Method for detecting software vulnerabilities by code injection
The method addresses the limitations of existing vulnerability detection methods by systematically identifying exploitable injections in software, ensuring comprehensive and efficient detection of vulnerabilities and enhancing software security.
Patent Information
- Application Number
- PCT/FR2024/051654
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-12
- Filing Date
- 2024-12-12
- Publication Date
- 2025-06-19
AI Technical Summary
Existing methods for detecting software vulnerabilities through code injection are either too costly in computational time and resources, or they are experimental and non-exhaustive, failing to provide formal proof of exhaustiveness and thus inadequate guarantees of software security.
A method is proposed that systematically and exhaustively identifies exploitable injections in an executable sequence by determining entry points, patterns, injection points, and associated constraints, and then determining the set of exploitable injections that respect the language and constraints.
This method allows for rigorous, exhaustive, and systematic identification of injections that can exploit vulnerabilities in software, thereby improving security guarantees by ensuring that all potential exploitable injections are identified without the need for exhaustive testing.
Smart Images

Figure FR2024051654_19062025_PF_FP_ABST
Abstract
Description
Description Title: Method for detecting software vulnerabilities by code injection Technical field
[0001] This disclosure relates to the field of computer security and more specifically to the identification of software vulnerabilities by code injection. Prior art
[0002] Security analysis is a crucial aspect to consider during the development and operation of software. For example, the specification, development, and testing phases of software may include a security analysis to identify and correct potential vulnerabilities in the software before it is put into operation. In another example, software that is being operated may be subject to security audits to detect and correct residual vulnerabilities in the software.
[0003] Identifying software vulnerabilities is particularly essential for IT systems in sensitive environments, such as the banking or energy sectors, or for detection and alert systems, for example. Such vulnerabilities can affect all types of software, such as applications, operating systems, or even networks.
[0004] Among the classes of software vulnerabilities, injection vulnerabilities pose a considerable risk to software. An injection vulnerability is a flaw, for example in the source code of a software program, which can be exploited by injecting portions of code, or scripts, during execution, in order to modify the software's behavior. Thus, exploiting an injection vulnerability in software can lead to distorting the execution of the underlying program. Typically, we can consider a web application including an authentication form allowing users to authenticate themselves using a username and password, for example to access a web page.When such an application is executed, a server retrieves the information (here, the username and password) entered on the authentication form in order to construct a query to query a database storing the data of (legitimate) users of the web page and to validate or not the authentication. A malicious individual could then, for example, provide a deliberately malformed password, so as to modify the semantics of the query allowing the database to be queried, and thus be authenticated without knowing a valid password, or even request actions from the database that are supposed to be forbidden, such as deleting data.
[0005] A software injection vulnerability can therefore be present when the software relies on data external to the source code, for example entered by the user, which is used to construct queries to another computer system (e.g., another program, software or even a remote database). This type of injection vulnerability is one of the most frequently encountered classes of vulnerabilities, particularly for web applications. In addition, injection vulnerabilities can potentially affect all programming languages and / or network protocols used, e.g. XSS, SQL, NoSQL, LDAP etc.
[0006] Solutions exist to detect or identify such software vulnerabilities through injections. Examples include vulnerability scanners that test software based on defined exploits, or databases that compile lists of known vulnerabilities based on the nature and version of each software.
[0007] In particular, two types of solutions are generally known: directory-based vulnerability detectors, which check the applicability of known vulnerabilities in an executable system (typically, a source code snippet). Since such detectors are based only on known vulnerabilities (which depend on the directory content), they are neither exhaustive nor proactive. Moreover, directory-based vulnerability detectors must be regularly updated to take into account newly discovered vulnerabilities. random data testing (or "fuzzing"), which tests an executable system by injecting random data into the system's user input. However, since the test space is considerable (or even virtually infinite), tests are in practice implemented on representative samples of the most critical cases.Such a solution is therefore not exhaustive, as potentially critical injections may have escaped testing. Furthermore, detecting vulnerabilities requires considerable execution times.
[0008] More generally, most known solutions for identifying vulnerabilities through software injections remain either too costly in computational time and resources to be implemented on most software, or experimental and non-exhaustive. Indeed, most existing solutions have a common approach of testing whether a particular injection (often already known or listed) can exploit the vulnerability of a software. Such an approach is never exhaustive in practice in that the set of potential injections in a program is infinite. Moreover, an exhaustive identification of potential injections in a program is made difficult in existing approaches by the diversity of natures, specificities and languages of software.Thus, existing approaches to identifying software vulnerabilities by injection do not provide formal proof of exhaustiveness and therefore a satisfactory guarantee of software security. Summary
[0009] This disclosure improves the situation.
[0010] A method implemented by computer means is provided for determining a set of exploitable injections in an executable sequence, the executable sequence using a language, the method comprising: a) determining at least one entry point of the executable sequence, b) determining, from said entry point, a pattern comprising at least one injection point of the executable sequence, the pattern respecting the language of the executable sequence, c) for each injection point, extracting a set of constraints associated with the injection point, d) determining, for each injection point, the set of exploitable injections from the pattern to which the injection point belongs, the associated set of constraints and the language.
[0011] Therefore, the proposed method advantageously allows a rigorous, exhaustive and systematic identification of injections that can exploit a vulnerability of an executable sequence, such as a program, a script or more generally a block of executable computer code. The method thus makes it possible to improve the security guarantees of a software with respect to vulnerabilities by injections.
[0012] Such identification is in particular possible regardless of the executable sequence considered and / or the language used by the executable sequence, in that the method proposes to formalize the executable sequence and its language, by considering the pattern, the injection points and the constraints in a formal manner.
[0013] To this end, the method proposes to determine injections that can be used to complete the executable sequence at its injection points, taking into account both the pattern, its syntax respecting the language and the constraints imposed on the data that can be entered at the injection points. The proposed method then performs an intersection of several elements allowing the systematic and exhaustive identification of exploitable injections at each injection point of the executable sequence.
[0014] The proposed method can therefore be advantageously launched (or tested) on an executable sequence by developers or code auditors, as part of a development, a test, or even a software audit. In particular, the proposed solution allows users who are not experts in cybersecurity to identify and list potential vulnerabilities by injections of the software they use or develop. In addition, such identification and listing can be exhaustive without having to test or compare given injections in the pattern in isolation.
[0015] An executable sequence may refer to a block or snippet of source code, or the contents of a script. The executable sequence may consist of a successive and logical sequence of lines of code. The executable sequence may be integrated into an executable system, which may refer to the system (hardware and / or software) implementing the execution of the executable sequence.
[0016] A language can refer to a programming language followed by the executable sequence, or to the programming language in which the executable sequence is coded. For example, the executable sequence can correspond to a portion of code written in Python and following the MySQL language. In particular, the language can also include an underlying grammar, allowing the definition of a language syntax.
[0017] An entry point may refer to a portion of the executable sequence in which input data of the executable sequence are present. In particular, these input data may occur in the entry point directly (in other words, the entry point directly integrates injection points) or indirectly (for example, by going up the tree of the executable system, the entry point calls variables and / or functions which themselves integrate injection points). When the input data satisfy the language and any other criteria, such as filtering constraints (for example, in terms of size, values, nature, etc.), such input data is said to be filtered and allows the completion of the pattern of the executable sequence.
[0018] An injection point may refer to a portion of the executable sequence that includes input data. Such input data may, for example, be data entered by a user of the executable sequence. Thus, when the executable sequence is executed, the input data—possibly after filtering—is integrated at the injection points so as to interpret the program corresponding to the executable sequence.
[0019] By a language-respecting pattern, it is understood that the pattern respects a syntax of the language, defined in particular by the grammar underlying the language. In other words, the elements (or symbols) forming the pattern are constructed and follow one another in such a way as to be syntactically interpretable in relation to the language.
[0020] A set of constraints associated with the injection point may refer to constraints or conditions imposed on the input data that can be validly injected at the injection point. Such constraints can be interpreted as filtering conditions (additional to the syntax conditions imposed by the language) allowing the data entered at each injection point to be filtered. For example, the set of constraints may include constraints related to the format, nature, size or value of the input data. For example, the data must be of type "string" or "float", the string entered at the injection point must not exceed a maximum number of characters, the numeric value entered at the injection point must belong to a given range of values, etc.Such conditions may be imposed by the developer via the contents of the executable sequence.
[0021] A set of exploitable injections refers to the set of data that can be injected at each injection point, so as to complete the pattern and allow the execution of the executable sequence. Such a set of exploitable injections thus includes injections (Le., in the form of data) that are both syntactically correct with respect to the language and that respect the proper filtering conditions of the executable sequence imposed by the set of constraints. In particular, the set of exploitable injections may correspond to an exhaustive set of injections allowing the injection point of the pattern to be completed by satisfying both the language and the constraints associated with the injection point.
[0022] The features set out in the following paragraphs may, optionally, be implemented, independently of each other or in combination with each other:
[0023] In one embodiment, the language is formed of symbols and constructed according to a set of rules, the pattern is constructed by a succession of symbols respecting said set of rules, and the set of exploitable injections is determined in step d) by the generation of generated symbols injectable into the pattern, said generated symbols being injectable at the level of the at least one injection point of the pattern, from the symbols of the pattern and respecting the set of rules and the set of constraints.
[0024] Therefore, the proposed method advantageously allows to formally construct, and therefore exhaustively, exploitable injections in the pattern in the form of a succession of symbols generated from the symbols of the pattern while respecting the language and the constraints. The proposed method thus allows to determine the nature of the exploitable injections while abstracting from the language used or the executable sequence considered: the method can be applied regardless of the language and the executable sequence. Indeed, the language is considered at a formal level, at the scale of the symbols composing it and the rules describing it.
[0025] A symbol refers to a unitary element of the language. For example, in the modern Latin alphabet, a symbol can correspond to each of the twenty-six letters of the alphabet. In another example, for an executable sequence conforming to XML language, the elements <objet>And< / objet>to define an object can be interpreted as symbols of the XML language. Alternatively, the elements <, o, b, j, e, t, >, <, / , o, b, j, e, t and > can be defined as symbols of the XML language. Any element of the language can then be obtained by a succession of symbols.
[0026] A set of rules refers to a set of syntactic constructions that generate the language (i.e., the set of elements or words of the language). In particular, the set of rules allows the definition of a grammar (or initial grammar) of the language. In other words, the set of rules allows the construction of successions of symbols that generate the entire language.
[0027] By pattern-injectable generated symbols, we mean language elements constructed to respect the set of rules and the set of constraints, so as to form an exploitable injection at the considered injection point. In other words, the set of exploitable injections determined by the proposed method corresponds to a set of injectable generated symbols formed from the rules and constraints of the injection point.
[0028] In one embodiment, the method further comprises, after step b): b1) determining, for the pattern of the executable sequence, a plurality of radicals corresponding to invariable portions of the pattern, two neighboring radicals of the pattern being separated by an injection point, and in which step d) comprises: d10) an iterative comparison of each radical of the plurality of radicals of the pattern with the set of rules.
[0029] Therefore, the proposed method advantageously allows to formalize (and therefore systematize) the determination of exploitable injections by relying on the formal structure of the pattern of the executable sequence. Indeed, independently of the language and the content of the executable sequence, radicals and injection points of the pattern can be identified. The method can therefore be implemented systematically on any executable sequence and for any language used, as long as a set of rules and symbols are associated with such a language.
[0030] In order to determine the set of exploitable injections, the method proposes to compare each radical of the pattern with the set of rules. Such an iterative comparison then makes it possible to specify and constrain the form of the exploitable injections with respect to each radical of the pattern. Indeed, an injection point being included between two radicals, an exploitable injection must at least respect the syntax of the language constrained by the radicals on either side of the injection point, so that the entire pattern formed by the succession of radicals and injection points is syntactically correct (i.e. respects at least the set of rules).
[0031] By an iterative comparison between a radical and the set of rules, it can be understood an intersection operation between each radical of the pattern and the set of rules considered. The radicals are notably considered in their order of appearance within the pattern. Such a comparison can be understood as a specification or specialization of the set of rules considered with respect to each radical compared. In particular, such a comparison can be implemented with the set of rules initially considered but also with other potential additional rules resulting from comparisons with the previous radicals. In other words, each comparison takes into account the results of the comparisons at the previous iterations (Le., with the radicals positioned upstream of the radical considered in the pattern).
[0032] In one embodiment, each iteration of step d 10) for each radical considered is associated with the generation of a set of generated symbols comprising the radical considered.
[0033] In one embodiment, the method further comprises: d11) iteratively specifying the set of rules to each radical of the plurality of radicals of the pattern resulting in sets of specified rules, each set of specified rules generating the set of generated symbols comprising the radical considered.
[0034] Therefore, at each comparison of a radical with the set of rules, the method advantageously proposes to specify the language used in relation to each radical (via an intersection operation), so as to generate (or engender) the elements of the language comprising this radical. In other words, at each iteration, the space of the language considered (i.e. the set of combinations of symbols that can be generated) decreases since the successions of acceptable symbols are specified (therefore constrained). Such a specification of the language is allowed by the addition (or specification) of specific rules defining the syntax of the language allowing the generation of the given radical. In other words, the set of rules specified for each radical grows in relation to the previous set of rules and integrates the latter. In other words, if Gj denotes a set of rules at the j-th iteration (j being a natural number greater than 1), we have dim(Gj) > dim (Gj-i), where dim corresponds to the dimension of the set of rules and Gj includes Gj-i.
[0035] By a specified rule set, we mean the initial rule set to which new rules are added to specialize the initial rule set to the radical under consideration. Thus, for each new iteration on a radical, the considered rule set can correspond to the rule set resulting from the previous iteration, so that the grammar of the language gradually specializes for the set of radicals in the pattern.
[0036] In one embodiment, the radicals of the pattern are formed by a succession of symbols, and step d) further comprises: d 101 ) associating the symbols of the radicals of the pattern with first labels, the symbols of the same radical being associated with the same first label, the respective symbols of the neighboring radicals being associated with successive first labels according to a first criterion, d 102) associating the symbols forming the set of rules with the same initial label, and in which each set of generated symbols and each set of specified rules are associated with generated labels depending on the first labels of the radicals of the pattern and the initial label of the set of rules, step d) comprising: d12) an iterative update of the generated labels according to a result of each iteration of steps d 10) and d11 ), d 13) an identification of the elements of the sets of generated symbols having generated labels different from the first labels.
[0037] Advantageously, the method allows one to navigate successive iterations by labeling the symbols of the manipulated elements (here, radicals and rules). Considering the generated labels associated with the generated symbols then allows one to systematically and formally identify (independently of the symbols in question) potential exploitable injections.
[0038] Identifying elements with generated labels different from the first labels advantageously allows identifying elements that do not correspond to the stems. Such elements can then include exploitable injections.
[0039] Labels can refer to values or annotations that allow us to identify, order, sequence or index the different symbols of an element (e.g., a pattern, a radical, a rule, a constraint, etc.). The values of labels can belong to an ordered set of objects. For example, labels can correspond to numerical indices, expressed for example by natural numbers. In another example, labels can correspond to letters of an alphabet. Successive labels according to a first criterion can then refer to values ordered according to a natural ordering. For example, if the labels correspond to numerical indices, two successive labels can correspond to two successive natural numbers. In another example, if the labels correspond to letters of the modern Latin alphabet, two successive labels can correspond to two letters following one another in this alphabet. The first criterion in these cases corresponds to a step of one between two successive indices. The first criterion can also be a parity criterion. For example, successive first labels according to a first criterion can correspond to successive first numerical indices of the same first parity. By first parity, we mean an element among the even character or the odd character. For example, if the first parity corresponds to the even character, successive indices of the same first parity designate the indices 0, 2, 4, 6..., 2n, , where n is a natural number.
[0040] In one embodiment, step d) comprises: d1) from the pattern and the language, determining a set of executable injections, d2) filtering the set of executable injections according to the associated set of constraints so as to obtain the set of exploitable injections, the set of exploitable injections being of smaller dimension than the set of executable injections.
[0041] Therefore, the method proposes to determine the exploitable injections by specifying and restricting the language with respect to both the syntax rules defining the language (step d1)) and the constraints imposed at each injection point (step d2)). Thus, step d1) makes it possible to obtain a first set of candidate injections, called executable injections, satisfying the sets of rules specified from the pattern and the language. Step d2) then makes it possible to filter these executable injections according to the set of constraints so as to obtain the set of injections that are actually exploitable.
[0042] By a set of executable injections, we mean a set of injections that are grammatically (or syntactically) correct with respect to the pattern and the language. In other words, an injection is an executable injection when its execution does not generate a syntax error in the executable sequence. However, such an executable injection does not always allow the execution of the executable sequence, in that it may not satisfy the developer's constraints. For example, an input data corresponding to an executable injection may be syntactically correct given the pattern stems, but its size or type may exceed the constraints imposed at the corresponding injection point, so that the injection is not actually exploitable.However, since an exploitable injection necessarily respects the rules of the language, the set of exploitable injections is necessarily included in the set of executable injections. Thus, the method advantageously makes it possible to exhaustively identify injections that are potential candidates for being exploitable injections in the executable sequence.
[0043] In one embodiment, the determination of the set of executable injections results from steps d10), d11), d101), d102), d12) and d13).
[0044] Therefore, the set of executable injections is the result of the intersection of all radicals in the pattern with the set of rules iteratively specialized for each radical. In other words, when all radicals in the pattern have been compared (or intersected) with the set of rules, the resulting set of specified rules allows generating the set of injections executable. In particular, at the end of steps d10), d11), d101), d102), d12) and d13), consideration of all constraints is not yet implemented, so that executable injections may not satisfy such constraints.
[0045] In one embodiment, the filtering of step d2) is implemented from an expanded set of constraints including said set of constraints.
[0046] Advantageously, the use of an expanded set of constraints allows for faster execution of the method. Indeed, taking into account the constraints imposed by the developer can sometimes lead to a large number of operations (elementary or not) in order to compare all executable injections with such constraints. Thus, a relaxation (or overestimation) of the constraints integrates the possibility of having false positives for the benefit of a better execution time. Nevertheless, such a relaxation of the constraints always excludes the presence of false negatives: the proposed method therefore remains advantageously systematic and exhaustive and the guarantee of security of the software with respect to code injections remains assured.
[0047] Such an expanded set of constraints may be associated with a second set of regular expressions that is over-approximate with respect to a first set of regular expressions associated with the set of constraints.
[0048] In one embodiment, the method comprises, after step c), : c1) converting, for each injection point, the associated set of constraints into a first set of regular expressions, and in which the set of exploitable injections corresponds to the executable injections filtered according to a set of regular expressions among the first set of regular expressions and a second set of regular expressions over-approximate with respect to the first set of regular expressions.
[0049] In one embodiment, each regular expression of the set of regular expressions is formed by a succession of symbols, and step d2) further comprises: d21) associating the symbols of each of the regular expressions of the set of regular expressions with second labels, the symbols of the same regular expression being associated with the same second label, the respective symbols of two neighboring regular expressions in the set of regular expressions being associated with successive second labels according to a second criterion, d22) comparing the second labels of each of the regular expressions of the set of regular expressions with the labels generated at the end of step d12), d23) determining the set of exploitable injections from the elements of the sets of generated symbols associated with the generated labels corresponding to the second labels.
[0050] Therefore, the set of exploitable injections can be advantageously obtained by a simple intersection between the regular expressions associated with the set of constraints and the set of generated symbols corresponding to the set of executable injections. such intersection is advantageously simple in that it is based on a comparison of the labels associated on the one hand with the injection points (via regular expressions) and on the other hand with the elements of the sets of symbols generated.
[0051] Thus, if step d13) makes it possible to determine the set of executable injections (Le., the sets of generated symbols whose generated labels do not correspond to the first labels), step d23) makes it possible to determine the set of exploitable injections (Le., the set of generated symbols whose generated labels correspond to the second labels), at the end of the constraint filtering step.
[0052] As defined above, labels may, for example, designate numerical indices. In particular, second successive labels according to a second criterion may correspond to second successive numerical indices of the same second parity. By second parity, we mean the remaining element among the even character or the odd character (not corresponding to the first parity). For example, if the second parity corresponds to the odd indices, successive indices of the same second parity designate the indices 1, 3, 5, 7,..., (2n+1), where n is a natural number.
[0053] In one embodiment, the method further comprises: e) from the set of exploitable injections, implementing a vulnerability analysis of the executable sequence.
[0054] Therefore, the proposed method advantageously makes it possible to exploit the set of exploitable injections determined in order to analyze or evaluate the executable sequence. Such an analysis or evaluation can for example be implemented in the context of an analysis of the risks of compromise of the executable sequence or of the software integrating such an executable sequence. Such a vulnerability analysis can advantageously be implemented by the developers at the time of development of the software integrating the executable sequence. The determination of the set of exploitable injections advantageously allows a developer or an individual who is not an expert in cybersecurity to identify potential vulnerabilities of the executable sequence.
[0055] Vulnerability analysis can, for example, be understood as a study or classification of all exploitable injections determined in the executable sequence according to given risk criteria. Vulnerability analysis can also be understood as an assessment of the degree of vulnerability of the software or system integrating the executable sequence, for example based on the number or nature of the exploitable injections determined.
[0056] In one embodiment, the method further comprises: e') assigning vulnerability indices respectively associated with each exploitable injection of the set of exploitable injections, said vulnerability index being determined from the pattern (M) completed by the exploitable injection at the corresponding injection point.
[0057] A vulnerability index refers to a value or measurement quantifying the degree of compromise of the system incorporating the executable sequence if the corresponding exploitable injection were to be actually exploited. Such a vulnerability index may, for example, be defined by the user of the process (the developer, a risk analyst, or another individual) according to a predefined risk scale. Such a vulnerability index may, for example, depend on the nature of the executable action resulting from the exploitable injection or the level of privilege associated with such an action. For example, an exploitable injection leading to the deletion or alteration of information or the distortion of a function may be considered to have a high vulnerability index, in that it distorts and diverts the result of the executable sequence.
[0058] According to another aspect, there is provided a computer program comprising instructions for implementing all or part of a method as defined above when this program is executed by a processor. According to another aspect, there is provided a non-transitory recording medium, readable by a computer, on which such a program is recorded. Brief description of the drawings
[0059] Other features, details and advantages will become apparent upon reading the detailed description below, and upon analyzing the attached drawings, in which: Fig. 1
[0060] [Fig. 1] shows a set of systems involved in the detection method proposed according to one embodiment. Fig. 2
[0061] [Fig. 2] shows steps of a detection method proposed according to one embodiment. Fig. 3
[0062] [Fig. 3] shows an example of injection into a database access query. Fig. 4
[0063] [Fig. 4] shows an example of a grammar defining a data format according to one embodiment. Fig. 5
[0064] [Fig. 5] shows an example of an executable sequence according to the embodiment of Figure 4. Description of the embodiments
[0065] Reference is made to Figure 1. Figure 1 illustrates a set of systems 1, 2, 3, 4 in the context of detecting code injections into an executable sequence. Such a set may include a detection system 1, a computer system 2, a user system 3 and optionally a database system 4.
[0066] The computer system 2 may correspond to an executable system (or environment), or runtime, configured to execute one or more programs. The computer system 2 may for example include software (for example, system or application), which comprises one or more programs. The computer system 2 may for example correspond to or be integrated into a server or a module as part of a service or an application. The functionalities and services of the computer system 2 are implemented by the execution of one or more predefined programs, in the form of scripts. Such programs include in particular one or more sequences of instructions, or executable sequences (eg, blocks of code). The execution and interpretation of these executable sequences allows the computer system 2 to implement one or more functionalities.Examples of execution and interpretation of executable sequences by the computer system 2 will be detailed later in the description. Such execution of one or more executable sequences may be planned, pre-programmed or pre-configured, for example as part of automatic and / or periodically executed functionalities. Such execution may also be punctually triggered, for example when data is entered into the computer system 2. In order to implement services and functionalities, the computer system 2 may comprise at least one processing unit 21 and one memory unit 22. The processing unit 21 may in particular correspond to the software unit executing executable sequences of one or more programs. The computer system 2 may also include a communication interface 20 configured to transmit and / or receive data with other computer systems and entities.For example, the computer system 2 may receive input data via the communication interface 20 and / or transmit data received and / or processed by the unit 21 to a database system 4. In FIG. 1, such a database system 4 is distinct from the computer system 2 but alternatively, such a database system 4 may be integrated into the computer system 2 (typically, in the case where the computer system 2 is a server).
[0067] The data entered into the computer system 2 may come from a user system 3. Such a user system 3 may, for example, correspond to a machine, a computer, a calculator, a mobile device, a terminal, a vehicle or any other computerized device supporting software and operating systems, which can be manipulated or accessed by a user. Such a user system 3 is connected to the computer system 2, for example by a wired, wireless connection and / or a communication network, so as to transmit data from the user system 3 to the computer system 2. For example, the user system 3 may transmit data which is manually entered into it by the user (eg, in the case of a user filling out an online form on his private computer), which are captured by the user system 3 (for example, in the case of a vehicle or machine capturing an ambient temperature) and / or which are generated by the user system 3 (eg, in the case of a motor vehicle generating system data such as vehicle temperature or remaining fuel level). The user system 3 may include at least one communication interface 30 configured to transmit the user data to the computer system 2.
[0068] The detection system 1 is a system proposed for implementing a method for detecting code injections in a computer system 2. In particular, the proposed detection system 1 is configured to determine a set of executable code injections for a given executable sequence of the computer system 2. For this, the detection system 1 may include an input interface 10, a processing unit 11, a memory unit 12 and an output interface 13. The input interface 10 may be configured to receive data, information, scripts from other systems such as the computer system 2 but also the user system 3 or the database system 4. In particular, the input interface 10 is configured to receive data relating to at least one executable sequence to be analyzed by the detection system 1. Such data will be detailed later in the description.The processing unit 11 is configured to process such data relating to the executable sequence, according to a detection method which will be detailed in FIG. 2, so as to identify a set of executable injections in such an executable sequence. The memory unit 12 is configured to store data and information from the input interface 10 and / or the processing unit 11. The processing unit 11 may comprise several processing modules 1 11 , 112, 113, 114, 115, 116 (or 1 11-116), the respective operations of which will be detailed in FIG. 2. Finally, the output interface 13 is configured to transmit the data at the end of processing by the detection system 1. Such a transmission, not shown in FIG. 1, may be intended for the computer system 2, the user system 3, the database system 4, one or more other third-party systems and / or a user via a human-machine interface.
[0069] Examples of execution and interpretation of executable sequences by the computer system 2 are now described. An executable sequence can be defined as an ordered succession of executable instructions allowing or contributing to the implementation of a program. Such an executable sequence typically corresponds to a block of code included in a source code of a program. Such an executable sequence is notably written in a given programming language, more simply referred to as language L. Such a language L can for example correspond to PHP, Python, Java, C, HTML or many other known languages for example. Such a language L can also correspond to a language specifically developed for a computer system 2 or a specific application. In particular, the programming language L used in the executable sequence is defined by a grammar.Such a grammar corresponds to symbols combined according to rules, called grammar rewriting rules, and which makes it possible to generate all the elements (or words) of the programming language L.
[0070] For example, with reference to Figure 3, there is illustrated an executable sequence of a computer system 2 written in PHP language which makes it possible to execute user requests (“Request” in Figure 2) in SQL language. Such an executable sequence is a program configured in particular to authenticate a user who enters a user name (login?) and a password (pass?) via a user system 3. The execution of the executable sequence consists of accessing a database 4 storing the user names (login) and passwords (pass) authenticated users and to return an error or a connected result to the user. In another example, with reference to Figure 5, an executable sequence written in Python is illustrated. Such an executable sequence is a program configured in particular to check data (representing id, temperature, fuel, failure parameters) at the input of the program and to store this data in a data format (in this case, the XML data format), in the form of a log file.
[0071] In the context of this description, the executable sequences considered involve variables that can integrate input data. In other words, the executable sequence includes entry points, which are portions of the executable sequence involving, directly or indirectly, variables corresponding to user input data. For example, with reference to the example in Figure 3, the instruction line reproduced below from the executable sequence is an entry point: $r = db_query(“select * from user where “login = '$login' and pass = '$pass'”)
[0072] Indeed, the db_query function includes login and pass variables that take user input data directly as input (which are replaced at the $login and $pass elements when the executable sequence is executed).
[0073] Referring to the example in Figure 5, lines 16, 17, 18, 19, 20, and 21 correspond to entry points of the executable sequence. Indeed, for example, the variables datai, data2, and data3 directly take user data as input (a temperature, fuel, and failure value respectively); the query object indirectly takes user data from the variables datai, data2, and data3 as input.
[0074] Therefore, the executable sequence may be potentially vulnerable to code injections at such entry points. Such code injections can be defined as elements, exploiting the programming language L used, allowing to distort the execution of the executable sequence or its result when exploited at the entry points of the executable sequence. With reference to Figure 3, examples of input data (client inputs) formed by pairs (login, pass) are illustrated. In the first example of Figure 3, the first pair (random, 123) returns an error result: such input data does not correspond to an authenticated user (according to the database system 4) and the authentication test is correctly executed.In the second example of Figure 3, the second pair (john, 12gh3) returns a connected result: such input data correspond to an authenticated user (according to database system 4) and the authentication test is correctly executed. However, in the third example of Figure 3, the third pair (yes, x' or '1 '='1) returns a connected result: such input data do not correspond to an authenticated user (according to database system 4). In the latter case, the input data embedded in the executable sequence forms the expression pass = 'x' or '1'=1 ', which is an assertion that is always true (it results in a boolean of type TRUE). Thus, the authentication is incorrectly validated. This last case illustrates an example of code injection. in the executable sequence of Figure 3, exploiting the grammar of the SQL programming language L of the query.
[0075] In another example, referring to Figure 5, line 17 [data2 = " <carburant> %s < / carburant> " % fuel] of the executable sequence involves an input data (which is a fuel data). A code injection could for example consist of replacing the element [%s] of line 17 by the set: [80 " " <panne> there is a breakdown< / panne> " #], the # element indicating that the rest of the instructions is commented out (and is therefore no longer part of the executable sequence). Such an injection would then modify the behavior of the executable sequence, which would then systematically consider a query data containing a temperature value, a fuel value = 80 and an indication "there is a fault" (even though with the temperature value and the fuel value = 80, no fault is normally reported). The operation of the computer system 2 is then distorted and the data stored in the XML format then has erroneous content.
[0076] The proposed detection system 1 and method of Figure 2 aim to detect all possible code injections for any executable sequence.
[0077] Reference is now made to Figure 2. Figure 2 details steps of a method for detecting code injections of an executable sequence, for example of a computer system 2, implemented by a detection system 1 as illustrated in Figure 1. An example based on Figures 4 and 5 will serve as a guide for the implementation of each step. In this example, the computer system 2 considered may be a server operating as part of a log service and configured to receive data from one (or more) user systems 3 corresponding to a vehicle. Such a user system 3 is configured to send (for example, periodically) engine temperature, fuel level and detected fault presence data to the computer system 2.The program (and in particular the executable sequence) considered aims to store this data in XML format in a database system 4 corresponding to a storage space. Such a storage space can in particular be stored at the level of the server, the vehicle or even a remote third-party entity.
[0078] The person skilled in the art will understand that this example is not intended to be limiting and that the steps of figure 2 apply to any (or a plurality of) computer systems 2 as shown diagrammatically in figure 1 and / or to any (or any plurality of) executable sequences of such a computer system 2.
[0079] At a step 200, data relating to the executable sequence is received, for example by the input interface 10. Such received data may include the elements of the executable sequence, for example the script containing the successive lines of instructions. In particular, the executable sequence is written in a programming language L defined at least by symbols and a (formal) grammar. For example, at step 200, the script of FIG. 5 detailing the executable sequence (which is a succession of instructions on lines 1 to 22) may be received. The data received at step 200 may include information relating to the programming language L used in the executable sequence, for example data relating to the set of symbols, to the rules of the grammar, such data allowing to define the way of generating the executable sequence. For example, at step 200, the script of figure 4 defining the grammar allowing to generate the set of data in XML format can be received.
[0080] At a step 210, from the data received at step 200, the executable sequence can be processed, for example by an executable sequence processing module 111. For this, an abstract interpretation method or even a symbolic execution can be implemented on the executable sequence, for the given programming language L. Such processing integrates a step 211 and a step 212.
[0081] In step 211, the elements making it possible to characterize and constrain the user data in the executable sequence are identified and interpreted. In other words, the constraints C on the variables involved in the executable sequence are identified. For example, the constraints C relating to variables of an executable sequence may relate to an input data type, an input data format, an input data size, a value or a set of values that can be taken by the input data and / or a minimum and / or maximum number of characters for example.
[0082] With reference to Figure 4, step 211 includes for example identifying that the variables of the executable sequence are id, temperature, fuel and failure and that these variables are respectively an integer, a float, an integer and a character string. In addition, with reference to Figure 5, the analysis of the executable sequence in step 211 includes identifying that constraints C are associated with the variables temperature, fuel, failure. In lines 3 and 13 of the executable sequence, a constraint C associated with the variable temperature is that its value is less than 200.0. In lines 7 and 14 of the executable sequence, a constraint C associated with the variable fuel is that its value is between 0 and 100. In lines 10 and 15 of the executable sequence, a constraint C is associated with the variable fuel: it is a character string strictly not exceeding 100 characters.
[0083] Step 211 may include translating the input variables of the executable sequence and their constraints C into a formal language (denoted L_x for an element x), for example in the form of rational expressions (also referred to as regular expressions) for the variables. For example, with reference to the example of Figures 4 and 5: the variable id may be translated by L_id = E* the variable temperature may be translated by L_temperature = ([0-9] | [1 -9][0-9] | 1 [0-9][0-9]) (• [0-9]+)? the variable fuel may be translated by L_carburant = [0-9] | [1 -9][0-9] | 100 the variable breakdown may be translated by L_panne = E{0,99}
[0084] At the end of step 211, the input variables of the executable sequence have been identified, characterized and formalized in formal language.
[0085] In step 212, at least one entry point of the executable sequence is identified. For this, the processing unit 11 can locate the intervention of the variables in the executable sequence and identify the lines directly or indirectly calling such variables. In particular, an entry point of the executable sequence can correspond to the elements forming a query in the executable sequence, such a query calling one or more identified variables. Such an entry point can correspond to the portion (or portions) of the executable sequence in which the input data are integrated during the execution and interpretation of the script. For example, with reference to the executable sequence of FIG. 5, lines 16 to 19 are entry points. Line 21 also includes an entry point. In the following and to simplify the description, the processing of a single entry point will be detailed.However, the rest of the process can be implemented on each entry point. The example of the entry point on line 21 is illustrated, in that the query element % (id, data) calls the four variables id, temperature, fuel and failure.
[0086] In step 212, the query entry point is identified and translated into formal language: L_request = <donnees> <noeud> <id> 0< / id> <noeud> <id> .* < / id> <temperature> .* < / temperature> <carburant> .* < / carburant> <panne> .* < / panne> < / noeud> < / noeud> < / donnees>
[0087] Indeed, the query element takes two variables: id (%d of line 20) and data (%s of line 20), the latter itself taking the values of the variables temperature, fuel and failure. The elements noted ".*" then correspond to the location, in the formal language, of the user data in the query entry point.
[0088] Step 212 then makes it possible to identify, in the entry point thus translated, a pattern M, corresponding to a succession of symbols ordered in the language L used by the executable sequence. The pattern M of the element requete corresponds to the symbols forming the element L_requete. Here, the pattern M comprises twenty-one symbols: <donnees>, < node>, <id> , 0, < / id> , <noeud> , <id> , .*, < / id> , <temperature> , .*, < / temperature> , <carburant> , .*, < / carburant> , <panne> , .*, < / panne> , < / noeud> , And< / donnees> . To simplify the rest of the description, the notation L_requete can designate the entry point requete and its pattern M in formal language.
[0089] At a step 220, the pattern M identified at step 212 is processed, for example by a pattern M processing module 112. Such processing integrates a step 221 and a step 222 making it possible to decompose the pattern M into radicals and injection points. For this, an abstract interpretation method or even a symbolic execution can be implemented on the pattern M, for the given programming language L.
[0090] At step 221, radicals are identified in the pattern M. Such radicals correspond to fixed parts of the pattern M, which cannot be distorted. In other words, such radicals are a succession of invariable symbols of the pattern M. For example, the pattern L_requete has five distinct radicals: Radical 1: A <donnees> <noeud> <id> 0< / id> <noeud> <id> Radical 2 :< / id> <temperature> Radical 3 :< / temperature> <carburant> Radical 4 :< / carburant> <panne> Radical 5 :< / panne> < / noeud> < / noeud> < / donnees> $
[0091] The marker " A » represents the beginning of the pattern M and the marker “$” represents the end of the pattern M. Such radicals are then ordered.
[0092] In step 222, injection points are identified in the pattern M. Such injection points correspond to the variable parts of the pattern M, which can potentially distort the content of the pattern M, in particular depending on the content of the injected input data. These are also the points at which code injections occur in the executable sequence. In other words, the injection points are the portions of the pattern M at which input data can complete the pattern M during the execution and interpretation of the executable sequence. In particular, the radicals correspond to the portions of the pattern M separated by the injection points. For example, the pattern L_requete has four injection points, represented by the markers “.*”. The four injection points represented in the pattern L_requete correspond in particular to the respective entry points of the data elements L_id, L_temperature, L_fuel and L_panne.
[0093] Thus, at the end of step 220, the identified pattern M is decomposed into radicals and injection points.
[0094] In a step 230, the executable sequence is processed at the symbol level, for example by a symbol processing module 113 of the processing unit 11. In particular, in step 230, the symbols combined in the executable sequence according to the grammar of the language L used can be associated with labels. In the remainder of the description, the labels will correspond to numerical indices (or indexes). However, in other embodiments, the labels can correspond to other values or elements (for example, letters of a given alphabet). In this case, the notion of successive labels can be defined according to a first predefined criterion (for example, a given step in the ordered set of values that can be taken by these labels). The symbols are said to be labeled with the numerical index, so that the detection system 1 is configured to record a mapping between a numerical index and a symbol.Such a step 230 includes a step 231 of processing the symbols of the radicals determined in step 221, a step 232 of processing the symbols of the expressions of the variables determined in step 211 and a step 233 of processing the symbols of the rewriting rules of the grammar used by the pattern M.
[0095] In particular, the treatment of symbols depends on the definition of symbols within the language L used by the executable sequence. A symbol corresponds to a unitary, basic element of the language L. In the example of figures 4 and 5, a tagged element (for example <temperature>or ) is considered a symbol. Thus, the grammar for generating the entire L language as illustrated by the rules in Figure 4 allows these symbols (and more precisely non-terminal symbols) to be rewritten. In another embodiment, a symbol could be <, / , d, o, n, e, s or even >. In this case, steps 230 and following are adapted to such granularity and to the definition of the symbol in the L language used.
[0096] In step 231, the symbols of the radicals are associated with numerical indices of the same first parity. For example, all the numerical indices associated with the symbols of the radicals are odd. In addition, for each radical considered in the pattern M, the symbols of the radical considered are associated with the same numerical indices. For example, each of the radicals in the pattern L_requete can be written with numerical indices (indexed at the level of each symbol of the radical) in the following way: Radical 1: A <donnees> 1 <noeud> 1 <id> 1 0 1 < / id> 1 <noeud> 1 <id> 1 Radical 2 :< / id> 3 <temperature> 3 Radical 3 :< / temperature> 5 <carburant> 5 Radical 4 :< / carburant> 7 <panne> 7 Radical 5 : < / panne 9 < / panne> < / noeud> 9 < / noeud> 9 < / donnees> 9 $
[0097] In step 232, the symbols of the regular expressions translating the variables and their constraints C in the executable sequence are associated with numerical indices of the same second parity (distinct from the first parity). For example, all the numerical indices associated with the symbols of the regular expressions of the variables are even. In addition, for each regular expression associated with a variable considered, the symbols of the regular expression considered are associated with the same numerical indices. For example, each of the radicals of the elements L_id, L_temperature, L_fuel and L_panne can be written with numerical indices (indexed at the level of each symbol of the element) in the following way: - LJd = Z 2 * - LJemperature = ([0-9] 4 | [1 -9] 4 [0-9] 4 | [1 -2] 4 [0-9] 4 [0-9] 4 ) ( 4 [0-9] 4 +)7 - L_fuel = [0-9] 6 1 [1 -9] 6 [0-9] 6 1 1 6 0 6 0 6 L_panne = E 8 {0.99}
[0098] In step 233, the symbols forming the rewriting rules of the grammar are also associated with numerical indices. In particular, such a grammar, which is the initial grammar making it possible to generate, among other things, the pattern M considered in the executable sequence in the format considered (in figures 4 and 5, the XML), is associated with the same initial numerical index, for example 1. Thus, with reference to the rules of the grammar illustrated in figure 4, step 233 makes it possible to write: - S — > <donnees> 1 N < / donnees> 1 - N — > <noeud> 1 <id> 1 entire 1 < / id> 1 N < / noeud> 1 - N — > <noeud> 1 <id> 1 entire 1 < / id> 1 L < / noeud> 1 - N ^ NN - L — > LL - D — > D - D — > <temperature> 1 floating 1 < / temperature> 1 - D — > <carburant> 1 entire 1 < / carburant> 1 - D — > <panne> 1 string 1 < / panne> 1 - D — > <x> 1 entire 1 < / x> 1
[0099] The grammar thus labeled at the end of step 233 corresponds to an initial grammar (Le., a set of initial rewriting rules), noted Gi, which in particular makes it possible to generate the elements of the executable sequence in XML format, for example the pattern L_requete.
[0100] At a step 240, a grammar for generating all the data in the format of the executable sequence that contain the radicals of the pattern M is determined. Such a grammar is called an executable grammar. In other words, step 240 aims to formally and exhaustively determine all the elements for completing the pattern M in a syntactically correct manner. Data from such an executable grammar then contains the radicals of the pattern M as well as the data used to complete this pattern M.
[0101] For this, the initial grammar is used and progressively specified with respect to each of the radicals of the pattern M. In other words, step 240 is a step formed by iterations on each of the radicals of the pattern M, so as to generate a succession of grammars G2, G3, Gj from the rewriting rules of the initial grammar G1 (eg, the rewriting rules of figure 4). Each grammar Gj is specified (thus by generating a grammar Gj+1) so that each word (or element) generated contains the first j radicals of the pattern M in their order of appearance in the pattern M, (where j is a natural integer greater than 1).
[0102] Thus, at step 240, each iteration j will add new rules and new symbols to the grammar Gj, so that the resulting generated grammar Gj+1 will integrate these new rules and these new symbols. By the construction of each grammar and the definition of the numerical indices at step 230, each element (or terminal) of the grammar Gj is associated with a numerical index between 1 and (2j-1 ).
[0103] In particular, at step 240, new symbols for the grammar Gj+i can be created and named as follows, for a symbol X of the grammar Gj considered: <rad:x>: new symbol rewritten according to the same rewriting rules as the symbol X, but generating the j ème radical, <sufk:x>: new symbol rewritten according to the same rewriting rules as the symbol X, but generating the suffix j ème radical, as a prefix in position k of the new symbol, <prfk :X> : new symbol rewritten according to the same rewriting rules as the symbol X, but generating the prefix of j ème radical, as a prefix in position k of the new symbol, <aft:x>: new symbol rewritten according to the same rewriting rules as the symbol X, but generating elements associated with a numerical index (2j+1).
[0104] Thus, taking into account the new symbols that can be created, each element of the grammar Gj+i is associated with a numerical index between 1 and (2j+ 1 ).
[0105] The first iteration implemented during step 240 is now detailed, allowing the transition from the initial grammar Gi to a second grammar G2. For this, we consider the first radical:
[0106] Radical 1: [ <donnees> 1 <noeud> 1 <id> 1 0 1 < / id> 1 <noeud> 1 <id> 1 ].
[0107] Such a first iteration is special in that Radical 1 appears as a prefix to the generated data. In particular, for such a first iteration, the new symbol <prfk:x>does not appear, by definition, since no element comes upstream of Radical 1.
[0108] The first iteration begins with the (non-terminal) element S of the initial grammar (line 1 of figure 4). Such an element S can generate combined elements (words) containing or not the Radical 1 . It is then a question of constructing a second grammar G2 in which the new symbol is rewritten according to the same rewriting rules as S but containing the Radical 1 , that is to say the new (non-terminal) symbol <rad:s>.
[0109] The element S derives the first radical as a prefix with the first rule of line 1 of Figure 4: [S— > <donnees> 1 N < / donnees> 1 ]. Such a rule contains a left part (S) and a right part ( <donnees> 1 N < / donnees> 1 ) of rewriting.
[0110] The element <donnees> 1 in the first position of the right part generates by equality the prefix <donnees> 1 of Radical 1. The element <donnees> 1 of the first rule being equal to the prefix <donnees> 1 from Radical 1, the new first rule translates this element into <donnees> 2 , by incrementing by one the numeric index associated with the element <donnees> 1 . The (non-terminal) element N in the second position of the right part generates the suffix <noeud> 1 <id> 1 0 1 < / id> 1 <noeud> 1 <id> 1 of Radical 1, after the first element (prefix) of this Radical 1. The new first rule therefore translates the element N into<suf1 :N> . The element< / id> < / noeud> < / noeud> < / donnees> 1 of the first rule being generated after the Radical 1, the new first rule translates such an element into< / donnees> 3 , by incrementing by two the numeric index associated with the element< / donnees> 1 .
[0111] The following new first rule is therefore generated:
[0112] [ <rad:s>— > <donnees> 2 <suf1 :N>< / donnees> 3 ]
[0113] Such a new first rule introduces the new (non-terminal) element<suf1 :N> , which must be constructed from the rules of N in the initial grammar Gi (i.e., lines 2 to 4 of Figure 4), so that<suf1 :N> generates the suffix <noeud> 1 <id> 1 0 1 < / id> 1 <noeud> 1 <id> 1 of Radical 1 in first position of<suf1 :N> .
[0114] The (non-terminal) element N generates the suffix <noeud> 1 <id> 1 0 1 < / id> 1 <noeud> 1 <id> 1 of Radical 1 with the second rule: [N— >NN] (line 4 of Figure 4). Such a second rule contains a left part (N) and a right part (NN) of rewriting.
[0115] The element N in the first position of the right part generates the suffix <noeud> 1 <id> 1 0 1 < / id> 1 <noeud> 1 <id> 1 of Radical 1. Thus, the new second generated rule translates such an element into<suf1 :N> . The element N in the second position of the right part being generated downstream of the Radical 1, the new second rule generated translates such an element into <aft:n>.
[0116] The following new second rule is therefore generated:
[0117] [<suf1 :N> -►<suf1 :N> <aft:n>]
[0118] Such a new second rule introduces the new (non-terminal) element <aft:n>, which is constructed according to the same rewriting rules as the symbol N (therefore generates the same language L as N), in the initial grammar Gi (i.e., lines 2 to 4 of figure 4), but by generating elements associated with a numerical index (2j+1, i.e., 3 here), and not with a numerical index 1 (as is the case for N in the initial grammar Gi).
[0119] The (non-terminal) element N generates the suffix <noeud> 1 <id> 1 0 1 < / id> 1 <noeud> 1 <id> 1 from Radical 1 with the second rule: [N— > <noeud> 1 <id> 1 entire 1 < / id> 1 N < / noeud> 1 ] (line 2 of Figure 4). Such a second rule contains a left part (N) and a right part ( <noeud> 1 <id> 1 entire 1 < / id> 1 N < / noeud> 1 ) of rewriting.
[0120] The element [ <noeud> 1 <id> 1 entire 1 < / id> 1 ], formed by the first four symbols of the right part generates the second to fifth symbols [ <noeud> 1 <id> 1 0 1 < / id> 1 ] forming Radical 1.
[0121] The element N corresponding to the fifth symbol on the right side, generates the suffix [ <noeud> 1 <id> 1 ] of Radical 1.
[0122] The following new third rule is therefore generated:
[0123] [<suf1 :N> — > <noeud> 2 <id> 2 0 2 < / id> 2 <suf6:n> <noeud> 3 ].
[0124] Such a new third rule introduces the new (non-terminal) element <suf6:n>, which must be constructed from the rules of N in the initial grammar Gi (i.e., lines 2 to 4 of Figure 4), so that <suf6:n>generates the suffix <noeud> 1 <id> 1 of Radical 1 in sixth position <suf6:n>(because up to the fifth position, the previous symbols are generated in particular by the symbols [ <noeud> 1 <id> 1 entire 1 < / id> 1 ] of the right part).
[0125] The (non-terminal) element N generates the suffix <noeud> 1 <id> 1 of Radical 1 with the fourth rule: [N— >NN] (line 4 of Figure 4). Such a second rule contains a left part (N) and a right part (NN) of rewriting.
[0126] The element N in the first position of the right part generates the suffix [ <noeud> 1 <id> 1 ] of Radical 1.
[0127] The following new fourth rule is therefore generated:
[0128] [ <suf6:n> -► <suf6:n> <aft:n>],
[0129] The element N generates the suffix [ <noeud> 1 <id> 1 ] of Radical 1 with the third rule: [N — > <noeud> 1 <id> 1 entire 1 < / id> 1 L < / noeud> 1 ] (line 3 of Figure 4). Such a third rule contains a left part (N) and a right part ( <noeud> 1 <id> 1 entire 1 < / id> 1 L < / noeud> 1 ) of rewriting.
[0130] The element [ <noeud> 1 <id> 1 ] forming the first two symbols of the right part generates the suffix [ <noeud> 1 <id> 1 ] of Radical 1.
[0131] The following new fifth rule is therefore generated:
[0132] [ <suf6:n>— > <noeud> 2 <id> 2 'entire 3 < / id> 3 <aft :L>< / noeud> 3 ].
[0133] The element N generates the suffix [ <noeud> 1 <id> 1 ] of Radical 1 with the second rule [N — > <noeud> 1 <id> 1 entire 1 < / id> 1 N < / noeud> 1 ] (line 2 of Figure 4). Such a second rule contains a left part (N) and a right part ( <noeud> 1 <id> 1 entire 1 < / id> 1 N < / noeud> 1 ) of rewriting.
[0134] The element [ <noeud> 1 <id> 1 ] forming the first two symbols of the right part, generates the suffix [ <noeud> 1 <id> 1 ] of Radical 1.
[0135] The following rule is therefore generated: [ <suf6:n>— > <noeud> 2 <id> 2 entire 3 < / id> 3 <aft :N>< / noeud> 3 ].
[0136] Such a first iteration then makes it possible to arrive at a grammar G2, specified in relation to the initial grammar G1 by taking into account the first radical. A set of new rules and new (non-terminal) symbols of forms <rad:s> , <sufk:n>(where k is a natural number), or <aft:n>are created.
[0137] The following iterations to arrive at the following grammars specified with respect to each of the radicals Radical 2, Radical 3, Radical 4 and Radical 5 are implemented on the same principle as the first iteration.
[0138] In order to facilitate the visualization of the result of each iteration on each radical, the general form of the elements obtained by rewriting with each new grammar iteratively determined is presented, on the basis of a formal example (in order to simplify the writing). To simplify the formal writing in the remainder of this step 240, we note (a, b, c, d) the radicals and (v, w, x, y, z) the injection points of a pattern M considered in an executable sequence. Step 240 implemented in such a formal example then aims to determine the set of injections executable in the different injection points (v, w, x, y, z). The order of the executable injections (corresponding to injectable data at the injection points) is then represented by the order of appearance in the general form shown below. The general form of the elements illustrated below makes it possible to successively represent the form of the words obtained by the corresponding grammar after each intersection with a radical, as well as their respective numerical indices, which serve as reference points in the implementation of the successive intersections. Each intersection is implemented as detailed previously for the first iteration (for Radical 1).
[0139] At the first iteration, the initial grammar Gi generates the following general form of elements associated respectively with numerical indices: v1 a1 w1 b1 x1 c1 y1 d1 z1 .
[0140] Thus, v1 gives the general form of the first injectable data, w1 that of the second injectable data, etc. Such a general form can include one or more terminals. For example, v1 corresponding to the first injectable data can be written in the form of one or more terminals. Such a form of the elements generated by the initial grammar Gi is then intersected, during the first iteration, with the radical a indexed with the value 1. The terminals upstream of the radical have a numerical index less than or equal to 1 and keep this mark. The terminals downstream of the radical have a numerical index equal to 1 which is updated with the value 3 (i.e., 2x1+1 ) at the end of the first iteration.
[0141] In other words, at the end of the first iteration on the radical a, the second grammar G2 obtained generates the general form of elements with the following numerical indices: v1 a2 w3 b3 x3 c3 y3 d3 z3.
[0142] Such a form of the elements generated by the second grammar G2 is then intersected, during the second iteration, with the radical b indexed with the value 3. The terminals upstream of the radical have a numerical index less than or equal to 3 and retain this mark. The terminals downstream of the radical have a numerical index equal to 3, which is updated with the value 5 (i.e., 2x2+1 ) at the end of the second iteration.
[0143] In other words, at the end of the second iteration on the radical b, the third grammar G3 obtained generates the general form of elements with the following numerical indices: v1 a2 w3 b4 x5 c5 y5 d5 z5.
[0144] Such a form of the elements generated by the third grammar G3 is then intersected, during the third iteration, with the radical c indexed with the value 5. The terminals upstream of the radical have a numerical index less than or equal to 5 and retain this mark. The terminals downstream of the radical have a numerical index equal to 5, which is updated with the value 7 (i.e., 2x3+1 ) at the end of the second iteration.
[0145] At the end of the third iteration on the radical c, the fourth grammar G4 obtained generates the general form of elements with the following numerical indices: v1 a2 w3 b4 x5 c6 y 7 d7 z7.
[0146] Such a form of the elements generated by the fourth grammar G4 is then intersected, during the fourth iteration, with the radical d indexed with the value 7. The terminals upstream of the radical have a numerical index less than or equal to 7 and retain this mark. The terminals downstream of the radical have a numerical index equal to 7, which is updated with the value 9 (Le., 2x4+1 ) at the end of the fourth iteration.
[0147] In other words, at the end of the fourth iteration on the last radical d, the fifth grammar G5 obtained generates the general form of elements with the following numerical indices: v1 a2 w3 b4 x5 c6 y 7 d8 z9.
[0148] Following the iteration of the index of the last radical, the grammar G5 corresponds to the language L of the data which allows the completion of the pattern M without generating a syntax error.
[0149] Consequently, step 240 makes it possible to determine a so-called executable grammar, which generates elements respecting the initial grammar G1 (and therefore the language L used by the executable sequence) and which contains the radicals of the pattern M. In the previous formal example, the grammar G5 corresponds to the executable grammar obtained at the end of step 240 and the set of executable injections is obtained from the general form associated with such a grammar G5.
[0150] In other words, step 240 makes it possible to formally determine the set of injections (Le., a general form of these injections, or words) which makes it possible to complete the pattern M at the injection points while respecting the grammar of the executable sequence.
[0151] At the end of step 240, the regular expressions corresponding to the constraints C on the variables (and therefore on the content of the injection points) are not yet taken into account. Step 240 makes it possible to formally arrive at an executable grammar (Le., a set of rewriting rules) making it possible to generate all the syntactically executable injections at the injection points (eg, at the level of the .* markers in the pattern M L_requete), with regard to the neighboring radicals upstream (Le., in prefix) and downstream (Le., in suffix) of these injection points in the pattern M. Nevertheless, the constraints C applied to the formalism of the content of these injection points is not yet formally taken into consideration at this stage.
[0152] At a step 250, an intersection, or filtering, of the elements generated by the executable grammar with the formal expressions of the constraints C on the input data is implemented. In the example of Figures 4 and 5, the formal expression of the elements generated by the executable grammar determined at step 240 is intersected with the expressions L_id, L_temperature, L_fuel and L_panne. For this, the expressions labeled at step 232 are used. The intersection of the formal expression of the elements generated by the grammar executable and regular expressions L_id, L_temperature, L_fuel and L_panne then relies on the numerical indices on the one hand, of the elements generated by the executable grammar as determined at the end of step 240, and on the other hand, of the regular expressions corresponding to the input data. Such an intersection then makes it possible to reduce the quantity and diversity of elements actually usable in each injection point compared to the elements that can be generated by the executable grammar.
[0153] For example, we consider on the one hand the general form of the elements generated by the executable grammar determined with the radicals of the L_requete pattern, that is to say:
[0154] v1 a2 w3 b4 x5 c6 y 7 d8 z9.
[0155] On the other hand, we consider the labeled regular expression associated with a given injection point of the L_requete pattern, for example, the second injection point associated with the L_tempe rature element:
[0156] LJemperature = ([0-9] 4 | [1-9] 4 [0-9] 4 | [1 -2] 4 [0-9] 4 [0-9] 4 ) ( 4 [0-9] 4 +)7
[0157] The index 4 of such a regular expression then makes it possible to filter the elements generated by the executable grammar which concern the second injection point. Such filtering can for example be implemented by any known filtering or intersection algorithm.
[0158] Step 250 then makes it possible to arrive at a final grammar, which generates all the injections syntactically executable in the pattern M and which meet the constraints C of the content of the injection points.
[0159] In one embodiment of step 250, such an intersection may require considerable execution time, due to the dimension of the spaces of the elements generated by the executable grammar on the one hand, and the degree of constraints C on the input data. In order to reduce such execution time, step 250 may include a relaxation (or over-approximation) of the constraints C on the input data, so as to reduce the degree of precision of the filtering in favor of a better execution time. In particular, such a relaxation of the constraints C may result in the consideration of modified regular expressions (called second regular expressions) associated with the relaxed constraints C with respect to the exact constraints C (associated with first regular expressions as determined in step 211). Second regular expressions associated with the elements LJd, L_temperature, L_fuel and L_panne may for example be expressed by: - LJd = Z 2 * - LJemperature = [0-9.] 4 * - L_fuel = [0-9] 6 {1,3} L_panne = E 8 {0.99}
[0160] In such an embodiment where the constraints C on the input data are over-approximate, the set of exploitable injections determined becomes broader in that it includes false positives (i.e., injections that validate the relaxed constraints C but which do not actually validate the exact constraints C of the input data). However, the detection of exploitable injections remains exhaustive in that the set of exploitable injections determined does not include false negatives (i.e., exploitable injections which would not be described by the general form of the elements generated by the final grammar).
[0161] Thus, at the end of step 250, the set of exploitable injections in the executable sequence (more precisely, in the or each pattern M of the executable sequence) is determined. Such a set of exploitable injections can be expressed by the set of symbols of the language L used in the exploitable sequence and the final grammar, or by a general form of the elements generated by such a final grammar. Consequently, for each injection point of the executable sequence, the proposed detection method makes it possible, at the end of step 250, to exhaustively identify all the combinations of symbols forming exploitable injections at this injection point.
[0162] In an optional step 260, the set of exploitable injections is analyzed or exploited, for example by an interpretation module 116 of the processing unit 11. Such a step 260 may for example be for the purposes of development, code analysis, code injection testing, risk analysis, vulnerability auditing. In one embodiment, the set of exploitable injections may be used to randomly generate inputs in the context of fuzzing, which makes it possible to improve the quality of the tested inputs, since they are selected from a set that is actually exploitable. The set of exploitable injections may also be used as a verification tool in the context of software development, in order to improve the robustness and / or reduce the degree of vulnerability of the source code.The set of exploitable injections may also be used for an analysis of the risk of compromise of the computer system 2, by analyzing the content of the exploitable injections. For example, step 260 may include the assignment of vulnerability indices or the prioritization of exploitable injections, which aims to quantify the degree of compromise of the executable sequence if such an injection were exploited. For example, in the SQL language, an exploitable injection including the symbol “DROP” may be considered as making the executable sequence vulnerable.
[0163] The proposed method thus makes it possible to formalize the detection of injections, so as to allow exhaustive, systematic detection applicable to any programming language L. Indeed, just like the content of the executable sequence, the language L used is an input to the detection system 1 since the latter implements processing according to a formal language (Le., by formalizing the symbols and the rewriting rules of the language L and its formal grammar). List of reference signs
[0164] 1: detection system 2: computer system 3: user system 4: database system 10, 20, 30, 13: communication interfaces (input and / or output) 11, 21: processing unit 11, 112, 113, 114, 115, 116: processing modules 12, 22: memory unit< / aft:n> < / sufk:n> < / rad:s> < / id> < / noeud> < / id> < / noeud> < / id> < / noeud> < / id> < / noeud> < / id> < / noeud> < / id> < / noeud> < / aft:n> < / suf6:n> < / suf6:n> < / id> < / noeud> < / id> < / noeud> < / noeud> < / id> < / noeud> < / noeud> < / suf6:n> < / noeud> < / id> < / noeud> < / noeud> < / noeud> < / id> < / noeud> < / noeud> < / aft:n> < / aft:n> < / aft:n> < / id> < / noeud> < / noeud> < / id> < / noeud> < / noeud> < / id> < / noeud> < / noeud> < / rad:s> < / donnees> < / donnees> < / donnees> < / rad:s> < / prfk:x> < / id> < / noeud> < / noeud> < / donnees> < / aft:x> < / sufk:x> < / rad:x> < / temperature>
Claims
Claims
1. Method implemented by computer means for determining a set of exploitable injections in an executable sequence, the executable sequence using a language (L), the method comprising: a) determining at least one entry point of the executable sequence, b) determining, from said entry point, a pattern (M) comprising at least one injection point of the executable sequence, the pattern respecting the language (L) of the executable sequence, c) for each injection point, extracting a set of constraints (C) associated with the injection point, d) determining, for each injection point, the set of exploitable injections from the pattern (M) to which the injection point belongs, the associated set of constraints (C) and the language (L).
2. Method according to claim 1, the language (L) being formed of symbols and constructed according to a set of rules (Gi), the pattern (M) being constructed by a succession of symbols respecting said set of rules (Gi), and in which the set of exploitable injections is determined in step d) by the generation of generated symbols injectable into the pattern (M), said generated symbols being injectable at the level of the at least one injection point of the pattern (M), from the symbols of the pattern (M) and respecting the set of rules (Gi) and the set of constraints (C).
3. Method according to claim 2 further comprising, after step b): b1) determining, for the pattern (M) of the executable sequence, a plurality of radicals corresponding to invariable portions of the pattern (M), two neighboring radicals of the pattern (M) being separated by an injection point, and in which step d) comprises: d10) an iterative comparison of each radical of the plurality of radicals of the pattern (M) with the set of rules (Gi).
4. Method according to claim 3, in which each iteration of step d10) for each radical considered is associated with the generation of a set of generated symbols comprising the radical considered.
5. A method according to claim 4 further comprising: d11) iteratively specifying the set of rules (Gi) to each radical of the plurality of radicals of the pattern (M) resulting in sets of specified rules (G2, G3,...), each set of specified rules (G2, G3,...) generating the set of generated symbols including the radical considered.
6. Method according to claim 5, the radicals of the pattern (M) being formed by a succession of symbols, and in which step d) further comprises: d101) associating the symbols of the radicals of the pattern (M) with first labels, the symbols of the same radical being associated with the same first label, the respective symbols of the neighboring radicals being associated with successive first labels according to a first criterion, d102) associating the symbols forming the set of rules (G1) with the same initial label, and in which each set of generated symbols and each set of specified rules (G2, G3,...) are associated with generated labels depending on the first labels of the radicals of the pattern (M) and the initial label of the set of rules (G1), step d) comprising: d12) an iterated update of the generated labels according to a result of each iteration of steps d 10) and d11), d 13) an identification of the elements of the sets of generated symbols having generated labels different from the first labels.
7. Method according to one of the preceding claims, in which step d) comprises: d 1 ) from the pattern (M) and the language (L), determining a set of executable injections, d2) filtering the set of executable injections according to the associated set of constraints (C) so as to obtain the set of exploitable injections, the set of exploitable injections being of smaller dimension than the set of executable injections.
8. Method according to claim 7 combined with claims 2 to 6, in which the determination of the set of executable injections results from steps d10), d11), d101), d102), d12) and d13).
9. Method according to claim 7 in which the filtering of step d2) is implemented from an expanded set of constraints (C') including said set of constraints (C).
10. Method according to one of claims 7 to 9 and further comprising, after step c),: c1) converting, for each injection point, the associated set of constraints (C) into a first set of regular expressions, and in which the set of exploitable injections corresponds to the executable injections filtered according to a set of regular expressions among the first set of regular expressions and a second set of regular expressions over-approximate with respect to the first set of regular expressions.
11. Method according to claim 10 taken in combination with claim 6, each regular expression of the set of regular expressions being formed by a succession of symbols, and in which step d2) further comprises: d21) associating the symbols of each of the regular expressions of the set of regular expressions with second labels, the symbols of the same regular expression being associated with the same second label, the respective symbols of two neighboring regular expressions in the set of regular expressions being associated with successive second labels according to a second criterion, d22) comparing the second labels of each of the regular expressions of the set of regular expressions with the labels generated at the end of step d 12), d23) determining the set of exploitable injections from the elements of the sets of symbols generated associated with the generated labels corresponding to the second labels.
12. Method according to one of the preceding claims and further comprising: e) from the set of exploitable injections, implementing a vulnerability analysis of the executable sequence.
13. Method according to one of the preceding claims and further comprising: e') assigning vulnerability indices respectively associated with each exploitable injection of the set of exploitable injections, said vulnerability index being determined from the pattern (M) completed by the exploitable injection at the corresponding injection point.
14. Computer program comprising instructions for implementing the method according to one of claims 1 to 13 when this program is executed by a processor.
15. Non-transitory recording medium readable by a computer on which is recorded a program for implementing the method according to one of claims 1 to 13 when this program is executed by a processor.
Citation Information
Cited By
Code execution method and device, equipment, medium and product
CN122450544A