Automatic tracing method for airborne software compiling process based on DSP (Digital Signal Processor)
By inserting assembly line number piles into C source code and using extended Bacos paradigm matching and syntax node precompilation, the problems of low testing efficiency and high labor cost during onboard software compilation are solved, and efficient and accurate automated traceability is achieved, suitable for DSP processors and software developed in C language.
Patent Information
- Application Number
- CN202510415872.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-04
AI Technical Summary
The prior art has problems such as low testing efficiency, high labor costs and insufficient traceability relationship establishment during the compilation of onboard software.
By inserting assembly line number piles in the C source code, combining extended Bacos paradigm matching and syntax node precompilation, intermediate codes are generated, and the longest common subsequence is calculated through dynamic programming algorithms, the accurate row-level correspondence between source code and assembly code and automated traceability are achieved.
It significantly improves the accuracy and efficiency of traceability, reduces labor costs, and is suitable for all software developed using DSP processors and C languages using the same compilation rules, enhancing the security of the software and the efficiency of airworthiness certification.
Smart Images

Figure CN120255902A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of software security verification, and particularly to an automatic tracing method for the compilation process of airborne software based on DSP. Background Art
[0002] In aircraft, DSP processors dominate in the fields of communication and radar due to their powerful computing capabilities. However, during the development process of airborne software, when the compiler converts the high-level language source code written by developers into the compiled object code for machine execution, it cannot guarantee 100% consistency between the object code and the source code, which poses a security risk. The aviation software development standard DO-178C requires that for software with a safety level of A, the tracing relationship of the compilation process must be established, and additional verification must be carried out for areas where tracing is not possible.
[0003] For example, a method for analyzing the statement coverage rate of the object code of safety-critical embedded software provided in the patent with the publication number CN106445803A. The method for analyzing the coverage rate of the object code of the safety-critical embedded software of the present invention compiles the program into object code with debugging information. Before the object code test, by selecting the functions to be tested, the monitoring module automatically inserts breakpoints at the entrances and exits of the functions. During the object code test, when running to the first breakpoint, the target machine records each executed object code, and when the next breakpoint is triggered, the monitoring module controls the target machine to send the executed object code back to the host computer through the communication link. When the host computer receives the feedback, it compares with the compiled binary object code and marks the executed binary object code. When the test ends, the statement coverage rate of the object code is obtained by calculating the object code execution rate;
[0004] A method for selecting the smallest code subset covering all source code structures provided in the patent with the publication number CN109783344A includes: a. classifying the source code of the software to be analyzed into high-level language, assembly language, and library functions to obtain a high-level language code set, an assembly code set, and a library function set; b. analyzing the syntax structure and the structure that may generate additional code during the compilation process for each function in the high-level language code set, and selecting a typical structure code subset that minimizes the workload of comparing the source code with the object code on the condition that each structure is covered at least once; and c. merging the typical structure code subset, the assembly code set, and the library function set to obtain the smallest code subset of the software to be analyzed;
[0005] An analysis method for the consistency between source code and object code in an airborne software for aviation, provided by the publication number CN107391368A, includes: analyzing the source code to obtain the language feature information of the source code; obtaining the typical code features of the source code according to the language feature information, and selecting the source code functions including the typical code features; disassembling the object code to obtain the disassembly code, and segmenting and identifying and labeling the disassembly code into functions to obtain the function list of the object code; establishing a mapping relationship between the source code functions and the functions in the function list of the object code; comparing whether the functions in the source code functions and the function list of the object code correspond; and labeling the source code functions in the function list that are not mapped to the object code and the functions in the function list of the object code that do not exist in the source code functions as inconsistent between the object code and the source code. The analysis method for the consistency between source code and object code in the airborne software for aviation provided by the present invention improves the analysis efficiency;
[0006] A method and system for comparing and analyzing the source code and object code of an airborne software for aviation, provided by the publication number CN110659200A, can identify the extra code generated by the compiler that cannot be traced back to the source code. A method for comparing and analyzing the source code and object code of an airborne software for aviation provided by the present invention includes the following steps: determining a general judgment criterion for the equivalence and inclusion relationships between syntax features; selecting the typical syntax structures and code subsets of the source code in the source file based on the general judgment criterion; compiling the code subsets to generate object files; and disassembling the object files to generate a cross-reference list of the source code and assembly code.
[0007] However, when the above tracing method is used, the following problems exist:
[0008] CN106445803A: The efficiency of coverage repeated testing is low;
[0009] CN109783344A: It is necessary to re-extract subsets for verification for each software;
[0010] CN107391368A: Only at the function level, it fails to penetrate into the specific code, resulting in an insufficiently in-depth establishment of the tracing relationship;
[0011] CN110659200A: The establishment of the linked list and the comparison relationship still relies on manual work, and no specific operation method is proposed.
[0012] In summary, the prior art has problems such as low testing efficiency, high labor cost, and insufficiently in-depth establishment of the tracing relationship. Summary of the Invention
[0013] The object of the present invention is to provide an automatic traceability method for the airborne software compilation process based on DSP, so as to solve the problems of low test efficiency, high labor cost and insufficient establishment of traceability relationships.
[0014] To achieve the above object, the present invention provides the following technical solution: An automatic traceability method for the airborne software compilation process based on DSP, including the following steps:
[0015] S1. Embedded line number marking: Perform preprocessing on the original C code and the compilation project file, and construct a line-level mapping relationship between the C source code lines and the basic blocks of the assembly code, including the following steps:
[0016] S11: Insert line number marking stakes containing assembly instructions line by line into the original C source code to be processed;
[0017] S12: Compile the C code of the marking stakes in step S11 to generate a project file, and generate assembly code with line number identification through a decompilation tool;
[0018] S13: Establish an associated mapping between the assembly code segment and the corresponding C source code line through the C source code line number information displayed in the assembly code in step S12;
[0019] S2. Extended Backus-Naur form matching: Perform lexical analysis on the C code preprocessed in step S1, and divide each line of code into specific syntax units, including:
[0020] S21: Use regular expressions to perform lexical analysis on the C code line, and extract the feature marking sequence;
[0021] S22: Formulate extended Backus-Naur form rules to classify the marking sequence into typical C language syntax structures;
[0022] S3. Syntax node precompilation: According to the syntax type of the C language statements identified in step S2 and the precompilation rules of different syntax nodes, precompile the C source code to generate intermediate code;
[0023] S4. Mapping relationship verification: Perform pattern matching on the intermediate code generated in step S3 and the assembly code with "line numbers" generated in step S1, and use the dynamic programming algorithm to calculate the longest common subsequence similarity. The specific implementation includes: converting the intermediate code into an uppercase ASCII string, extracting the assembly operation code to generate an uppercase string, constructing a two-dimensional state matrix and filling it through dynamic programming, and determining the maximum matching sequence.
[0024] Preferably, the pre-compilation rules for syntax nodes in step S3 include the comparison pre-compilation rules for assignment nodes, operation nodes, function call nodes, function declaration nodes, function return nodes, If-else judgment nodes, Switch judgment nodes, For loop nodes, While loop nodes, loop exit nodes, and input nodes.
[0025] Preferably, the construction rule for the dynamic programming matrix in step S4 is as follows:
[0026] If text1[i - 1] == text2[j - 1], the matrix element satisfies dp[i][j] = dp[i - 1][j - 1] + 1;
[0027] If text1[i - 1] != text2[j - 1], then the matrix element takes dp[i][j] = max(dp[i - 1][j], dp[i][j - 1]);
[0028] dp[i][j] represents the length of the longest common subsequence of the first i characters of string text1 and the first j characters of string text2. Therefore, dp[m][n] is the matching rate obtained after LCS processing.
[0029] Preferably, finding the longest common subsequence in the two pieces of code in step S4 is specifically as follows:
[0030] Starting from dp[m][n], if text1[i - 1] == text2[j - 1], then the current character is part of the longest common subsequence, and continue to trace back upward to dp[i - 1][j - 1];
[0031] If text1[i - 1] != text2[j - 1], then compare the values of dp[i - 1][j] and dp[i][j - 1], and choose the larger one to continue tracing back;
[0032] During the backtracking process, add the characters to the result string from back to front.
[0033] Compared with the prior art, the beneficial effects of the present invention are:
[0034] Through the embedded line number piling technology, this invention inserts assembly line number piles line by line into the source code. After compilation and disassembly processing, an accurate line-level correspondence relationship between the source code lines and the basic blocks of the assembly code is established, realizing the basis for automatic traceability. Through predefined syntax node pre-compilation rules, different syntax structures in the source code are pre-compiled into intermediate codes, which are then matched and compared with the assembly operation codes obtained by disassembly, and the longest common subsequence is extracted to obtain the traceability result. This innovative step not only avoids the tediousness and high cost of manual comparison, but also significantly improves the accuracy and efficiency of traceability. Brief Description of the Drawings
[0035] Figure 1 It is a block diagram of the overall structural flow of an automatic traceability method for the compilation process of airborne software based on DSP in this invention. Specific Embodiments
[0036] Next, the technical solutions in the embodiments of this invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this invention. Obviously, the described embodiments are only a part of the embodiments of this invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this invention without creative efforts shall fall within the protection scope of this invention.
[0037] Please refer to Figure 1 , this invention provides a technical solution: an automatic traceability method for the compilation process of airborne software based on DSP, including the following steps:
[0038] S1. Embedded line number marking: Perform preprocessing on the original C code and the compilation project file, and construct a line-level mapping relationship between the C source code lines and the basic blocks of the assembly code. The specific operations are as follows:
[0039] S11: Use embedded assembly instructions to insert "assembly line number piles" line by line into the C source code to be processed;
[0040] S12: Compile the C source code containing "assembly line number piles" into a target project file, and use the disassembly tool dis2000 for DSP to generate assembly code with "line numbers";
[0041] S13: Through the "line numbers" of the C source code shown in the assembly code, establish the correspondence relationship between the assembly code segment under the "line number" and the same-line code of the C source code.
[0042] S2. Extended Backus-Naur Form Matching: Perform lexical analysis on the C source code preprocessed in step S1, extract the token stream using regular expressions, and parse and identify each predefined component in the syntax rules through the Extended Backus-Naur Form (EBNF). Divide each line of C source code into different C language statements, specifically as follows:
[0043] S21: Perform lexical analysis on each line of C source code and extract the token stream using regular expressions;
[0044] S22: Define EBNF rules, identify the token stream as 12 representative C language types, and the judgment rules are shown in Table 1:
[0045] Table 1: Definition of EBNF rule table
[0046]
[0047] Among them, the meanings of the symbols used in the EBNF rules in the table are as follows:
[0048] Terminal symbol: Represents a specific value or lexical unit of a syntax rule, usually enclosed in quotes (” or "");
[0049] Non-terminal symbol: Expressed in Chinese words, represents a part or placeholder of a syntax rule;
[0050] Definition or specification: Use “::=” to represent the definition or specification of a syntax rule;
[0051] Choice: Use the vertical bar (|) to represent choice, that is, a certain part of the rule can be one of multiple options;
[0052] Optional: Use square brackets ([]) to represent optional, that is, the content within the square brackets can appear or not appear;
[0053] Repetition: Use the asterisk (*) symbol to indicate that a certain symbol or symbol sequence can appear zero or more times in the production.
[0054] S3. Precompilation of Syntax Nodes: According to the syntax types of the C language statements identified in step S2 and the precompilation rules of different syntax nodes, precompile the C source code to generate intermediate code. The precompilation rules are as follows:
[0055] (1) Precompilation rules for assignment nodes:
[0056] (1-1) Assignments made during variable definition are not shown in the code segment. Assignments used in function arguments and algebraic operations follow the following rules:
[0057] a) When the parameter is of short or int type, 1 word instruction is used for assignment;
[0058] b) When the parameter is of float, double, or long int type, 1 double-word instruction is used for assignment;
[0059] (1 - 2) The assignment of pointer types is compiled using the following rules:
[0060] a) 1 assignment instruction initializes the pointer;
[0061] b) 1 subtraction instruction modifies the address register to the target address;
[0062] c) 1 assignment instruction makes the pointer point to the target address;
[0063] d) Refer to the assignment rules in (1 - 1) to store the variable value into the target address.
[0064] (2) Pre-compilation rules for operation node mapping:
[0065] (2 - 1) The "+" operator: 1 addition instruction is used, and its special usage is as follows:
[0066] a) If the "++" operator is before the variable in the source code, an addition instruction is used and compiled before other syntax nodes in the same code line;
[0067] b) If the "++" operator is after the variable in the source code, an addition instruction is used and compiled after other syntax nodes in the same code line;
[0068] (2 - 2) The "-" operator: 1 subtraction instruction is used, and its special usage is as follows:
[0069] a) If the "--" operator is before the variable in the source code, a subtraction instruction is used and compiled before other syntax nodes in the same code line;
[0070] b) If the "--" operator is after the variable in the source code, a subtraction instruction is used and compiled after other syntax nodes in the same code line;
[0071] (2 - 3) The "*" operator: 1 multiplication instruction is used;
[0072] (2 - 4) The " / " operator and the "%" operator: Compiled in the form of built-in functions, and the specific operation instructions are as follows:
[0073] a) Use 1 function call instruction to enable the built-in division function;
[0074] b) Use 1 remainder instruction to perform the division remainder operation to obtain the remainder result;
[0075] c) If it is a division operation, an additional assignment instruction is used to save the quotient of the division calculation;
[0076] (2-5) "<< operator": Use 1 left shift instruction;
[0077] (2-6) ">> operator": Use 1 right shift instruction;
[0078] (2-7) "& operator": Use 1 AND instruction;
[0079] (2-8) "^ operator": Use 1 exclusive OR instruction;
[0080] (2-9) "| operator": Use 1 OR instruction.
[0081] (3) Precompilation rules for function call nodes:
[0082] (3-1) 1 function call start instruction;
[0083] (3-2) Several parameter passing assignment instructions:
[0084] a) When there are no parameters passed, there are no assignment instructions;
[0085] b) When there are parameters passed, the compilation rules refer to (1) assignment nodes;
[0086] c) When there are multiple parameters passed, multiple assignment instructions are used, and the object of each assignment instruction is the address pointer of the n bits after the first parameter passed.
[0087] (3-3) 1 jump instruction (to jump and execute the function body content).
[0088] (4) Precompilation rules for function declaration nodes:
[0089] (4-1) Function declarations are omitted during compilation;
[0090] (4-2) Compile the called underlying function body first according to the nested call relationship.
[0091] (5) Precompilation rules for function return nodes:
[0092] (5-1) The return node is located at the end of the called function;
[0093] (5-2) 1 jump back instruction to jump back to the upper layer of the function nesting relationship.
[0094] (6) Precompilation rules for If-else judgment nodes:
[0095] (6-1) 1 judgment and comparison instruction;
[0096] (6 - 2) One result jump instruction, and the specific judgment operands are as follows:
[0097] a) The operand for judging equality;
[0098] b) The operand for judging greater than;
[0099] c) The operand for judging less than;
[0100] d) The operand for judging greater than or equal to;
[0101] e) The operand for judging less than or equal to;
[0102] f) The operand for judging not equal to;
[0103] (6 - 3) There are logical expressions associated with multiple Boolean expressions:
[0104] a) NOT gate: Reverse - compile the judgment instruction (for example, compile the judgment instruction of the greater - than condition into the less - than - or - equal - to operand);
[0105] b) OR gate: Compile in the order of the Boolean expressions before and after (perform judgments in sequence, and if it meets the condition, jump to execute);
[0106] c) AND gate: Reverse - compile all except the last one in the order of the Boolean expressions before and after (perform judgments in sequence, and if one of the previous conditions does not meet, skip to execute).
[0107] (7) Pre - compilation rules for the Switch judgment node comparison:
[0108] (7 - 1) Pre - compilation rules for judgment code:
[0109] Similar to the OR - gate judgment process in (4 - 3 - b), use the comparison instruction and the jump instruction SB with the result of equality, and compile each case's judgment condition one by one in the order of the cases before and after. If it meets the condition, jump to execute the corresponding code - segment branch;
[0110] The specific operation instructions for the switch judgment code are as follows:
[0111] a) One assignment instruction to read the value of the judgment variable into the register;
[0112] b) In ascending order, each case judgment uses two instructions in (6) (one judgment - comparison instruction and one result - jump instruction);
[0113] (7 - 2) Pre - compilation rules for code segments
[0114] The branch of the pre-compiled code segment starts from the default segment (if the judgment conditions of the previous cases are not satisfied, the default code segment is executed), and is compiled in the reverse order of the case numbers; the specific operation instructions for each switch code segment are as follows:
[0115] a) Compile the syntax nodes normally in the program running order;
[0116] b) In the case of a break, compile 1 jump instruction at the end to jump out of the subsequent cases, which is a special usage of the break node in (10));
[0117] (8) Pre-compilation rules corresponding to the For loop node:
[0118] (8-1) Before compiling the internal code segment of the loop:
[0119] Compile the arithmetic instructions before the loop, reverse-compile the judgment instructions to judge whether the condition for entering the loop is satisfied, and if not, jump to the loop exit;
[0120] The specific operation instructions for entering the loop judgment are as follows:
[0121] a) 2 assignment instructions (select single-word and double-word assignments according to the variable type) for initializing registers and variable assignments;
[0122] b) 1 comparison instruction;
[0123] c) 1 jump instruction according to the result;
[0124] (8-2) After compiling the internal code segment of the loop:
[0125] Compile the arithmetic instructions after the loop, normally compile the judgment instructions to judge whether the condition for entering the loop is satisfied, and if satisfied, jump to the loop entry.
[0126] The specific operation instructions for exiting the loop judgment are as follows:
[0127] a) 1 instruction for calculating the judgment condition;
[0128] b) 1 assignment instruction;
[0129] c) 1 comparison instruction;
[0130] d) 1 jump instruction according to the result;
[0131] (9) Pre-compilation rules corresponding to the While loop node:
[0132] (9-1) Normal while loop:
[0133] It is the same as the pre-compilation rule of the for loop node in (8). Two judgments are made before and after the loop body code to determine whether to skip the loop body and whether to execute the loop body again;
[0134] (9-2) do-while loop:
[0135] Only one assignment instruction is compiled before the loop body code (single-word and double-word assignments are selected according to the variable type) to save the loop count;
[0136] After the loop body code, it is judged whether to execute the loop body content again. The specific operation instructions are the same as in (8-2);
[0137] (10) Pre-compilation rules for the loop exit node comparison:
[0138] (10-1) break node:
[0139] The code segment for judging whether to jump to the loop exit is compiled separately. The specific operation instructions are as follows:
[0140] a) One instruction for calculating the judgment condition;
[0141] b) One assignment instruction;
[0142] c) One comparison instruction;
[0143] d) One jump instruction according to the result;
[0144] c) One jump instruction to exit the loop;
[0145] (10-2) continue node:
[0146] Similar to the first four operation instructions of the break node, one SB jump instruction is added before the loop body to exit the current loop and enter the next loop.
[0147] (11) Pre-compilation rules for the input node comparison:
[0148] (11-1) Use three assignment instructions to store the input target address in the register, assign the register content to the pointer, and save the initial input address;
[0149] (11-2) Use one function call instruction to enable the built-in input function;
[0150] (11-3) Use one subtraction instruction to adjust the pointer register to point to the next save address;
[0151] (11-4) Use two assignment instructions to store the input content in the register and then store it from the register to the save address;
[0152] (11 - 5) Use one jump instruction to exit the input function of the built - in function and wait for the next call.
[0153] (12) Pre - compilation rules for output node comparison:
[0154] (12 - 1) Use one assignment instruction to store the starting address of the output target in a register;
[0155] (12 - 2) Use one function call instruction to enable the built - in output function;
[0156] (12 - 3) Use one assignment instruction to directly store the content at the register - corresponding address into the output position. If the output is a variable, one more assignment instruction is required to output the variable name;
[0157] (12 - 4) Use one jump instruction to exit the output function of the built - in function and wait for the next call;
[0158] S4. Mapping relationship verification: Match and compare the intermediate code generated in step S3 with the assembly code with "line numbers" generated in step S1. Find the longest common subsequence (LCS) between the two pieces of code through dynamic programming and evaluate their similarity. Specifically:
[0159] S41: Convert the intermediate code of the operation instructions obtained by pre - compilation into a string test1 in uppercase ASCII code format line by line;
[0160] S42: Find the corresponding line code segment in the assembly code processed in step one, remove the operands, and only keep the operation codes, then convert them into a string text2 in uppercase ASCII code format;
[0161] S43: Use a two - dimensional array dp[i][j] to store the state and fill the dynamic programming array in a loop. The specific rules are as follows:
[0162] (1) If text1[i - 1] == text2[j - 1], then the matrix element satisfies dp[i][j] = dp[i - 1][j - 1] + 1;
[0164] (2) If text1[i - 1]!= text2[j - 1], then the matrix element takes dp[i][j] = max(dp[i - 1][j], dp[i][j - 1]);
[0165] (3) dp[i][j] represents the length of the longest common subsequence of the first i characters of the string text1 and the first j characters of the string text2. Therefore, dp[m][n] (where m and n are the lengths of text1 and text2 respectively) is the matching rate obtained after LCS processing.
[0166] S44: Backtrack to extract the longest common subsequence, which is the opcode for successful matching and tracing. Specifically:
[0167] (1) Starting from dp[m][n], if text1[i - 1] == text2[j - 1], then the current character is part of the longest common subsequence, and continue backtracking upward to dp[i - 1][j - 1];
[0168] (2) If text1[i - 1] != text2[j - 1], then compare the values of dp[i - 1][j] and dp[i][j - 1], and choose the larger one to continue backtracking;
[0169] (3) During backtracking, add characters to the result string from back to front.
[0170] The present invention can effectively solve the complexity and inefficiency problems in the process of tracing from C source code to the compiled target code of a DSP processor. By embedding line number stubs, inserting assembly line number stubs line by line in the source code, and then through compilation and disassembly processing, an accurate line-level correspondence between the source code lines and the basic blocks of the assembly code is established, realizing the basis for automated tracing. Through predefined syntax node precompilation rules, different syntax structures in the source code are precompiled into intermediate codes, and compared and matched with the assembly opcodes obtained by disassembly. This innovative step not only avoids the tediousness and high cost of manual comparison, but also significantly improves the accuracy and efficiency of tracing.
[0171] The present invention can automatically establish the trace relationship between the codes of typical syntaxes and automatically compare and verify line by line for all software developed by a processor (DSP) with the same compilation rules and a high-level language (C language). This method not only greatly reduces the labor cost of tracing from source code to target code, improves the efficiency and accuracy of tracing, but also enhances the security of the software. Specifically, the present invention realizes the comprehensive tracing of the compilation process through steps such as embedding line number stubs, precompiling syntax nodes, and comparison and matching. Embedding line number stubs establishes the line-level correspondence between the source code lines and the basic blocks of the assembly code, laying the foundation for subsequent tracing work. The precompilation of syntax nodes precompiles the identified syntax nodes into intermediate codes according to the precompilation rules of different syntax nodes, facilitating the comparison and matching with the assembly opcodes obtained by disassembly. Finally, by calculating the longest common subsequence, the similarity of the two pieces of code is evaluated, thereby realizing the tracing of the compilation process.
[0172] For the places where the matching fails, only a small amount of manual supplementary verification is required to obtain the complete tracing result, greatly improving the efficiency of airworthiness certification. In summary, with the advantages of automatic tracing, high efficiency and accuracy, low cost, and wide applicability, this technical solution has brought a revolutionary change to the verification of the airborne DSP software compilation process.
[0173] The present invention not only greatly improves the accuracy and efficiency of traceability, but also significantly reduces the labor cost. It is applicable to all DSP processors using the same compilation rules and software developed in C language, with wide applicability and generality. In addition, it can also improve the efficiency of airworthiness certification and provide a strong technical guarantee for the safety verification of civil aviation airborne Class A software.
[0174] The present invention also demonstrates an embodiment of an automatic traceability method for the compilation process of airborne software based on DSP, including the following:
[0175] First, the code of a C language software is shown. This software will be compiled and traced in the following demonstration, and its source code is as follows:
[0176]
[0177]
[0178] S1. Embedded line number marking:
[0179] S11: Use embedded assembly instructions to insert "assembly line number stubs" line by line in the source C code to be processed. Specifically:
[0180] (1) Obtain the code file: Traverse the given file directory. For each file in the directory, check whether the file name ends with.c to ensure that only C language source files are processed;
[0181] (2) Set in-place editing: For each eligible file, set in-place editing after the file is opened;
[0182] (3) Insert embedded instructions: Traverse each line of each.c file and determine whether to insert the embedded assembly instruction code for displaying line numbers after this line according to the following three conditions (to ensure that the code does not affect normal compilation and operation after code stubbing):
[0183] (3-1) If the line ends with a semicolon;
[0184] (3-2) If the line contains a left curly brace {;
[0185] (3-3) If the line contains a right curly brace};
[0186] (4) Maintain the code format: Remove the blank characters (such as spaces and line breaks) at the end of each line, and then print it out (for the lines that need to insert the embedded assembly instruction code, a new line and assembly code will be added after removing the blanks). The final result is as follows:
[0187]
[0188]
[0189]
[0190] S12: Compile the C source code containing "assembly line number stubs" into a target project file, and use the disassembly tool dis2000 for DSP to generate assembly code with "line numbers". Specifically:
[0191] (1) Compile the source file: Call the compiler cl2000 for DSP to compile the C source code with embedded instructions into a target code project file and save it in the specified directory.
[0192] (2) Read the project file: Check the ".o" project file in the specified directory and record its address.
[0193] (3) Call dis2000: Call the disassembly tool dis2000 for DSP at the address where the project file is located to convert the target code in the project file into assembly code.
[0194] (4) Save the assembly code: Output the converted assembly code in txt format and save it. The obtained assembly code is as follows:
[0195]
[0196]
[0197]
[0198]
[0199]
[0200] S13: Establish the correspondence between the assembly code segment under the "line number" and the same-line code in the C source code through the "line number" of the C source code displayed in the assembly code. Specifically:
[0201] (1) Parse the line numbers in the assembly: For the assembly code obtained above, identify and extract all labels by retrieving ":" in the assembly code, and merge adjacent labels into the leading label.
[0202] (2) Split the code by labels: Split the code between labels into assembly code segments and establish the correspondence between the code segments and the labels. The result is shown in Table 2:
[0203] Table 2: Label-Split Code Table
[0204]
[0205]
[0206]
[0207]
[0208]
[0209]
[0210] (3) Extract the assembly operation code: Remove the information such as "assembly code line", "original object code", "operand", etc. from the assembly code segment, and only retain the assembly operation code;
[0211] (4) Create a line number comparison table: Similarly, split the C language source code into labeled code lines by tags, and establish a corresponding relationship with the assembly operation code. The result is shown in Table 3:
[0212] Table 3: Line number comparison table
[0213]
[0214]
[0215]
[0216]
[0217] S2. Extended Backus-Naur form matching:
[0218] S21: Perform lexical analysis on each line of C source code, and use regular expressions to extract the token stream. Specifically:
[0219] (1) Extract the code line by line
[0220] Regularly read the code between the tags, and use the function name tag to mark the function declaration additionally. The separate "{" line and the line above it use the same line number tag;
[0221] (2) Split the characters
[0222] Split the characters of each line of the source code read according to spaces, line breaks ("\n") and special symbols (), and obtain the basic characters. The result is shown in Table 4:
[0223] Table 4: Source code character split table
[0224]
[0225]
[0226]
[0227] S22: Define the EBNF rules to recognize the token stream as 12 representative C language categories. The judgment rules are as follows in the table, specifically:
[0228] (1) Recognize the lexeme types of characters
[0229] Recognize the code between tags as different types of lexemes, such as types, identifiers, etc.;
[0230] (2) Recognize the corresponding EBNF expressions to classify the grammar
[0231] Read the lexeme token, and classify each line of source code into the corresponding grammar category by referring to the EBNF statement types defined in Step 2 of the technical solution. The results are shown in Table 5:
[0232] Table 5: Classification Table of Source Code Grammar Categories
[0233]
[0234]
[0235]
[0236]
[0237]
[0238] S3. Pre-compilation of grammar nodes:
[0239] The pre-compilation principle of some lines is as follows:
[0240] line12:
[0241] Use MOV - type assignment instructions to initialize and assign the register storing the loop count; use the CMPB instruction to compare the initial loop count with the loop upper limit;
[0242] If the judgment result is greater than or equal to, use the SB instruction to skip the loop body.
[0243] line14:
[0244] After entering the loop body, execute the printf grammar node, which is a built - in function;
[0245] Use 1 MOV assignment instruction to store the starting address of the output target in the register and use the SPM instruction to start the function call;
[0246] Use several MOV assignment instructions to store the register address corresponding to the output content into a pointer; use the LCR instruction to exit the output function of the built-in function.
[0247] line16:
[0248] After one execution of the loop body, use the ADDB instruction to increment the loop count by one;
[0249] Use the MOV register to transfer the loop count to a register for easy judgment;
[0250] Use the CMPB instruction to compare the loop count with the loop upper limit;
[0251] If the judgment result is less than, use the SB instruction to jump back to the loop body and repeat the execution. The overall pre-compilation process of the example software is shown in Table VI:
[0252] Table VI: Pre-compilation process table
[0253]
[0254]
[0255]
[0256]
[0257]
[0258] S4. Mapping relationship verification:
[0259] Match and compare the intermediate code obtained by pre-compiling the example software in S3 with the assembly operation code obtained by disassembling in S1, and obtain the matching percentage by calculating the longest common subsequence. The results are shown in Table VII:
[0260] Table VII: Longest common subsequence matching percentage table
[0261]
[0262]
[0263]
[0264] It can be seen that the overall LCS statement matching rate has reached 95.52%. Among them, the compilation rules of the main function declaration and return statement are different from those of the called function, and the assignment instructions have increased, resulting in a decrease in the matching rate. Subsequently, manual analysis will be added to the places where the matching fails.
[0265] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0266] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An automatic traceability method for the airborne software compilation process based on DSP, characterized in that: Including the following steps: S1. Embedded line number marking: Perform preprocessing on the original C code and the compilation project file, and construct a line-level mapping relationship between the C source code lines and the basic blocks of the assembly code, including the following steps: S11: Insert line number marking stakes containing assembly instructions line by line into the original C source code to be processed; S12: Compile the C code of the marking stakes in step S11 to generate a project file, and generate assembly code with line number identification through a decompilation tool; S13: Establish an associated mapping between the assembly code segment and the corresponding C source code line through the C source code line number information displayed in the assembly code in step S12; S2. Extended Backus-Naur Form matching: Perform lexical analysis on the C code preprocessed in step S1, and divide each line of code into specific syntax units, including: S21: Use regular expressions to perform lexical analysis on the C code line and extract the feature marking sequence; S22: Formulate extended Backus-Naur Form rules to classify the marking sequence into typical C language syntax structures; S3. Syntax node precompilation: According to the syntax type of the C language statements identified in step S2 and the precompilation rules of different syntax nodes, precompile the C source code to generate intermediate code; S4. Mapping relationship verification: Perform pattern matching on the intermediate code generated in step S3 and the assembly code with "line numbers" generated in step S1, and use the dynamic programming algorithm to calculate the longest common subsequence similarity. The specific implementation includes: converting the intermediate code into a capital ASCII string, extracting the assembly operation code to generate a capital string, constructing a two-dimensional state matrix and filling it through dynamic programming, and determining the maximum matching sequence.
2. The automatic traceability method for the airborne software compilation process based on DSP according to claim 1, characterized in that: The syntax node precompilation rules in step S3 include the comparison precompilation rules for assignment nodes, operation nodes, function call nodes, function declaration nodes, function return nodes, If-else judgment nodes, Switch judgment nodes, For loop nodes, While loop nodes, loop exit nodes, and input nodes.
3. The automatic traceability method for the airborne software compilation process based on DSP according to claim 2, wherein: The construction rule of the dynamic programming matrix in step S4 is: If text1[i - 1] == text2[j - 1], the matrix element satisfies dp[i][j] = dp[i - 1][j - 1] + 1; If text1[i - 1] != text2[j - 1], then the matrix element takes dp[i][j] = max(dp[i - 1][j], dp[i][j - 1]); dp[i][j] represents the length of the longest common subsequence of the first i characters of the string text1 and the first j characters of the string text2. Therefore, dp[m][n]n is the matching rate obtained after LCS processing.
4. The automatic traceability method for the airborne software compilation process based on DSP according to claim 3, characterized in that: In step S4, finding the longest common subsequence in the two pieces of code is specifically as follows: Starting from dp[m][n], if text1[i - 1] == text2[j - 1], then the current character is part of the longest common subsequence, and continue to trace back upward to dp[i - 1][j - 1]; If text1[i - 1] != text2[j - 1], then compare the values of dp[i - 1][j] and dp[i][j - 1], and choose the larger one to continue backtracking; During the backtracking process, add the characters to the result string from the back to the front.
Citation Information
Patent Citations
Security-critical embedded software target code coverage analysis method
CN106445803A
Method for analyzing consistency of source codes and target codes in airborne software
CN107391368A
A method of selecting a minimum code subset covering all source code structures
CN109783344A
Source code and target code contrastive analysis method and system for aviation airborne software
CN110659200A