Method and apparatus for processing data associated with a development tool

The integration of a machine learning system in development tools for computing devices optimizes input data management, reducing redundant processing and accelerating software builds by predicting and managing data versions effectively.

DE102024200827A1Pending Publication Date: 2025-07-31ROBERT BOSCH GMBH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
DE102024200827
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-30
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing development tools for generating program code for computing devices lack efficiency and effectiveness in managing input data versions, leading to redundant processing and resource wastage.

Method used

Employing a machine learning (ML) system to ascertain and predict the appropriate version of input data for development tools, such as code generators and compilers, and optimize their execution using techniques like parallelization and caching to reduce redundant processing.

Benefits of technology

Enhances the efficiency of software builds by minimizing redundant executions, optimizing resource utilization, and accelerating the development process through intelligent data management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A method, for example a computer-implemented method, for processing data associated with a development tool for creating program code for a computing device, for example a microcontroller, comprising: determining first information characterizing a version of input data for the development tool by means of a machine learning (ML)-based ML system, using the first information for the development tool.
Need to check novelty before this filing date? Find Prior Art

Description

State of the art

[0001] The disclosure relates to a method for processing data associated with a development tool for creating program code for a computing device.

[0002] The disclosure relates to an apparatus for processing data associated with a development tool for creating program code for a computing device. Disclosure of the invention

[0003] Some embodiments relate to a method, for example, a computer-implemented method, for processing data associated with a development tool for creating program code for a computing device, for example, a microcontroller, comprising: determining first information characterizing a version of input data for the development tool using an ML system based on machine learning (ML); using the first information for the development tool. In some embodiments, this enables the specification of input data to be used for the development tool using the ML system, whereby, in some embodiments, for example, specific versions of the input data to be used can be predicted using the ML system.

[0004] In some embodiments, the version of the input data characterizes a particular version of the input data, for example a particular development stage and / or a particular variant of the input data.

[0005] In some embodiments, the method is thus provided to comprise: selecting, based on the first information, which version of the input data is to be used by the development tool.

[0006] In some embodiments, the development tool comprises or is at least one of the following elements: a) code generator, for example for generating source code associated with the computing device, wherein, for example, the input data is at least partially provided for the code generator, or b) compiler for generating code associated with the computing device, for example executable by the computing device, for example machine code and / or bytecode, wherein, for example, the input data is at least partially provided for the compiler.

[0007] In some embodiments, the development tool represents, for example, a software build toolchain or the development tool is at least part of a software build toolchain. In some embodiments, the software build toolchain can be viewed as a collection of tools, e.g., software tools, that are used, e.g., in a specific development process, for example, for creating the program code for the computing device (e.g., "software build"). Optionally, in some embodiments, linkers, libraries, debuggers, and other tools can also be part of the software build toolchain.

[0008] In some embodiments, the method comprises: predicting, by means of the ML system, which version of the input data should be used by the development tool, supplying a corresponding version of the input data to the development tool, and, optionally, processing the input data with the corresponding version by the development tool.

[0009] In some embodiments, the supplying comprises at least one of the following elements: a) supplying first input data, for example characterizing specification information, for a or the code generator of the development tool, and / or b) supplying second input data, for example characterizing source code, for example source code associated with the computing device, for example source code created by means of the code generator, for a or the compiler of the development tool.

[0010] In some embodiments, the method comprises storing previously used input data, and, optionally, selectively using, e.g., reusing, at least a portion of the previously used input data, e.g., based on the first information.

[0011] In some embodiments, the method comprises providing an interface, for example a user interface, for receiving second information, and, optionally, using the second information to influence the first information. In some embodiments, this enables, for example, manual correction or modification of the first information, for example by human experts. In some embodiments, the first information can also be influenced via the interface by means of another system, for example an expert system. In some embodiments, combinations of human and automatic, for example machine, influencing of the first information are possible.

[0012] In some embodiments, the method comprises evaluating, for example using the ML system, a change in input data for a code generator or the code generator to determine whether the change in the input data leads to significantly changed output data according to at least one predefinable criterion, and, optionally, specifying the input data for processing by the code generator based on the evaluation, or, optionally, omitting execution of the code generator with respect to the input data. In some embodiments, this can, for example, prevent a renewed, i.e. repeated, execution of the code generator for input data whose changes would, for example, result in no or no significant change in the output data.

[0013] In some embodiments, the method comprises evaluating, for example, using the ML system, a change in input data for a compiler or the compiler to determine whether the change in the input data leads to significantly changed output data according to at least one predefinable criterion, and, optionally, specifying the input data for processing by the compiler based on the evaluation, or, optionally, omitting execution of the compiler with respect to the input data. In some embodiments, this can, for example, prevent a renewed, i.e., repeated, execution of the compiler for input data whose changes would, for example, result in no or no significant change in the output data.

[0014] In some embodiments, the method comprises: determining, for example by means of the ML system, whether calls to components of the development tool, for example with regard to predefinable input data, are parallelizable, and, optionally, temporally overlapping, for example in parallel, calling of the components of the development tool, for example with regard to the predefinable input data.

[0015] In some embodiments, the method comprises: executing a first software build, for example without using the first information, for example to create the program code for the computing device, for example for reference purposes, executing at least one further software build using the first information, and, optionally, using a cache memory, for example a shared cache memory, for at least temporarily storing data associated with at least a) the first software build or b) the at least one further software build.

[0016] In some embodiments, the method comprises extracting knowledge regarding previous software builds, for example by means of the development tool, deriving, based at least on the knowledge, at least one hypothesis regarding an effect of a change in input data, for example source code, for a software build by means of the development tool, and, optionally, using the at least one hypothesis for determining the first information.

[0017] In some embodiments, the method comprises: using at least one large language model, LLM, for the ML system, for example comprising using a first LLM, for example of the BERT (Bidirectional Encoder Representations from Transformers) type, for example for a prediction of whether or to what extent output data of a software build are affected by a change in the input data, for example comprising using a second LLM, for example of the Decoder-Transformer type, for example Bidirectional Tokenizer-Transformer 5, ByT5, for generative tasks, for example for predicting diagnostic information relating to a software build, providing at least one prediction for the development tool by means of the at least one LLM, and, optionally, using the at least one prediction by the development tool.

[0018] In some embodiments, the method comprises: providing a bit vector whose individual bit positions are each assigned to a component of the development tool, wherein a value of a respective bit position of the bit vector indicates whether the associated component of the development tool is to be used for a software build or whether it can be omitted, for example, and using the bit vector for the software build.

[0019] Further embodiments relate to an apparatus for carrying out the method according to the embodiments.

[0020] Further embodiments relate to a system for developing program code for a computing device, for example a microcontroller, comprising a device according to the embodiments, optionally the development tool, optionally the ML system.

[0021] Further embodiments relate to a computer-readable storage medium comprising instructions which, when executed by a computer, cause the computer to carry out the method according to the embodiments.

[0022] Further embodiments relate to a computer program comprising instructions which, when executed by a computer, cause the computer to carry out the method according to the embodiments.

[0023] Further embodiments relate to a data carrier signal that transmits and / or characterizes the computer program according to the embodiments.

[0024] Further embodiments relate to a use of the method according to the embodiments and / or the device according to the embodiments and / or the system according to the embodiments and / or the computer-readable storage medium according to the embodiments and / or the computer program according to the embodiments and / or the data carrier signal according to the embodiments for at least one of the following elements: a) specifying the input data for the development tool, b) predicting at least one aspect of a software build using the development tool, c) reusing existing, for example previously generated, output data of at least one component of the development tool, d) omitting an execution of at least one component of the development tool, for example for a software build using the development tool,e) separating dependencies of input data for at least one component of the development tool, f) accelerating software builds using the development tool, g) reusing input data for at least one component of the development tool.

[0025] Further features, possible applications, and advantages of the invention will become apparent from the following description of exemplary embodiments of the invention, which are illustrated in the figures of the drawing. All described or illustrated features, individually or in any combination, constitute the subject matter of the invention, regardless of their summary in the claims or their references, as well as regardless of their wording or representation in the description or in the drawing.

[0026] The drawing shows: Fig. 1 schematically shows a simplified flow diagram according to some embodiments, Fig. 2 schematically shows a simplified block diagram according to some embodiments, Fig. 3 schematically shows a simplified flow diagram according to some embodiments, Fig. 4 schematically shows a simplified flow diagram according to some embodiments, Fig. 5 schematically shows a simplified flow diagram according to some embodiments, Fig. 6 schematically shows a simplified flow diagram according to some embodiments, Fig. 7 schematically shows a simplified flow diagram according to some embodiments, Fig. 8 schematically shows a simplified flow diagram according to some embodiments, Fig. 9 schematically shows a simplified flow diagram according to some embodiments, Fig. 10 schematically shows a simplified flow diagram according to some embodiments, Fig. 11 schematically shows a simplified flow diagram according to some embodiments, Fig. 12 schematically shows a simplified block diagram according to some embodiments, Fig. 13 schematically shows a simplified flow diagram according to some embodiments, Fig. 14 schematically shows a simplified block diagram according to some embodiments, Fig. 15 schematically shows a simplified flow diagram according to some embodiments, Fig. 16 schematically shows a simplified flow diagram according to some embodiments, Fig. 17 schematically shows a simplified flow diagram according to some embodiments, Fig. 18 schematically shows a simplified flow diagram according to some embodiments, Fig. 19 schematically shows a simplified block diagram according to some embodiments, Fig. 20 schematically shows a simplified flow diagram according to some embodiments, Fig. 21 schematically shows a simplified flow diagram according to some embodiments, Fig. 22 schematically illustrates aspects of uses according to some embodiments.

[0027] Some embodiments, Fig. 1, Fig. 2, relate to a method, for example a computer-implemented method, for processing data generated by a development tool 10 ( Fig. 2) for creating program code for a computing device, for example a microcontroller, 20 associated data, comprising: determining 100 first information I-1, which characterizes a version ED-VERS of input data ED for the development tool 10, by means of an ML system 300 based on machine learning, ML, using 102 the first information I-1 for the development tool 10. This enables, in some embodiments, the specification of input data ED to be used for the development tool 10 using the ML system 300, whereby, in some embodiments, for example, certain versions ED-VERS of the input data ED to be used can be predicted by means of the ML system 300.

[0028] In some embodiments, the version ED-VERS of the input data ED characterizes a specific version of the input data ED, for example a specific development status and / or a specific variant of the input data ED.

[0029] In some embodiments, Fig. 1, it is thus provided that the method comprises: selecting 102a, based on the first information I-1, which version of the input data ED is to be used by the development tool 10.

[0030] In some embodiments, Fig. 2, the development tool 10 has or is at least one of the following elements: a) code generator 10-1, for example for generating source code SRC associated with the computing device 20, wherein, for example, the input data ED is at least partially provided for the code generator 10-1, or b) compiler 10-2 for generating a code associated with the computing device 20, for example executable by the computing device 20, EXE for example machine code and / or bytecode, wherein, for example, the input data ED is at least partially provided for the compiler 10-2.

[0031] In some embodiments, Fig. 2, the development tool 10 represents, for example, a software build toolchain or the development tool 10 is at least part of a software build toolchain. In some embodiments, the software build toolchain can be viewed as a collection of tools, e.g., software tools, that are used, for example, in a specific development process, for example, for creating the program code for the computing device 20. Optionally, in some embodiments, linkers, libraries, debuggers, and other tools can also be part of the software build toolchain or the development tool 10.

[0032] In some embodiments, Fig. 3, the method comprises: Predicting 110, by means of the ML system 300, which version ED-VERS of the input data ED is to be used by the development tool 10, Supplying 112 a corresponding version ED-VERS of the input data to the development tool 10 ( Fig. 2), and, optionally, processing 114 of the input data ED with the corresponding version ED-VERS by the development tool 10.

[0033] In some embodiments, Fig. 3, the supply 112 comprises at least one of the following elements: a) supply 112a of first input data, for example characterizing specification information, for a or the code generator 10-1 of the development tool 10, and / or b) supply 112b of second input data, for example characterizing source code SRC, for example source code associated with the computing device 20, for example source code created by means of the code generator 10-1, for a or the compiler 10-2 of the development tool 10.

[0034] In some embodiments, Fig. 4, the method comprises: storing 120 previously used input data ED', and, optionally, selectively using 122, for example reusing 122a, at least a part ED'' of the previously used input data ED', for example based on the first information 1-1.

[0035] In some embodiments, the previously used input data ED' comprises, for example, at least one of the following elements: a) input data previously used for the code generator 10-1, on the basis of which, for example, a specific version of a source code SRC was determined, or b) input data previously used for the compiler 10-2, on the basis of which, for example, a specific version of an executable code EXE was determined.

[0036] In some embodiments, Fig. 5, the method comprises: providing 130 an interface, for example user interface, 302 (see also Fig. 2) for receiving second information I-2, and, optionally, using 132 the second information to influence 132a the first information I-1. In some embodiments, this enables, for example, manual correction or modification of the first information I-1, for example by human experts. In some embodiments, the first information I-1 can also be influenced via the interface 302 by means of another system, e.g., an expert system. In some embodiments, combinations of human and automatic, e.g., machine, influencing of the first information I-1 are possible.

[0037] In some embodiments, Fig. 6, the method comprises: evaluating 140, for example by means of the ML system 300, a change CHG-ED-10-1 in input data ED-10-1 for a code generator 10-1 or the code generator 10-1 to determine whether the change in the input data leads to output data AD-10-1 that has been significantly changed according to at least one predefinable criterion KRIT1, and, optionally, specifying 142 the input data ED-10-1 for processing by the code generator 10-1 based on the evaluation 140, or, optionally, omitting 144 an execution of the code generator 10-1 with regard to the input data ED-10-1. In some embodiments, this can, for example, avoid a renewed, i.e. repeated, execution of the code generator 10-1 for input data ED-10-1 whose changes would, for example, cause no or no significant change in the output data AD-10-1.

[0038] The optional block 142a according to Fig. 6 symbolizes an optional specification of the changed input data, e.g., if the corresponding change in the input data leads to significantly changed output data AD-10-1 according to the criterion KRIT1.

[0039] The optional block 142b according to Fig. 6 symbolizes an optional non-specification of the changed input data, e.g. a failure to execute the code generator 10-1 with regard to the changed input data, e.g. if the corresponding change in the input data leads to output data AD-10-1 that is not significantly changed according to the criterion KRIT1.

[0040] In some embodiments, Fig. 7, the method comprises: evaluating 150, for example by means of the ML system 300, a change CHG-ED-10-2 of input data ED-10-2 for a compiler 10-2, respectively, to determine whether the change in the input data leads to output data AD-10-2 that has been significantly changed according to at least one predefinable criterion KRIT2, and, optionally, specifying 152 the input data ED-10-2 for processing by the compiler 10-2 based on the evaluation 150, or, optionally, omitting 154 an execution of the compiler 10-2 with respect to the input data ED-10-2. In this way, in some embodiments, for example, a renewed, i.e. repeated, execution of the compiler 10-2 for such input data ED-10-2 can be avoided, the changes of which would, for example, cause no or no significant change in the output data. The optional blocks 152a, 152b according to Fig. 7 correspond to the functionality of the optional blocks 142a, 142b according to Fig. 6.

[0041] In some embodiments, Fig. 8, the method comprises: Determining 160, for example by means of the ML system 300 ( Fig. 2), whether calls to components 10-1, 10-2 of the development tool 10, for example with respect to predefinable input data ED-10-1, ED-10-2, can be parallelized, and, optionally, temporally overlapping, for example parallel, calls 162 of the components 10-1, 10-2 of the development tool 10, for example with respect to the predefinable input data ED-10-1, ED-10-2.

[0042] In some embodiments, Fig. 9, the method comprises: executing 170 a first software build SB-1, for example without using the first information I-1, for example to create the program code for the computing device 20, for example for reference purposes, executing 172 at least one further software build SB-2, ... using the first information I-1, and, optionally, using 174 a, for example shared (i.e., e.g., jointly usable), cache memory CACHE for at least temporarily storing data which are associated with at least a) the first software build SB-1 or b) the at least one further software build SB-2, ...

[0043] In some embodiments, Fig. 10, the method comprises: Extracting 180 of knowledge KN-SB-PREV regarding previous software builds SB-PREV, see also Fig. 12, for example using the development tool 10, deriving 182 ( Fig. 10), based at least on the knowledge KN-SB-PREV, at least one hypothesis HYP-COD-CHG regarding an effect of a change in input data ED, for example source code SRC, for a software build by means of the development tool 10, and, optionally, using 184 the at least one hypothesis HYP-COD-CHG for determining the first information 1-1.

[0044] In some embodiments, Fig. 11, Fig. 12, the method comprises: using 190 at least one large language model, LLM, for the ML system 300, for example comprising using 190a a first LLM LLM-1, for example of the BERT, Bidirectional Encoder Representations from Transformers, type, for example for a prediction of whether or to what extent output data of a software build are affected by a change in the input data ED, for example comprising using 190b a second LLM LLM-2, for example of the Decoder-Transformer type, for example Bidirectional Tokenizer-Transformer 5, ByT5, for generative tasks, for example for predicting diagnostic information regarding a software build, providing 192 at least one prediction PRED for the development tool 10 by means of the at least one LLM LLM-1, LLM-2, and, optionally, using 194 the at least one prediction PRED by the development tool 10, for example for selecting, e.g.Versions of, input data ED.

[0045] In some embodiments, Fig. 13, the method comprises: providing 195 a bit vector BV, the individual bit positions of which are each assigned to a component 10-1, 10-2 of the development tool 10, wherein a value of a respective bit position of the bit vector BV indicates whether the associated component of the development tool 10 is to be used for a software build or whether it can be omitted, for example (i.e., for example, should not be used for a specific software build), and using 197 the bit vector BV for the software build SB.

[0046] Further embodiments, Fig. 14, refer to an apparatus 200 for carrying out the method according to the embodiments.

[0047] In further exemplary embodiments, it is provided that the device 200 comprises: a computing device (“computer”) 202 having at least one computing core, a memory device 204 assigned to the computing device 202 for at least temporarily storing at least one of the following elements: a) data DAT (e.g. the input data ED, e.g. in different versions, and / or output data of the components 10-1, 10-2, and / or data associated with the ML system 300), b) computer program PRG, for example for executing the method according to the embodiments.

[0048] In further exemplary embodiments, the memory device 204 comprises a volatile memory (e.g., random access memory (RAM)) 204a, and / or a non-volatile (NVM) memory (e.g., flash EEPROM) 204b, or a combination thereof or with other memory types not explicitly mentioned.

[0049] Further exemplary embodiments relate to a computer-readable storage medium SM comprising instructions PRG' which, when executed by a computer 202, cause the computer to carry out the method according to the embodiments.

[0050] Further exemplary embodiments relate to a computer program PRG comprising instructions which, when the program PRG is executed by a computer 202, cause the computer 202 to carry out the method according to the embodiments.

[0051] Further exemplary embodiments relate to a data carrier signal DCS that characterizes and / or transmits the computer program PRG according to the embodiments. The data carrier signal DCS can be transmitted (sent and / or received), for example, via an optional data interface 206 of the device 200.

[0052] Further embodiments, Fig. 2, refer to a system 1000 for developing program code for a computing device, for example a microcontroller, 20 comprising a device 200 according to the embodiments, optionally the development tool 10, optionally the ML system 300.

[0053] In some embodiments, the ML system 300 can be implemented, e.g., at least in part, by means of the device 200.

[0054] In some embodiments, the development tool 10 can be realized, e.g., at least partially, by means of the device 200 (not shown).

[0055] Fig. 15 schematically shows a simplified flow diagram according to some embodiments. It depicts calls (e.g., "calls") CG-1, CG-2, ..., CG-N-1 of at least one code generator 10-1 ( Fig. 2) and calls COMP-1, COMP-2, COMP-3, ..., COMP-M ( Fig. 15) (where, e.g. M=N) at least one compiler 10-2 ( Fig. 2), which e.g. at least part of the development tool 10 ( Fig. 2). The calls CG-1, CG-2, ..., CG-N-1 are assigned as input data, for example, specification information or data S0, S1, S2, ..., S N-1 (e.g. in the form of XML, e.g. ARXML, files) can be supplied, on the basis of which the calls CG-1, CG-2, ..., CG-N-1 e.g. source codes, e.g. in the form of source code files .c0, .c1, .c2, ..., .c N-1 , (e.g., without loss of generality, formulated in the computer language C). Optionally, include files .h0, .h1, .h2, ..., .h N-1 , eg comprising declarations, which in some embodiments, eg with a respective source code file, in a respective include directory incl0, incl1, incl2, ..., incl N be stored and made available in this form, for example, for the calls COMP-1, ...

[0056] In some embodiments, several, e.g. even different, code generators can be used, e.g. a code generator for a runtime environment (e.g. “RTE”), a code generator for a configuration of the computing device or a control unit, etc.

[0057] A first area B1 in Fig. 15, for example, symbolizes an area assigned to the input data S0, .c0, h0, a second area B2 symbolizes aspects of code generation, a third area B3 symbolizes aspects of compilation, and a fourth area B4 symbolizes aspects of further tools (not shown) such as linkers and / or post-processing tools.

[0058] In some versions, the example configuration of Fig. 15 by means of the ML system 300, e.g. using the first information I-1 ( Fig. 2), it can be specified which input data or versions of input data are to be used for the different calls CG-1, CG-2, ..., CG-N-1 of the code generators and / or calls COMP-1, COMP-2, COMP-3, ..., COMP-M of the compiler(s). This allows, for example, some embodiments to eliminate at least some of the calls or, for example, to execute them in parallel.

[0059] Fig. 16 schematically shows aspects of an i-th call CG-i of a code generator according to some embodiments. The switches S1, S2, ..., Si symbolize an optional supply of different versions of input data ED ( Fig. 2) for the i-th call CG-i of the code generator according to Fig. 15. For example, S0 symbolizes specification information or data of a first version, e.g. version, and S0prev symbolizes another, e.g. earlier, version, e.g. version of the specification information, which is symbolically represented by the block z -1 , a symbol representing a time delay in discrete-time signal processing theory. Using switch S1, controlled by a signal CTRL-S1, it is possible to select which version S0, S0prev which can be fed to the i-th call CG-i of the code generator. Comparable selection mechanisms are available for the other input data S1, ..., S i-1 or their different versions S1, S1prev, ..., S i , Siprev or versions are given by the additional switches S2, ..., Si or corresponding signals CTRL-S2, ..., CTRL-Si.

[0060] In some embodiments, one or more signals CTRL-S1, CTRL-S2, ..., CTRL-Si are, for example, based on or in the form of the first information 1-1 ( Fig. 1, Fig. 2) can be specified, so that the respective versions S1, S1prev, ..., S i , Siprev or versions can be specified as input data for the i-th call CG-i of the code generator, e.g. using the ML system 300.

[0061] In some embodiments, the switches S1, S2, ..., Si can also be regarded as symbolizing a conditional data flow for selecting a version of the respective input data.

[0062] In some embodiments, the first information I-1 can characterize a prediction that can be determined by means of the ML system 300, which respective version of the input data, e.g. in the sense of the switches S1, S2, ..., Si, is to be selected for the exemplary i-th call of the code generator according to Fig. 15.

[0063] In some embodiments, this is illustrated by way of example with reference to Fig. 16 applies to one or more calls to code generators and / or compilers and / or other components of the development tool.

[0064] In some embodiments, the first information can be present, for example, in the form of the already described bit vector BV, for example to control the use of a respective component of the development tool, for example a call to a code generator or compiler or the like, and / or to select a respective variant of the relevant input data, see the symbolic switches according to Fig. 16.

[0065] In some embodiments, dependencies of, for example, the i-th call of the code generator to previous input data can be maintained or broken in this way.

[0066] Fig. 17 schematically shows aspects of a call COMP-kI of a compiler according to some embodiments. A first signal crp, which can be predicted, for example, by means of the ML system 300 k,l controls the selection (see switch S1') of a specific version .c k,1 or .ck,lprev a source code QC-1 as input data for the compiler call COMP-kI. In some embodiments, further signals hrp, for example, predictable by the ML system 300, are k,l,0 , ..., crp k,l,N which allows a selection (see switches S2') of specific versions of include files, collectively in Fig. 17 designated by the reference symbol INCL-1, as further input data for the compiler's call COMP-kI.

[0067] In some embodiments, at least some of the signals crp k,l , hrp k,l,0 , ..., crp k,l,Ncharacterized or formed by the first information I-1, for example, organized in the form of the bit vector BV. For example, a value of "1" for a respective bit position of the bit vector BV can indicate that calls from the code generator and / or the compiler can be parallelized, e.g., because the value of "1" for a respective bit position of the bit vector BV indicates that an earlier version of input data should be used (e.g., value of the signal crp k,l of “1” controls switch S1' in Fig. 17 to the right to transfer data flow from .ck,lprev as input data for the compiler's COMP-kI call).

[0068] In some embodiments, Fig. 16, calls to the code generator can be omitted if all switches S1, S2, Si associated with the code generator indicate that earlier versions of corresponding input data for the code generator are to be used, because in this case previously determined output data of the code generator are reusable, e.g. retrievable from an optional cache.

[0069] In some embodiments, Fig. 17, calls to the compiler can be omitted if all switches S1', S2' associated with the code generator indicate that earlier versions of corresponding input data should be used for the compiler, because in this case previously determined output data .o k,l of the compiler are reusable, e.g. retrievable from an optional cache.

[0070] In some exemplary embodiments, a size of the bit vector BV, which in some embodiments may also be referred to as a prediction vector, may be determined based on |srp| + |crp| + |hrp| = N*(N+1) / 2 + (N+2)*(Z), where |srp| characterizes, for example, a number of the signals CTRL-S1, CTRL-S2, ..., CTRL-Si, where |crp| characterizes, for example, a number of the signals crp k,l where |hrp| is, for example, a number of signals hrp k,l,0 , ..., hrp k,l where, for example, N characterizes a number of include directories.

[0071] In some example embodiments, the bit vector may be provided in the form of a JSON file, for example.

[0072] In some embodiments, the exemplary configuration according to Fig. 17 a selective breaking of dependencies, e.g., regarding the compiler's call COMP-kI. For example, a prediction regarding a reuse of the source code QC-1 can be used by determining, depending on the signal crp k,l the current version .c k,l of the source code is used for compiling using the COMP-kI call or the, e.g. previous, version .ck,lprev. In some embodiments, something similar also applies, for example, to the selection of the include files INCL-1.

[0073] In some embodiments, Fig. 17, the source code or source code file is .c k,l for example, an I-th source code file (e.g. comprising source code in the computer language C), which has been generated, e.g., by a k-th code generator or k-th call of a code generator, or, e.g., manually created, e.g., in the case of k=0.

[0074] In some embodiments, Fig. 17, a reuse of include files, e.g. header files, can be achieved with a granularity at the directory level (e.g. the respective include directories incl k-1 , incl k , incl k+1 , ...), e.g. because the header files h k generated all at once, e.g., after executing the kth call of the code generator. In other words, in some embodiments, the switch S2' can be used to select, for example, a specific include directory and thus the include files contained therein for a COMP-kI call of the compiler.

[0075] Fig. 18 schematically shows a simplified flow diagram according to some embodiments, in which the ML system is symbolized by block 300. Similar to Fig. 16 symbolize the blocks CG-1, CG-2, ..., CG-N of Fig. 18 calls of at least one code generator, and the blocks COMP-0, COMP-1, COMP-2, COMP-N symbolize calls of at least one compiler. The blocks X0, X1, X2, ..., Xm symbolize input data for the code generator or the corresponding calls of the code generator. The arrows a1 symbolize, by way of example, changes to the input data X0. The arrows a2 symbolize an information flow from the ML system 300, e.g., in the sense of at least a portion of the first information I-1, for example, to resolve or break dependencies between the input data for the code generator and / or the compiler that are no longer required, e.g., based on the changes X0.

[0076] Fig. Figure 19 shows a comparison between a software build pipeline according to some conventional approaches, with chronologically successive calls CGC of a code generator and a subsequent compilation COMPC of the code (an imaginary time axis runs in Fig. 19 vertically downwards), and a software build pipeline according to some embodiments, in which at least some calls CGC', COMPC' of a code generator and / or compiler are at least partially parallelizable using the principle according to the embodiments, which can result in a time saving ZE in some exemplary embodiments.

[0077] Fig. 20 shows aspects of a use of the principle according to the embodiments, wherein block 300' symbolizes a variant of the ML system 300 according to exemplary embodiments. In some embodiments, the ML system 300' has a, for example, deep, for example, artificial, neural network NN, which is designed to output the first information I-1, e.g., in the form of a bit vector BV, e.g., prediction vector. Block 302 symbolizes a processing device that receives information a3 characterizing changes in the input data ED for the development tool. An optional coding device 304 transforms the information describing the changes obtained by means of block 302, e.g., into a first input data NN-ED-1 suitable for the neural network NN.An optional device 306 for evaluating dependencies with respect to the input data ED for the development tool creates - for example in cooperation with one or more further coding devices 307a, 307b - further input data NN-ED-2 for the neural network NN.

[0078] Block 308 symbolizes an optional decoder that decodes or evaluates output data of the neural network and outputs it, for example, via an optional demultiplexer 309.

[0079] In some embodiments, the demultiplexer is controllable by block 306, for example.

[0080] Block 310 symbolizes a training mechanism, e.g. based on a backpropagation method, for training the neural network NN.

[0081] In some embodiments, a learning vector LV for the training mechanism 310 is obtained, for example, from a comparison VERG between a software build SB-REF intended for reference purposes and one or more, e.g., predictive, software builds SB-PRED, e.g., executed using the first information I-1.

[0082] In some embodiments, a shared cache memory CACHE is provided, which enables an exchange of information between the different builds SB-REF, SB-PRED.

[0083] In some embodiments, the neural network NN can be used to determine whether, based on given changes a3 of the input data, at least some dependencies between processing steps or corresponding input data for a software build can be resolved or broken, wherein, for example, a result of these determinations can be provided in the form of the first information I-1, e.g., in order to execute at least one, for example, predictive, software build based on the result.

[0084] In some embodiments, information characterizing the dependencies to be checked between the processing steps can be supplied to the neural network, for example by means of the coding devices 307a, 307b, wherein, for example, the coding device 307a outputs such information to the neural network NN that characterizes a starting point for a dependency to be checked, wherein, for example, the coding device 307b outputs such information to the neural network NN that characterizes a target point for the dependency to be checked.

[0085] In some embodiments, various input data NN-ED-1, NN-ED-2 can be fed to the neural network NN, which data characterize both the changes a3 and aspects relating to at least one dependency to be checked.

[0086] In some embodiments, the neural network NN is configured to output a numerical value, e.g., a scalar, e.g., a floating-point number, which characterizes whether or not the at least one dependency described by the input data NN-ED-2 is resolved in view of the change(s) described by the input data NN-ED-1.

[0087] In some embodiments, the block 308 transforms the numerical value output by the neural network NN, for example, into a binary value, eg truth value, eg for controlling the demultiplexer 309, which in some embodiments is designed, for example, to form the bit vector BV.

[0088] In some embodiments, using the principle according to the embodiments, e.g. based on the first information I-1, a plurality of, e.g. predictive, software builds SB-PRED can be executed, e.g. each with at least one different parameter or parameter value for the neural network. For example, in some embodiments, a first software build can be executed based on the first information I-1, wherein the decoder 308 operates with a first threshold value that, e.g., enables a comparatively conservative prediction with regard to an analysis of dependencies in the first software build, and a second software build can be executed based on the first information I-1, wherein the decoder 308 operates with a second threshold value that is different from the first threshold value and that, e.g., enables a comparatively "risky" prediction with regard to the analysis of dependencies in the second software build.In this way, in some embodiments, different software builds can be carried out, each with different strategies regarding a predictive evaluation of dependencies on input data or possibly to be broken.

[0089] Processing steps for or in the development tool used 10.

[0090] In some embodiments, correctly formed output data (e.g. of the code generator and / or the compiler and / or at least one other component of the development tool 10) can already be temporarily stored in the cache memory CACHE, e.g. for future use or reuse for at least one software build, which can be accelerated, e.g.

[0091] In some embodiments, Fig. 20, the ML system 300' may also have a different structure and / or topology, e.g., instead of the neural network NN shown as an example. For example, the ML system 300' may also have one or more of the above-described examples with reference to Fig. 12, eg at least one first LLM LLM-1, or several LLMs LLM-1, LLM-2, eg a combination of a first LLM LLM-1 of the BERT type with a second LLM LLM-2, for example of the ByT5 type.

[0092] Fig. 21 shows exemplary aspects of data processing relating to the ML system 300, 300' according to some embodiments. Element E1 symbolizes a collection of data, for example output data (e.g., generated source code and / or generated executable code, etc.), which is obtained during software builds (e.g., conventional and / or based on the principle according to the embodiments), e.g., using the development tool 10. In some embodiments, the collection E1 is executed repeatedly, e.g., periodically with a first period of e.g., 1 day, for example, under the control of at least one script, e.g., automatically (e.g., without user interaction).

[0093] Element E2 symbolizes an update of a database E2a based on the collection E1. In some embodiments, the update E2 is repeated, e.g., executed periodically with a second period, for example, under the control of at least one script, i.e., e.g., automatically (e.g., without user interaction). For example, the second period can also be equal to the first period, e.g., 1 day. In some embodiments, information characterizing changes, which can be determined, e.g., in the collection E1 and / or update E2, can be used to train the ML system 300, 300'.

[0094] Element E3 symbolizes an update of "delta knowledge," e.g., knowledge regarding changes obtained, for example, from blocks E1, E2, which is repeated, e.g., periodically with a third period duration, e.g., under the control of at least one script, e.g., automatically (e.g., without user interaction). For example, the third period duration is longer than the first and / or second period duration; for example, the third period duration is 1 week.

[0095] In some embodiments, a possible information gain associated with the delta knowledge is evaluated in block E3, and if the information gain exceeds 80%, for example, the corresponding delta knowledge is accepted, e.g., used for training at least one component of the ML system 300, 300', which in some embodiments can also be done in block E3, for example.

[0096] Element E4 symbolizes aspects of the ML system 300, 300', e.g., a model based on a combination of a BERT-LLM LLM-1 and a ByT5-LLM LLM-2. In some embodiments, the model can be pre-trained, e.g., using the delta knowledge from element E3; see also the block arrow. In some embodiments, the model can be trained using at least some information obtained using at least one of the elements E1, E2, E2a, E3, e.g., even during operation of the model.

[0097] In some embodiments, the model of element E4 may be specifically adapted to the development system 10 or at least one component thereof, e.g., with respect to at least one of the following aspects: a) semantic investigations, e.g. calculations, b) semantic grouping, c) segmentation, d) checksums, e) Rules.

[0098] In some embodiments, the ML system 300, 300' or its model can be operated, for example, as a service that can be called, for example, via a network and that outputs, for example, the first information I-1 when called, see block arrow A5.

[0099] In some embodiments, the principle according to the embodiments can be used to create a system specialized for a specific area of ​​software development, for example an ML-based platform for deep learning, e.g. optimized for a specific software development area.

[0100] In some embodiments, the LLMs LLM-1, LLM-2 are optimized, for example, for a specific software development domain, e.g., for deeply embedded source code, such as can be used for control units, e.g., for vehicles.

[0101] Further exemplary aspects and embodiments are described below, which in some embodiments can be combined individually or in any combination with at least one of the above aspects.

[0102] In some embodiments, knowledge is extracted from previous software development runs, and from this, for example, hypotheses about the effects of source code changes are derived, preferably automatically (without human interaction).

[0103] In some embodiments, a rapid availability of, for example, high-quality predictions (e.g., regarding dependencies of individual aspects of a software build) is provided, which can, for example, represent valuable information for software developers, e.g., if the software build tool chain is a comparatively long-running one.

[0104] In some embodiments, the system 1000 or at least one component thereof is optimized, for example, to the needs of a particular software development domain, e.g., by supporting various types of source code formats, such as domain-specific input files from code generators such as AUTOSAR XML, proprietary configuration files for scripts, etc.

[0105] In some embodiments, for example, pre-structured training data for the model LLM-1, LLM-2 (or for a combined model, see element E4 of Fig. 21), which enables the model to learn efficiently.

[0106] In some embodiments, the system 1000 or the model tracks updates in the software architecture and / or in the processing tools (e.g., code generator and / or compiler and / or other tools) and closes knowledge gaps, e.g., by performing dynamic training.

[0107] In some embodiments, the ML system 300, 300' comprises a combination of transformer-based, e.g., pre-trained Large Language Models (LLMs) as described above.

[0108] In some embodiments, pure encoder-transformer models, e.g., belonging to the BERT family, are tuned for classification tasks, such as predicting whether the final software build output artifacts will be affected by a software change or not.

[0109] In some embodiments, decoder-transformer models such as ByT5 are tuned for generative tasks, such as the creation of a list of predicted diagnostic messages that are triggered during software development.

[0110] In some embodiments, an overall prediction is derived from the results of these two types of models LLM-1, LLM-2.

[0111] In some embodiments, e.g. for decoder models, an output format is proposed which can be referred to as a "rule binary string". In some embodiments, a sequence of the two characters "0" and "1" is trained for this output format, wherein a position index of each character indicates, e.g., whether a property is true or false, e.g., at least similar to the bit vector BV already described. For example, in some embodiments, if the model produces a '0' at a certain index position, the corresponding processing tool can be skipped, while a '1' at this index position means that the corresponding processing tool is to be used.

[0112] In some embodiments, a tokenizer of the models is adapted, for example, to use special segmentations according to the domain-specific file types such as C source file, C header file, AUTOSAR XML file and generic XML file.

[0113] In some embodiments, historical and current results, e.g. from continuous integration pipelines, are used for conditioning the models as follows: - Raw data from continuous integration pipelines is fetched from an artifactory (e.g. repository manager, e.g. tool for dependency management in software development), ie the source files, configuration files and generated files e.g. of each software build run. - Semantic filtering, e.g. specific for each file type belonging to a software build run, is applied to the content of e.g. each file, such as removing artifacts / tags / text between keywords etc. that are not needed for classification / generation tasks. - Filtered file contents are inserted into columns of a database. In some embodiments, the database schema is optimized so that training data for the models can be easily generated. - For example, pairs of software build runs are identified in the database, and for each pair, the deltas between the files belonging to the corresponding software build runs are calculated. The filtered file contents are cut into chunks of a specified size, e.g., 10 MB each, to speed up processing time. Checksums are calculated for the chunks in parallel. - After initial fine-tuning and evaluation, the models are trained and evaluated regularly, e.g., weekly, using deltas from the database. Deltas are only accepted for delta training, e.g., if the information gain exceeds a threshold (e.g., 80%), e.g., to avoid overtraining the model.

[0114] Further embodiments, Fig.22, relate to a use 400 of the method according to the embodiments and / or the device 200 according to the embodiments and / or the system 1000 according to the embodiments and / or the computer-readable storage medium SM according to the embodiments and / or the computer program PRG, PRG' according to the embodiments and / or the data carrier signal DCS according to the embodiments for at least one of the following elements: a) specifying 401 the input data ED for the development tool 10, b) predicting 402 at least one aspect of a software build by means of the development tool 10, c) reusing 403 existing, for example previously generated, output data of at least one component of the development tool 10, d) omitting 404 an execution of at least one component of the development tool 10, for example for a software build by means of the development tool 10,e) separating 405 dependencies of input data ED for at least one component 10-1, 10-2 of the development tool 10, f) accelerating 406 software builds using the development tool 10, g) reusing 407 input data ED for at least one component of the development tool.

Claims

[1] Method, for example a computer-implemented method, for processing data associated with a development tool (10) for creating program code for a computing device, for example a microcontroller (20), comprising: determining (100) first information (I-1) characterizing a version (ED-VERS) of input data (ED) for the development tool (10) by means of an ML system (300) based on machine learning, ML, using (102) the first information (I-1) for the development tool (10). [2] The method of claim 1, comprising: selecting (102a), based on the first information (I-1), which version of the input data (ED) is to be used by the development tool (10). [3] Method according to at least one of the preceding claims, wherein the development tool (10) comprises and / or is at least one of the following elements: a) code generator (10-1), for example for generating source code (SRC) associated with the computing device (20), wherein, for example, the input data (ED) are at least partially provided for the code generator (10-1), or b) compiler (10-2) for generating a code (EXE) associated with the computing device (20) and executable by the computing device (20), for example machine code and / or byte code, wherein, for example, the input data (ED) are at least partially provided for the compiler (10-2). [4] Method according to at least one of the preceding claims, comprising: predicting (110), by means of the ML system (300), which version of the input data (ED) is to be used by the development tool (10), supplying (112) a corresponding version of the input data (ED) to the development tool (10), and, optionally, processing (114) the input data (ED) with the corresponding version by the development tool (10). [5] Method according to claim 4, wherein the supplying (112) comprises at least one of the following elements: a) supplying (112a) first input data, for example characterizing specification information, for a or the code generator (10-1) of the development tool (10), and / or b) supplying (112b) second input data, for example characterizing source code, for example source code (SRC) associated with the computing device (20), for example source code created by means of the code generator (10-1), for a or the compiler (10-2) of the development tool (10). [6] Method according to at least one of the preceding claims, comprising: storing (120) previously used input data (ED'), and, optionally, selectively using (122), for example reusing (122a) at least a part (ED'') of the previously used input data (ED'), for example based on the first information (I-1). [7] Method according to at least one of the preceding claims, comprising: providing (130) an interface, for example a user interface, (302) for receiving second information (I-2), and, optionally, using (132) the second information (I-2) for influencing (132a) the first information (I-1). [8] Method according to at least one of the preceding claims, comprising: evaluating (140), for example by means of the ML system (300), a change (CHG-ED-10-1) of input data (ED-10-1) for a or the code generator (10-1) as to whether the change (CHG-ED-10-1) of the input data (ED-10-1) leads to output data (AD-10-1) that are significantly changed according to at least one predefinable criterion (KRIT1), and, optionally, specifying (142) the input data (ED-10-1) for processing by the code generator (10-1) based on the evaluation (140), or, optionally, omitting (144) an execution of the code generator (10-1) with regard to the input data (ED-10-1). [9] Method according to at least one of the preceding claims, comprising: evaluating (150), for example by means of the ML system (300), a change (CHG-ED-10-2) of input data (ED-10-2) for a or the compiler (10-2) as to whether the change (CHG-ED-10-2) of the input data (ED-10-2) leads to output data (AD-10-2) that have been significantly changed according to at least one predefinable criterion (KRIT2), and, optionally, specifying (152) the input data (ED-10-1-2) for processing by the compiler (10-2) based on the evaluation (150), or, optionally, omitting (154) an execution of the compiler (10-1) with regard to the input data (ED-10-2). [10] Method according to at least one of the preceding claims, comprising: determining (160), for example by means of the ML system (300), whether calls to components (10-1, 10-2) of the development tool (10), for example with regard to predefinable input data (ED-10-1, ED-10-2), can be parallelized, and, optionally, temporally overlapping, for example in parallel, calling (162) of the components (10-1, 10-2) of the development tool, for example with regard to the predefinable input data (ED-10-1, ED-10-2). [11] Method according to at least one of the preceding claims, comprising: executing (170) a first software build (SB-1), for example without using the first information (I-1), for example to create the program code for the computing device, for example for reference purposes, executing (172) at least one further software build (SB-2, ...) using the first information (I-1), and, optionally, using (174) a, for example shared, cache memory (CACHE) for at least temporarily storing data associated with at least a) the first software build (SB-1) or b) the at least one further software build (SB-2, ...). [12] Method according to at least one of the preceding claims, comprising: extracting (180) knowledge (KN-SB-PREV) regarding previous software builds (SB-PREV), for example by means of the development tool (10), deriving (182), based at least on the knowledge (KN-SB-PREV), at least one hypothesis (HYP-COD-CHG) regarding an effect of a change in input data, for example source code, for a software build by means of the development tool (10), and, optionally, using (184) the at least one hypothesis (HYP-COD-CHG) for determining (100) the first information (I-1). [13] Method according to at least one of the preceding claims, comprising: using (190) at least one large language model, LLM, (LLM-1, LLM-2) for the ML system (300), for example comprising using (190a) a first LLM (LLM-1), for example of the BERT, Bidirectional Encoder Representations from Transformers, type, for example for a prediction as to whether or notto what extent output data of a software build are affected by a change in the input data (ED), for example comprising using (190b) a second LLM (LLM-2), for example of the decoder-transformer type, for example Bidirectional Tokenizer-Transformer 5, ByT5, for generative tasks, for example for predicting diagnostic information relating to a software build, providing (192) at least one prediction (PRED) for the development tool (10) by means of the at least one LLM (LLM-1, LLM-2), and, optionally, using (194) the at least one prediction (PRED) by the development tool (10). [14] Method according to at least one of the preceding claims, comprising: providing (195) a bit vector (BV), the individual bit positions of which are each assigned to a component (10-1, 10-2) of the development tool (10), wherein a value of a respective bit position of the bit vector (BV) indicates whether the associated component (10-1, 10-2) of the development tool (10) is to be used for a software build (SB) or whether it can be omitted, for example, and using (197) the bit vector (BV) for the software build (SB). [15] Device (200) for carrying out the method according to at least one of the preceding claims. [16] System (1000) for developing program code for a computing device, for example a microcontroller, comprising a device (200) according to claim 15, optionally the development tool (10), optionally the ML system (300). [17] Computer-readable storage medium (SM) comprising instructions (PRG) which, when executed by a computer (202), cause the computer to carry out the method according to at least one of claims 1 to 14. [18] Computer program (PRG) comprising instructions which, when the program (PRG) is executed by a computer (202), cause the computer (202) to carry out the method according to at least one of claims 1 to 14. [19] Data carrier signal (DCS) which transmits and / or characterises the computer program (PRG) according to claim 18. [20] Use (400) of the method according to at least one of claims 1 to 14 and / or the device (200) according to claim 15 and / or the system (1000) according to claim 16 and / or the computer-readable storage medium (SM) according to claim 17 and / or the computer program (PRG) according to claim 18 and / or the data carrier signal (DCS) according to claim 19 for at least one of the following elements: a) specifying (401) the input data (ED) for the development tool (10), b) predicting (402) at least one aspect of a software build (SB) by means of the development tool (10), c) reusing (403) existing, for example previously generated, output data of at least one component (10-1, 10-2) of the development tool (10), d) omitting (404) an execution of at least one component (10-1, 10-2) of the development tool (10), for example for a software build (SB) using the development tool (10),e) separating (405) dependencies of input data (ED) for at least one component (10-1, 10-2) of the development tool (10), f) accelerating (406) software builds (SB) by means of the development tool (10), g) reusing (407) input data (ED) for at least one component (10-1, 10-2) of the development tool (10).

Citation Information

Patent Citations

  • Flow based fault testing

    US20140047275A1

  • Using comments of a program to provide optimizations

    US20190146764A1