A computer-implemented method for software system modernization

The method uses LLMs to analyze, decompose, and rearrange legacy software components, enabling automated and efficient modernization by preserving functionality and reorganizing them into a coherent system in a different technology.

WO2025196200A1PCT designated stage Publication Date: 2025-09-25CGI SUOMI OY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/057662
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-20
Filing Date
2025-03-20
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing software modernization methods lack automation and efficiency in transforming legacy systems using large language models (LLM) to create modernized systems without losing functionality.

Method used

A computer-implemented method utilizing LLMs to analyze, decompose, and rearrange components of legacy software systems, generating new building blocks that are causally linked to form a modernized system, with stages including as-is codebase analysis, to-be design generation, and code generation.

Benefits of technology

Enables automated and efficient modernization of software systems by preserving functionality and reorganizing components in an ordered manner, producing a coherent system in a different technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025057662_25092025_PF_FP_ABST
    Figure EP2025057662_25092025_PF_FP_ABST
Patent Text Reader

Abstract

A computer-implemented method for software system modernization comprising utilizing a large language model (LLM) to process one or more source files, wherein the processing comprises, as one or separate steps, reading the one or more source files with the LLM to analyze the function of the one or more source files, utilizing the same or another LLM, processing one or more source files to discover public definitions and declarations in the one or more source files, wherein said discovered definitions may be expressed with the declared identifiers and context of the declaration, such as namespace, and then semantically indexed, detect references to external identifiers made in the one or more source files, wherein said references may be expressed with the referenced identifiers and context of the reference, match a definition for each discovered reference, wherein a candidate definition expression is formed based on the reference, and a semantically indexed form for said matching is produced with an LLM from the candidate definition expression, for each matched definition, extract the semantics of the reference from the context of the source code having the reference declaration, to a referenced functionality expression, process each source file into one or more components, wherein each component comprises one or more functionally inherent parts present in a source file, generate a design description for each component using the LLM to extract a design summary of the functionality present in the source code for each component so that the summary includes the information, which information may enable redesigning the same functionality in different technology, providing a target design generation configuration and mapping configuration, and mapping components in reference to said target design generation configuration and mapping configuration, based on the mapping, linking the components into system structures and the referenced functionality links across the components, arranging components in an optimal generation order within each system structure so that the most isolated components are ordered first, and that the most dependent components are ordered last, utilizing the same or another LLM, in view of the target design generation configuration and mapping configuration, generate a codebase from the system structures and components.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] A COMPUTER-IMPLEMENTED METHOD FOR SOFTWARE SYSTEM MODERNIZATION

[0002] FIELD OF THE INVENTION

[0003] The present invention relates to computer-implemented methods. In particular, however not exclusively, the present invention pertains to a computer-implemented method for software system analysis and modernization.

[0004] BACKGROUND

[0005] In the advent of generative Al, generating source code by human prompt has been gaining traction. The process of generating prompts themselves can be automated as well with generative Al, for instance extracting design specifications from source code. Combination of these capabilities opens up new avenues in transforming source code, like a human programmer would be able to conduct with manual work. Known prior solutions range from translating a set of instructions in a first programming language into second instructions in a second programming language such as in EP2979176B1 with a model driven approach. Other known solutions include monolith database to distributed database transformation such as US11615076B2 or a system and method for optimized modernization of applications as is the case in US20230401038A1.

[0006] SUMMARY OF THE INVENTION

[0007] The objective of the present disclosure is to at least alleviate the problems described hereinabove and to provide a solution allowing for example alteration of a legacy software system. Another objective of the present disclosure is to provide a solution that can analyze and understand functions of an existing software system and make changes to the software system. The objective is generally achieved with a computer-implemented method, and computer program product in accordance with the present disclosure.

[0008] The objective technical problem of the current patent application is how to provide an automated and autonomous computer-implemented method for modernizing a software system with the use of language model(s). The technical effect of the current patent application is achieved through the automated and sequential computer-implemented method which uses LLM and foundational building blocks of a software system to create new building blocks that are gathered and arranged autonomously by the method to causally create a modernized system of the old system.

[0009] Accordingly, in one aspect of the present disclosure a computer-implemented method for software system modernization comprising utilizing a large language model (LLM) to process one or more source files, wherein the processing comprises, as one or separate steps,

[0010] • reading the one or more source files with the LLM to analyze the function of the one or more source files, utilizing the same or another LLM, processing one or more source files to

[0011] • discover public definitions and declarations in the one or more source files, wherein said discovered definitions may be expressed with the declared identifiers and context of the declaration, such as namespace, and then semantically indexed,

[0012] • detect references to external identifiers made in the one or more source files, wherein said references may be expressed with the referenced identifiers and context of the reference,

[0013] • match a definition for each discovered reference, wherein a candidate definition expression is formed based on the reference, and a semantically indexed form for said matching is produced with an LLM from the candidate definition expression,

[0014] • for each matched definition, extract the semantics of the reference from the context of the source code having the reference declaration, to a referenced functionality expression,

[0015] • process each source file into one or more components, wherein each component comprises one or more functionally inherent parts present in a source file,

[0016] • generate a design description for each component using the LLM to extract a design summary of the functionality present in the source code for each component so that the summary includes the information, which information may enable redesigning the same functionality in different technology, providing a target design generation configuration and mapping configuration, and mapping components in reference to said target design generation configuration and mapping configuration, based on the mapping, linking the components into system structures and the referenced functionality links across the components, arranging components in an optimal generation order within each system structure so that the most isolated components are ordered first, and that the most dependent components are ordered last, utilizing the same or another LLM,

[0017] • in view of the target design generation configuration and mapping configuration, generate a codebase from the system structures and components.

[0018] Accordingly, to another aspect of the present disclosure a data processing system comprising means for carrying out the method of any one of claims 1-15.

[0019] Accordingly, to another aspect of the present disclosure a computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any one of claims 1-15.

[0020] Accordingly, to another aspect of the present disclosure a computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of any one of claims 1-15.

[0021] The solution produces a technical effect by splitting a (legacy) source file into components which may be converted to different technology without losing their functionalities and are linked across the high-level system structures in an ordered manner to produce a coherent system in different technology. The technical effect is used to solve a technical problem of how to modernize a software system by automation and use of large language models.

[0022] Different embodiments of the present disclosure are disclosed in the dependent claims.

[0023] The expression “a number of’ refers herein to any positive integer starting from one (1), e.g. to one, two, or three.

[0024] The expression “a plurality of’ refers herein to any positive integer starting from two (2), e.g. to two, three, or four.

[0025] BRIEF DESCRIPTION OF THE RELATED DRAWINGS Next the invention is described in more detail with reference to the appended drawings in which

[0026] Fig. 1 illustrates a flow diagram depicting key entities of a 3-stage process in accordance with the present disclosure,

[0027] Fig. 2 illustrates a flow diagram depicting a summary of the 3-stage modernization configurations and inputs over the stages in accordance with the present disclosure, Fig. 3 illustrates a flow diagram depicting a summary of the technological clustering process to derive the contents for the legacy architecture configuration in accordance with the present disclosure.

[0028] DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] In accordance with one embodiment of the present disclosure, the solution may be roughly split into two entities: a language model and automation environment. The first entity, the language model environment comprises a language model(s) which may be an inference model, such as OpenAI GPT-4o, and an embedding language model, such as OpenAI text-embedding-ada-002. The solution may access a third- party API interface, which is used to call the language models by the automated modernization process. A skilled person will recognize that these models evolve rapidly, and a variety of language models may be suitable for the purpose of the present disclosure.

[0030] The second entity, the automation environment comprises at least a container image, a configuration file(s), a relational database, and source files. The container image operates in isolation and consists of the executable program code, i.e. the container image holds the automated modernization process. The automated modernization process executes the actual modernization work on the source files and is also able to call the API interfaces. The configuration file stores the required system architecture and overall requirements for the modernization of the system, including the API interface configurations. The configuration file is read as the automated modernization process is initiated. The relational database stores intermediate state data of the automation process. Further, the automation process reads the state of the constructional model for the next phase of the automation process. Source files used in the modernization process are included as references in the configuration file and are read as the automated modernization process is initiated. To accomplish the described automated modernization some initial settings from a human user may be required. The user prepares the integration and / or connections from the automation environment to the language model API interfaces. Used configuration file is adjusted to be suitable for the desired modernization, for example, the target modernization programming language is set to a desired language, such as COBOL, Java or any other desired programming language. The user executes the automated modernization process which will then trigger the set of automated events and API calls. Finally, the user stores the automatically created modernized code files and documentation.

[0031] The method for a software system modernization of the present disclosure, which software system may be for example a legacy codebase, comprises at least three different stages: as-is codebase analysis (110), to-be design generation (124) and code generation (140). Each stage is split into a varying number of steps, some of which are optional and as such may be used for optimization or customize default values. Although software system modernization is discussed, the present disclosure may be used to analyze and alter a software system so that the technology may be altered, and the code of the software system may be changed whereat actual “modernization” of the technology needs not be the objective, and the objective may be to simply alter the analyzed codebase and / or produce new code so that the analyzed software system is altered. Alternatively, it is also possible that no alteration is done and the present disclosure may be used only for analysis, documentation or other purposes that do not require alteration of the technology. Further below, although legacy system is explicitly discussed, the expression “legacy system” may herein pertain to any existing software system. Likewise, the expression “target system” is used to refer to a software that the legacy system is modified to.

[0032] The first stage “as-is codebase analysis” (110) comprises a number of steps. As an input, the as-is codebase analysis (110) accepts a codebase (106) that includes source code and / or as-is architecture configuration (AIAC) (108) i.e. architecture configuration of the legacy system. A default configuration may also be used. As- is components (AIC) (116), i.e. components of the legacy system, are the principal and atomic entities for decomposing the functionality of the as-is codebase (106), i.e. codebase of the legacy system. The source code, which may be legacy source code, is processed to produce as-is design (AID) documentation (112), i.e. design documentation of the legacy system. However, the as-is codebase analysis (110) is not necessarily limited only to accepting the above examples and may accept other forms of code or textual data processable by the system. Below, the steps are explained in more detail. Some of the steps may be optional, and the order and / or number of steps per as-is codebase analysis (110) is necessarily not defined only by the steps below.

[0033] In technology analysis, which may be an optional step, each source file (302) is processed with a large language model (LLM) to produce an as-is technology summary (304) for each, i.e. what high level technologies are observed to be leveraged in the file and what technical purpose the file serves. A source file (302) comprises at least software source code. Each as-is technology summary (306), i.e. technology summary of the legacy system, is then semantically indexed with an LLM, i.e. an embedding (308) is generated. The embedding vectors (310) are then processed using a clustering algorithm (312) with optimal clustering composition (314) iterated to with silhouette analysis (316, 318), a measure of similarity for a data point within a cluster compared to other clusters. This process automatically discovers the technical clusters present in the as-is codebase with the respective source files and their technology summaries. It can be used to provide information about the nature of an unknown as-is codebase, to for instance detail the Al AC (320) with as-is codebase specific instructions in terms of introducing suitable tags.

[0034] In preprocessing, which may be an optional step, each source file (102) is amended with the content from source files (102) that are statically included into it, if any, so that the source files (102) would contain full information relevant to the implementation built into the source file (102). A source file (102) comprises at least software source code. The amending is done based on AIAC (108) with configurable implementation (LLM based or parsing based), with included source files (102) discarded from further processing as their content has been already included into other source files (102, 104).

[0035] Tagging, which may be an optional step, is a process where each source file (102) is tagged with LLM to 0..n tags that have their matching criteria defined in the AIAC (108) in natural language. The tags are used as markers for mapping each file (210) to as-is design generation configuration (208) in AIAC (108) in a dynamic manner so that the contents and composition of AID (112) can be optimized and adjusted beyond default form, if necessary.

[0036] Cross-references definitions discovery uses an AIAC-configured implementation, where each source file (102) is processed to discover the public definitions and declarations in the source file (102). The discovered definitions are expressed with the declared identifiers and context of the declaration, such as namespace, and then semantically indexed (i.e. embedded with an LLM) to a modernization database (MD). The processing and definitions discovery is by default conducted with an LLM with a series of prompts that first discover the relevant namespace of the code and subsequently supported with that information obtain the definitions and declarations with the LLM and expressed in natural language. Alternatively, the definitions discovery may be supported by a technology specific programming language parser when configured to be used with source files having selected tags configured in the AIAC. The resulting parse tree is processed for extracting the programming language specific definition entities and their context for semantically indexing them to a modernization database.

[0037] Cross-references references discovery uses an AIAC-configured implementation, where each source file (102) is processed to detect the references to external identifiers made in the source file (102). The references are expressed with the referenced identifiers and context of the reference, and stored to the MD. The processing and references discovery is by default conducted with an LLM with a series of prompts that first discover the relevant namespace of the code and subsequently supported with that information obtain the external references with the LLM and expressed in natural language. Alternatively, the references discovery may be supported by a technology specific programming language parser when configured to be used with source files having selected tags configured in the AIAC. The resulting parse tree is processed for extracting the programming language specific reference entities and their context for semantically indexing them to a modernization database.

[0038] In cross-referencing each reference that is discovered in Cross-references references discovery for each source file (102), a corresponding definition match is attempted: a candidate definition expression is formed based on the reference, and a semantically indexed form (i.e. an embedding) for said matching is produced with an LLM from the candidate definition expression. A cosine similarity search is performed with this embedding over the definition expressions embeddings created in prior cross-references definitions discovery step and stored in the MD. In case a match is found where the similarity is over a configured threshold, the matched reference-to-definition mapping is stored in the MD.

[0039] In cross-referencing referenced functionalities (122), the semantics of the reference of each matched reference-to-definition mapping are extracted using an LLM from the context of the source code having the reference (120) declaration, to a natural language referenced functionality expression, and these referenced functionalities (122) are stored in the MD along with the reference-to-definition mappings.

[0040] For as-is components decomposition, i.e. components decomposition of the legacy system, each source file (102) is processed into L.n AICs (116) based on the implementation configured into use in the AIAC (108). The processing is by default conducted using LLM to decompose the source file (102) into functionally inherent parts by using the LLM to discover the functionalities present in the code file (102) and expressing them as distinct modules, along with the declared references (120) and definitions (118) from the respective source code carried over to each AIC (116). Thereby each source file (102) is decomposed into functionally logical smaller AIC (116) parts using the LLM. Alternatively, the decomposition into functionally inherent parts may be supported by a technology specific programming language parser when configured to be used with source files having selected tags configured in the AIAC. The resulting parse tree is processed for discovering the programming language specific logical structure of the source code, and then continuing with the LLM based approach to discover the functionalities of each structural part as their own module.

[0041] In as-is components design, i.e. components design of the legacy system, an AID (114) is generated for each AIC (116) using LLM to extract a design summary of the functionality present in the source code (102) for each AIC (116) so that the summary includes the information to redesign the same functionality in different technology in the subsequent stage. The extraction process is controlled by dynamic prompting where the design is produced from source code (102) for each AIC (116) for its functionality and using the previously associated tags to optionally refine / adjust the design for specific types of source code components, and / or design documentation objectives, using tag-based control prompt configurations (214) in the AIAC (108). Each AID (114) is stored into the MD along with the AIC (116) and its referenced functionalities (122).

[0042] In as-is design documentation gathering all AIDs (114) are combined to one documentation base and interlinked based on the referenced functionalities to form the as-is design documentation (112). The second stage, also called the to-be design generation (124), i.e. design generation of the target system, of the process uses the as-is components (116) for generating the to-be designs (TBD) (132), i.e. designs of the target system, based on the as-is components (116), their designs (114) and referenced functionalities (122). These entities are used together with the to-be architecture configuration (TBAC) (128), i.e. architecture configuration of the target system, for specifying details for the to-be state. Relations between as-is components (116) and to-be components (TBC) (130), i.e. components of the target system, are described in a granular fashion in natural language where to-be components (130) are the principal and atomic entity for specifying the to-be state design. Some of the steps may be optional, and the order and / or number of steps per to-be design generation is necessarily not defined only by the steps below.

[0043] In solution building block (SBB) composition AICs (116) are mapped to system structures, SBB(s) (126), using the to-be design generation configuration (228), i.e. design generation configuration of the target system, and the mapping configuration per system structure SBB (126) in the TBAC (128). SBB(s) (126) represent the highest level of grouping to reorganize the functionality embodied in the AICs (116) for realizing the modernization case specific architecture for the to-be state, in other words, solution building blocks (126) are the aggregating entities to enable architectural transformations of the system. In this step the AICs (116) are so called to-be AICs (116), i.e. legacy components of the target system, and they retain the referenced functionality links of AICs (116) between them.

[0044] In solution building block linking, the to-be AICs (116) of each SBB (126) are linked and ordered according to the new SBB (126) structure and the referenced functionality links across the to-be AICs (116): the to-be SBBs (224), i.e. SBB or system structure of the target system, are arranged in an optimal generation order with the most isolated SBBs (126) ordered first, and the most dependent SBBs (126) ordered last.

[0045] In to-be components design, i.e. components design of the target system, for each ordered SBB (126), each to-be AIC (116) mapped to the SBB (126) is processed with an implementation configured in the TBAC (128), resulting in the TBCs (130) of the SBB (126). The default LLM based implementation uses the configuration of the SBB (126) in TBAC (128), including natural language specifications of the technology stack, technology components summary, and inter-SBB integration, as an input together with the input AID (114) and the inbound / outbound referenced functionalities (138) to reason with LLM to produce the to-be design (TBD) (132), design of the target state, for the TBC (130). The TBD (132) includes functional requirements extracted from the AID (114) by the LLM and corresponding nonfunctional requirements designed according to the specified TBAC (128) by the LLM. In case the TBC (130) is referenced from outside the SBB (126), a new interface TBC (130) is generated which creates a new interface according to the inter-SBB integration specification to act as an interface that is exposed for the TBC(s) (130) of the dependent SBB (126) to connect to the component and obtain access to the referenced functionality (138). Conversely the relations of the referenced functionalities (138) are modified to accommodate the interface TBC (130), so that due to the ordering of the SBB (126) the interface TBCs (130) are generated before the dependent TBCs (130) in the dependent SBBs (126) are generated with the updated referenced functionalities (138). This effectively allows to reorganize the fundamental organization embodied in the AICs (116) and where needed, amend new functional interfaces to enable new references. TBCs (130) with their TBDs (132) are stored into the MD.

[0046] In to-be design documentation, i.e. design documentation of the target system, gathering all TBCs (130) are combined to one documentation base and interlinked based on the referenced functionalities (138) to form the to-be design documentation (134).

[0047] The third stage, the code generation module (140), is when the to-be code will be generated. For the input the codebase generation (140) uses the designs (132) and referenced functionalities (138) of the to-be components (130) of the SBBs (126). Codebases (154) are generated “inside-out” starting from independent components with reference relations being progressively fulfilled from the core to the edges. Reference relations are fulfilled by designing the externally visible identifier declarations of each dependent upon component, and which are distributed as available interfaces for the dependent components design. Codebases (154) are realized by top-down planning starting from technology summarization and fundamental structures, and incrementally detailed with structure and code files to realize each component in their dependency order from reference relations. Some of the steps may be optional, and the order and / or number of steps per code generation is necessarily not defined only by the steps below.

[0048] In technology solution realization for each SBB (126) a technology solution (142) is generated for realizing its TBCs (130), using the to-be design generation configured in the TBAC (128), and the configured technology solution generation implementation. In the default technology solution generation implementation, for each technology solution the following steps are executed, and results stored into the MD.

[0049] In codebase technology planning (CTP) all TBCs (130) of the SBB (126) with their respective TBDs (132) are iteratively summarized by LLM with the configuration in TBAC (128), including natural language specifications of the technology stack and technology components summary, to produce a codebase technology plan (144) for realizing the codebase for the technology solution (142). The CTP (144) includes a high-level mapping of the TBCs (130) to the provided technology components summary that guides the eventual realization of the TBCs (130) to match with the targeted to-be architecture.

[0050] For codebase structure planning using the CTP (144) and the configuration in TBAC (128), including natural language specifications of the technology stack and technology components summary, LLM is used to plan the file system level structure of the upcoming codebase of the technology solution. The LLM-planned codebase structure plan (146) (CSP) is used as a basis for applying further LLM- planned changes per TBC (130) to be implemented into the codebase.

[0051] In to-be component ordering, i.e. component ordering of the target system, all TBCs (130) in the SBB (126) are ordered so that given the incoming and outgoing referenced functionalities (138) of the TBC (130), the most independent TBCs (130) (no outgoing referenced functionalities) come first and the most dependent TBCs (130) come last. Using this ordering each TBC (130) is processed as follows.

[0052] In to-be component structure planning, i.e. component structure planning of the target system, for each TBC (130) a codebase component structure plan (CCSP) (148) is produced using LLM based on the TBD (132), CSP (146) and the referenced functionalities (138). The CCSP (148) includes all of the folders, files and purposes for each for realizing the TBD (132), and it is applied on top of the CSP (146).

[0053] In to-be component declared identifiers planning, i.e. component declared identifiers planning of the target system, for each TBC (130) LLM is used to produce a codebase component identifiers plan (CCIP) (150) based on the TBD (132), codebase component structure plan (CCSP) (148) and the incoming referenced functionalities (138). The CCIP (150) is a detailed plan based on the CCSP (148) for capturing as planned by LLM what identifier signatures need to be made visible by the implementation of each file towards the other files so that the referenced functionalities (138) can be implemented by leveraging these visible identifiers.

[0054] In to-be component implementation generation, i.e. component implementation generation of the target system, for each file planned to be implemented in the CCSP (148) for the TBC (130), LLM is used to produce the content of the file. The LLM is provided with the technology stack details from TBAC (128), the TBD (132), the file changes as planned in CCSP (148), and the identifiers as planned in CCIP (150) and supplied via the TBC (130) ordering the identifiers that have been planned from the TBCs (130) providing the referenced functionalities (138) of this TBC (130). Thus, the LLM is supplied with the full context to generate code content: via the TBD (132) we inform the functional and non-functional requirements and provide the necessary contextual information via the previously computed plans on how the specific file is positioned in the codebase and how it interacts with the other entities in the codebase. After all TBCs (130) have been generated the remaining bootstrap files that may have been planned in the CSP (146) are generated as individual files.

[0055] In codebase generation files generated in the process and stored to the MD are written to the local file system in the folder structure planned, thus realizing a complete, generated codebase (154).

[0056] The method of the present disclosure is provided as a computer implemented method, which may be carried out on a computer, computer network or the like computing means. The LLM and discussed algorithms may be part of a single software system or some tasks or processes may be run external to the system, so that the whole method may be realized as a software on a terminal or server, or some functions or LLMs may be run on a separate server, cloud, or the like, and accessed thereof.

[0057] The one or more LLMs utilized by the method may be tuned using training data but commonly available LLMs, which can read files, understand source code and text, and provide code as output may also be used. In other words, training of the LLM is not mandatory and open versions of LLM may be used directly. Fig. 1 depicts a flow chart of key entities of the three stages. In the figure, boxes marked with an inner square in the upper-right corner indicate a step of the process, or entities that are formed or generated fully or partially based on LLM output. First, in to-be code input context (101), i.e. code input context of the legacy system, the as-is codebase (102) file may be preprocessed (104) to disregard any content that has been already included into other source files. Each in-scope as-is codebase file (102), i.e. codebase file of the legacy system, is processed into one or more as- is components (116) and designs (114) are generated for each based on the as-is architecture analysis configuration (108). The as-is codebase (106) and the as-is architecture configuration (108) are then used by the as-is codebase analysis module (110) to generate an as-is documentation (112). Additionally, each source file (102) is processed into l..n as-is components (116) using LLM based on the implementation configured into use in the as-is architecture configuration (108). In as-is design generation context (113), i.e. design generation context of the legacy system, as-is component design (114), i.e. component design of the legacy system, is done for each as-is component (116), where an as-is component design (114) is generated automatically using the LLM by extracting a design summary of the functionality present in the source code for each as-is component (116). The extraction process is controlled by automatic and or dynamic prompting where the design is produced from as-is component code definition (118), i.e. component code definition of the legacy system, and previously associated tags or references that include as-is component code reference (120), i.e. component code reference of the legacy system, and as-is component referenced functionality (122), i.e. component referenced functionality of the legacy system. The formed as-is component designs (114) are combined to one documentation base and interlinked based on the referenced functionalities to form the as-is design documentation (112). The to-be design generation module (124) maps as-is components (116) to the solution building block (126) in to-be design generation context (131), i.e. design generation context of the target system, where solution building blocks (126) represent the highest level of grouping for reorganizing the functioning of the as-is components (116) for the to-be state, i.e. the modernized version or target version of the system. At this point the as-is components (116) are called “to-be as-is components” and they retain their referenced functionalities. The solution building blocks (126) are then linked and ordered with the most isolated solution building blocks (126) ordered first and the most dependent solution building blocks (126) ordered last. The newly ordered solution building blocks (126) are then processed using an implementation configured to the to-be architecture configuration (128) to dynamically build to-be components (130) of the solution building block (126). Each as- is component (116) is designed to to-be components (130) in one or more solution building blocks (126) based on the to-be architecture configuration (128). To-be component design (132) is generated using an LLM-based implementation. The to-be component design (132) includes functional requirements extracted from the as-is component design (114) and corresponding non-functional requirements designed according to the specified to-be architecture configuration (128) by the LLM. The to-be documentation (134), i.e. documentation of the target system, is created by combining all the to-be components (130) to one documentation base and interlinked based on the to-be component references (136), i.e. component references of the target system, and to-be component referenced functionalities (138). In the code generation module (140), for each solution building block (126) a technology solution (142) is generated for realizing the to-be components (130) of the solution building block (126) using the to-be architecture configuration (128) which includes the to-be design generation, i.e. design generation of the target system. Each solution building block (126) is realized as a technology solution (142), each generating a codebase (152) and code files based on the to-be components (130). A codebase technology plan (144) is produced from the to-be architecture configuration (128) using LLM. The codebase technology plan (144) is used for realizing the codebase (152) for the technology solution (142). Based on the codebase technology plan (144) and the to-be architecture configuration (128) a codebase structure plan (146), a file system level structure of the upcoming codebase of the technology solution (142) may be generated dynamically using the LLM. Using the to-be component designs (132), i.e. component designs of the target system, the codebase structure plan (146) and the referenced functionalities, the LLM creates a codebase component structure plan (148). An LLM is also used to produce a codebase component identifiers plan (150) for each to-be component (130) based on the to-be component design (132), codebase component structure plan (148) and the incoming referenced functionalities. Using all the files generated in the final process, generated codebase file (152) can be created. All these files are also stored to the modernization database and are written to the local file system realizing a complete generated codebase (154).

[0058] Eig. 2 depicts a flow chart of a summary of configurations and inputs over the three stages. With prompt-based design generation (202) the configuration consists of the prompts to generate the design based on the source based on the as-is architecture configuration (216), i.e. architecture configuration of the legacy system. Supported prompt-based design generation strategies (204) may be used by as-is design generation strategy (206), i.e. design generation strategy of the legacy system, which can be compiled to the as-is design generation configuration (208), i.e. design generation configuration of the legacy system. Matching of tags (210) is done using as-is source tags (212), i.e. source tags of the legacy system, which are matched using tag-matching prompts (214). As-is architecture configuration (216) is formed from as-is architecture building block configuration (218), i.e. architecture building block configuration of the legacy system, which consists of as-is design generation strategy (206), as-is design generation configuration (208), matching as-is source tags (210), and the as-is source tags (212). As-is components design (220), i.e. components design for the legacy system, is generated for each as- is component using the LLM to extract a design summary of the functionality present in the source code for each as-is component. The design summary can be used to redesign the same functionality in different technology in the subsequent stage. In the second stage, the target architecture configuration (222) is used to specify the details for the to-be state, which is done by defining the to-be solution building block configuration (224), i.e. solution building block configuration of the target system. As-is components (226) are mapped to the to-be solution building block configuration (224) by using the to-be design generation configuration (228), i.e. design generation configuration of the target system. The ordered solution building blocks are processed with an implementation configured in the to-be architecture configuration (222) resulting in to-be components of the solution building block (224). The to-be design generation strategy (230), i.e. design generation strategy of the target system, is generated 1-1 from as-is designs with prompt-based planning (232) for the to-be specifications based on the as-is designs and the to-be technology and architecture specifications. Other supported to-be design generation strategies (234), i.e. design generation strategies of the target system, are the decoupled to-be strategy (236), i.e. decoupled strategy of the target system, and as-is 1-1 customized override strategy (238), i.e. 1-1 customized override strategy of the legacy system. However, the given examples are not definitive and other supported to-be design generation strategies (234) may exist. For each ordered solution building block, each to-be as-is component mapped to the solution building block is processed with an implementation configured in the to-be architecture configuration (222), which results in the to-be components design (240), i.e. components design of the target system. To-be architecture building blocks may or may not be 1-1 to the as-is architecture building blocks as it is determined by the to-be design generation strategy (230). The generated to-be design descriptions contain both functional and technology independent technical design descriptions. These are transformed to technology solutions with the technology solution configuration of the following solution building block. In the final stage where the output solutions (242) and to-be components (244), i.e. components of the target system, are created, a default technology solution generation strategy (246) is used. The default technology solution generation strategy (246) is an ordered generation strategy based on a component dependency graph and it contains an iterative planning process per component which considers baseline repository planning, needed repository structure changes for each component, interface identifier planning, file content and distribution of declared interfaces. A supported technology solution generation strategy (248) is used as a to-be technology solution generation strategy (250), i.e. technology solution generation strategy of the target system, which is associated with the to-be technology solution generation configuration (252), i.e. technology solution generation configuration of the target system. The to-be solution building block source component (254), i.e. solution building block source component of the target system, is received from the previous stage and forms the technology solution (256) together with the to-be technology solution generation configuration (252) and the to-be technology solution generation strategy (250).

[0059] Fig. 3 depicts a summary of an optional technology clustering process to help in deriving the contents for the legacy architecture configuration. First, a code file (302) is processed by LLM, which summarizes various source file technologies (304) from which a summary of detected technologies in the code file (306) is created. The LLM creates a semantic representation of the code file technology summary (308), which code file technology summary is vectorized into an embedding (310). K- means clustering is used for the technology summary embedding vectors (312) after which technology clusters present in the codebase are detected (314). Based on the k- means clustering method, clusters silhouette analysis and cluster numeral optimization (316) is done. As part of the optional process, silhouettes (318) may be formed, and manual changes, such as optimization or alterations can be made to the as-is architecture configuration (320), i.e. architecture configuration of the legacy system. The described process is not constricted by the described steps and the number of steps may alternate, decrease, increase or repeat.

[0060] The given examples are not exhaustive and alternative embodiments are possible. The scenarios given for each figure may not be definitive and the number and order of steps may change.

[0061] Consequently, a skilled person may on the basis of this disclosure and general knowledge apply the provided teachings in order to implement the scope of the present invention as defined by the appended claims in each particular use case with necessary modifications, deletions, and additions.

Claims

CLAIMS1. A computer- implemented method for software system modernization comprising utilizing a large language model (LLM) to process one or more source files, wherein the processing comprises, as one or separate steps,• reading the one or more source files with the LLM to analyze the function of the one or more source files, utilizing the same or another LLM, processing one or more source files to• discover public definitions and declarations in the one or more source files, wherein said discovered definitions may be expressed with the declared identifiers and context of the declaration, such as namespace, and then semantically indexed,• detect references to external identifiers made in the one or more source files, wherein said references may be expressed with the referenced identifiers and context of the reference,• match a definition for each discovered reference, wherein a candidate definition expression is formed based on the reference, and a semantically indexed form for said matching is produced with an LLM from the candidate definition expression,• for each matched definition, extract the semantics of the reference from the context of the source code having the reference declaration, to a referenced functionality expression,• process each source file into one or more components, wherein each component comprises one or more functionally inherent parts present in a source file,• generate a design description for each component using the LLM to extract a design summary of the functionality present in the source code for each component so that the summary includes the information, which information may enable redesigning the same functionality in different technology, acquiring a target design generation configuration and mapping configuration, and mapping components in reference to said target design generation configuration and mapping configuration, based on the mapping, linking the components into system structures and the referenced functionality links across the components,arranging components in an optimal generation order within each system structure so that the most isolated components are ordered first, and that the most dependent components are ordered last, utilizing the same or another LLM,• in view of the target design generation configuration and mapping configuration, generate a codebase from the system structures and components.

2. The computer- implemented method of claim 1 comprising utilizing a large language model (LLM) to process one or more source files, wherein the processing further comprises, as one or separate steps, formulating a summary of the functionality of each of the one or more source files.

3. The computer- implemented method of claim 2 comprising utilizing a large language model (LLM) to process one or more source files, wherein the processing further comprises, as one or separate steps, indexing the summary in order to generate embedding vectors.

4. The computer-implemented method of any preceding claim comprising preprocessing wherein each source file is amended with the content from source files that are statically included into it, if any, so that each source file contains full information relevant to the intended functionality of the source file.

5. The computer- implemented method of claim 4 wherein the preprocessing is carried out by parsing, optionally utilizing an LLM.

6. The computer-implemented method of claim 3 wherein utilizing a cluster algorithm to process the embedding vectors to discover technical clusters present in the one or more source files.

7. The computer-implemented method of any preceding claim wherein further utilizing the same or different LLM for tagging each source file wherein tags may be used as markers for mapping each file.

8. The computer- implemented method of any preceding claim wherein processing one or more source files to discover public definitions and declarations in the source file is conducted with one or more of prompts that first discover the relevant namespace of the code, and subsequently supported with thatinformation obtain the definitions and declarations with the LLM and expressed in natural language.

9. The computer- implemented method of any preceding claim wherein processing one or more source files to references discovery is conducted with one or more prompts that first discover the relevant namespace of the code, and subsequently supported with that information obtain the external references with the LLM and expressed in natural language.

10. The computer- implemented method of any preceding claim wherein processing one or more source files to match a definition for each discovered reference comprises a cosine similarity search with the semantically indexed form over the definition of the semantically indexed form created in discovery of the discover public definitions and declarations.

11. The computer- implemented method of any preceding claim wherein processing the one or more source files into components comprises decomposing the source file into functionally inherent parts by using the LLM to discover the functionalities present in the source code and expressing them as distinct modules, along with the declared references and definitions from the respective source code.

12. The computer-implemented method of any preceding claim wherein extraction process is controlled by dynamic prompting where the design is produced from source code for each component for its functionality and using the previously associated tags to optionally refine or adjust the design for specific types of source code components, and / or design documentation objectives, using tag- based control prompt configurations.

13. The computer- implemented method of any preceding claim wherein a codebase component structure plan using target design and referenced functionalities is created, wherein the codebase component structure plan, target design and incoming referenced functionalities are used to plan target component declared identifiers, where the declared identifiers are made visible by the implementation of each file towards the other files so that the referenced functionalities can be implemented, and where the same or different LLM is used to produce the content of a codebase file using target architectureconfiguration, target design, codebase component structure plan and declared identifiers.

14. The computer- implemented method of any preceding claim wherein the design descriptions are combined to one documentation base and interlinked based on the referenced functionalities.

15. The computer- implemented method of any preceding claim wherein the components mapped in reference to a target system structure are combined to one documentation base and interlinked based on the referenced functionalities.

16. A data processing system comprising means for carrying out the method of any one of claims 1-15.

17. A computer program comprising instructions which, when the program is executed by a computer, cause the computer to carry out the method of any one of claims 1-15.

18. A computer-readable medium comprising instructions which, when executed by a computer, cause the computer to carry out the method of any one of claims 1-15.

Citation Information

Patent Citations

  • Method for transforming first code instructions in a first programming language into second code instructions in a second programming language

    EP2979176B1

  • Monolith database to distributed database transformation

    US11615076B2

  • System and method for optimized modernization of applications

    US20230401038A1