A software system refactoring scheme generation method and device, and electronic equipment

By parsing IT software system transformation tasks using a large language model, generating sub-task instructions, and performing multi-objective optimization, the problem of reliance on human experience in existing technologies is solved, and efficient and reliable transformation scheme generation is achieved.

CN122507352APending Publication Date: 2026-08-04中移信息技术有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
中移信息技术有限公司
Filing Date
2026-05-07
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

The existing intelligent generation of IT software system transformation solutions relies too heavily on human experience, making it impossible to achieve multi-objective optimization under multiple constraints, resulting in low efficiency and poor reliability in solution generation.

Method used

The system uses a pre-defined large language model to perform intent parsing, extract entity information and constraints, generate sub-task instructions, construct a decision space through multi-objective optimization, select structured transformation schemes, and use the large language model to generate transformation schemes expressed in natural language.

Benefits of technology

It achieves intelligent transformation from natural language requirements to structured component combination solutions, breaking through the reliance on human experience and improving the efficiency and reliability of solution generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122507352A_ABST
    Figure CN122507352A_ABST
Patent Text Reader

Abstract

The application discloses a software system reconstruction scheme generation method and device and electronic equipment. Based on the reconstruction task description text of the software system to be reconstructed, a preset large language model is used for intent analysis to obtain entity information and constraint conditions, and subtask instructions determined according to the entity information and the constraint conditions; the query entity information and / or constraint information included in each subtask instruction are used for retrieval to determine a candidate component set corresponding to each subtask instruction; a decision space is constructed based on the candidate component set corresponding to each subtask instruction, and each candidate reconstruction scheme in the decision space is subjected to multi-objective optimization solving under the constraint conditions to screen and determine a target reconstruction scheme of a structured expression of the software system to be reconstructed; and based on the target reconstruction scheme of the structured expression, a large language model is used to generate a target reconstruction scheme expressed in natural language. The method breaks through the dependence on artificial experience, realizes multi-objective optimization under multiple constraints, and improves efficiency and reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus and electronic device for generating software system modification schemes. Background Technology

[0002] As enterprise information technology (IT) system architectures become increasingly complex, their upgrade and transformation needs are characterized by high frequency, diversity, and high constraints. These needs involve not only technical aspects such as architectural upgrades and compatibility adaptations, but also business-level aspects such as process refactoring and performance optimization. Against this backdrop, traditional solution design models relying on human experience are facing efficiency bottlenecks. Meanwhile, the maturity of technologies such as Artificial Intelligence (AI), Large Language Models (LLM), and vector retrieval has made it possible to build data-driven and logically traceable intelligent decision-making frameworks, becoming a core trend in solving the challenges of generating enterprise IT software system transformation solutions.

[0003] Currently, the intelligent generation of existing IT software system transformation solutions relies excessively on human experience, failing to achieve intelligent transformation from natural language requirements to structured component combination solutions, and failing to achieve multi-objective optimization under multiple constraints, resulting in low solution generation efficiency and poor reliability. Summary of the Invention

[0004] This invention provides a method, apparatus, and electronic device for generating software system transformation schemes, in order to solve the problem that the intelligent generation of existing IT software system transformation schemes relies too much on human experience, making it difficult to achieve multi-objective optimization under multiple constraints, resulting in low scheme generation efficiency and poor reliability.

[0005] According to one aspect of the present invention, a method for generating a software system modification scheme is provided, comprising: Based on the description text of the modification task of the software system to be modified, the intent is parsed using a preset large language model to obtain entity information and constraints, as well as sub-task instructions determined according to the entity information and constraints. The candidate component set corresponding to each subtask instruction is determined by retrieving the query entity information and / or constraint information included in each subtask instruction. A decision space is constructed based on the candidate component set corresponding to each subtask instruction. Under the constraints, a multi-objective optimization solution is performed on each candidate modification scheme in the decision space to screen and determine the target modification scheme of the structured expression of the software system to be modified. Based on the structured representation of the target transformation scheme, a large language model is used to generate a target transformation scheme expressed in natural language.

[0006] According to another aspect of the present invention, a software system modification scheme generation apparatus is provided, comprising: The text parsing module is used to perform intent parsing based on the modification task description text of the software system to be modified, using a preset large language model to obtain entity information and constraints, as well as sub-task instructions determined based on the entity information and constraints. The data retrieval module is used to retrieve data based on the query entity information and / or constraint information included in each subtask instruction in order to determine the candidate component set corresponding to each subtask instruction. The intelligent decision-making module is used to construct a decision space based on the candidate component set corresponding to each sub-task instruction, and to perform multi-objective optimization on each candidate modification scheme in the decision space under the constraints, so as to screen and determine the target modification scheme of the structured expression of the software system to be modified. The Natural Language Scheme Generation Module is used to generate target transformation schemes expressed in natural language based on structured representations and large language models.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which is then executed by the at least one processor to enable the at least one processor to execute the software system modification scheme generation method according to any embodiment of the present invention.

[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the software system modification scheme generation method according to any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the software system modification scheme generation method described in any embodiment of the present invention.

[0010] The technical solution of this invention involves: first, parsing the modification task description text of the software system to be modified using a preset large language model to obtain entity information and constraints, as well as sub-task instructions determined based on the entity information and constraints; second, retrieving the query entity information and / or constraint information included in each sub-task instruction to determine the candidate component set corresponding to each sub-task instruction; third, constructing a decision space based on the candidate component set corresponding to each sub-task instruction; fourth, performing multi-objective optimization on each candidate modification scheme in the decision space under the constraints to filter and determine the target modification scheme of the structured expression of the software system to be modified; and finally, generating the target modification scheme expressed in natural language using a large language model based on the target modification scheme of the structured expression. By parsing the intent of the transformation task description text using a pre-set large language model, entity information and constraints are automatically extracted and sub-task instructions are generated, realizing the intelligent transformation of natural language requirements into structured task instructions. Then, based on the sub-task instructions, a candidate component set is determined and a decision space is constructed. Under constraints, multi-objective optimization solutions are performed on each candidate transformation scheme in the decision space, automatically selecting a structured component combination scheme that takes into account multiple constraints and multi-dimensional optimization objectives. This breaks through the dependence on human experience and realizes the intelligent transformation from natural language requirements to structured component combination schemes and multi-objective optimization decision-making under multiple constraints, significantly improving the efficiency and reliability of scheme generation.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a flowchart of a method for generating a software system modification scheme, provided in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram illustrating the process of generating a software system modification scheme according to Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of a software system modification scheme generation device provided in Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] Example 1 Figure 1 This is a flowchart illustrating a method for generating a software system modification scheme according to Embodiment 1 of the present invention. This embodiment is applicable to situations involving the generation of software system modification schemes. This method can be executed by a software system modification scheme generation device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes: S110. Based on the modification task description text of the software system to be modified, use a preset large language model to perform intent parsing to obtain entity information and constraints, as well as sub-task instructions determined according to the entity information and constraints.

[0017] In this embodiment, the task description text is the original text information input by the user in natural language, used to describe the current status and modification requirements of the software system to be modified. The preset large language model is a domain-adaptive large language model (LLM). In this embodiment, the preset large language model is based on a general base model (such as Qwen2.5-72B), and is a domain-specific model trained using LoRA (Low-Rank Adaptation) and other parameter-efficient fine-tuning techniques and software system modification domain instruction data. It possesses professional knowledge and chain-of-thought reasoning capabilities within the domain of the software system to be modified. Entity information consists of the core objects constituting the software system extracted from the task description text, such as hardware entities, software entities (operating systems, databases, etc.), and their attribute parameters. Constraints are explicit limitations extracted from the task description text, including but not limited to time, budget, and compatibility requirements. Subtask instructions can be understood as structured executable units automatically generated by the preset large language model through reasoning, decomposing complex modification tasks. Each subtask instruction can include a subtask identifier, task name, execution priority, dependencies, subtask description, query entity information, constraints, etc., to guide subsequent multimodal retrieval and decision optimization.

[0018] Specifically, based on the user-input description text of the modification task for the software system to be modified, a pre-defined large language model is used to perform intent parsing and reasoning. During this process, the pre-defined large language model utilizes its attention mechanism to accurately identify entity information and constraints. Then, based on the identified entity information and constraints, the task complexity (simple or complex) is intrinsically determined through thought chain reasoning, generating a structured instruction object containing task complexity, entity information, constraints, and a list of subtasks. When the task complexity is determined to be complex, the subtask list contains several subtask instructions decomposed based on entity information and constraints; when the task complexity is determined to be simple, the subtask list is empty, and a simple query is directly performed based on the instruction to return the answer, achieving short-circuit optimization.

[0019] Optionally, after generating the structured instruction object, the format and content can be validated by a parsing validator.

[0020] For example, the following is a structured instruction object (JSON format): {"task_id":"T-20240520-001","complexity":"complex","entities":[{"entity_name":"x86 architecture server cluster","entity_type":"existing hardware","description":"a business operation cluster consisting of 30 units"}],"constraints":[{"constraint_type":"time","constraint_key":"renovation cycle","constraint_value":6,"unit":"months","description"}],constraints":[{"constraint_type":"time","constraint_key":"renovation cycle","constraint_value":6,"unit":"months","description"}] ":"Complete the transformation within 6 months"}],"sub_tasks":[{"sub_task_id":"ST-001","sub_task_name":"Hardware adaptation solution","priority":"high","dependencies":[]}],"original_input":"30 x86 servers to be domestically upgraded to Kunpeng 920 platform within 6 months, budget within 5 million, compatible with Oracle12c","parsing_status":"success","parsing_time":"2025-12-01T14:30:22.156Z"}; Here, `task_id` represents the task identifier, used for system tracking and log association; `complexity` represents the task complexity (simple / complex), determining the subsequent processing flow; `entities` represents the core objects in the requirements (such as hardware and software), including name, type, and brief description, i.e., entity information; `constraints` represents the requirements' constraints (such as time and budget), clarifying the task execution boundaries, i.e., constraints; `sub_tasks` represents the specific executable subtasks after the complex task is decomposed, including priority and dependencies, i.e., a list of subtasks; `original_input` represents the user's original input content, used for backtracking verification; `parsing_status` represents the parsing result status (success / failure); and `parsing_time` represents the timestamp of parsing completion, used for statistics and traceability.

[0021] S120. Retrieve based on the query entity information and / or constraint information included in each subtask instruction to determine the candidate component set corresponding to each subtask instruction.

[0022] In this embodiment, query entity information is used to determine the category and identifier of the target object to be retrieved (such as retrieving specific component types like CPUs and databases, and their associated entities); constraint information is used to determine the boundary filtering conditions for the retrieval, including but not limited to time requirements, budget limits, and compatibility requirements. The candidate component set can be understood as a set of components that may be applicable to a subtask instruction, obtained through retrieval. This candidate component set may include several alternative components and their attribute information (such as specifications, compatibility relationships, etc.), serving as the basic input for subsequent construction of the decision space and scheme selection.

[0023] Specifically, for each subtask instruction, the query entity information and / or constraint information in the current subtask instruction are converted into a representation that the knowledge base can recognize, and then input into the knowledge base for retrieval to obtain several candidate components and their attribute information that match the current subtask, thereby determining the candidate component set corresponding to the subtask instruction.

[0024] S130. Construct a decision space based on the candidate component set corresponding to each subtask instruction, and perform multi-objective optimization on each candidate modification scheme in the decision space under the constraints, so as to screen and determine the target modification scheme of the structured expression of the software system to be modified.

[0025] In this embodiment, the decision space is the set of all possible modification schemes formed by combining candidate component sets corresponding to each subtask instruction. A candidate modification scheme is a single solution in the decision space, a possible complete system modification configuration formed by combining specific components selected from the candidate component sets corresponding to each subtask instruction. Multi-objective optimization is the process of simultaneously optimizing multiple independent and potentially conflicting objective functions (such as maximizing performance matching, maximizing ecological compatibility, minimizing modification costs, and minimizing implementation risks) within the decision space. Constraints include hard constraints and soft constraints. Hard constraints can be understood as technical and physical limitations that must be strictly met (such as physical compatibility between components and architectural feasibility); if they are not met, the scheme is not feasible. Soft constraints are business preferences that can be met as much as possible but can be weighed (such as budget limits and performance bottom lines); if they are not met, they are still acceptable but require corresponding costs. The structured description of the target modification scheme can be understood as the recommended modification scheme determined after multi-objective optimization and screening, described in a machine-readable structured data format (such as JSON).

[0026] Specifically, firstly, the components in the candidate component set corresponding to each subtask instruction are combined according to the technical domain category to construct a decision space containing all possible component combinations, where each component combination is a candidate modification scheme; then, under the constraints (including hard constraints that must be strictly satisfied and soft constraints that can be traded off), multi-objective optimization is performed on each candidate modification scheme in the decision space; and finally, the target modification scheme with a structured expression is determined through comprehensive evaluation and screening.

[0027] S140. A target transformation scheme based on structured representation: using a large language model to generate a target transformation scheme expressed in natural language.

[0028] In this embodiment, the large language model is a domain-adaptive fine-tuned model, which can be the same as the preset large language model. Through prompt word engineering, the target modification scheme of the structured expression is converted into a target modification scheme of natural language expression. The target modification scheme of natural language expression is a modification scheme presented in human-readable text form (such as a technical solution report), typically including sections such as a modification overview, component selection details, evaluation analysis, implementation suggestions, and supporting evidence.

[0029] Specifically, to facilitate user understanding and review, the structured target transformation plan is organized into an input format that can be understood by the large language model. The large language model is then used to generate the target transformation plan expressed in natural language.

[0030] The technical solution provided in Embodiment 1 of this invention involves: based on the description text of the modification task of the software system to be modified, using a preset large language model for intent parsing to obtain entity information and constraints, as well as sub-task instructions determined based on the entity information and constraints; retrieving the query entity information and / or constraint information included in each sub-task instruction to determine the candidate component set corresponding to each sub-task instruction; constructing a decision space based on the candidate component set corresponding to each sub-task instruction; performing multi-objective optimization on each candidate modification scheme in the decision space under the constraints to screen and determine the target modification scheme of the structured expression of the software system to be modified; and generating the target modification scheme expressed in natural language based on the target modification scheme of the structured expression using a large language model. By parsing the intent of the transformation task description text using a pre-set large language model, entity information and constraints are automatically extracted and sub-task instructions are generated, realizing the intelligent transformation of natural language requirements into structured task instructions. Then, based on the sub-task instructions, a candidate component set is determined and a decision space is constructed. Under constraints, multi-objective optimization solutions are performed on each candidate transformation scheme in the decision space, automatically selecting a structured component combination scheme that takes into account multiple constraints and multi-dimensional optimization objectives. This breaks through the dependence on human experience and realizes the intelligent transformation from natural language requirements to structured component combination schemes and multi-objective optimization decision-making under multiple constraints, significantly improving the efficiency and reliability of scheme generation.

[0031] In some embodiments, the step of retrieving the candidate component set corresponding to each subtask instruction based on the query entity information and / or constraint information included in each subtask instruction includes: for each subtask instruction, generating a graph query statement for a preset knowledge graph and a semantic query vector for a preset vector database based on the query entity information and / or constraint information included in the current subtask instruction; inputting the graph query statement into the preset knowledge graph for querying to obtain a first retrieval list, and inputting the semantic query vector into the preset vector database for similarity search to obtain a second retrieval list; and fusing and reordering the first retrieval list and the second retrieval list to obtain the candidate component set corresponding to the subtask instruction.

[0032] In this embodiment, the preset knowledge graph is a pre-constructed graph database used to store structured knowledge of the domain to which the software system to be modified belongs. This knowledge graph extracts precise information such as entity relationships, parameter attributes, and compatibility rules of system components from structured or semi-structured data sources such as official product documentation, compatibility certification lists, and test reports. Using entity recognition and relationship extraction technologies, it constructs a graph network with component entities (such as CPU, OS, and database) as nodes and edges such as "compatible with," "substitute for," and "depends on," storing entity attributes (such as clock speed and version number) as node attributes. After construction, the Cypher graph query language is used for efficient relationship queries, providing the precise factual basis required for decision-making. The preset vector database is a pre-constructed database used to store unstructured text data of the domain to which the software system to be modified belongs. This database aims to store semantic representations of textual data such as technical white papers, migration cases, community Q&A, and fault reports. It uses embedding models (such as BGE) to convert the above text into high-dimensional vectors and store them. It creates efficient indexes based on the approximate nearest neighbor algorithm, supports millisecond-level semantic similarity search, and provides empirical and procedural soft knowledge support for decision-making.

[0033] The graph query statement is a structured query language (such as Cypher) executable by the preset knowledge graph. The semantic query vector is a high-dimensional dense vector generated using an embedding model based on the query entity information and constraint information. It is used to perform semantic similarity matching in a preset vector database to retrieve technical documents and cases related to the entity and constraints. The first search list can be understood as a structured result list obtained after inputting the graph query statement into the preset knowledge graph, including matched component entities, entity attribute parameters, relationship paths, and confidence scores. The second search list can be understood as a list of document fragments obtained after inputting the semantic query vector into the preset vector database for similarity search, sorted by semantic similarity, including unstructured text such as technical white papers, migration cases, and community Q&A.

[0034] Specifically, for each subtask instruction, the query entity information and / or constraint information based on the current subtask instruction are converted into a graph query statement executable by a preset knowledge graph; simultaneously, they are converted into semantic query vectors executable by a preset vector database; then, the graph query statement is input into the preset knowledge graph for querying, resulting in a first search list containing matching component entities, attribute parameters, relationship paths, and confidence scores. The semantic query vector is input into a preset vector database for similarity search, resulting in a second search list sorted by cosine similarity and containing document fragments such as technical white papers and migration case studies. The first and second search lists are then merged and reordered: a fused feature vector is constructed for the search results in each search list, which includes at least knowledge graph features and vector retrieval features; subsequently, a learning ranking algorithm (such as the LambdaMART (Learning to Rank) model) is used to predict the relevance score of each fused feature vector, and the optimal ranking list is generated by arranging the results in descending order of the scores. This refers to the set of candidate components corresponding to the current subtask instruction.

[0035] For example, the system receives a description of the modification task of the software system to be modified from user input. Suppose the description of the modification task is "Our existing system is running on a source processor with a first instruction set architecture. We plan to migrate the system to a target processor with a second instruction set architecture, requiring compatibility with the target processor and performance no less than that of the source processor". After using the preset large language model to parse the intent, four sub-task instructions can be obtained, and their corresponding query statements are as follows: "task_distribution": {"multimodal_retrieval": ["Retrieve compatibility with the target processor and performance no less than that of the source processor", "Retrieve performance cases of the target processor replacing the source processor", "Retrieve the software compatibility list and certified operating system, database and middleware information of the target processor platform", "Retrieve migration guidelines, best practices and common problem solutions for migrating from the first instruction set architecture platform to the second instruction set architecture platform"].

[0036] For each subtask instruction, two retrieval processes are initiated in parallel based on the query entity information and / or constraint information of the current subtask instruction: 1. Preset knowledge graph retrieval The query entity information and constraint information are converted into a graph query statement that can be executed by a predefined knowledge graph. For example, the query statement "compatible with the target processor and with performance no less than that of the source processor" can be converted into a Cypher statement: MATCH (source:CPU {role: 'source processor'}) MATCH (target:CPU {architecture: 'target processor'}) WHERE target.performance_score >= source.performance_score RETURN target.name, target.performance_score After the search, a structured first search list is output. It includes entities and relationships, with confidence scores.

[0037] 2. Preset vector database retrieval The query entity information and constraint information of the subtask are converted into semantic query vectors. For example, "retrieving performance cases of the target processor replacing the source processor" can be converted into a semantic query vector. Search the vector database to find matching pairs. Output a second search list sorted by the document fragment with the highest cosine similarity. .

[0038] Furthermore, the LambdaMART model is used to fuse the first and second search lists. This method abandons heuristic rules such as simple weighted averages and automatically learns the contribution weights of different modalities and features to the final relevance ranking through a machine learning model, thereby achieving intelligent, accurate, and adaptive result fusion. The specific implementation steps are as follows: (1) Feature Engineering For each structured result in the first search list and each document fragment in the second search list, construct a unified fusion feature vector. This feature vector fuses heterogeneous information from two channels: 1) Characteristics of knowledge graphs confidence_kg: Relation confidence level of the graph query results; node_degree: The degree (number of connections) of the resulting entity in the knowledge graph, which measures its importance; path_length: The length of the path between the query entity and the result entity. The shorter the length, the stronger the correlation may be. entity_centrality: The graph centrality index of the resulting entity, indicating its betweenness centrality and isograph centrality.

[0039] 2) Vector retrieval features cosine_similarity: The cosine similarity score between the query vector and the document vector; document_authority: The authority weight of the document source (preset, official documentation > technical blogs > community Q&A); document_freshness: The freshness of the document (timestamp); bm25_score: The traditional text matching score (BM25) between the query and the document, as a supplement to semantic similarity.

[0040] 3) General features original_rank_kg: The original rank of this result in the list returned by the knowledge graph; original_rank_vec: The original rank of this result in the list returned by the vector database; normalized_score_kg: The normalized confidence score of the knowledge graph; normalized_score_vec: The vector similarity score after normalization.

[0041] Then, each fused feature vector is input into the loaded LambdaMART model, which predicts a final, uniform relevance score for each fused feature vector. Furthermore, based on the correlation scores of each fused feature vector... The corresponding search results are sorted in descending order to generate an optimal sorted list. .

[0042] The candidate component set determined after fusion and reordering has the following characteristics: its top entries collectively outline the potential solutions for the subtask. For example, for migration path planning, the top-ranked results might include a "list of compatible target processor models" from the first search list, a "best practice guide for migration" from the second search list, and "general principles for transitioning from x86 to ARM architecture." Each candidate entry in this set is supported by high-confidence information from both engines: the first search list provides precise information (such as specific compatibility relationships, component attribute parameters, version dependency rules, etc.), constituting the hard facts and quantitative basis required for decision-making; the second search list provides experiential information (such as successful migration cases, detailed implementation steps, community-validated solutions, troubleshooting records, etc.), constituting the procedural knowledge and contextual support required for decision-making.

[0043] Ultimately, the output is a structured set of candidate components that has been intelligently fused and sorted. It not only indicates which candidate items are possible, but also answers why these items are feasible and worth recommending through multi-source heterogeneous information, providing comprehensive and reliable input for the subsequent construction of the decision space.

[0044] Optionally, the LambdaMART model needs to be trained before fusion and ranking. The training process is as follows: 1) Constructing training data Domain experts manually annotated a large number of historical search lists, assigning a relevance level label (e.g., 0: irrelevant, 1: moderately relevant, 2: relevant, 3: highly relevant) to each query-document pair, thus forming the training set. Where q represents a query-document pair, To fuse feature vectors, Relevance level labels assigned to domain experts.

[0045] 2) Model Training The LambdaMART algorithm, which combines gradient boosting decision trees and the LambdaRank concept, gradually reduces prediction error by stacking multiple weak decision tree models. Unlike directly fitting relevance scores, LambdaMART defines a gradient λ for each document, the magnitude of which represents the strength and direction of the adjustment that document should undergo in the ranking process. This gradient is calculated from the impact of swapping the positions of two documents on the overall performance metric NDCG. ; in: It is the change in the NDCG index after swapping the positions of document i and document j. It is the score that the model currently predicts for documents i and j.

[0046] This process directly uses NDCG as the optimization objective, achieving end-to-end optimization from feature fusion to the final ranking result, exhibiting stronger robustness than heuristic methods. Furthermore, the LambdaMART model enables adaptive fusion: through training, it automatically learns the optimal weight allocation between knowledge graph features and vector retrieval features for different query types. For example, for precise queries such as "software version number," the model automatically assigns higher weights to knowledge graph features; for fuzzy queries such as "transfer implementation experience," it automatically assigns higher weights to vector database features, thus overcoming the limitation of fixed rules being unable to adapt to changes in query type.

[0047] In some embodiments, the step of constructing a decision space based on the candidate component set corresponding to each subtask instruction, and performing multi-objective optimization on each candidate modification scheme in the decision space under the constraints, includes: extracting candidate components from the candidate component set corresponding to each subtask instruction, combining them according to technical domain categories to construct a set of candidate modification schemes to form a decision space; under the constraints, using performance matching degree, ecological compatibility, modification cost, and implementation risk as optimization objectives, performing multi-objective optimization on the decision space to obtain a Pareto optimal solution set; and determining the target modification scheme of the structured representation of the software system to be modified from the Pareto optimal solution set based on a multi-criteria decision analysis algorithm.

[0048] In this embodiment, the technology domain category refers to the component categories in the software system architecture, divided according to functional layers. These include the hardware layer (such as servers and CPUs), the system software layer (such as operating systems), and the basic software layer (such as databases and middleware). This ensures the systematic nature of cross-domain component combinations and the integrity of the technology stack when constructing the decision space. The candidate modification scheme set includes at least one candidate modification scheme. Each candidate modification scheme is a complete system modification configuration composed of specific components selected from the candidate component set of each subtask instruction according to the technology domain category. It represents a single solution in the decision space, including the selection and combination of components such as servers, CPUs, operating systems, and databases. The Pareto optimal solution set can be understood as the set of non-dominated solutions obtained in multi-objective optimization where no solution can improve any objective without worsening other objectives. It represents the optimal trade-off set among performance, compatibility, cost, and risk objectives. The multi-criteria decision analysis algorithm can be understood as a decision-making method that selects and determines the final recommended solution from the Pareto optimal solution set, ranking and selecting the best solution by calculating the closeness of each solution to the ideal solution.

[0049] Specifically, the decision space is first constructed: candidate components are extracted from the candidate component set corresponding to each subtask instruction (such as CPU options for the hardware adaptation subtask and database options for the database migration subtask), and cross-domain combinations are performed according to the technology domain category to construct a set of candidate modification schemes containing all possible component combinations, forming a multi-objective optimization decision space X (i.e., solution space). Each solution x∈X represents a candidate modification scheme, including the selection of components such as server, CPU, operating system, and database, where x = {comp1, comp2, ..., comp...} n}

[0050] Furthermore, a multi-objective optimization model is established: under constraints (including hard constraints such as physical compatibility between components and technology stack integrity, and soft constraints such as budget constraints and performance baseline constraints), a mathematical model is established to simultaneously optimize the four major objectives: 1) Performance matching degree: Maximize the performance ratio between the new solution and the original system. ; in, The performance metrics of the new scheme under parameter x. These serve as benchmark values ​​for the performance indicators of the existing system.

[0051] 2) Ecosystem compatibility: Maximize the geometric mean of the compatibility confidence among components in the following formula. ; Where n is the number of components in the system. For the i-th and j-th components in the system, The confidence level of compatibility between two components typically ranges from [0,1], with values ​​closer to 1 indicating stronger compatibility.

[0052] 3) Renovation cost: Minimize the total investment cost as shown in the following formula.

[0053] ; in, The cost of purchasing, developing, or deploying the i-th component. The costs incurred for system migration (such as data migration, service migration, etc.).

[0054] 4) Implementation risk: Minimize the overall risk score of the following scheme.

[0055] ; Where n is the number of components in the system. For the i-th component in the system, This represents the risk coefficient of the i-th component during implementation. Its value range is usually [0,1]. The closer the value is to 1, the higher the risk of the component.

[0056] Then, perform multi-objective optimization: solve the above multi-objective problem in the decision space to obtain a set of Pareto optimal solutions in which there is no absolute superiority or inferiority among the four objectives. Each solution represents a different trade-off between performance, cost, compatibility and risk.

[0057] Furthermore, a multi-criteria decision analysis algorithm (such as the TOPSIS method based on entropy weight to calculate target weight and Euclidean distance to calculate proximity to the ideal solution) is used to select the target modification scheme with the comprehensive optimal structure from the Pareto optimal solution set. This scheme includes the selected component combination and multi-objective evaluation quantitative indicators.

[0058] The above technical solutions enable a leap from subjective experience-based judgment to multi-objective quantitative optimization decision-making in software system transformation. Under multiple constraints such as performance matching, ecosystem compatibility, transformation cost, and implementation risks, a Pareto optimal solution set that takes into account the trade-offs of each objective is generated, and a comprehensive optimal structured objective transformation scheme is determined. This provides a structured decision-making basis for the subsequent generation of interpretable, data-driven professional system transformation schemes.

[0059] In some embodiments, the step of performing multi-objective optimization on the decision space to obtain a Pareto optimal solution set includes: determining an initial population within the decision space and constraining the generation range of the initial population based on component compatibility relationship data stored in the preset knowledge graph; performing genetic operations within the feasible solution space that satisfies the compatibility relationship data constraints, guiding candidate modification schemes to crossover and mutation based on component attribute constraints of the preset knowledge graph to generate new candidate solutions; and performing Pareto sorting and genetic iteration on the new candidate solutions until a termination condition is met to obtain the Pareto optimal solution set.

[0060] Specifically, the improved NSGA-II algorithm is used to seek a Pareto optimal solution set (i.e., a set of non-dominated solutions where no solution can improve any objective without worsening other objectives) within the decision space, including: (1) Heuristic population initialization: An initial population P0 (each individual represents a candidate modification scheme) is generated based on the candidate component set corresponding to each subtask instruction. The generation range of the initial population is constrained by the compatibility relationship data between components stored in the preset knowledge graph (such as "compatible with", "dependent on" relationship and confidence level): Candidate components with high confidence are directly injected into the initial population, and the remaining individuals are randomly generated only within the compatible components verified by the preset knowledge graph to ensure that the population has a high starting point and excellent quality.

[0061] (2) Calculate the Pareto front rank of each solution in the population. ; in, To solve The Pareto front hierarchy is used, with lower hierarchy values ​​being better (1 being the highest). P represents the current population. express Dominate ,Right now Not inferior to in any goal And it is better to be stricter on at least one objective.

[0062] (3) Crowding calculation, environmental selection and genetic manipulation First, calculate the crowding distance to maintain the diversity of the solution set: ; in, To solve The crowding distance is used to measure the solution. The density of its neighboring solutions in the objective space; M is the total number of objective functions (e.g., M=4). The solutions are as follows: after sorting by the m-th objective, Objective function values ​​of adjacent solutions; Let be the maximum and minimum values ​​of the m-th objective function in the current population.

[0063] Furthermore, environment selection is performed based on Pareto front ranking (preferring solutions with lower rankings) and crowding distance (preferring solutions with larger distances) to retain high-quality solutions. The quantitative standard for high-quality solutions is: the smaller the Pareto ranking (the earlier the solution is on the non-dominated front), the better it performs in the four optimization objectives (…). maximize, maximize, minimize, The more difficult it is for a solution to be dominated by other solutions (i.e., no other solution can improve another objective without worsening one objective), the higher the ranking quantification directly corresponds to the non-dominated ranking result. A larger crowding distance indicates that other solutions around the solution are sparser, better preserving the diversity of the objective space and avoiding optimization blind spots caused by the clustering of high-quality solutions. In summary, solutions with the best Pareto ranking and a large crowding distance are preferred, ensuring both dominance and distribution diversity.

[0064] Subsequently, genetic operations (simulating binary crossover and polynomial mutation) are performed within the feasible solution space that satisfies the data constraints of compatibility relationships, and candidate modification schemes are guided to undergo crossover and mutation based on component attribute constraints (such as performance parameters, cost data, and risk coefficients) of a pre-defined knowledge graph. ; ; in, Let k be the k-th dimension variable of the two offspring individuals generated by the crossover; Let k be the k-th dimension variable of the two parent individuals; It is a random number that is uniformly distributed in the interval [0,1]. The crossover index is used to control for the similarity between offspring and parents. The larger the value, the closer the offspring is to the parent. Its typical value range is: Genetic operations prioritize the combination and replacement of candidate components that satisfy compatibility constraints to improve the proportion of feasible solutions.

[0065] In this embodiment, an initial population is constructed based on the candidate component set corresponding to each subtask instruction. High-confidence candidate components are directly injected into the initial population, and the remaining individuals are randomly generated only within compatible components verified by the preset knowledge graph to improve the quality of the initial solution and accelerate convergence. Furthermore, genetic operations are guided by the compatibility relationships and component attribute data between components stored in the preset knowledge graph, prioritizing the selection of candidate components that meet compatibility constraints for combination and replacement to increase the proportion of feasible solutions and guide the search direction. Simultaneously, based on the crowding calculation, a diversity metric dimension based on the information coverage of the candidate component set supporting the candidate modification scheme is introduced (the information coverage is determined based on the precise information provided by the first search list and the empirical information provided by the second search list) to balance the diversity of the solution set distribution in the target space and the feasibility of engineering practice.

[0066] (4) Determine the target modification scheme: Through genetic iteration until the termination condition is met, a Pareto optimal solution set is obtained, which contains multiple non-dominated solutions. Each solution represents a different trade-off between objectives such as performance, compatibility, cost, and risk. There is no absolutely optimal unique solution. Therefore, it is necessary to select a final recommended scheme from this Pareto optimal solution set based on the company's specific preferences (such as whether it values ​​cost or performance more) to generate a clear and executable modification plan.

[0067] In this embodiment, a multi-criteria decision analysis algorithm (such as calculating the target weight based on entropy weight and calculating the closeness to the ideal solution based on Euclidean distance) is used to screen and determine the target modification scheme with a structured representation from the Pareto optimal solution set, including: 1) Constraint violation priority: The constraint dominance principle is adopted to prioritize the selection of solutions that satisfy all hard constraints; 2) TOPSIS decision-making based on entropy weights: First, calculate the information entropy weights of each objective: ; in, Let be the information entropy of the j-th objective function, which measures the uncertainty of that objective in decision-making; The elements of the standardized decision matrix are: n is the number of alternatives (size of the Pareto optimal solution set), and m is the number of objective functions.

[0068] Then, the relative closeness of each candidate modification scheme to the ideal solution is calculated: ; in, and Let represent the Euclidean distances between the candidate modification scheme and the positive ideal solution and the negative ideal solution, respectively.

[0069] Furthermore, the candidate modification scheme with the highest degree of similarity is selected as the target modification scheme in the structured expression, realizing intelligent decision-making from multi-objective trade-offs to a single executable solution.

[0070] In some embodiments, after generating a target modification scheme expressed in natural language, the method further includes: collecting user feedback data on the target modification scheme; optimizing and updating the preset knowledge graph and / or the preset large language model based on the feedback data; wherein the optimization and updating of the preset knowledge graph includes at least one of the following: dynamically adjusting the confidence of relationships between entities in the preset knowledge graph using a relation confidence weighted algorithm; identifying component co-occurrence patterns based on an association rule mining algorithm, and adding the new compatibility relationship represented by the component co-occurrence pattern to the preset knowledge graph when the support and confidence of the component co-occurrence pattern meet a preset threshold.

[0071] In this embodiment, the feedback data includes explicit feedback data and implicit feedback data; wherein, explicit feedback data is the user's direct evaluation signal, including but not limited to solution rating. Likes / dislikes Manually corrected records; implicit feedback data consists of potential preference signals mined from user behavior logs, including but not limited to solution adoption rates. Component click popularity and browsing time weighting Etc. The relation confidence weighted algorithm can be understood as a quantitative update algorithm that dynamically adjusts the confidence of relationships between entities in a pre-defined knowledge graph based on feedback data. The association rule mining algorithm can be understood as a statistical learning method that discovers potential compatibility relationships in a pre-defined knowledge graph based on historical candidate modification scheme data. The component co-occurrence pattern can be understood as the pattern of combinations of different system components (such as a specific model of CPU and database) being used simultaneously in the same candidate modification scheme or historical candidate modification schemes.

[0072] Specifically, after generating the target improvement plan expressed in natural language, feedback learning and closed-loop optimization can be performed, including: Collect user feedback data on the target renovation plan, including explicit feedback data (plan rating). And likes / dislikes ) and implicit feedback data (solution adoption rate) And browsing time weight Furthermore, based on the feedback data, the credibility of the feedback is calculated: ; in, These are the weighting coefficients for each feedback signal.

[0073] Based on feedback data, the pre-defined knowledge graph is dynamically evolved and enhanced, including at least one of the following: 1) Relationship Confidence Adjustment: The confidence of relationships between entities in the preset knowledge graph is dynamically adjusted using a relation confidence weighted algorithm. ; in, As the confidence level weight, Weighting for the strength of new evidence. It is a relationship The original confidence level, It is the strength of new evidence from the feedback.

[0074] 2) New Relationship Discovery: Based on association rule mining algorithms, identify component co-occurrence patterns (i.e., the combination patterns of different components appearing simultaneously in the same candidate modification scheme, which can be used to discover new compatibility relationships): ; ; Among them, Support is the support level of the rule, which measures the generality of the rule; Confidence is the confidence level of the rule, which measures the reliability of the rule. The number of times A and B appear simultaneously (in the context of software system transformation, this refers to the number of solutions where both components are used at the same time). This represents the total number of transactions (number of candidate modification schemes). Count(A) represents the number of transactions that simultaneously contain both A and B; Count(B) represents the number of transactions containing A; and Count(B) represents the number of transactions containing B. When the support and confidence exceed the thresholds, the new compatibility relationships represented by the component co-occurrence patterns are added to the preset knowledge graph.

[0075] Optionally, by detecting conflicts between feedback data and existing knowledge (i.e., relationship confidence, compatibility relationships, etc. already stored in the preset knowledge graph), when the quantitatively calculated conflict score exceeds a preset threshold, it indicates that there may be risks in the automated update. At this time, a manual review process is triggered, with domain experts intervening to make a judgment, so as to avoid erroneous information from polluting the preset knowledge graph and ensure the accuracy and reliability of the knowledge base update.

[0076] In addition, the preset large language model is updated based on feedback data, including: 1) Training data augmentation: Transforming high-quality feedback data into training samples: ; in, These are labels adjusted based on feedback data; The quality score is calculated based on feedback data; This is the quality threshold, such as 0.8.

[0077] 2) Course learning strategy: Training is performed in tiers based on sample quality, and training samples are selected using sampling probability based on feedback reliability. ; Where T is a temperature parameter that controls the sharpness of the selection strategy.

[0078] 3) Model retraining strategy: Set conditions to trigger retraining, such as the amount of new training data reaching a threshold, the performance degradation of the preset large language model exceeding the limit, or time-triggered periodic retraining, to incrementally fine-tune the preset large language model and enhance its domain adaptability.

[0079] Optionally, based on the above, to further improve the quality and efficiency of feedback learning, a meta-learning mechanism is established to monitor and optimize the process, including: 1) learning performance evaluation indicators. Calculate the improvement score to quantify the optimization effect. ; Where K represents the number of evaluation indicators; The value of the k-th indicator before optimization. This represents the optimized k-th metric value.

[0080] 2) Hyperparameter adaptive adjustment Dynamically adjust parameters such as the learning rate based on learning performance: ; in, To adjust the amplitude coefficient, exp ensures that the learning rate is always positive.

[0081] 3) Feedback quality filtering Establish a feedback quality assessment model and train a binary classifier to determine whether the feedback is reliable: ; The MLP structure consists of a 2-3 layer fully connected network, with the output layer activated by sigmoid. To score the proposals, To determine the adoption rate of the proposal, The browsing time is weighted, and UserCredibility represents user credibility, calculated based on historical feedback quality.

[0082] Alternatively, the LambdaMART model can be optimized based on feedback data: Set the incremental learning objective function: ; in, For learning rate, The number of new feedback samples. Sample weights (set based on feedback credibility). For loss function, For the ranking model of samples The predicted score, The relevance labels are adjusted based on feedback. The retrieval ranking model is incrementally updated using the FTRL (Follow-the-Regularized-Leader) optimizer. ; in, This represents the cumulative gradient from round 1 to round t. The learning rate is the scheduling parameter; This is the L1 regularization coefficient, which controls the sparsity of the model. These are the model parameters for the s-th round.

[0083] In some embodiments, the modification task description text based on the software system to be modified is used to perform intent parsing using a preset large language model to obtain entity information and constraints, as well as sub-task instructions determined based on the entity information and constraints. This includes: preprocessing and standardizing the terminology of the modification task description text to obtain a standardized task description text; generating task description prompts based on a first preset prompt word template and the standardized task description text, and inputting them into the preset large language model to guide the preset large language model to perform reasoning to obtain entity information and constraints, as well as sub-task instructions determined based on the entity information and constraints.

[0084] In this embodiment, preprocessing involves noise reduction of the task description text to provide a clean base text for subsequent semantic parsing. This preprocessing includes, but is not limited to, text cleaning, word segmentation, and format standardization. Terminology standardization involves standardizing and replacing abbreviations, aliases, and colloquialisms in the preprocessed text (e.g., replacing "model X" with "model X processor") to eliminate domain terminology ambiguity and ensure consistent model understanding. The first preset prompt word module is a structured instruction framework designed for the preset large language model, used to guide the model in performing domain-specific multi-task reasoning. It includes at least role definitions (e.g., "as an IT system transformation solution architect"), task definitions (entity recognition, constraint extraction, etc.), output format constraints (e.g., forced JSON output), and sample examples. The standardized task description text is standardized text obtained after preprocessing and terminology standardization, conforming to domain terminology norms and possessing a clean structure. It serves as the standardized data input to the preset large language model for deep semantic parsing.

[0085] Specifically, the task description text is preprocessed and terminology is standardized to eliminate input noise and domain terminology ambiguity, resulting in a standardized task description text. Then, task description prompts are generated based on the first preset prompt word template and the standardized task description text, and input into the preset large language model. Guided by the task description prompts, the preset large language model performs deep semantic reasoning to identify entity information (system components and their attributes) and constraints (boundary limiting conditions), and performs complexity determination based on the entity information and constraints. When the complexity is determined, the task is decomposed to obtain multiple sub-task instructions, realizing the conversion from natural language requirements to structured machine instructions.

[0086] The above technical solutions enable unambiguous conversion from user natural language requirements to structured machine instructions, significantly reducing reliance on human experience and improving the automation level and decision-making efficiency of IT system transformation solution generation.

[0087] In some embodiments, the preprocessing and terminology standardization of the modification task description text to obtain a standardized task description text includes: preprocessing the modification task description text to obtain a preprocessed requirement description, wherein the preprocessing includes text cleaning and word segmentation; and standardizing and replacing term variants in the preprocessed requirement description based on a preset domain terminology dictionary to obtain the standardized task description text; wherein the preset domain terminology dictionary includes a mapping relationship between standard terms in the domain to which the software system to be modified belongs and term variants corresponding to the standard terms.

[0088] In this embodiment, text cleaning refers to the data purification operation performed on the text describing the transformation task, including removing irrelevant characters (such as special symbols and formatting marks), correcting spelling errors, and unifying encoding formats. Word segmentation involves dividing continuous text into lexical units with independent semantic meaning, transforming the text into a computationally understandable sequence structure, providing a foundation for subsequent semantic analysis. The pre-built domain terminology dictionary refers to a pre-constructed knowledge base mapping terms specific to the domain of the software system to be transformed. It adopts a core structure of "variant-standard term-attribute," establishing a mapping relationship between standard terms and their variants (such as abbreviations, aliases, and colloquialisms). This pre-built domain terminology dictionary is constructed through automated collection of vendor documents, industry standards, and technical community data, and is manually verified. It supports dynamic iterative updates to adapt to the dynamic changes in domain terminology.

[0089] Specifically, the task description text undergoes preprocessing such as text cleaning and word segmentation to obtain a preprocessed requirement description. Then, based on a pre-defined domain terminology dictionary, the term variations in the preprocessed requirement description are standardized and replaced. Through rule matching or confidence calculation, the term variations in the text are replaced with corresponding standard terms, resulting in the standardized task description text.

[0090] The above technical solutions have enabled the precise standardization of domain terminology, effectively solving the problems of non-standard and ambiguous user input terms in software system transformation scenarios, and significantly improving the accuracy and reliability of subsequent large language model intent parsing.

[0091] In some embodiments, the first preset prompt word template includes at least a role definition, a task definition, output format constraints, and sample examples. The role definition guides the preset large language model to perform intent parsing on the modification task description text using a professional role within the domain of the software system to be modified. The task definition includes entity recognition, constraint extraction, complexity judgment, and task decomposition. Entity recognition guides the preset large language model to identify entity information from the modification task description text. Constraint extraction guides the preset large language model to identify constraints from the modification task description text. Complexity judgment guides the preset large language model to determine task complexity based on the entity information and the constraints. Task decomposition guides the preset large language model to decompose the modification task description text into multiple sub-task instructions when the task complexity is complex. The output format constraints limit the output format of the preset large language model. The sample examples provide a standard reference for intent parsing for the preset large language model.

[0092] In this embodiment, the first preset prompt word template establishes a domain-specific perspective through role definition, constructs a step-by-step reasoning chain through task definition (entity recognition, constraint extraction, complexity judgment, and task decomposition), ensures structured output through output format constraints, and provides reference standards through sample examples. These four elements work together to stimulate the thinking chain reasoning ability of the preset large language model, achieving accurate conversion from natural language to structured instructions.

[0093] In some embodiments, the target modification scheme based on structured representation utilizes a large language model to generate a target modification scheme expressed in natural language. This includes: determining structured input data based on the target modification scheme based on structured representation, the candidate component set, and the evaluation results generated during the multi-objective optimization process; generating prompts based on the structured input data and a second preset prompt word template, and inputting these prompts into the large language model to guide the large language model in generating a target modification scheme expressed in natural language. During the decoding process, the large language model references facts and values ​​in the structured input data through a pointer network mechanism and performs real-time fact consistency verification on the generated content. Simultaneously, it employs a bundle search strategy and temperature parameter adjustment to optimize the generation quality.

[0094] In this embodiment, the evaluation results can be understood as quantitative scores, Pareto frontier levels, and TOPSIS relative closeness of each candidate transformation scheme in the multi-objective optimization process across four dimensions: performance matching degree, ecological compatibility, transformation cost, and implementation risk. Structured input data refers to input data that organizes the structured target transformation schemes, candidate component sets, and evaluation results generated during the multi-objective optimization process into a unified machine-readable format (such as JSON), used to provide factual basis for the large language model to generate reports. The second preset prompt word template refers to a structured instruction framework designed for the report generation task, used to guide the large language model to convert structured data into natural language text, including role definitions (such as "senior IT system transformation architect"), report structure requirements (overview, scheme details, evaluation analysis, implementation suggestions, etc.), content specifications (highlighting advantages and supporting evidence, accurately citing technical parameters, objectively explaining risks and mitigation suggestions), and format constraints. The pointer network mechanism refers to the mechanism that calculates the probability of copying words from the input data during the decoding process to directly reference facts and values ​​in the structured input data. The bundle search strategy refers to a heuristic search strategy that retains the k candidate sequences with the highest probabilities at each generation step in order to balance generation quality and diversity.

[0095] Specifically, the structured description of the target transformation plan, candidate component set, and evaluation results generated during the multi-objective optimization process are organized into a unified JSON format to obtain structured input data. Then, based on the structured input data and a second preset prompt template, a report with generated prompts is generated. For example, the generated prompts include the following elements: You are a senior IT system transformation architect. Please generate a technical solution report based on the following JSON data: {instruction_template}. Requirements: 1. Use a standard technical solution format, including sections such as overview, solution details, evaluation analysis, and implementation recommendations; 2. Highlight the advantages of the solution and supporting evidence, accurately citing the provided technical parameters; 3. Objectively explain risk factors and propose mitigation suggestions; 4. Use professional language, clear logic, and accurate data. Input data: {json_data}.

[0096] The generated prompts are then input into a large language model, which generates report content based on an attention mechanism. To further enhance the accuracy of the report content, accuracy is ensured from two dimensions: the credibility of the content source and the consistency of logical expression. 1) During the decoding process, the large language model references facts and values ​​in the structured input data through a pointer network mechanism, specifically by calculating the replication probability: ; in, The copy probability represents the probability that a word will be copied from the input data in the current step. This refers to the hidden state of the Decoder at time step t. This is the encoding of the i-th word in the input sequence by the Encoder; For learnable parameter matrix, This means concatenating two vectors; It is the sigmoid function, which maps the score to the [0,1] interval. This method can establish a clear association between the generated content and authoritative, authentic external knowledge sources. Through pointer-like references or anchoring, every conclusion and data has a traceable basis, fundamentally avoiding the generation of unfounded information and ensuring the factuality and credibility of the content.

[0097] 2) Perform real-time factual consistency verification on the generated content by comparing whether the values ​​and names in the generated text match the structured input data to avoid illusions and ensure logical closure.

[0098] Furthermore, a beam search strategy and temperature parameter adjustment are employed to optimize the generated quality.

[0099] The beam search algorithm is employed to balance generation quality and diversity. Beam search is a heuristic graph search algorithm that retains the k sequences with the highest probabilities (k being the beam width) at each generation step, thereby balancing generation quality and diversity.

[0100] ; in, sequence arrive The cumulative log probability score; To generate words given the preceding context and input X. The probability of.

[0101] Temperature parameter adjustment: The generated creativity is controlled by the temperature parameter τ.

[0102] ; in, The probability distribution after temperature adjustment; To adjust the probability distribution before temperature; τ represents all candidate words in the vocabulary; τ is a temperature parameter that controls the smoothness of the distribution. When τ=1, the original distribution remains unchanged. When τ<1, the distribution becomes sharper, reducing randomness and increasing certainty. When τ>1, the distribution becomes smoother, increasing diversity and enhancing creativity. Here, we set τ=0.7 to ensure a balance between professionalism and appropriate creativity.

[0103] The above technical solutions enable a reliable conversion of machine decision-making results into human-readable solutions. The pointer network mechanism and fact consistency verification ensure the factual accuracy of the content. The bundle search strategy and temperature parameter adjustment balance the professionalism and creativity of the generated solutions, significantly improving the interpretability and delivery quality of the transformation solutions.

[0104] For example, Figure 2 This is a schematic diagram illustrating the generation process of a software system modification scheme provided in Embodiment 2 of the present invention. Figure 2 As shown, the system mainly includes a preparation phase and an online service phase. The preparation phase, which forms the foundation for system operation, is completed before system launch and is updated regularly. It includes the construction of a knowledge graph (i.e., a pre-defined knowledge graph), a vector knowledge base (i.e., a pre-defined vector knowledge base), and a domain-adaptive large model (i.e., a pre-defined large language model). The knowledge graph construction process includes: schema design for data related to system transformation, data preprocessing, knowledge extraction and entity alignment, knowledge fusion and graph construction to form a knowledge graph; vector knowledge base construction process: document slicing for data related to system transformation, vectorization and indexing to form a vector knowledge base; and domain-adaptive large model training process: based on instruction data, Qwen2.5-72B is selected as the base model, and LoRA fine-tuning training is used to output a domain-adaptive large model that is proficient in the knowledge of software system transformation.

[0105] The online service phase is the actual processing of user queries, receiving user inquiries related to system modification (i.e., a description of the modification task for the software system to be modified). The specific process is as follows: Invoke the domain-adaptive large model, and perform intent parsing and thought chain reasoning under the guidance of specific prompt words: identify entities, extract constraints, judge complexity, and decompose complex problems into several sub-tasks, and output structured query instructions; Based on query commands, the system performs parallel retrieval of the knowledge graph (executes graph queries to obtain precise product specifications and compatibility relationships) and retrieval of the vector knowledge base (executes semantic search to obtain similar migration cases and technical document fragments). After feature engineering processing of the multimodal retrieval results, the system calls the trained LambdaMart ranking model in real time to provide ranking results, and filters to obtain a multimodal result set (a set of candidate components for each subtask command) that integrates structured knowledge and unstructured experience. Based on the multimodal result set, multi-objective optimization is performed (considering performance, cost, compatibility, and risk). The NSGA-II algorithm is used to find the Pareto optimal solution set and conflict resolution is performed. Finally, the optimal solution is determined through TOPSIS scheme recommendation. Based on the optimal solution using structured representation, the domain-adaptive large model is invoked to transform the solution data in machine language format into a logically clear, evidence-rich, and easily human-understandable software system modification solution, which is then delivered to the user.

[0106] Furthermore, to achieve continuous improvement of models and other technologies, explicit and implicit user feedback information (i.e., feedback data) is collected: Three optimization paths are executed based on feedback data: Optimize search ranking: Train the LambdaMart ranking model using feedback data to make the ranking results more in line with user preferences; Update the knowledge graph: Based on information confirmed or refuted by users, adjust the confidence of relationships in the knowledge graph or add new relationships; Incremental fine-tuning: Select high-quality feedback data as new samples to incrementally train the domain-adaptive large model and continuously enhance the model's capabilities.

[0107] Example 2 Figure 3 This is a schematic diagram of a software system modification scheme generation device provided in Embodiment 2 of the present invention. Figure 3 As shown, the device includes: The text parsing module 21 is used to perform intent parsing based on the modification task description text of the software system to be modified, using a preset large language model to obtain entity information and constraints, as well as sub-task instructions determined according to the entity information and constraints. Data retrieval module 22 is used to retrieve data based on the query entity information and / or constraint information included in each subtask instruction in order to determine the candidate component set corresponding to each subtask instruction; The intelligent decision module 23 is used to construct a decision space based on the candidate component set corresponding to each sub-task instruction, and to perform multi-objective optimization on each candidate modification scheme in the decision space under the constraints, so as to screen and determine the target modification scheme of the structured expression of the software system to be modified. Natural language scheme generation module 24 is used to generate target transformation schemes expressed in natural language based on structured representations and utilizes a large language model.

[0108] The technical solution provided in Embodiment 2 of this invention performs intent parsing on the description text of the transformation task using a preset large language model, automatically extracts entity information and constraints, and generates sub-task instructions, thus realizing the intelligent transformation of natural language requirements into structured task instructions. Furthermore, based on the sub-task instructions, a candidate component set is retrieved and a decision space is constructed. Under constraints, multi-objective optimization solutions are performed on each candidate transformation scheme within the decision space, automatically selecting a structured component combination scheme that balances multiple constraints and multi-dimensional optimization objectives. This overcomes the reliance on human experience, achieving intelligent transformation from natural language requirements to structured component combination schemes and multi-objective optimization decision-making under multiple constraints, significantly improving the efficiency and reliability of scheme generation.

[0109] Optionally, the data retrieval module 22 includes: The query generation unit is used to generate, for each subtask instruction, a graph query statement for a preset knowledge graph and a semantic query vector for a preset vector database, based on the query entity information and / or constraint information included in the current subtask instruction. The parallel retrieval unit is used to input the graph query statement into a preset knowledge graph for querying to obtain a first retrieval list, and input the semantic query vector into a preset vector database for similarity search to obtain a second retrieval list; The fusion and reordering unit is used to fuse and reorder the first search list and the second search list to obtain the candidate component set corresponding to the subtask instruction.

[0110] Optionally, the intelligent decision-making module 23 includes: The spatial construction unit is used to extract candidate components from the candidate component set corresponding to each subtask instruction, combine them according to the technical domain category, and construct a set of candidate modification schemes to form a decision space; The spatial solution unit is used to perform multi-objective optimization on the decision space under the constraints, with performance matching degree, ecological compatibility, transformation cost and implementation risk as optimization objectives, to obtain the Pareto optimal solution set. The structured scheme determination unit is used to determine the target modification scheme of the structured representation of the software system to be modified from the Pareto optimal solution set based on a multi-criteria decision analysis algorithm.

[0111] Optionally, the spatial solution element includes: The population initialization subunit is used to determine the initial population within the decision space and constrain the generation range of the initial population based on the component compatibility relationship data stored in the preset knowledge graph. The genetic operation subunit is used to perform genetic operations within the feasible solution space that satisfies the compatibility relationship data constraints, and guides the crossover and mutation of candidate modification schemes based on the component attribute constraints of the preset knowledge graph to generate new candidate solutions. The iterative optimization subunit is used to perform Pareto sorting and genetic iteration on the new candidate solutions until the termination condition is met, thereby obtaining the Pareto optimal solution set.

[0112] Optionally, the device further includes: The feedback data acquisition module is used to collect user feedback data on the target modification plan; The optimization and update module is used to optimize and update the preset knowledge graph and / or the preset large language model based on the feedback data. The optimization and updating of the preset knowledge graph includes at least one of the following: The confidence level of relationships between entities in the preset knowledge graph is dynamically adjusted using a relationship confidence weighting algorithm. Based on the association rule mining algorithm, component co-occurrence patterns are identified. When the support and confidence of the component co-occurrence pattern meet the preset threshold, the new compatibility relationship represented by the component co-occurrence pattern is added to the preset knowledge graph.

[0113] Optionally, the text parsing module 21 includes: The text preprocessing unit is used to preprocess and standardize the terminology of the transformation task description text to obtain a standardized task description text. The prompt word generation unit is used to generate task description prompt words based on the first preset prompt word template and the standardized task description text, and input them into the preset large language model to guide the preset large language model to perform reasoning, obtain entity information and constraints, and sub-task instructions determined based on the entity information and constraints.

[0114] Optionally, the text preprocessing unit includes: The preprocessing subunit is used to preprocess the text describing the transformation task to obtain a preprocessed requirement description, wherein the preprocessing includes text cleaning and word segmentation. The terminology mapping subunit is used to standardize and replace term variations in the preprocessed requirement description based on a preset domain terminology dictionary to obtain the standardized task description text; wherein, the preset domain terminology dictionary includes the mapping relationship between standard terms in the domain to which the software system to be modified belongs and the term variations corresponding to the standard terms.

[0115] Optionally, the first preset prompt word template includes at least a role definition, a task definition, output format constraints, and sample examples. The role definition guides the preset large language model to perform intent parsing on the modification task description text using a professional role within the domain of the software system to be modified. The task definition includes entity recognition, constraint extraction, complexity judgment, and task decomposition. Entity recognition guides the preset large language model to identify entity information from the modification task description text. Constraint extraction guides the preset large language model to identify constraints from the modification task description text. Complexity judgment guides the preset large language model to determine task complexity based on the entity information and the constraints. Task decomposition guides the preset large language model to decompose the modification task description text into multiple sub-task instructions when the task complexity is complex. The output format constraints limit the output format of the preset large language model. The sample examples provide a standard reference for intent parsing for the preset large language model.

[0116] Optionally, the natural language scheme generation module 24 includes: The input data determination unit is used to determine structured input data based on the target modification scheme, the candidate component set, and the evaluation results generated during the multi-objective optimization solution process, according to the structured representation. The natural language scheme generation unit is used to generate a report and generate prompt words based on the structured input data and the second preset prompt word template, and input them into the large language model to guide the large language model to generate a target modification scheme expressed in natural language. During the decoding process, the large language model references the facts and values ​​in the structured input data through a pointer network mechanism and performs real-time fact consistency verification on the generated content. At the same time, it adopts a bundle search strategy and temperature parameter adjustment to optimize the generation quality.

[0117] The software system modification scheme generation device provided in this embodiment of the invention can execute the software system modification scheme generation method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0118] Example 3 Figure 4 This is a schematic diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0119] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0120] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0121] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as software system modification scheme generation methods.

[0122] In some embodiments, the software system modification scheme generation method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the software system modification scheme generation method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the software system modification scheme generation method by any other suitable means (e.g., by means of firmware).

[0123] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0124] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0125] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0126] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0127] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0128] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0129] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0130] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

[0131] This invention also provides a computer program product, including a computer program and / or instructions, which, when executed by a processor, implements the software system modification scheme generation method provided in any embodiment of this application.

[0132] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0133] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. A method for generating a software system modification scheme, characterized in that, include: Based on the description text of the modification task of the software system to be modified, the intent is parsed using a preset large language model to obtain entity information and constraints, as well as sub-task instructions determined according to the entity information and constraints. The candidate component set corresponding to each subtask instruction is determined by retrieving the query entity information and / or constraint information included in each subtask instruction. A decision space is constructed based on the candidate component set corresponding to each subtask instruction. Under the constraints, a multi-objective optimization solution is performed on each candidate modification scheme in the decision space to screen and determine the target modification scheme of the structured expression of the software system to be modified. Based on the structured representation of the target transformation scheme, a large language model is used to generate a target transformation scheme expressed in natural language.

2. The method according to claim 1, characterized in that, The step of retrieving the candidate component set corresponding to each subtask instruction based on the query entity information and / or constraint information included in each subtask instruction includes: For each subtask instruction, based on the query entity information and / or constraint information included in the current subtask instruction, a graph query statement for a preset knowledge graph and a semantic query vector for a preset vector database are generated for the subtask instruction. The graph query statement is input into a preset knowledge graph for querying to obtain a first search list. The semantic query vector is input into a preset vector database for similarity search to obtain a second search list. The first search list and the second search list are merged and reordered to obtain the candidate component set corresponding to the subtask instruction.

3. The method according to claim 2, characterized in that, The process involves constructing a decision space based on the candidate component set corresponding to each subtask instruction, and then performing multi-objective optimization on each candidate modification scheme in the decision space under the constraints, including: Candidate components are extracted from the candidate component set corresponding to each subtask instruction, and combined according to the technical domain category to construct a set of candidate modification schemes to form a decision space; Under the constraints, with performance matching degree, ecological compatibility, transformation cost and implementation risk as optimization objectives, the decision space is solved by multi-objective optimization to obtain the Pareto optimal solution set. Based on a multi-criteria decision analysis algorithm, the target modification scheme for the structured representation of the software system to be modified is determined from the Pareto optimal solution set.

4. The method according to claim 3, characterized in that, The process of performing multi-objective optimization on the decision space to obtain the Pareto optimal solution set includes: Within the decision space, an initial population is determined, and the generation range of the initial population is constrained based on the inter-component compatibility relationship data stored in the preset knowledge graph. Genetic operations are performed within the feasible solution space that satisfies the compatibility relationship data constraints. Based on the component attribute constraints of the preset knowledge graph, candidate modification schemes are guided to crossover and mutation to generate new candidate solutions. The new candidate solutions are subjected to Pareto sorting and genetic iteration until the termination condition is met, thus obtaining the Pareto optimal solution set.

5. The method according to claim 2, characterized in that, After generating the target modification scheme expressed in natural language, the method further includes: Collect user feedback data on the target modification plan; Based on the feedback data, the preset knowledge graph and / or the preset large language model are optimized and updated; The optimization and updating of the preset knowledge graph includes at least one of the following: The confidence level of relationships between entities in the preset knowledge graph is dynamically adjusted using a relationship confidence weighting algorithm. Based on the association rule mining algorithm, component co-occurrence patterns are identified. When the support and confidence of the component co-occurrence pattern meet the preset threshold, the new compatibility relationship represented by the component co-occurrence pattern is added to the preset knowledge graph.

6. The method according to claim 1, characterized in that, The modification task description text based on the software system to be modified is parsed using a preset large language model to obtain entity information and constraints, as well as sub-task instructions determined based on the entity information and constraints, including: The task description text is preprocessed and terminology is standardized to obtain a standardized task description text. Based on the first preset prompt word template and the standardized task description text, a task description prompt word is generated and input into the preset large language model to guide the preset large language model to perform reasoning, obtain entity information and constraints, and sub-task instructions determined based on the entity information and constraints.

7. The method according to claim 6, characterized in that, The preprocessing and terminology standardization of the task description text to obtain a standardized task description text includes: The text describing the transformation task is preprocessed to obtain a preprocessed requirement description, wherein the preprocessing includes text cleaning and word segmentation. Based on a pre-defined domain terminology dictionary, the term variations in the pre-processed requirement description are standardized and replaced to obtain the standardized task description text; wherein, the pre-defined domain terminology dictionary includes the mapping relationship between standard terms in the domain to which the software system to be modified belongs and the term variations corresponding to the standard terms.

8. The method according to claim 6 or 7, characterized in that, The first preset prompt word template includes at least a role definition, a task definition, output format constraints, and sample examples. The role definition guides the preset large language model to perform intent parsing on the modification task description text using a professional role within the domain of the software system to be modified. The task definition includes entity recognition, constraint extraction, complexity judgment, and task decomposition. Entity recognition guides the preset large language model to identify entity information from the modification task description text. Constraint extraction guides the preset large language model to identify constraints from the modification task description text. Complexity judgment guides the preset large language model to determine task complexity based on the entity information and constraints. Task decomposition guides the preset large language model to decompose the modification task description text into multiple sub-task instructions when the task complexity is complex. The output format constraints limit the output format of the preset large language model. The sample examples provide a standard reference for intent parsing for the preset large language model.

9. The method according to claim 1, characterized in that, The target modification scheme based on structured representation utilizes a large language model to generate a target modification scheme expressed in natural language, including: Based on the target modification scheme described in the structured representation, the candidate component set, and the evaluation results generated during the multi-objective optimization solution process, the structured input data is determined. The system generates prompts based on the structured input data and the second preset prompt word template, and inputs them into the large language model to guide the large language model to generate a target modification scheme expressed in natural language. During the decoding process, the large language model references the facts and values ​​in the structured input data through a pointer network mechanism and performs real-time fact consistency verification on the generated content. At the same time, it uses a bundle search strategy and temperature parameter adjustment to optimize the generation quality.

10. A software system modification scheme generation device, characterized in that, include: The text parsing module is used to perform intent parsing based on the modification task description text of the software system to be modified, using a preset large language model to obtain entity information and constraints, as well as sub-task instructions determined based on the entity information and constraints. The data retrieval module is used to retrieve data based on the query entity information and / or constraint information included in each subtask instruction in order to determine the candidate component set corresponding to each subtask instruction. The intelligent decision-making module is used to construct a decision space based on the candidate component set corresponding to each sub-task instruction, and to perform multi-objective optimization on each candidate modification scheme in the decision space under the constraints, so as to screen and determine the target modification scheme of the structured expression of the software system to be modified. The Natural Language Scheme Generation Module is used to generate target transformation schemes expressed in natural language based on structured representations and large language models.

11. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the software system modification scheme generation method according to any one of claims 1-9.