Software engineering cost function point estimation method based on large language model
By combining large language models and knowledge graphs, we build correlation maps and identify software functional modules, the deviation and subjectivity problems of traditional software functional point estimation methods are solved, and higher estimation accuracy and dynamic response capabilities are achieved.
Patent Information
- Application Number
- CN202510000752.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional software functional point estimation methods have problems with deviations in understanding requirements, subjectivity of manual estimation, and the inability to accurately reflect the dynamic changes of the system.
By combining the natural language understanding ability of the large language model with the semantic relationship reasoning ability of the knowledge graph, a business domain-application domain-data domain correlation map is built, a software functional module is identified, a software function module is judged whether it is constructed repeatedly, and a list of data and transaction functions corresponding to the functional points is output.
It improves the accuracy of functional point recognition, reduces the deviation of manual participation, and reflects the dynamic changes of the system in real time, optimizing the evaluation process of software engineering cost.
Smart Images

Figure BDA0005225004340000066 
Figure BDA0005225004340000072 
Figure BDA0005225004340000081
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of engineering cost, and in particular relates to a software engineering cost function point estimation method based on a large language model. Background Art
[0002] In the field of software engineering, function point estimation has been widely used as an important tool to measure the scale and complexity of software systems. Traditional function point estimation methods rely on software requirement documents and the experience of developers to evaluate the workload by identifying functional elements in the system (such as input, output, query, file, and interface). However, as the complexity and scale of software systems continue to grow, traditional function point estimation methods face many challenges, including deviations in the understanding of requirements, the subjectivity of manual estimation, and the inability to accurately reflect the dynamic changes of the system.
[0003] In recent years, the breakthrough of big language models in the field of natural language processing has brought new possibilities for automated and intelligent analysis. Big language models have powerful language understanding and generation capabilities, and can automatically process large amounts of text data and extract valuable information on this basis. This provides a potential solution for improving software function point estimation. At the same time, knowledge graphs, as a structured data representation method, model entities and relationships in a graph manner, which can better organize and reason about complex information. Knowledge graphs can not only provide more accurate contextual relationships for function point estimation, but also further enhance the accuracy and consistency of information based on the generation of big language models.
[0004] In the field of software engineering cost, how to combine large language models and knowledge graphs to improve the automation level and accuracy of function point estimation has become an urgent problem to be solved. Summary of the invention
[0005] The purpose of the present invention is to provide a software engineering cost function point estimation method based on a large language model. By combining the natural language understanding ability of the large language model with the semantic relationship reasoning ability of the knowledge graph, a more intelligent and accurate software function point estimation method is provided, thereby optimizing the software engineering cost evaluation process to solve the problems raised in the above background technology.
[0006] To achieve the above object, the present invention adopts the following technical solution: a method for estimating function points of software engineering cost based on a large language model, comprising the following steps:
[0007] Build a business domain-application domain-data domain association graph; use a large language model combined with pre-designed instruction templates to identify the business domain, application domain and content of each functional module corresponding to the software; determine whether the functional modules of the software are duplicated, and return the data function and transaction function list of the most similar historical functional modules for the duplicated functional modules; retrieve and identify the data function and transaction function corresponding to each functional module; and output the data and transaction function list corresponding to the functional point based on the large language model and knowledge graph.
[0008] Preferably, the association graph includes business, application and data table, the business and application are in a containment relationship, and the application and data table are in a reference relationship.
[0009] Preferably, the use of a large language model in combination with a pre-designed instruction template to identify the business domain, application domain and content of each functional module corresponding to the software includes: using regular expressions to locate chapters of the software feasibility study report, and identifying the business domain, application domain and content of each functional module corresponding to the software from the chapters through a large language model.
[0010] Preferably, the determination of whether the functional modules of the software are duplicated includes: retrieving software reports of the same business domain in a database to form a comparison library, and calculating the semantic similarity between the current software functional module and the functional modules in the comparison library; if the similarity exceeds a threshold, the functional module is considered to be duplicated.
[0011] Preferably, the method returns a list of data functions and transaction functions of the most similar historical function modules to the duplicated function modules, including directly returning the data functions and transaction functions corresponding to the historical function modules for the function modules identified as duplicated construction; wherein the data functions include internal logic files and external interface files, and the transaction functions include external input, external output and external query.
[0012] Preferably, the method further comprises: retrieving data tables associated with the software, wherein the retrieving data tables associated with the software is based on the previously identified business domain and application domain, and querying all associated data tables in the graph.
[0013] Preferably, the retrieving and identifying the data functions and transaction functions corresponding to each functional module includes: identifying the data functions and transaction functions corresponding to each functional module includes fine-tuning data set construction based on self-reflection, large language model fine-tuning, and data function and transaction function identification based on the large language model.
[0014] Preferably, the self-reflection-based fine-tuning dataset construction includes: extracting relevant information of functional modules in the comparison library to form a dataset, and guiding the large language model to learn and self-correct through instruction templates until the correct result is output.
[0015] Preferably, the method further includes fine-tuning of large language model instructions for function point estimation; the fine-tuning of large language model instructions for function point estimation includes: using a positive sample data set to perform supervised fine-tuning on the large language model to obtain a preliminary strategy, and using a feedback-based reinforcement learning optimization reward model for function points that are not successfully identified.
[0016] Preferably, the data and transaction function list corresponding to the function point output based on the large language model and knowledge graph includes: for the description of new software function point, after identifying its business domain and application domain, finding the associated data table set through graph query, using the large model under the optimized strategy to output candidate answers, and selecting the candidate answer with the highest reward value as the final answer to determine the workload of the function point.
[0017] Technical effects and advantages of the present invention: Compared with the prior art, the software engineering cost function point estimation method based on a large language model proposed by the present invention has the following advantages:
[0018] The present invention combines the natural language understanding ability of the large language model with the semantic relationship reasoning ability of the knowledge graph to provide a more intelligent and accurate software function point estimation method, thereby optimizing the software engineering cost assessment process. This method can not only greatly improve the accuracy of function point identification, but also reduce the deviation of manual participation and reflect the dynamic changes of the system in real time. It provides a new technical means for cost estimation and resource allocation of complex software projects, which helps to improve the efficiency and quality of software development projects. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 The present invention is a flow chart of the method for estimating function points of software engineering cost based on a large language model. DETAILED DESCRIPTION
[0020] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0021] The present invention provides a method for estimating function points of software engineering cost based on a large language model, which can not only greatly improve the accuracy of function point identification, but also reduce the deviation caused by manual participation and reflect the dynamic changes of the system in real time.
[0022] Before describing the present implementation in detail, the abbreviations and key terms in this application are defined as follows:
[0023] Knowledge Graph: Knowledge Graph is a structured data representation method that displays entities (such as people, places, things, etc.) and their relationships in the form of a graph. It uses nodes to represent entities and edges to represent relationships, which can help machines better understand semantics and support information query and reasoning.
[0024] Large Language Model (LLM): Large Language Model is a type of AI model based on deep learning. It can learn a large amount of language knowledge and general domain knowledge through large-scale pre-training on massive unlabeled data. It also has general multi-task solving capabilities through key technologies such as instruction fine-tuning and reinforcement learning based on human feedback. Typical large language models include the GPT-4 series, Tongyi Qianwen series, and Wenxin Yiyan series.
[0025] Engineering Cost (EngineeringCostEstimation): Engineering cost refers to the total cost estimated and predicted during the planning and implementation phase of an engineering project. It covers materials, labor, equipment, construction management and other related expenses. Engineering cost usually helps project managers and decision makers effectively control budgets and costs by comprehensively evaluating factors such as project design drawings, contracts, and market prices.
[0026] Function Point Estimation: Function Point Estimation is a method used to evaluate the scale and complexity of software development projects. It determines the workload of a project by calculating the function points in the software system (such as input, output, query, file, interface, etc.).
[0027] Five-point method: The five-point method is a software engineering cost estimation method. The main steps include determining the function point analysis type, identifying the system boundary, identifying the function point elements, determining the relevant factors, and calculating the function point scale. The core is to identify the function point elements of the software, including data functions (i.e. internal logic file I LF and external interface file EI F) and transaction functions (external input EI, external output EO and external query EQ), a total of five points, so it is called the "five-point method".
[0028] like Figure 1 As shown, a method for estimating function points of software engineering cost based on a large language model in this embodiment includes the following steps:
[0029] Step 1. Build a business domain-application domain-data domain association graph
[0030] Based on the one-to-many relationship between the business domain and the application domain, and the one-to-many relationship between the application domain and the data domain, a business domain-application domain-data domain association graph G is constructed, in which the entity types include "business", "application" and "data table". The relationship between "business" and "application" is an "inclusion" relationship, and the relationship between "application" and "data table" is an "involvement" relationship.
[0031] Step 2. Identify the business domain and application domain corresponding to the software, as well as the content of each functional module
[0032] Based on regular expressions, the "Business Architecture" and "Application Architecture" sections of the software feasibility study report are located. A large language model combined with a pre-designed instruction template is used to identify the business domain corresponding to the software from the "Business Architecture" section, and to identify the application domain corresponding to the software and the contents of each functional module from the "Application Architecture" section.
[0033] Step 3. Determine whether the functional modules of the software are duplicated
[0034] First, search the database for software feasibility study reports that belong to the same business domain as the current software, and then identify the contents of their various functional modules according to the method in step 2 to form a comparison library D f ={f1,f2,…,f n}, where f i is the content of the i-th functional module, and n represents the comparison library D f Finally, for each functional module of the current software, the similarity between it and each functional module in the comparison library is calculated using semantic similarity. If the similarity exceeds the threshold, the functional module is considered to be duplicated, see step 4. Otherwise, the functional module is considered not to be duplicated, see step 5.
[0035] Step 4. Return the most similar historical function module
[0036] For a functional module F of the current software, use the semantic similarity method to calculate its similarity with each functional module f in the comparison library. i The similarity of i , let the historical function module with the largest similarity be f, and its corresponding similarity be s max , if s max If it is greater than the preset threshold α, it is considered that the current function module F and the historical function module f are duplicated, and the data function (i.e., internal logic file I LF and external interface file E IF) and transaction function (external input EI, external output EO and external query EQ) corresponding to f are directly returned.
[0037] Step 5. Retrieve the data table associated with the software
[0038] Based on the business domain and application domain identified in step 2, Query all data tables associated with it T = {t1, t2, ..., t m}, where t i It represents the i-th associated data table, which consists of the description of each table and the name and description of each field. m represents the number of all associated data tables.
[0039] Step 6. Identify the data functions and transaction functions corresponding to each functional module
[0040] In step 6, this solution uses the large language model to identify the corresponding data functions and transaction functions according to the content description of each functional module and the data table associated with the software, specifically including step 6.1 fine-tuning dataset construction based on self-reflection, step 6.2 large language model fine-tuning, and step 6.3 data function and transaction function identification based on the large language model.
[0041] Step 6.1 Self-reflection based fine-tuning dataset construction
[0042] For comparison library D f Each functional module f i , extract the associated data table set T i , corresponding data function and transaction function list y i ={ILF i ,EIF i ,EI i ,EO i ,EQ i}, forming a data set where d i ={f i ,T i ,y i}.
[0043] Design instruction template p1, requiring the large language model to be based on the description of each functional module f i and the associated data table set T i Output in Including the data function list and transaction function list identified by LLM, as shown below:
[0044]
[0045] Design instruction template p2, requiring large language model comparison y i and judge Are the data function list and transaction function list included consistent with y iand give reflections i If y i and If the i Should contain successful experience. If not, then r i Lessons from failure should be included.
[0046] Design instruction template p3 for y in step (3) i and Inconsistent questions require large language models to reflect on i Re-answer. If the answer is still incorrect, repeat this step until the answer is correct. Therefore, the question-answering rules for round t are as follows.
[0047]
[0048] in, Represents the reflections of rounds t-1, t-2, …, and 0.
[0049] According to steps (2) to (4), when using the large language model to identify the data function and transaction function list corresponding to each functional module, the output may be similar to y i Matching or not matching Here it will match Together with reflection i Recorded as All Composition of positive sample data set will not match Together with reflection i Recorded as All Composing negative sample dataset
[0050] Step 6.2 Fine-tuning the large language model instructions for function point estimation
[0051] From the positive sample data set Find the function point samples that are successfully recognized by the large language model for the first time, and use these samples to perform supervised fine-tuning on the large language model to obtain the strategy π SFT .
[0052] For those function points that were not successfully recognized by the large language model for the first time, and There are corresponding samples in each of them. These samples are combined to perform feedback-based reinforcement learning on the large language model, and the optimal reward model r is obtained by minimizing the following loss function. θ .
[0053]
[0054] Using the strategy π in step (1) SFT Initialize the large language model. Given any software function point description f and the associated data table set T, use the large language model to output the corresponding data function and transaction function list, and then use the reward model r in step (2) θ Give rewards or penalties to the output, and finally obtain the optimal strategy by maximizing the overall reward Represents the optimal model. The loss function is shown below.
[0055]
[0056] Where β is the equilibrium and π SFT The adjustment coefficient.
[0057] Step 7. Output the data and transaction function list corresponding to the function point based on the large language model and knowledge graph
[0058] For a new software function point description f, identify the business domain and application domain corresponding to the software, and then use the graph to Query to find the data table set T associated with it, and then use The large model under the strategy outputs N candidate answers Finally, use the reward model r in step 6.2 θ Calculate the reward for each candidate answer, select the candidate answer with the highest reward value as the final answer, and substitute the list of data functions and transaction functions contained in the final answer into the cost formula to calculate the workload corresponding to the function point.
[0059] When compiling software engineering estimates, the present invention can automatically match similar projects in historical records and find out duplicate function points based on semantic similarity. For duplicate function points, a list of data functions and transaction functions of the function points can be automatically generated based on historical records. For newly constructed software function points, a set of data tables can be automatically associated based on the knowledge graph, and then the data table and specific operation type to be operated by the function point can be identified in combination with a large language model, thereby intelligently generating a list of data functions and transaction functions of the software.
[0060] In summary, the present invention combines the knowledge graph of planning task project library, enterprise architecture, and historical feasibility report function points with the large language model to identify whether there is duplicate construction in the new content of the information system project feasibility report and prompt historical related duplicate construction content, and automatically identify the internal logic files, external interface files, query, input, and output operations according to the five-point method.
[0061] In addition, by combining the natural language understanding ability of the large language model with the semantic relationship reasoning ability of the knowledge graph, a more intelligent and accurate software function point estimation method is provided, thereby optimizing the evaluation process of software engineering cost. This method can not only greatly improve the accuracy of function point identification, but also reduce the deviation of manual participation and reflect the dynamic changes of the system in real time. It provides a new technical means for cost estimation and resource allocation of complex software projects, which helps to improve the efficiency and quality of software development projects.
[0062] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A software engineering cost function point estimation method based on a large language model, characterized in that: The following steps are involved: Build a business domain-application domain-data domain association map; Use a large language model combined with pre-designed instruction templates to identify the business domain, application domain and content of each functional module corresponding to the software; Determine whether the functional modules of the software are duplicated, and return the data function and transaction function list of the most similar historical functional modules for the duplicated functional modules; Retrieve and identify the data functions and transaction functions corresponding to each functional module; and Output the data and transaction function list corresponding to the function points based on the large language model and knowledge graph.
2. The software engineering cost function point estimation method based on a large language model according to claim 1 is characterized in that: The association graph includes business, application and data table. The business and application are in a containment relationship, and the application and data table are in a reference relationship.
3. The software engineering cost function point estimation method based on a large language model according to claim 1 is characterized in that: The use of a large language model in combination with a pre-designed instruction template to identify the business domain, application domain and content of each functional module corresponding to the software includes: Use regular expressions to locate the chapters of the software feasibility study report, and use the large language model to identify the corresponding business domain, application domain and functional module content of the software from the chapters.
4. The method for estimating function points of software engineering cost based on a large language model according to claim 3 is characterized in that: The step of determining whether the functional modules of the software are repeatedly constructed includes: Retrieve software reports of the same business domain from the database to form a comparison library, and calculate the semantic similarity between the current software function module and the function module in the comparison library. If the similarity exceeds the threshold, the function module is considered to be duplicated.
5. The method for estimating function points of software engineering cost based on a large language model according to claim 4, characterized in that: The data function and transaction function list of the most similar historical function module is returned to the duplicated function module, including For the function modules identified as duplicated, directly return the data functions and transaction functions corresponding to the historical function modules; The data functions include internal logic files and external interface files, and the transaction functions include external input, external output and external query.
6. The method for estimating function points of software engineering cost based on a large language model according to claim 5, characterized in that: The method further comprises: The data tables associated with the retrieval software are queried in the graph based on the previously identified business domain and application domain to find all associated data tables.
7. The method for estimating function points of software engineering cost based on a large language model according to claim 6, characterized in that: The retrieving and identifying the data function and transaction function corresponding to each functional module includes: Identifying the data functions and transaction functions corresponding to each functional module includes fine-tuning dataset construction based on self-reflection, large language model fine-tuning, and data function and transaction function identification based on the large language model.
8. The method for estimating function points of software engineering cost based on a large language model according to claim 7, characterized in that: The self-reflection-based fine-tuning dataset construction includes: The relevant information of the functional modules in the comparison library is extracted to form a data set, and the large language model is guided to learn and self-correct through instruction templates until the correct results are output.
9. The method for estimating function points of software engineering cost based on a large language model according to claim 8, characterized in that: The method also includes fine-tuning of large language model instructions for function point estimation; The large language model instruction fine-tuning for function point estimation includes: A large language model is fine-tuned in a supervised manner using a positive sample dataset to obtain a preliminary strategy, and a feedback-based reinforcement learning optimization reward model is used for function points that are not successfully identified.
10. A method for estimating function points of software engineering cost based on a large language model according to claim 9, characterized in that: The data and transaction function list corresponding to the function points output based on the large language model and knowledge graph includes: For the description of new software function points, after identifying its business domain and application domain, the associated data table set is found through graph query, the large model under the optimized strategy is used to output candidate answers, and the candidate answer with the highest reward value is selected as the final answer to determine the workload of the function point.
Citation Information
Patent Citations
Intelligent software cost evaluation method and system capable of enabling large language model
CN117635243A
Software workload assessment method and device, electronic equipment and storage medium
CN118093376A
Vertical type government affair large model service method and system based on interactive learning
CN118820448A