Method for measuring and selecting ts open source software based on static analysis and llm
Patent Information
- Application Number
- CN202610984232.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-29
AI Technical Summary
[0005]为解决背景技术中提到的问题,本发明提出一种方法,可以解决现有TypeScript代码种类繁多、质量参差不齐、选型难度大的问题,本发明通过以下技术方案来实现
本发明采用了面向 TypeScript 开源软件的多粒度文档生成与功能关键词检索相结合的技术方案,相较于现有技术缺乏检索手段、仅依靠人工筛选候选软件的技术方案,具有快速定位功能匹配的开源软件、大幅缩短软件选型前期准备周期的效果;
Smart Images

Figure CN122837848A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of software engineering technology, specifically to a method for measuring and selecting open-source software based on static analysis and LLM. Background Technology
[0002] TypeScript (TS) is a superset of JavaScript. Since its release by Microsoft in 2012, it has become a mainstream programming language for front-end, full-stack, and even back-end development thanks to its core advantages of static typing and full JavaScript compatibility. In the 2023 Stack Overflow Developer Survey, TypeScript became the third most popular general-purpose programming language after JavaScript and Python. Its open-source ecosystem has experienced explosive growth, with TS open-source projects on platforms such as GitHub and GitLab covering multiple areas including web frameworks, tool libraries, and cloud-native components, forming a vast and complex ecosystem.
[0003] At the same time, the rapid development of TypeScript has also brought significant industry pain points: First, TypeScript still lacks a formal language specification, its syntax structure continues to evolve rapidly, and some new features lack standardized usage guidance, making it difficult for developers to master them systematically; Second, there is a large amount of redundancy in the TS open source ecosystem with similar functions, and the usage volume of different libraries varies greatly, making it difficult for developers to quickly select libraries that meet their needs, and also making it difficult to efficiently understand the overall architecture and business logic of large open source projects.
[0004] While existing TypeScript parsing tools (such as TypeDoc, ts-morph, Babel+acorn, etc.) have achieved mature single-file code parsing and AST generation capabilities, they generally suffer from core limitations: they all focus on the syntax parsing of a single TS code file, failing to achieve overall architecture analysis at the code repository level, full call relationships, and dependency analysis. Furthermore, their functional extensibility is insufficient, making it difficult to meet developers' needs for systematic analysis of large TS projects. Therefore, there is an urgent need for a TypeScript-specific measurement and selection tool to fill this technological gap and address the industry's core pain points. Summary of the Invention
[0005] To address the problems mentioned in the background section, this invention proposes a method that can solve the problems of numerous types of existing TypeScript code, inconsistent quality, and difficulty in selection. This invention achieves this through the following technical solution.
[0006] A method for measuring and selecting TS open-source software based on static analysis and LLM, characterized by the following steps: S1 receives the TypeScript file path and the result output path, and performs validity checks on the parameters and the target file. S2, extract the metadata of the TypeScript file to be analyzed and construct the metadata object; S3: Traverse the abstract syntax tree of the source file, extract the core attributes of functions, classes, and class methods, generate unique identifiers, and store them in a structured list. At the same time, extract the valid call relationship between function calls and class instantiation, and store it in a structured list. The core attributes include at least name, starting line number, export status, and number of parameters. S4 integrates the extracted metadata, function list, class list, and call relationship list into an analysis result object, serializes it into JSON format, and then constructs a directed graph of class relationships and a directed graph of function calls based on networkx, outputting structured analysis results; S5 extracts multi-dimensional evaluation indicators from static analysis results and retrieval platform data, uses a large language model to construct a selection evaluation system, and completes the quantitative and qualitative evaluation of the candidate warehouse. S6 performs in-depth validation on the top-ranked candidate warehouses using a large language model, and generates a visual analysis report and final selection recommendations by combining the static analysis results.
[0007] Furthermore, step S1 specifically includes: S11 extracts the path to the TypeScript file to be analyzed and the output path of the JSON format analysis results from the command line parameters, verifies the completeness of the parameters, and outputs an error message and terminates the process if any parameter is missing. S12, verify the physical existence of the TypeScript file to be analyzed. If the file does not exist, output an error message and terminate the process. S13, verify that the object pointed to by the path to be analyzed is a file type. If it is a directory, output an error message and terminate the process. Furthermore, step S2 specifically includes: S21, create a ts-morph Project instance, add the TypeScript file to be analyzed as the project source file, and complete the parsing environment initialization; S22: Count the number of lines of code in the file to be analyzed by reading the file and splitting it by newline character. If a reading error occurs, the line count is assigned to 0. S23 extracts the full file path and basic file name from the source file, and combines them with the count of lines of code to construct a file metadata object containing the path, name, and line count.
[0008] Furthermore, the core attributes extracted from functions, classes, and class methods in step S3 specifically include: S31, traverse the ordinary function nodes in the source file, generate a unique ID for each ordinary function consisting of the function name / anonymous identifier and the file base name, extract the core attributes of function name, starting line number, export status, and number of parameters, and store them in a structured list of functions; S32, traverse the class nodes in the source file, generate a unique ID for each class consisting of the class name / anonymous identifier and the file base name, adapt to different versions of ts-morph to get the parent class name list, extract the class name, starting line number, export status core attributes, and store them in a structured manner in the class list; S33. Traverse the method nodes of each class, generate a unique ID for each class method consisting of class name.method name / anonymous identifier and file base name, extract the core attributes of method name, class name, starting line number, export status, and number of parameters, and store them in a structured way in the function list.
[0009] Furthermore, the extraction of valid call relationships in step S3 specifically includes: S34, traverse and identify function call nodes in the abstract syntax tree, obtain the function / method to which the call node belongs and generate its unique ID, extract the name of the called function and generate its unique ID, filter out the valid call relationships of the called functions in the current file, and record the caller ID, callee ID, and call line number and store them in a structured manner. S35, traverse the attribute declaration nodes in the class nodes, extract the type identifier of each attribute and call the type cleaning function to remove generic parameters and union type modifiers, match the cleaned type names with the extracted class name set, filter out the valid class association relationship of the attribute type corresponding to the class in the current file, record the source class ID, target class ID and the line number of the attribute and store it in a structured manner.
[0010] Furthermore, step S4 specifically includes: S41, integrate the extracted file metadata, functions, classes, and call relationships into a unified analysis result object and serialize it into an indented JSON string, and write it into the result output path specified in step S1 in UTF-8 encoding; S42 outputs the parsing statistics to the console, including the number of extracted functions, the number of classes, and the number of valid call relationships; S43, construct a directed graph of class relationships, add class nodes and function nodes in the call relationship, configure node attributes for the function nodes, and add inheritance relationship edges and instantiation relationship edges; S44, construct a directed graph of function calls, add function / method nodes and configure node attributes for the nodes, add call relationship edges and configure call line number attributes for the call relationship edges.
[0011] Furthermore, the selection and evaluation system in step S5 includes: S51, parse the software comparison request parameters, load the list of names of the software to be compared, the structured evaluation result data, the callback notification address and the unique request identifier, perform format verification and legality verification on the evaluation result data, and deserialize the evaluation result data into a standard evaluation result object; S52: Extract all external API documents from the evaluation result data of the software to be compared, parse and generate their respective API name sets, calculate the intersection ratio of the API name sets and compare it with a preset threshold to determine whether they are different versions of the same software. S53, if they are different versions of the same software, then distinguish between the new and old versions, extract the new APIs and deprecated APIs, and call the large language model to generate a comparison report between the new and old versions; S54, if they are different functional software, then the repository documents and all module documents of each software are combined, and a large language model is called to generate a comparative analysis report that includes functional points, applicable scenarios and selection suggestions.
[0012] Furthermore, when the comparison process in step S52 is executed successfully, the result data containing the unique request identifier, comparison report, and success status is encapsulated and the comparison result is returned to the requester through the callback interface; when the comparison process is executed abnormally, various error information in parsing, model calling, and network transmission is captured, detailed error logs are recorded, the failure status and error details are encapsulated and returned through the callback interface.
[0013] Furthermore, the deep verification in step S6 specifically includes: S61 receives the user's input function requirements and deserializes them into a structured object that matches the parsing result, thus completing the parsing and structured verification of the user input. S62, calculate the similarity between the user's input functional requirements and the candidate repository functions, calculate Jaccard similarity based on API name, calculate TF-IDF cosine similarity based on repository documents, and calculate text similarity based on module documents to initially screen out repositories whose functions and requirements are relatively well matched. S63. Input the repository-level and module-level functional documents of the software repository into the large language model and score them from three aspects: functional matching degree, document completeness, and repository usability. The functional matching degree evaluates whether the functions, modules, and APIs provided by the repository meet the core indicators of user needs. The document completeness evaluates whether the repository's documentation, module introductions, and functional descriptions are complete, clear, and easy to use. The repository usability evaluates the adaptability of the repository for actual project implementation, including ease of use, scalability, and compatibility. S64: Based on the user-defined rating weights and minimum thresholds for individual ratings, normalize the scores of each candidate warehouse, and then perform a secondary screening based on user needs for warehouses with scores below the threshold. S65: Sum the individual scores of each dimension according to the configuration weight to obtain the comprehensive selection score of the candidate warehouse, sort them from high to low, and output the Top-N high-quality candidate warehouse list. S66 outputs final selection recommendations, identifies the optimal and alternative repositories, and provides implementation guidance on integration steps, dependency handling, and performance optimization.
[0014] Furthermore, the content uniformly stored in step S7 includes: user-input keywords and technical constraints, matching records from the retrieval platform, batch static analysis JSON data of candidate warehouses, GEXF files of class relationship diagrams and function call diagrams, multi-dimensional evaluation scoring tables, fusion verification results, in-depth analysis reports, and selection recommendations; all data are stored in categories according to warehouse name and support retrieval by keywords and indicators.
[0015] The present invention adopts the above technical solution and has the following beneficial effects: This invention adopts a technical solution that combines multi-granularity document generation and functional keyword retrieval for TypeScript open-source software. Compared with the existing technical solution that lacks retrieval methods and relies solely on manual screening of candidate software, it has the effect of quickly locating open-source software with matching functions and significantly shortening the preparation cycle for software selection. This invention adopts a technical solution based on the collaborative evaluation of quantitative data from static code analysis and large language models. Compared with the existing technical solution that relies solely on development experience and repository descriptions for subjective judgment, this invention provides objective and quantifiable criteria for software selection and effectively improves the scientific nature and accuracy of selection decisions. This invention adopts a technical solution of integrating candidate software with target project analysis and pre-integration risk detection. Compared with the existing technical solution that exposes compatibility issues only after the selection is completed, it has the effect of avoiding dependency conflicts and module coupling risks in advance and significantly reducing the cost of selection and implementation. This invention adopts a technical solution that adapts to the full-process measurement and selection support of TypeScript syntax features. In view of the characteristics of TypeScript's rapid iteration of new features and syntax and lack of unified code standards, it updates the syntax library in a timely manner and lowers the error threshold. Compared with the existing general analysis solutions that cannot fit the characteristics of the TypeScript ecosystem, this invention has the effect of strong compatibility and improving the success rate and adaptability of TypeScript project selection. Attached Figure Description
[0016] Figure 1 The flowchart illustrates the overall process of the TS open-source software measurement and selection method based on static analysis and LLM, as described in the specific implementation.
[0017] Figure 2The flowchart illustrates the static analysis and graph construction phase as described in the specific implementation method.
[0018] Figure 3 This is a flowchart illustrating the multi-dimensional selection and evaluation phase of open-source software as described in a specific implementation method.
[0019] Figure 4 This is a flowchart illustrating the candidate repository in-depth verification and selection suggestion generation stage as described in a specific implementation method.
[0020] Figure 5 This is a flowchart illustrating the data storage and traceability process described in a specific implementation. Detailed Implementation
[0021] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] This embodiment provides a method for measuring and selecting TS open-source software based on static analysis and LLM. (See also...) Figure 1 The method in this embodiment specifically includes the following steps: S1 receives the TypeScript file path and the output path, and performs validation of the parameters and target file; see [link / reference] Figure 2 , Figure 2 This demonstrates the entire preprocessing logic of TypeScript code, from parameter and file validity validation, ts-morph environment initialization, Abstract Syntax Tree (AST) generation, core entity extraction of functions / classes, call relationship identification, to the construction of directed graphs of function calls and class relationships. Through structured data extraction and visualization graph construction, it provides core technical quantitative indicators and code structure basis for subsequent open-source software selection. Step 1) specifically includes the following steps: S11 extracts the path to the TypeScript file to be analyzed and the output path of the JSON format analysis results from the command line parameters, verifies the completeness of the parameters, and outputs an error message and terminates the process if any parameter is missing.
[0023] S12 verifies the physical existence of the TypeScript file to be analyzed. If the file does not exist, an error message is output and the process is terminated.
[0024] S13, verify that the object pointed to by the path to be analyzed is a file type. If it is a directory, output an error message and terminate the process.
[0025] S2, extract the metadata of the TypeScript file to be analyzed and construct the metadata object; Step 2) specifically includes the following steps: S21, create a ts-morph Project instance, add the TypeScript file to be analyzed as the project source file, and complete the parsing environment initialization.
[0026] S22 counts the number of lines of code in the file to be analyzed by splitting the data by newline character during file reading, and assigns the line count to 0 when a reading error occurs.
[0027] S23 extracts the full file path and basic file name from the source file, and combines them with the count of lines of code to construct a file metadata object containing the path, name, and line count.
[0028] S3 involves traversing the abstract syntax tree of the source file, extracting the core attributes of functions, classes, and class methods, generating unique identifiers, and storing them in a structured manner. It also extracts the valid call relationships between function calls and class instantiations and stores them in a structured manner. The core attributes must include at least a list, a quantity, and a node. The specific steps in S3 for extracting the core attributes of functions, classes, and class methods include: S31: Traverse the ordinary function nodes in the source file, generate a unique ID for each ordinary function consisting of the function name / anonymous identifier and the file base name, extract the core attributes of function name, starting line number, export status, and number of parameters, and store them in a structured way in the function list.
[0029] S32 iterates through the class nodes in the source file, generates a unique ID for each class consisting of the class name / anonymous identifier and the file base name, adapts to different versions of ts-morph to extract the parent class name list using the parent class retrieval interface, extracts the class name, starting line number, and export status core attributes, and stores them in a structured manner in the class list.
[0030] S33. Traverse the method nodes of each class, generate a unique ID for each class method consisting of class name.method name / anonymous identifier and file base name, extract the core attributes of method name, class name, starting line number, export status, and number of parameters, and store them in a structured way in the function list.
[0031] The specific steps for extracting valid call relationships in step S3 include: S34: Traverse and identify function call nodes in the abstract syntax tree, obtain the function / method to which the call node belongs and generate its unique ID, extract the name of the called function and generate its unique ID, filter out the valid call relationships of the called functions in the current file, and record the caller ID, callee ID, and call line number and store them in a structured manner.
[0032] S35, traverse the attribute declaration nodes in the class nodes, extract the type identifier of each attribute and call the type cleaning function to remove generic parameters and union type modifiers, match the cleaned type names with the extracted class name set, filter out the valid class association relationship of the attribute type corresponding to the class in the current file, record the source class ID, target class ID and the line number of the attribute and store it in a structured manner.
[0033] S4, integrate the extracted metadata, lists, class lists, and call relationship lists into an analysis result object, serialize it into JSON format, and then construct a directed graph of class relationships and a directed graph of function calls based on NetworkX, outputting structured analysis results; Step 4) specifically includes the following steps: S41, construct a unified analysis result object containing file metadata, a list of functions, a list of classes, and a list of call relationships; serialize the analysis result object into an indented JSON string and write it to the output path specified in step S1 in UTF-8 encoding.
[0034] S42 outputs the parsing statistics to the console, including the number of extracted functions, the number of classes, and the number of valid call relationships.
[0035] S43, construct a directed graph of class relationships, add class nodes and function nodes in the call relationship, configure the node with type, name, file path, and line number attributes, add inheritance relationship edges based on the parent class information of the class, and add instantiation relationship edges based on the class instantiation relationship.
[0036] S44, construct a directed graph of function calls, add function / method nodes and configure the node with attributes such as type, name, class name, export status, and number of parameters, add call relationship edges based on valid function call relationships, and configure the call line number attribute for the edges.
[0037] S5 extracts multi-dimensional evaluation indicators from static analysis results and retrieval platform data, and uses a large language model to construct a selection evaluation system to complete the quantitative and qualitative evaluation of the candidate repository; see reference Figure 3 , Figure 3 This demonstrates a complete evaluation chain, from candidate repository matching on the search platform, extraction of multi-dimensional technology / ecosystem indicators, construction of a five-dimensional selection evaluation system, to quantitative scoring and intelligent ranking of candidate repositories. It replaces traditional subjective selection experience with objective and quantifiable evaluation standards, enabling the scientific screening of TypeScript open-source software and providing data support for selection decisions. Step S5 specifically includes:
[0038] S51 parses the software comparison request parameters, loads the list of software names to be compared, structured evaluation result data, callback notification address and unique request identifier, performs format validation and legality verification on the evaluation result data, and deserializes it into a standard evaluation result object.
[0039] S52. Extract all external API documents from the evaluation results of the software to be compared, parse and generate their respective API name sets, calculate the intersection ratio of each pair, compare the ratio with a preset threshold, and determine whether the two software are different versions of the same software. If so, proceed to step S53; otherwise, proceed to step S54.
[0040] S53. If the two versions are determined to be different versions of the same software, the old and new versions are distinguished according to the total number of APIs in the two software versions. The API documents of the old and new versions are traversed to filter out the new APIs that only exist in the new version and the obsolete APIs that only exist in the old version. The core document information such as the name and function description of the corresponding APIs is extracted and organized. Based on the integrated version comparison reference document set, a standardized comparison report of the old and new versions containing functional differences, API changes and optimization points is generated by calling the large model.
[0041] S54. If the software is determined to be different functional software, all software to be compared are traversed, and the repository documents and all module documents of each software are extracted and spliced in turn to form a set of software function description documents in a unified format containing complete functional information and module information. Based on this set, a large language model is called to generate a standardized comparative analysis report containing functional points, applicable scenarios and selection suggestions.
[0042] When the comparison process in step S52 is executed successfully, the result data, including the unique request identifier, comparison report, and success status, is encapsulated and returned to the requester through a callback interface. When the comparison process encounters an error, various error information such as parsing, model invocation, and network transmission is captured, detailed error logs are recorded, and the failure status and error details are encapsulated and returned through a callback interface.
[0043] S6 performs deep validation on the top-ranked candidate warehouses using a large language model, and generates a visual analysis report and final selection recommendations based on the static analysis results; see reference Figure 4 , Figure 4 This demonstrates the entire process from integrating top-scoring candidate repositories with the target project, detecting coupling / dependency conflict risks, generating in-depth analysis reports, to providing final selection recommendations and integration guidance. By proactively detecting integration risks, the rework costs of project implementation are significantly reduced, providing clear implementation guidelines for project integration. Step S6 specifically includes the following steps:
[0044] S61 receives the user's input function requirements and deserializes them into a structured object that matches the parsing result, thus completing the parsing and structured verification of the user input.
[0045] S62, calculate the similarity between user input and candidate repository functions, calculate Jaccard similarity based on API name, calculate TF-IDF cosine similarity based on repository documents, and calculate text similarity based on module documents, to initially screen out repositories whose functions and requirements are relatively well matched.
[0046] S63 inputs the repository-level and module-level functional documentation of the software repository into the large model, and scores it from three aspects: functional matching degree, documentation completeness, and repository usability. Functional matching degree evaluates whether the functions, modules, and APIs provided by the repository meet the core indicators of user needs. Documentation completeness evaluates whether the repository's documentation, module introductions, and functional descriptions are complete, clear, and easy to use. Repository usability evaluates the adaptability of the repository for actual project use, including ease of use, scalability, and compatibility, with a maximum score of 100 and a minimum score of 0.
[0047] In S64, users set the scoring weights for each dimension and the minimum threshold for a single item's score according to their needs. The tool normalizes the scores of each candidate repository and then filters the repositories with scores below the threshold based on user needs.
[0048] S65, Comprehensive Score and Ranking Output: The individual scores of each dimension are weighted and summed according to the configuration weight to obtain the comprehensive selection score of the candidate warehouse. The candidate warehouses are sorted from high to low and the Top-N high-quality candidate warehouse list is output to provide a priority basis for subsequent in-depth verification.
[0049] S66, based on the in-depth analysis report, outputs the final selection recommendation, identifies the optimal and alternative repositories, and provides implementation guidance such as integration steps, dependency handling, and performance optimization.
[0050] S7 stores all analysis and selection data in a unified manner, enabling traceability and retrieval of selection results. (See also...) Figure 5 , Figure 5 This demonstrates the logic for data collection, classification, archiving, multi-dimensional index construction, persistent storage, and retrieval traceability throughout the entire selection process. This ensures the reproducibility and traceability of the entire selection process, providing historical data support and a reusable foundation for subsequent selection tasks of similar TypeScript open-source software. Step S7 uniformly stores the following content: user-input keywords and technical constraints, matching records from the retrieval platform, batch static analysis JSON data of candidate repositories, class relationship diagrams and function call diagrams in GEXF files, multi-dimensional evaluation scoring tables, fusion verification results, in-depth analysis reports, and selection recommendations. Data is stored categorized by repository name and supports retrieval by keywords and indicators.
[0051] The above descriptions are embodiments of the present invention, but the specific embodiments described herein are merely illustrative and not intended to limit the invention. Any omissions, modifications, or equivalent substitutions made within the scope of the claims of this invention without departing from the principles and spirit of the invention should be included within the scope of protection of the claims.
Claims
1. A method for measuring and selecting TS open-source software based on static analysis and LLM, characterized in that, The method includes the following steps: S1 receives the TypeScript file path and the result output path, and performs validity checks on the parameters and the target file. S2, extract the metadata of the TypeScript file to be analyzed and construct the metadata object; S3: Traverse the abstract syntax tree of the source file, extract the core attributes of functions, classes, and class methods, generate unique identifiers, and store them in a structured list. At the same time, extract the valid call relationship between function calls and class instantiation, and store it in a structured list. The core attributes include at least name, starting line number, export status, and number of parameters. S4 integrates the extracted metadata, function list, class list, and call relationship list into an analysis result object, serializes it into JSONL format, and then constructs a directed graph of class relationships and a directed graph of function calls based on networkx, outputting structured analysis results; S5 extracts multi-dimensional evaluation indicators from static analysis results and retrieval platform data, uses a large language model to construct a selection evaluation system, and completes the quantitative and qualitative evaluation of the candidate warehouse. S6 uses a large language model to perform in-depth validation on the candidate warehouses with the highest scores, and combines the static analysis results to generate a visual analysis report and final selection recommendations. S7 stores all process analysis and selection data in a unified manner, enabling traceability and retrieval of selection results.
2. The method for measuring and selecting TS open-source software based on static analysis and LLM as described in claim 1, characterized in that, Step S1 specifically includes: S11 extracts the path to the TypeScript file to be analyzed and the output path of the JSON format analysis results from the command line parameters, verifies the completeness of the parameters, and outputs an error message and terminates the process if any parameter is missing. S12, verify the physical existence of the TypeScript file to be analyzed. If the file does not exist, output an error message and terminate the process. S13, verify that the object pointed to by the path to be analyzed is a file type. If it is a directory, output an error message and terminate the process.
3. The method for measuring and selecting TS open-source software based on static analysis and LLM as described in claim 1, characterized in that, Step S2 specifically includes: S21, create a ts-morph Project instance, add the TypeScript file to be analyzed as the project source file, and complete the parsing environment initialization; S22: Count the number of lines of code in the file to be analyzed by reading the file and splitting it by newline character. If a reading error occurs, the line count is assigned to 0. S23 extracts the full file path and basic file name from the source file, and combines them with the count of lines of code to construct a file metadata object containing the path, name, and line count.
4. The method for measuring and selecting TS open-source software based on static analysis and LLM as described in claim 1, characterized in that, The core attributes extracted from functions, classes, and class methods in step S3 specifically include: S31, traverse the ordinary function nodes in the source file, generate a unique ID for each ordinary function consisting of the function name / anonymous identifier and the file base name, extract the core attributes of function name, starting line number, export status, and number of parameters, and store them in a structured list of functions; S32, traverse the class nodes in the source file, generate a unique ID for each class consisting of the class name / anonymous identifier and the file base name, adapt to different versions of ts-morph to get the parent class name list, extract the class name, starting line number, export status core attributes, and store them in a structured manner in the class list; S33. Traverse the method nodes of each class, generate a unique ID for each class method consisting of class name.method name / anonymous identifier and file base name, extract the core attributes of method name, class name, starting line number, export status, and number of parameters, and store them in a structured way in the function list.
5. The method for measuring and selecting TS open-source software based on static analysis and LLM as described in claim 1, characterized in that, The extraction of valid call relationships in step S3 specifically includes: S34, traverse and identify function call nodes in the abstract syntax tree, obtain the function / method to which the call node belongs and generate its unique ID, extract the name of the called function and generate its unique ID, filter out the valid call relationships of the called functions in the current file, and record the caller ID, callee ID, and call line number and store them in a structured manner. S35, traverse the attribute declaration nodes in the class nodes, extract the type identifier of each attribute and call the type cleaning function to remove generic parameters and union type modifiers, match the cleaned type names with the extracted class name set, filter out the valid class association relationship of the attribute type corresponding to the class in the current file, record the source class ID, target class ID and the line number of the attribute and store it in a structured way; S35: Traverse and identify class instantiation nodes in the abstract syntax tree, obtain the function / method to which the class instantiation node belongs and generate its unique ID, extract the name of the instantiated class and generate its unique ID, filter out the valid instantiation relationships of the instantiated class in the current file, and record the caller ID, the instantiated class ID, the call line number and store them in a structured manner.
6. The method for measuring and selecting TS open-source software based on static analysis and LLM as described in claim 1, characterized in that, Step S4 specifically includes: S41, integrate the extracted file metadata, functions, classes, and call relationships into a unified analysis result object and serialize it into an indented JSON string, and write it into the result output path specified in step S1 in UTF-8 encoding; S42 outputs the parsing statistics to the console, including the number of extracted functions, the number of classes, and the number of valid call relationships; S43, construct a directed graph of class relationships, add class nodes and function nodes in the call relationship, configure node attributes for the function nodes, and add inheritance relationship edges and instantiation relationship edges; S44, construct a directed graph of function calls, add function / method nodes and configure node attributes for the nodes, add call relationship edges and configure call line number attributes for the call relationship edges.
7. The method for measuring and selecting TS open-source software based on static analysis and LLM as described in claim 1, characterized in that, The selection and evaluation system in step S5 includes: S51, parse the software comparison request parameters, load the list of names of the software to be compared, the structured evaluation result data, the callback notification address and the unique request identifier, perform format verification and legality verification on the evaluation result data, and deserialize the evaluation result data into a standard evaluation result object; S52: Extract all external API documents from the evaluation result data of the software to be compared, parse and generate their respective API name sets, calculate the intersection ratio of the API name sets and compare it with a preset threshold to determine whether they are different versions of the same software. S53, if they are different versions of the same software, then distinguish between the new and old versions, extract the new APIs and deprecated APIs, and call the large language model to generate a comparison report between the new and old versions; S54, if they are different functional software, then the repository documents and all module documents of each software are combined, and a large language model is called to generate a comparative analysis report that includes functional points, applicable scenarios and selection suggestions.
8. The method for measuring and selecting TS open-source software based on static analysis and LLM as described in claim 7, characterized in that, When the comparison process in step S52 is executed successfully, the result data containing the unique request identifier, comparison report, and success status is encapsulated and the comparison result is returned to the requester through the callback interface. When the comparison process is executed abnormally, various error information in parsing, model calling, and network transmission is captured, detailed error logs are recorded, the failure status and error details are encapsulated and returned through the callback interface.
9. The method for measuring and selecting TS open-source software based on static analysis and LLM as described in claim 1, characterized in that, The deep verification in step S6 specifically includes: S61 receives the user's input function requirements and deserializes them into a structured object that matches the parsing result, thus completing the parsing and structured verification of the user input. S62, calculate the similarity between the user's input functional requirements and the candidate repository functions, calculate Jaccard similarity based on API name, calculate TF-IDF cosine similarity based on repository documents, and calculate text similarity based on module documents to initially screen out repositories whose functions and requirements are relatively well matched. S63. Input the repository-level and module-level functional documents of the software repository into the large language model and score them from three aspects: functional matching degree, document completeness, and repository usability. The functional matching degree evaluates whether the functions, modules, and APIs provided by the repository meet the core indicators of user needs. The document completeness evaluates whether the repository's documentation, module introductions, and functional descriptions are complete, clear, and easy to use. The repository usability evaluates the adaptability of the repository for actual project implementation, including ease of use, scalability, and compatibility. S64: Based on the user-defined rating weights and minimum thresholds for individual ratings, normalize the scores of each candidate warehouse, and then perform a secondary screening based on user needs for warehouses with scores below the threshold. S65: Sum the individual scores of each dimension according to the configuration weight to obtain the comprehensive selection score of the candidate warehouse, sort them from high to low, and output the Top-N high-quality candidate warehouse list. S66 outputs final selection recommendations, identifies the optimal and alternative repositories, and provides implementation guidance on integration steps, dependency handling, and performance optimization.
10. The method for measuring and selecting TS open-source software based on static analysis and LLM as described in claim 1, characterized in that, The unified storage content in step S7 includes: user-input keywords and technical constraints, matching records from the retrieval platform, batch static analysis JSON data of candidate warehouses, GEXF files of class relationship diagrams and function call diagrams, multi-dimensional evaluation scoring tables, fusion verification results, in-depth analysis reports, and selection recommendations; all data is stored in categories according to warehouse name and supports retrieval by keywords and indicators.