Software project quality defect analysis method and device, equipment and medium
By analyzing multi-source heterogeneous data and using multi-dimensional models, combined with historical defect data and interaction correlation analysis, the problem of low accuracy in software project quality defect analysis has been solved. This enables a comprehensive and systematic assessment of the health of software projects and risk warning, thereby improving the effectiveness and predictability of project quality management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies have low accuracy in software project quality defect analysis and cannot comprehensively and systematically reflect the true health status of the project, resulting in low coverage of evaluation results and a lack of comprehensive data collection and evaluation across multiple dimensions such as architecture design, test completeness, documentation quality, team collaboration efficiency, and security.
By acquiring multi-source heterogeneous project data, code quality and development feature analysis is performed. A pre-built multi-dimensional quality defect analysis model is used to generate a quality defect index. Combined with historical defect data, trend analysis and interaction correlation analysis are conducted to identify quality influencing factors, generate quality alarm information, and provide governance decision-making suggestions.
It significantly expands the coverage of technology health assessment, improves the accuracy and comprehensiveness of assessment results, can identify risk points of accelerated health deterioration in advance, enhances the initiative and foresight of project quality management, and provides objective and quantitative data support.
Smart Images

Figure CN121935129A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cloud security technology, and in particular to a method, apparatus, equipment and medium for analyzing quality defects in software projects. Background Technology
[0002] As the complexity and iteration speed of software systems continue to increase, various technical problems introduced during the long-term evolution of software due to development decisions, time pressure, or failure to follow best practices will continue to accumulate, leading to a decline in system maintainability, scalability, and stability. This phenomenon is often referred to in the industry as technical health decay or technical debt.
[0003] In the healthcare field, software has become the infrastructure supporting diagnosis, treatment, patient data management, and medical research. This type of software involves complex business processes, stringent data compliance requirements, and extremely high reliability standards. A continuous decline in the health of this software technology can directly lead to disruptions in the treatment process, risks to patient data security, or errors in clinical decision support, ultimately impacting patient safety and the quality of healthcare.
[0004] In the fintech sector, software such as mobile payments, core banking systems, quantitative trading platforms, and risk management engines are central to financial services. These systems have stringent requirements for transaction consistency, real-time response capabilities, security, and regulatory compliance. Deterioration in the internal technical health of these systems can lead to transaction failures, fund security vulnerabilities, systemic risks, and even compliance penalties, directly impacting the sound operation of financial institutions.
[0005] Currently, the industry primarily relies on static code analysis tools (such as SonarQube), code quality platforms, and human architecture assessment meetings to identify and manage such issues. However, static analysis lacks a comprehensive data collection and assessment model covering multiple dimensions such as architecture design, test completeness, documentation quality, team collaboration efficiency, and security. This results in assessment reports that fail to comprehensively and systematically reflect the true technical health of the project, leading to low coverage of assessment results and consequently, lower accuracy in analyzing software project quality defects. Summary of the Invention
[0006] This invention provides a method, apparatus, equipment, and medium for analyzing quality defects in software projects, in order to solve the technical problem of low accuracy in analyzing quality defects in software projects.
[0007] Firstly, a method for analyzing quality defects in software projects is provided, including: Obtain multi-source heterogeneous project data of the target software project, analyze the code quality in the multi-source heterogeneous project data to obtain first analysis data, and analyze the project development characteristics in the multi-source heterogeneous project data to obtain second analysis data; Using a pre-built multi-dimensional quality defect analysis model, the quality score corresponding to each quality dimension in the target software project is analyzed based on the first analysis data and the second analysis data. The quality score corresponding to each quality dimension is then integrated into the quality defect index corresponding to the target software project. Obtain historical defect data of the target software project, analyze the evolution trend of the quality defect index based on the historical defect data, and generate quality alarm information for the target software project based on the evolution trend. Interactive correlation analysis is performed on the first analysis data and the second analysis data to obtain the correlation events in the multi-source heterogeneous project data, and the quality influencing factors corresponding to the target software project in the correlation events are identified. Analyze the quality defects of the target software project based on the quality influencing factors and the quality alarm information.
[0008] Secondly, a software project quality defect analysis device is provided, comprising: The data analysis module is used to acquire multi-source heterogeneous project data of the target software project, analyze the code quality in the multi-source heterogeneous project data to obtain first analysis data, and analyze the project development characteristics in the multi-source heterogeneous project data to obtain second analysis data. The quality defect index analysis module is used to analyze the quality score corresponding to each quality dimension of the target software project based on the first analysis data and the second analysis data using a pre-built multi-dimensional quality defect analysis model, and to merge the quality score corresponding to each quality dimension into the quality defect index corresponding to the target software project. The quality alarm information generation module is used to acquire historical defect data of the target software project, analyze the evolution trend of the quality defect index based on the historical defect data, and generate quality alarm information of the target software project based on the evolution trend. The quality influencing factor identification module is used to perform interactive correlation analysis on the first analysis data and the second analysis data to obtain the correlation events in the multi-source heterogeneous project data, and identify the quality influencing factors corresponding to the target software project in the correlation events. The quality defect analysis module is used to analyze the quality defects of the target software project based on the quality influencing factors and the quality alarm information.
[0009] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the quality defect analysis method for the aforementioned software project.
[0010] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the quality defect analysis method for the aforementioned software project.
[0011] In the aforementioned software project quality defect analysis method, apparatus, equipment, and medium, the following scheme can be implemented: Multi-source heterogeneous project data of the target software project can be obtained through a client; code quality in the multi-source heterogeneous project data is analyzed to obtain first analysis data; project development characteristics in the multi-source heterogeneous project data are analyzed to obtain second analysis data; a pre-built multi-dimensional quality defect analysis model is used to analyze the quality score corresponding to each quality dimension of the target software project based on the first and second analysis data; the quality scores corresponding to each quality dimension are merged into a quality defect index corresponding to the target software project; historical defect data of the target software project is obtained; the evolution trend of the quality defect index is analyzed based on the historical defect data; and quality alarm information of the target software project is generated based on the evolution trend; the first and second analysis data are then analyzed. This invention analyzes data through interactive correlation analysis to obtain related events in multi-source heterogeneous project data and identifies the quality influencing factors corresponding to the target software project within these related events. Based on these quality influencing factors and quality alarm information, it analyzes the quality defects of the target software project and feeds these defects back to the client. By establishing a quantitative evaluation model covering multiple dimensions such as code, architecture, testing, documentation, collaboration, and security, and integrating static code data with dynamic process data, the scope of technical health assessment is significantly expanded, enabling the assessment results to more comprehensively and systematically reflect the true health status of the software project. Using time series analysis technology to decompose and predict historical quality defect indices, it is possible to identify risk points of accelerated deterioration in technical health in advance, generate quality alarm information, and improve the initiative and foresight of project quality management. Interactive correlation analysis and causal inference of multi-source data can trace from surface technical symptoms to deeper root causes such as development processes and team collaboration, improving the effectiveness and return on investment of governance measures. By generating correlation maps of quality influencing factors, governance decision-making recommendations for software project quality defects are generated, providing objective and quantitative data support for project managers' resource allocation decisions, thereby solving the problem of low accuracy in analyzing software project quality defects. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a schematic diagram of an application environment for a software project quality defect analysis method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a method for analyzing quality defects in software projects according to an embodiment of the present invention; Figure 3 yes Figure 2 A flowchart illustrating a specific implementation method of step S1; Figure 4 yes Figure 2 A flowchart illustrating a specific implementation method of step S3; Figure 5 This is a schematic diagram of a software project quality defect analysis device according to an embodiment of the present invention; Figure 6 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 7 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0014] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0015] The software project quality defect analysis method provided in this invention can be applied to, for example... Figure 1In this application environment, the client communicates with the server via a network. The server can obtain multi-source heterogeneous project data of the target software project from the client, analyze the code quality in the multi-source heterogeneous project data to obtain first analysis data, and analyze the project development characteristics in the multi-source heterogeneous project data to obtain second analysis data. Using a pre-built multi-dimensional quality defect analysis model, the server analyzes the quality score corresponding to each quality dimension of the target software project based on the first and second analysis data, and merges the quality scores corresponding to each quality dimension into a quality defect index corresponding to the target software project. The server obtains historical defect data of the target software project, analyzes the evolution trend of the quality defect index based on the historical defect data, and generates quality alarm information for the target software project based on the evolution trend. The server performs interactive correlation analysis on the first and second analysis data to obtain... This invention identifies correlated events in multi-source heterogeneous project data and determines the quality influencing factors corresponding to the target software project within these events. Based on these quality influencing factors and quality alarm information, it analyzes quality defects in the target software project and feeds these defects back to the client. By establishing a quantitative evaluation model covering multiple dimensions such as code, architecture, testing, documentation, collaboration, and security, and integrating static code data with dynamic process data, the scope of technical health assessment is significantly expanded, enabling the assessment results to more comprehensively and systematically reflect the true health status of the software project. Utilizing time series analysis technology to decompose and predict historical quality defect indices allows for the early identification of risk points leading to accelerated deterioration of technical health, generating quality alarm information and enhancing the initiative and foresight of project quality management. Interactive correlation analysis and causal inference of multi-source data allow for tracing from surface technical symptoms to deeper root causes such as development processes and team collaboration, improving the effectiveness and return on investment of governance measures. Through the correlation graph of quality influencing factors, governance decision-making suggestions for software project quality defects are generated, providing objective and quantitative data support for project managers' resource allocation decisions, thereby solving the problem of low accuracy in software project quality defect analysis. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.
[0016] Please see Figure 2 As shown, Figure 2 A flowchart illustrating a software project quality defect analysis method provided in this embodiment of the invention includes the following steps: S1. Obtain multi-source heterogeneous project data of the target software project, analyze the code quality in the multi-source heterogeneous project data to obtain first analysis data, and analyze the project development characteristics in the multi-source heterogeneous project data to obtain second analysis data.
[0017] In this embodiment of the invention, multi-source heterogeneous project data refers to a data set with differences in structure and semantics collected from different tools and platforms throughout the entire lifecycle of software project development and operation.
[0018] In detail, multi-source heterogeneous project data of the target software project can be obtained from a pre-stored storage area through computer statements with data scraping capabilities (such as Java statements, Python statements, etc.), where the stored data includes, but is not limited to, databases and blockchains.
[0019] In this embodiment of the invention, the first analytical data refers to the feature data that has been cleaned and standardized and is used to characterize the intrinsic quality of the code.
[0020] In this embodiment of the invention, the step of analyzing the code quality in the multi-source heterogeneous project data to obtain first analysis data includes: The set of original code metrics in the code repository of the multi-source heterogeneous project data is extracted based on the preset data interface specifications. The original code metric set is transformed according to a preset code quality feature mapping table to obtain a standardized code feature vector; Noise data points are detected in the standardized code feature vector, the noise data points are cleaned to form a target code feature vector, and the target code feature vector is used as the first analysis data.
[0021] In detail, through predefined application programming interface (API) call specifications or database connection protocols, raw, unprocessed sets of metrics—i.e., raw code metric sets—are automatically extracted from repositories of version control systems such as Git and SVN, and from the output of static code analysis tools such as SonarQube and Checkstyle. These raw code metric sets refer to the basic data set directly extracted from the code repository that can quantify code attributes. For example, in analyzing the core transaction module of a fintech system, the hash values, authors, and timestamps of the most recent 1000 commits can be extracted from the Git repository, and raw data such as the cyclomatic complexity, number of duplicate lines of code, and non-compliance with coding standards for each Java file can be extracted from the SonarQube scan report. In a healthcare scenario, for an electronic medical record system, basic data such as the number of lines of code in the medical record data storage module, the duplication rate of diagnostic logic code, and the coupling degree of the interface with medical devices can be extracted.
[0022] Specifically, a code quality feature mapping table refers to a pre-established set of mapping rules used to correlate raw code metrics with standardized code quality features. Using this predefined mapping rule table, raw metrics from different sources and with different dimensions are converted into unified, comparable numerical features, i.e., standardized code feature vectors. This mapping table defines rules such as "cyclomatic complexity greater than 20 maps to feature value 0.8" and "0 security vulnerabilities per thousand lines of code maps to feature value 1.0." In a fintech scenario, the original loop complexity of a fintech system is 18, which, after transformation by the mapping table, results in a standardized loop complexity of 0.72. The original code duplication rate is 25%, directly mapped to a standardized feature of 0.25, ultimately forming a standardized vector containing this type of feature. In a healthcare scenario, the original function nesting depth of a remote diagnostic system is 6, which, after transformation, results in a standardized nesting depth of 0.6. The original annotation rate is 15%, mapped to a standardized feature of 0.15, constructing the corresponding feature vector.
[0023] Furthermore, based on statistical methods (such as outlier detection based on Z-score) or machine learning models (such as the Isolation Forest algorithm), data points deviating from the normal range in the feature vector are identified. These noise points may be caused by false positives from analysis tools, extreme edge cases, or data collection errors. For example, if the detection finds that the annotation density feature value of a file is 0 (possibly because the tool failed to parse the annotations), this is identified as a noise point. When cleaning noise points, this value can be replaced with the average annotation density of the entire module, or the file can be directly excluded from the current round of analysis. After processing, the target code feature vector is obtained, which serves as the final first analysis data. This solves the problem of distortion in analysis results caused by noise and outliers in the original data, significantly improving the accuracy and robustness of subsequent quality assessments.
[0024] In this embodiment of the invention, the second analytical data refers to refined and correlated feature data used to characterize the dynamics of the development process and the behavioral patterns of the team.
[0025] In this embodiment of the invention, reference is made to Figure 3 As shown, the analysis of project development characteristics in the multi-source heterogeneous project data to obtain second analysis data includes: S31. Parse the version control logs and project management records in the multi-source heterogeneous project data; S32. Extract the time sequence and metadata of development activities in the target software project based on the version control log and the project management record; S33. Identify key pattern indicators of the development activities based on the time series and the metadata; S34. Associate the key pattern indicators with the context information of the project phase in the development activity to generate second analysis data.
[0026] In detail, version control logs refer to log data that records the history of code version changes in a target software project, including information such as code committer, commit time, changes, changed files, and commit notes. Project management records refer to record data generated in project management tools that are related to the project development process, including information such as requirement and task assignments, task completion status, iteration cycle planning, defect fixing records, and team member changes. The process involves calling the log reading interface of version control tools (such as Git and SVN) to obtain the complete version control logs of the target software project and organizing them into structured data in chronological order. Then, using the open API of project management tools (such as Jira and Trello), key fields such as requirement number, task description, responsible person, start time, end time, defect level, and fix duration are extracted from the project management records. The two types of data are then formatted and aligned to form a unified structured dataset. In fintech scenarios, for a securities trading system, the Git logs can be used to extract the transaction logic changes and commit times for each code commit, while Jira can be used to extract the allocation records of transaction function requirements and defect repair records. In healthcare scenarios, the version control logs of a vaccination management system can provide code change information for the vaccination data statistics module, and the project management records can be used to extract the task progress and related defect repair status for vaccine process optimization requirements.
[0027] Specifically, the temporal sequence of development activities refers to the sequence data formed by arranging various activities in the development process according to their chronological order. Meta-information refers to the basic information describing the core attributes of development activities, including activity type, participating entities, duration, related requirements or defects, etc. The parsed structured data is sorted and organized chronologically to form an event sequence with rich meta-information. For example, an event sequence could be generated as follows: Time T1: Developer A submitted code on module M1 (number of lines added / deleted: 200, 50); Time T2: Jira ticket BUG-101 was marked as resolved, with a resolution time of 48 hours; Time T3: The automated build for module M1 failed, thus achieving dynamic and temporal modeling of the development process. Quantitative indicators that characterize development efficiency, collaboration patterns, and process health can be extracted from the event sequence. For example, this includes calculating the average code submission frequency, mean time to repair (MTTR) of defects, build failure rate, and the average time from requirement submission to closure. For instance, in the analysis of a healthcare information system, it might be found that the average repair time of the image processing module is significantly longer than that of other modules. The calculated metrics are correlated with the specific stage of the project (such as requirements analysis, centralized development, pre-deployment testing, and operations), taking into account context such as team structure and technology stack. For example, high code commit frequency may be normal during centralized development, while during operations, it may indicate frequent patching and system instability. After correlation, a set of metrics with context labels is formed, i.e., the second set of analytical data. This solves the problem of misjudgment that may result from understanding process metrics out of context, making the process analysis results more interpretable and practical.
[0028] Furthermore, the first and second analysis data describe the state of the software project from two complementary perspectives: static outputs and dynamic processes. Neither type of data alone can comprehensively assess the decline in technical health. Therefore, it is necessary to input these two heterogeneous data types into a unified model for comprehensive analysis and quantitative fusion to generate a global, comparable health evaluation index.
[0029] S2. Using a pre-built multi-dimensional quality defect analysis model, analyze the quality score corresponding to each quality dimension in the target software project based on the first analysis data and the second analysis data, and merge the quality score corresponding to each quality dimension into the quality defect index corresponding to the target software project.
[0030] In this embodiment of the invention, the multi-dimensional quality defect analysis model is a comprehensive evaluation system covering six dimensions: code quality, architecture health, test completeness, documentation quality, team collaboration efficiency, and security. This enables a comprehensive quantitative assessment of technical debt in software projects, overcoming the limitations of traditional single-dimensional code quality assessment. The quality score refers to the numerical value obtained through the multi-dimensional quality defect analysis model, used to quantitatively evaluate the health status of each quality dimension.
[0031] In this embodiment of the invention, the step of using a pre-constructed multi-dimensional quality defect analysis model to analyze the quality score corresponding to each quality dimension in the target software project based on the first analysis data and the second analysis data includes: Identify the quality dimension data in the first analysis data and the second analysis data; According to the preset mapping rules, the quality dimension data are respectively input to the input layer of each quality dimension of the multi-dimensional quality defect analysis model, and the dimension score of each quality dimension data is output. The adaptive weight configuration algorithm in the multi-dimensional quality defect analysis model is invoked based on the attribute tags of the target software project. The adaptive weight configuration algorithm is used to dynamically adjust the contribution weights of each quality dimension input layer to the quality dimension. The original quality score for each quality dimension in the target software project is obtained by weighting the contribution weight and the dimension score.
[0032] In detail, based on the quality dimensions defined in the model (such as code quality, architecture health, test completeness, documentation quality, team collaboration efficiency, and security), a subset of features directly or indirectly related to each dimension is selected from the first and second analysis data. For example, code complexity is a feature belonging to the code quality dimension; build failure rate may be related to both test completeness and team collaboration efficiency. For instance, in a fintech scenario, the quality dimension data of a cryptocurrency trading system includes data such as the standardized value of loop complexity and code duplication rate for the code quality dimension, the compliance data of the encryption module code for the security dimension, and the data such as the completion time of cross-departmental development tasks for the team collaboration efficiency dimension. In a healthcare scenario, the quality dimension data of a pathology analysis system includes data such as the interface coupling degree for the architecture health dimension, the frequency of reporting defects related to clinical data testing for the test completeness dimension, and the data related to the completeness of pathology analysis algorithm documentation for the documentation quality dimension.
[0033] Specifically, the multi-dimensional quality defect analysis model employs an independent computational submodule (input layer and subsequent processing layer) for each quality dimension. Each submodule receives a corresponding subset of features and outputs a score between 0 and 1 through its internal computational logic (such as a small neural network or a set of rule functions), representing the health score of that dimension. For example, the code quality submodule receives code feature vectors and outputs a score of 0.65; the architecture health submodule receives features related to dependencies and coupling and outputs a score of 0.8. Attribute labels include project type (e.g., financial transaction system, medical imaging system), business criticality level (e.g., high availability requirements, security and compliance requirements), and current stage (e.g., startup phase, rapid growth phase, maintenance phase). For example, for a medical imaging AI analysis project in a rapid growth phase, the weights of architecture health and scalability may need to be increased; for a financial payment system in a strict compliance phase, the weights of security and documentation completeness are crucial. The adaptive weight configuration algorithm selects a set of weight values from a predefined weight configuration strategy library or calculates them in real time based on the attribute labels. ,in The total number of quality dimensions, and For example, the weights assigned to the aforementioned medical imaging project might be: code quality (0.15), architecture health (0.25), test completeness (0.20), documentation quality (0.10), team collaboration efficiency (0.20), and security (0.10). Multiplying the score output by each dimension's sub-module by its contribution weight yields a weighted score, reflecting the relative importance of different dimensions within the current project context.
[0034] In this embodiment of the invention, the quality defect index is a scalar value used to quantify the severity of the overall technical health decline of the target software project. The higher the value, the worse the health and the heavier the technical debt.
[0035] In detail, the weighted raw quality scores for each dimension are synthesized using an aggregation function (such as weighted summation or Euclidean distance-based aggregation) to ultimately calculate a single Quality Defect Index (TDI). For example, using weighted summation: ,in It is the first The raw quality scores for each dimension. According to this formula, if the scores for each dimension are high (poor health), the TDI value will also be high. This index provides a global, quantitative snapshot of health.
[0036] Furthermore, through the analysis of multi-source heterogeneous project data, we obtained primary analytical data reflecting code quality and secondary analytical data reflecting development activity characteristics. This data provides a comprehensive foundation for quality assessment, but the data itself cannot directly reflect the project's quality defects. Therefore, it is necessary to construct a multi-dimensional quality defect analysis model to conduct professional and multi-dimensional analysis of this data, transforming the scattered data into scores and indices that can quantify the quality status. This will enable a comprehensive assessment of project quality defects and provide core evidence for subsequent risk warning and root cause analysis.
[0037] S3. Obtain historical defect data of the target software project, analyze the evolution trend of the quality defect index based on the historical defect data, and generate quality alarm information for the target software project based on the evolution trend.
[0038] In this embodiment of the invention, historical defect data refers to a collection of quality defect indices and their related metadata calculated and stored at multiple points in time (such as daily or weekly).
[0039] In detail, historical defect data of a target software project can be retrieved from a pre-stored storage area using computer statements (such as Java statements, Python statements, etc.) with data scraping capabilities, where the stored data includes, but is not limited to, databases and blockchains.
[0040] In this embodiment of the invention, the evolution trend refers to the changing pattern and development direction of the quality defect index at different time stages, including upward trend, downward trend, stable trend or fluctuating trend, etc.
[0041] In this embodiment of the invention, reference is made to Figure 4 As shown, the step of analyzing the evolution trend of the quality defect index based on the historical defect data includes: S41. Identify the historical quality defect index in the historical defect data and extract the generation timestamp corresponding to the historical quality defect index; S42. Construct time series data based on the historical quality defect index and the generated timestamp; S43. Perform trend decomposition on the quality defect index based on the time series data to obtain long-term trend components, periodic components, and residual components; S44. Determine the evolution trend of the quality defect index based on the long-term trend component, the periodic component, and the residual component.
[0042] In detail, historical data is identified and extracted to construct, such as The time series, where For timestamps, This refers to the quality defect index at the corresponding time. For example, in a fintech scenario, the historical defect data of a bank's credit card management system shows a quality defect index of 35 generated at 10:30 AM on January 15, 2024, and 42 generated at 2:15 PM on February 20, 2024. These indices and their corresponding timestamps are extracted and correlated. In a healthcare scenario, the historical defect data of a rehabilitation assessment system shows a quality defect index of 28 generated at 9:45 AM on March 5, 2024, and 33 generated at 11:20 AM on April 10, 2024. The correlation between the index and the timestamp is completed. Trend decomposition is performed using algorithms such as STL (Seasonalization) or Hodrick-Prescott filtering, which decomposes the original TDI sequence into three parts: long-term trend component... (A smooth curve reflecting the long-term rise or fall of TDI), cyclical components (Reflecting regular fluctuations such as version release cycles) and residual components (Reflecting the impact of random noise and sudden events). The long-term trend components are fitted using a locally weighted regression scatter smoothing method to capture the overall direction of the index's change over a long period; periodic components are extracted based on Fourier transform to identify the fixed period and amplitude of index fluctuations; the residual components are obtained by subtracting the long-term trend components and periodic components from the original time series data.
[0043] Specifically, by calculating the slope (first derivative) and rate of change of the slope (second derivative) of the trend line within the recent time window, it can be determined whether the health decline is accelerating, decelerating, or remaining stable. Analyzing the direction and rate of change of the long-term trend component, if the long-term trend component continues to rise and the rate of increase is greater than a preset threshold, it indicates a significant upward trend in the quality defect index, with accelerated accumulation of technical debt; if the long-term trend component continues to decline and the rate of decline is greater than a preset threshold, it indicates a significant downward trend in the index, with continuous improvement in quality; if the fluctuation amplitude of the long-term trend component is less than a preset threshold, it is a stable trend. Combining the fluctuation cycle and amplitude of the cyclical component, the changing pattern of the index within each cycle is analyzed, for example, whether the index shows a regular increase at the end of the iteration cycle. Referring to the fluctuation range of the residual component, if the fluctuation amplitude of the residual component is within a reasonable range, the trend analysis results are reliable; if the residual component exhibits large abnormal fluctuations, supplementary analysis is required based on the specific project situation at each time point. Combining the analysis results of the three components, the overall evolution trend of the quality defect index is determined. For example, in the historical data analysis of a financial risk control system, it may be found that the long-term trend component of its TDI has a continuously positive slope and an increasing value over the past three months, indicating that the technical health is deteriorating at an accelerated pace.
[0044] In this embodiment of the invention, quality alarm information refers to notification information generated based on the evolution trend of the quality defect index and in combination with preset alarm rules, used to alert project quality risks, including alarm level, risk description, potential impact, key time points, etc.
[0045] In detail, based on the analysis of evolutionary trends and combined with preset threshold rules, different levels of alarms are automatically triggered. For example, the rule is set as follows: if the recent average slope of the trend component exceeds the threshold θ1, and there are consecutive sudden increases in the residual component exceeding the threshold θ2, then a red alarm is triggered.
[0046] Furthermore, the current quality defect index of the target software project only reflects the current quality status, but it cannot reflect the development and change patterns of quality defects, nor can it predict future quality risks. Historical defect data, however, contains past quality defect indices and related information. By analyzing this data, we can uncover the evolution trend of the quality defect index, and then judge potential future quality risks based on the trend, generating corresponding quality alarm information, thus extending from current status assessment to risk warning. S4. Perform interactive correlation analysis on the first analysis data and the second analysis data to obtain the correlation events in the multi-source heterogeneous project data, and identify the quality influencing factors corresponding to the target software project in the correlation events.
[0047] In this embodiment of the invention, a related event refers to a combination pattern of specific code feature changes and specific development activities that are highly correlated in time and frequently co-occur.
[0048] In this embodiment of the invention, the step of performing interactive correlation analysis on the first analysis data and the second analysis data to obtain the correlation events in the multi-source heterogeneous project data includes: Align and match the code feature change points in the first analysis data with the development activity events in the second analysis data on the timeline to obtain target matching data; In the target matching data, scan for code feature changes and development activity events with a frequency higher than a preset occurrence frequency; Determine whether the scanned code feature changes and development activity events constitute the same co-occurrence pattern; When the scanned code feature changes and development activity events form the same co-occurrence pattern, the co-occurrence pattern is identified as a related event in the multi-source heterogeneous project data.
[0049] In detail, code feature change points refer to specific data points in the first analysis data that reflect significant changes in code quality characteristics, such as a sudden increase in loop complexity or a significant decrease in code duplication rate. Development activity events refer to specific development activities recorded in the second analysis data, such as emergency code submissions, requirement changes, team member adjustments, and shortened testing cycles. The process involves extracting the occurrence time of each code feature change point from the first analysis data and the occurrence time of each development activity event from the second analysis data; constructing a unified timeline and marking the code feature change points and development activity events on the timeline according to their occurrence time; setting a time matching threshold (e.g., 24 hours) and associating code feature change points on the timeline with development activity events at intervals not exceeding this threshold; and compiling successfully matched combinations into target matching data, with each data point containing information on the code feature change point, development activity event, and the time interval between them.
[0050] Specifically, the frequency of occurrence refers to a pre-set minimum threshold for selecting statistically significant code feature changes and development activity events. This threshold is determined based on the project's development cycle and data volume. The process involves statistically analyzing the frequency of occurrence of each code feature change and development activity event in the target matching data; comparing this frequency with the preset frequency, and selecting code feature changes and development activity events with frequencies exceeding the preset threshold. Using association rule mining algorithms (such as the Apriori algorithm), patterns with high support (frequency of occurrence) and confidence (conditional probability) are searched among all aligned matching data. The co-occurrence relationships between the selected high-frequency code feature changes and development activity events are analyzed; the support and confidence of the combination of code feature changes and development activity events are calculated. Support represents the probability of this combination appearing in the target matching data, and confidence represents the probability of the corresponding code feature change appearing when a development activity event occurs. If both support and confidence are higher than the preset threshold (e.g., support ≥ 20%, confidence ≥ 70%), then the two are determined to be the same co-occurrence pattern. For example, a high-frequency pattern might be observed: within one week of closing an urgent deployment type Jira ticket, the probability of code duplication in the relevant modules increasing exceeds 70%. When code feature changes and development activity events are determined to be the same co-occurrence pattern, this combination pattern is formally identified as a related event. The related event is then named and its characteristics described, reflecting the core code feature changes and development activity events in the combination. The characteristic description includes information such as co-occurrence frequency, time interval patterns, and the project modules involved.
[0051] Furthermore, causal inference analysis is performed on the identified related events to screen out the leading events that have statistical significance with the fluctuation of the quality defect index as quality influencing factors. The quality influencing factors refer to the related events or their development activities that have been statistically analyzed and confirmed to have a significant causal relationship with the fluctuation of the quality defect index.
[0052] Furthermore, causal inference techniques (such as Granger causality tests and causal graph-based algorithms) are used to analyze causal directions based on correlation. Propensity score matching is used to perform causal inference analysis on each associated event to determine whether development activities are the cause of code feature changes, such as analyzing whether urgent requirement changes lead to increased loop complexity. The correlation coefficient between each associated event and the fluctuation of the quality defect index is calculated, and hypothesis testing is performed to determine whether the relationship is statistically significant. Components of associated events with statistical significance (P-value ≤ 0.05) that are leading events in the causal relationship are selected as quality influencing factors. For example, the analysis found that an urgent deployment (cause) led to a decrease in code quality and an increase in TDI (effect), not the other way around. Events confirmed as causes (such as urgent deployments, module refactoring by a specific team, and knowledge gaps after key personnel leave) are identified as quality influencing factors.
[0053] Furthermore, historical defect data analysis revealed the evolution trend of the quality defect index and generated alarm information, but the root cause of this trend and the alarms was not clearly identified. Through cross-correlation analysis and causal inference, the core influencing factors leading to quality defects can be identified, explaining the fundamental reasons for the fluctuations in the quality defect index.
[0054] S5. Analyze the quality defects of the target software project based on the quality influencing factors and the quality alarm information.
[0055] In this embodiment of the invention, quality defects refer to the specific manifestations, causes, scope of impact, and development trends of quality defects in the target software project, which are comprehensively analyzed by combining quality influencing factors and quality alarm information.
[0056] In this embodiment of the invention, the step of analyzing the quality defects of the target software project based on the quality influencing factors and the quality alarm information includes: The depth and breadth of the traceability analysis of the quality influencing factors are determined based on the alarm level of the quality alarm information. Based on the depth and breadth of the traceability analysis, construct a correlation graph between the quality influencing factors and the dimensional scores of each quality dimension data in the multi-dimensional quality defect analysis model; Based on the correlation graph, locate the health decay dimension and health decay path of the target software project; The quality defects of the target software project are determined based on the health decay dimension and the health decay path.
[0057] In detail, alarm level refers to the preset alarm level based on the evolution trend and risk level of the quality defect index, including Level 1 alarm (high risk), Level 2 alarm (medium risk), and Level 3 alarm (low risk). The depth of tracing analysis refers to the hierarchical depth of tracing the causes of quality influencing factors; for example, Level 1 depth traces to the direct cause, Level 2 depth traces to the indirect cause, and Level 3 depth traces to the root cause. The breadth of tracing analysis refers to the range of quality dimensions covered by the tracing analysis; for example, Level 1 breadth covers all quality dimensions, Level 2 breadth covers core quality dimensions, and Level 3 breadth covers directly related quality dimensions. The alarm level for quality alarm information is clearly defined. This level is determined based on the evolution trend of the quality defect index. For example, if the quality defect index increases by 2% per month and exceeds the safety threshold, it is classified as a Level 1 alarm. The depth and breadth of the traceability analysis are determined according to the level-depth / breadth correspondence rules. For example, a Level 1 alarm corresponds to a Level 3 traceability depth (root cause) and a Level 1 traceability breadth (all quality dimensions); a Level 2 alarm corresponds to a Level 2 traceability depth (indirect cause) and a Level 2 traceability breadth (core quality dimensions); and a Level 3 alarm corresponds to a Level 1 traceability depth (direct cause) and a Level 3 traceability breadth (directly related quality dimensions). For instance, in a fintech scenario, a quality alarm in a bank's core business system is a Level 1 alarm, and the traceability analysis depth is determined to be Level 3, while the traceability analysis breadth is Level 1. In a healthcare scenario, a quality alarm in an intensive care unit is a Level 2 alarm, and the traceability analysis depth and breadth are determined to be Level 2.
[0058] Specifically, a correlation graph is a graphical tool that visually presents the relationships between quality influencing factors and the dimensional scores of each quality dimension. It includes nodes (quality influencing factors, quality dimensions, and dimensional scores) and edges (correlation relationships and influence strength). Based on the depth and breadth of the retrospective analysis, the quality influencing factors and quality dimensions participating in the construction of the correlation graph are determined. The influence strength of the correlation relationship is determined by calculating the correlation coefficient between the quality influencing factors and the dimensional scores of each quality dimension; the larger the absolute value of the correlation coefficient, the stronger the influence. Using graph visualization technology, the quality influencing factors are used as starting nodes, quality dimensions as intermediate nodes, and dimensional scores as target nodes. Edges of varying thicknesses represent influence strength, thus constructing the correlation graph. Analyze the changes in the dimension scores of each quality dimension in the correlation graph. If the dimension score of a certain quality dimension is lower than the preset health threshold and has a strong correlation with the quality influencing factors (absolute value of correlation coefficient ≥ 0.7), then the quality dimension is determined as the health decay dimension. Based on the node connection relationship in the correlation graph, organize the transmission process of quality influencing factors causing quality defects, which in turn leads to a decrease in the dimension score of the health decay dimension and an increase in the overall quality defect index, forming a health decay path.
[0059] Furthermore, the quality defect analysis content refers to the analysis report content formed after organizing the core information of quality defects, including health decay dimensions, decay degree, health decay path, quality influencing factors, and potential impact of defects. Based on the health decay dimensions and decay path, combined with quality influencing factors and quality alarm information, the key information of quality defects is comprehensively organized; the current score, health threshold, and decay range of each health decay dimension are clearly defined; each link of the health decay path is described in detail; the specific mechanism of action of quality influencing factors is analyzed; and the potential impact of defects on project functions, performance, security, etc., is assessed; this information is then structured and organized to form the quality defect analysis content.
[0060] Furthermore, by integrating the analysis results of the correlation graph with the governance cost estimation model, governance decision recommendations are generated, including specific governance actions, priority ranking, and expected resource consumption. The governance cost estimation model is based on historical data and industry benchmarks, used to estimate the workload (person-days) required to fix specific technical issues and the potential impact on business continuity (such as downtime). By comparing the fix benefits (expected TDI reduction and risk reduction) with the fix costs, a rough ROI can be calculated for each identified issue. Finally, a governance decision recommendation is generated, for example: prioritize refactoring the core class X of the payment module (estimated time 5 person-days, expected TDI reduction of 0.1, high ROI); secondly, introduce a mandatory code review checkpoint process for team A (estimated time 2 person-days for configuration, designed to prevent recurrence of similar issues).
[0061] As can be seen, the above solution significantly expands the scope of technical health assessment by establishing a quantitative evaluation model covering multiple dimensions such as code, architecture, testing, documentation, collaboration, and security, and integrating static code data with dynamic process data. This allows the assessment results to more comprehensively and systematically reflect the true health status of the software project. Utilizing time series analysis technology to decompose and predict historical quality defect indices enables early identification of risk points that could accelerate the deterioration of technical health, generating quality alerts and enhancing the initiative and foresight of project quality management. Interactive correlation analysis and causal inference of multi-source data allow tracing from surface technical symptoms to deeper root causes such as development processes and team collaboration, improving the effectiveness and return on investment of governance measures. By generating correlation graphs of quality influencing factors, governance decision-making recommendations for software project quality defects are generated, providing objective and quantitative data support for project managers' resource allocation decisions, thereby addressing the problem of low accuracy in software project quality defect analysis.
[0062] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0063] In one embodiment, a software project quality defect analysis apparatus is provided, which corresponds one-to-one with the software project quality defect analysis method described in the above embodiments. For example... Figure 5 As shown, the quality defect analysis device for this software project includes a data analysis module 101, a quality defect index analysis module 102, a quality alarm information generation module 103, a quality influencing factor identification module 104, and a quality defect analysis module 105. Detailed descriptions of each functional module are as follows: Data analysis module 101 is used to acquire multi-source heterogeneous project data of the target software project, analyze the code quality in the multi-source heterogeneous project data to obtain first analysis data, and analyze the project development characteristics in the multi-source heterogeneous project data to obtain second analysis data. The quality defect index analysis module 102 is used to analyze the quality score corresponding to each quality dimension of the target software project based on the first analysis data and the second analysis data using a pre-built multi-dimensional quality defect analysis model, and to integrate the quality score corresponding to each quality dimension into the quality defect index corresponding to the target software project. The quality alarm information generation module 103 is used to acquire historical defect data of the target software project, analyze the evolution trend of the quality defect index based on the historical defect data, and generate quality alarm information of the target software project based on the evolution trend. The quality influencing factor identification module 104 is used to perform interactive correlation analysis on the first analysis data and the second analysis data to obtain the correlation events in the multi-source heterogeneous project data, and identify the quality influencing factors corresponding to the target software project in the correlation events. The quality defect analysis module 105 is used to analyze the quality defects of the target software project based on the quality influencing factors and the quality alarm information.
[0064] In one embodiment, the data analysis module 101, when performing code quality analysis on the multi-source heterogeneous project data to obtain first analysis data, is used to: The set of original code metrics in the code repository of the multi-source heterogeneous project data is extracted based on the preset data interface specifications. The original code metric set is transformed according to a preset code quality feature mapping table to obtain a standardized code feature vector; Noise data points are detected in the standardized code feature vector, the noise data points are cleaned to form a target code feature vector, and the target code feature vector is used as the first analysis data.
[0065] In one embodiment, when the data analysis module 101 performs analysis on the project development characteristics in the multi-source heterogeneous project data to obtain second analysis data, it is further configured to: Analyze the version control logs and project management records in the multi-source heterogeneous project data; Extract the time sequence and metadata of development activities in the target software project based on the version control log and the project management record; Key pattern indicators of the development activities are identified based on the time series and the metadata. The key pattern indicators are correlated with the contextual information of the project phase in the development activity to generate second analysis data.
[0066] In one embodiment, the quality defect index analysis module 102, when executing a pre-built multi-dimensional quality defect analysis model to analyze the quality score corresponding to each quality dimension in the target software project based on the first analysis data and the second analysis data, is used to: Identify the quality dimension data in the first analysis data and the second analysis data; According to the preset mapping rules, the quality dimension data are respectively input to the input layer of each quality dimension of the multi-dimensional quality defect analysis model, and the dimension score of each quality dimension data is output. The adaptive weight configuration algorithm in the multi-dimensional quality defect analysis model is invoked based on the attribute tags of the target software project. The adaptive weight configuration algorithm is used to dynamically adjust the contribution weights of each quality dimension input layer to the quality dimension. The original quality score for each quality dimension in the target software project is obtained by weighting the contribution weight and the dimension score.
[0067] In one embodiment, the quality alarm information generation module 103, when performing the analysis of the evolution trend of the quality defect index based on the historical defect data, is used to: Identify the historical quality defect index in the historical defect data and extract the generation timestamp corresponding to the historical quality defect index; Time series data is constructed based on the historical quality defect index and the generated timestamp; The quality defect index is decomposed based on the time series data to obtain long-term trend components, periodic components, and residual components. The evolution trend of the quality defect index is determined based on the long-term trend component, the periodic component, and the residual component.
[0068] In one embodiment, the quality influencing factor identification module 104, when performing interactive correlation analysis on the first analysis data and the second analysis data to obtain the correlation events in the multi-source heterogeneous project data, is used to: Align and match the code feature change points in the first analysis data with the development activity events in the second analysis data on the timeline to obtain target matching data; In the target matching data, scan for code feature changes and development activity events with a frequency higher than a preset occurrence frequency; Determine whether the scanned code feature changes and development activity events constitute the same co-occurrence pattern; When the scanned code feature changes and development activity events form the same co-occurrence pattern, the co-occurrence pattern is identified as a related event in the multi-source heterogeneous project data.
[0069] In one embodiment, the quality defect analysis module 105, when performing the analysis of quality defects in the target software project based on the quality influencing factors and the quality alarm information, is used to: The depth and breadth of the traceability analysis of the quality influencing factors are determined based on the alarm level of the quality alarm information. Based on the depth and breadth of the traceability analysis, construct a correlation graph between the quality influencing factors and the dimensional scores of each quality dimension data in the multi-dimensional quality defect analysis model; Based on the correlation graph, locate the health decay dimension and health decay path of the target software project; The quality defects of the target software project are determined based on the health decay dimension and the health decay path.
[0070] This invention provides a software project quality defect analysis device. By establishing a quantitative evaluation model covering multiple dimensions such as code, architecture, testing, documentation, collaboration, and security, and integrating static code data and dynamic process data, it significantly expands the coverage of technical health assessment, enabling the evaluation results to more comprehensively and systematically reflect the true health status of the software project. Utilizing time series analysis technology to decompose and predict historical quality defect indices, it can identify risk points of accelerated deterioration in technical health in advance, generating quality alarm information and improving the initiative and foresight of project quality management. Interactive correlation analysis and causal inference of multi-source data can trace from surface technical symptoms to deeper root causes such as development processes and team collaboration, improving the effectiveness and return on investment of governance measures. Through the correlation graph of quality influencing factors, it generates governance decision-making suggestions for software project quality defects, providing objective and quantitative data support for project managers' resource allocation decisions, thereby solving the problem of low accuracy in software project quality defect analysis.
[0071] Specific limitations regarding the software project quality defect analysis device can be found in the limitations of the software project quality defect analysis method described above, and will not be repeated here. Each module in the aforementioned software project quality defect analysis device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0072] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a quality defect analysis method for a software project on the server side.
[0073] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 7As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the client-side functions or steps of a quality defect analysis method for a software project.
[0074] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain multi-source heterogeneous project data of the target software project, analyze the code quality in the multi-source heterogeneous project data to obtain first analysis data, and analyze the project development characteristics in the multi-source heterogeneous project data to obtain second analysis data; Using a pre-built multi-dimensional quality defect analysis model, the quality score corresponding to each quality dimension in the target software project is analyzed based on the first analysis data and the second analysis data. The quality score corresponding to each quality dimension is then integrated into the quality defect index corresponding to the target software project. Obtain historical defect data of the target software project, analyze the evolution trend of the quality defect index based on the historical defect data, and generate quality alarm information for the target software project based on the evolution trend. Interactive correlation analysis is performed on the first analysis data and the second analysis data to obtain the correlation events in the multi-source heterogeneous project data, and the quality influencing factors corresponding to the target software project in the correlation events are identified. Analyze the quality defects of the target software project based on the quality influencing factors and the quality alarm information.
[0075] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Obtain multi-source heterogeneous project data of the target software project, analyze the code quality in the multi-source heterogeneous project data to obtain first analysis data, and analyze the project development characteristics in the multi-source heterogeneous project data to obtain second analysis data; Using a pre-built multi-dimensional quality defect analysis model, the quality score corresponding to each quality dimension in the target software project is analyzed based on the first analysis data and the second analysis data. The quality score corresponding to each quality dimension is then integrated into the quality defect index corresponding to the target software project. Obtain historical defect data of the target software project, analyze the evolution trend of the quality defect index based on the historical defect data, and generate quality alarm information for the target software project based on the evolution trend. Interactive correlation analysis is performed on the first analysis data and the second analysis data to obtain the correlation events in the multi-source heterogeneous project data, and the quality influencing factors corresponding to the target software project in the correlation events are identified. Analyze the quality defects of the target software project based on the quality influencing factors and the quality alarm information.
[0076] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0077] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0078] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0079] It should be noted that if any software tools or components not belonging to our company appear in the embodiments of this application, they are merely for illustrative purposes and do not represent actual use.
[0080] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for analyzing quality defects in software projects, characterized in that, include: Obtain multi-source heterogeneous project data of the target software project, analyze the code quality in the multi-source heterogeneous project data to obtain first analysis data, and analyze the project development characteristics in the multi-source heterogeneous project data to obtain second analysis data; Using a pre-built multi-dimensional quality defect analysis model, the quality score corresponding to each quality dimension in the target software project is analyzed based on the first analysis data and the second analysis data. The quality score corresponding to each quality dimension is then integrated into the quality defect index corresponding to the target software project. Obtain historical defect data of the target software project, analyze the evolution trend of the quality defect index based on the historical defect data, and generate quality alarm information for the target software project based on the evolution trend. Interactive correlation analysis is performed on the first analysis data and the second analysis data to obtain the correlation events in the multi-source heterogeneous project data, and the quality influencing factors corresponding to the target software project in the correlation events are identified. Analyze the quality defects of the target software project based on the quality influencing factors and the quality alarm information.
2. The software project quality defect analysis method as described in claim 1, characterized in that, The analysis of code quality in the multi-source heterogeneous project data yields first analysis data, including: The set of original code metrics in the code repository of the multi-source heterogeneous project data is extracted based on the preset data interface specifications. The original code metric set is transformed according to a preset code quality feature mapping table to obtain a standardized code feature vector; Noise data points are detected in the standardized code feature vector, the noise data points are cleaned to form a target code feature vector, and the target code feature vector is used as the first analysis data.
3. The software project quality defect analysis method as described in claim 1, characterized in that, The analysis of project development characteristics in the multi-source heterogeneous project data yields second analysis data, including: Analyze the version control logs and project management records in the multi-source heterogeneous project data; Extract the time sequence and metadata of development activities in the target software project based on the version control log and the project management record; Key pattern indicators of the development activities are identified based on the time series and the metadata. The key pattern indicators are correlated with the contextual information of the project phase in the development activity to generate second analysis data.
4. The software project quality defect analysis method as described in claim 1, characterized in that, The step of using a pre-built multi-dimensional quality defect analysis model to analyze the quality score corresponding to each quality dimension in the target software project based on the first analysis data and the second analysis data includes: Identify the quality dimension data in the first analysis data and the second analysis data; According to the preset mapping rules, the quality dimension data are respectively input to the input layer of each quality dimension of the multi-dimensional quality defect analysis model, and the dimension score of each quality dimension data is output. The adaptive weight configuration algorithm in the multi-dimensional quality defect analysis model is invoked based on the attribute tags of the target software project. The adaptive weight configuration algorithm is used to dynamically adjust the contribution weights of each quality dimension input layer to the quality dimension. The original quality score for each quality dimension in the target software project is obtained by weighting the contribution weight and the dimension score.
5. The software project quality defect analysis method as described in claim 1, characterized in that, The step of analyzing the evolution trend of the quality defect index based on the historical defect data includes: Identify the historical quality defect index in the historical defect data and extract the generation timestamp corresponding to the historical quality defect index; Time series data is constructed based on the historical quality defect index and the generated timestamp; The quality defect index is decomposed based on the time series data to obtain long-term trend components, periodic components, and residual components. The evolution trend of the quality defect index is determined based on the long-term trend component, the periodic component, and the residual component.
6. The software project quality defect analysis method as described in claim 1, characterized in that, The step of performing interactive correlation analysis on the first analysis data and the second analysis data to obtain the correlation events in the multi-source heterogeneous project data includes: Align and match the code feature change points in the first analysis data with the development activity events in the second analysis data on the timeline to obtain target matching data; In the target matching data, scan for code feature changes and development activity events with a frequency higher than a preset occurrence frequency; Determine whether the scanned code feature changes and development activity events constitute the same co-occurrence pattern; When the scanned code feature changes and development activity events form the same co-occurrence pattern, the co-occurrence pattern is identified as a related event in the multi-source heterogeneous project data.
7. The software project quality defect analysis method as described in claim 1, characterized in that, The analysis of quality defects in the target software project based on the quality influencing factors and the quality alarm information includes: The depth and breadth of the traceability analysis of the quality influencing factors are determined based on the alarm level of the quality alarm information. Based on the depth and breadth of the traceability analysis, construct a correlation graph between the quality influencing factors and the dimensional scores of each quality dimension data in the multi-dimensional quality defect analysis model; Based on the correlation graph, locate the health decay dimension and health decay path of the target software project; The quality defects of the target software project are determined based on the health decay dimension and the health decay path.
8. A software project quality defect analysis device, characterized in that, include: The data analysis module is used to acquire multi-source heterogeneous project data of the target software project, analyze the code quality in the multi-source heterogeneous project data to obtain first analysis data, and analyze the project development characteristics in the multi-source heterogeneous project data to obtain second analysis data. The quality defect index analysis module is used to analyze the quality score corresponding to each quality dimension of the target software project using a pre-built multi-dimensional quality defect analysis model for the first analysis data and the second analysis data, and to merge the quality score corresponding to each quality dimension into the quality defect index corresponding to the target software project. The quality alarm information generation module is used to acquire historical defect data of the target software project, analyze the evolution trend of the quality defect index based on the historical defect data, and generate quality alarm information of the target software project based on the evolution trend. The quality influencing factor identification module is used to perform interactive correlation analysis on the first analysis data and the second analysis data to obtain the correlation events in the multi-source heterogeneous project data, and identify the quality influencing factors corresponding to the target software project in the correlation events. The quality defect analysis module is used to analyze the quality defects of the target software project based on the quality influencing factors and the quality alarm information.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for analyzing quality defects in a software project as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for analyzing quality defects in a software project as described in any one of claims 1 to 7.