An AI-based method and system for evaluating the information-based operation of software.
By employing an AI-based information-based evaluation method for software operation, combined with multi-dimensional and adaptive analysis, the limitations of existing software quality evaluation methods, such as their singularity and one-sidedness, have been addressed. This approach enables comprehensive and accurate evaluation of software in complex environments, thereby enhancing the scientific rigor and effectiveness of the evaluation.
Patent Information
- Application Number
- CN202411449818.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-17
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-10-17
AI Technical Summary
Existing software quality assessment methods focus on single attributes or local functions, neglecting the software's adaptability in complex environments and its interaction performance with other software, resulting in partial and one-sided assessment results.
An artificial intelligence-based approach is adopted to determine the rigidity and flexibility quality standards of the target software, combine multiple random evaluation scenarios, obtain the homography matrix and evaluate the index intervals, combine the evaluation of adaptive adjustment capability and risk resistance capability, conduct independent verification and mutual verification, and finally calculate the mean to generate the target evaluation result.
It enables comprehensive, real-time, and objective evaluation of software operation, ensuring the completeness and accuracy of evaluation results, while taking into account the software's adaptability and resilience in changing environments, thus improving the scientific nature and effectiveness of software quality management.
Smart Images

Figure CN119645824B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital data processing technology, specifically to an information-based evaluation method and system for software operation based on artificial intelligence. Background Technology
[0002] Significant progress has been made in artificial intelligence technology, particularly in areas such as machine learning, natural language processing, and computer vision, which has driven the digital transformation of various industries and provided powerful tools and technical support for software quality assessment.
[0003] With the widespread application of software across various industries, especially its increasingly prominent role in critical business areas, the demand for evaluating software performance quality is becoming increasingly stringent. Evaluations not only need to cover multiple dimensions such as functionality, performance, security, and compatibility, but also emphasize in-depth examination of the software's adaptability, stability, and collaborative effectiveness in actual operating environments. Existing software performance evaluation methods focus on single attributes or partial functions of the software, lacking a comprehensive assessment of software operation, and particularly neglecting the software's adaptability in complex environments and its interaction performance with other software.
[0004] In summary, existing technologies suffer from the technical problems of relying on a single indicator for software quality assessment and providing partial and one-sided results for information-based evaluation of software operation. Summary of the Invention
[0005] This application provides an artificial intelligence-based method and system for information-based evaluation of software operation, aiming to solve the technical problems of existing software quality evaluation standards having a single indicator and the results of information-based evaluation of software operation being partial and one-sided.
[0006] In view of the above problems, the technical solution to achieve the present application is as follows:
[0007] This application provides an information-based evaluation method for software operation based on artificial intelligence. The method includes: determining the quality standards of the target software, wherein the quality standards include rigid quality standards and flexible quality standards, and the quality measurement dimensions include single-item evaluation dimensions and interactive evaluation dimensions.
[0008] Determine the evaluation scenarios based on the target software and determine the valid scenario operation data, wherein the evaluation scenarios are at least two random scenarios;
[0009] Connect to the software database, and based on the quality standard, perform an evaluation analysis based on the evaluation scenario in the intelligent evaluation module, including obtaining the homography matrix and evaluating the index interval based on the scenario interval, to determine the first evaluation result;
[0010] Based on the software database, the target software is evaluated for its adaptive adjustment capability and risk resistance capability, and a second evaluation result is determined. The second evaluation result is marked with an elastic relaxation coefficient.
[0011] The first evaluation result and the second evaluation result are verified, including independent verification and mutual verification, and the verification result is determined.
[0012] If the verification result is qualified, the average of the first evaluation result and the second evaluation result is calculated and used as the target evaluation result.
[0013] In another aspect, this application provides an information-based evaluation system for software operation based on artificial intelligence, wherein the system includes: a standard determination module for determining the quality standards of the target software, wherein the quality standards include rigid quality standards and flexible quality standards, and the quality measurement dimensions include single-item evaluation dimensions and interactive evaluation dimensions;
[0014] The running data determination module is used to determine the evaluation scenario based on the target software and determine the valid scenario running data, wherein the evaluation scenario is at least two random scenarios;
[0015] The evaluation and analysis module is used to connect to the software database and, based on the quality standard, perform evaluation and analysis based on the evaluation scenario in the intelligent evaluation module, including obtaining the homography matrix and evaluating the index intervals based on the scenario intervals, to determine the first evaluation result.
[0016] The capability assessment module is used to combine the software database to assess the target software's adaptive adjustment capability and risk resistance capability, and determine the second assessment result, which is marked with an elastic relaxation coefficient.
[0017] The verification module is used to verify the first evaluation result and the second evaluation result, including independent verification and mutual verification, and to determine the verification result.
[0018] The mean calculation module is used to calculate the mean of the first evaluation result and the second evaluation result if the verification result is qualified, and use it as the target evaluation result.
[0019] In summary, one or more technical solutions provided in this application address the technical problems of single indicators in software quality assessment standards and partial and one-sided results in information-based assessment of software operation. They enable the use of effective scenario operation data to avoid assessment bias caused by the limitations of test samples. Based on real-time and objective software assessment, through in-depth mining and analysis of software operation data, a flexible relaxation coefficient identifier is determined to ensure the software's adaptability and flexibility under changing environments and demand conditions. This balances the software's basic operational performance with its ability to self-optimize and resist potential threats in constantly changing environments. Based on a comprehensive analysis of multiple dimensions, adaptability, and scale randomness, the technical effect ensures the completeness of the assessment and the accuracy of the results. Attached Figure Description
[0020] Figure 1 This application provides a flowchart illustrating an information-based evaluation method for software operation based on artificial intelligence.
[0021] Figure 2 This application provides a flowchart illustrating the process of verifying the qualification of software operation information evaluation results based on artificial intelligence;
[0022] Figure 3 This application provides a schematic diagram of the structure of an information-based evaluation system for software operation based on artificial intelligence.
[0023] Explanation of reference numerals in the attached diagram: Standard determination module M100, Operational data determination module M200, Evaluation and analysis module M300, Capability assessment module M400, Verification module M500, Mean calculation module M600. Detailed Implementation
[0024] Example 1
[0025] The present application will now be described in detail with reference to the accompanying drawings, such as... Figure 1 As shown, this application provides an artificial intelligence-based information evaluation method for software operation, wherein the method includes:
[0026] S1: Determine the quality standards for the target software, wherein the quality standards include rigid quality standards and flexible quality standards, and the quality measurement dimensions include individual evaluation dimensions and interactive evaluation dimensions;
[0027] S2: Determine the evaluation scenarios based on the target software and determine the valid scenario operation data, wherein the evaluation scenarios are at least two random scenarios;
[0028] Existing software operation information assessment methods focus on a single attribute or partial function of the software. Furthermore, based on preset and limited test scenarios, they are difficult to simulate the complexity and diversity of the real world, resulting in assessment results that cannot accurately reflect the software's performance in actual operation. In addition, they are insufficient in identifying potential software risks and assessing its risk resistance capabilities, and only take remedial measures after problems occur, rather than providing early warnings and proactive responses.
[0029] Based on this, this application constructs a comprehensive, real-time, accurate, and automated information-based evaluation system for software operation. Verification has shown that this application takes into account both the basic operational performance of software and its ability to self-optimize and resist potential threats in a constantly changing environment, effectively addressing the complex evaluation needs in modern software development and operation environments, and improving the scientificity and effectiveness of software quality management.
[0030] Specifically, the quality standards include rigid quality standards, which include: functionality (the target software meets all predetermined functional requirements, including core functions, auxiliary functions, and interface compatibility with external systems); reliability (evaluating the target software's ability to perform its expected functions under specified conditions and within a specified time, including indicators such as mean time to repair (MTTR), failure rate, etc.); performance efficiency (measuring the target software's performance in terms of request processing, response time, and resource utilization (CPU, memory, disk, network); security (auditing the target software's compliance and protection capabilities in terms of data protection, access control, encryption mechanisms, and vulnerability prevention); and compatibility (checking the target software's adaptability to different operating systems, hardware configurations, browser versions, third-party components, etc.).
[0031] The quality standards include flexible quality standards, which include: maintainability (assessing the clarity of the target software architecture, code readability, modularity, documentation completeness, and ease of error diagnosis and repair), scalability (examining the target software's ability to easily add new features, support more users, and handle larger loads when facing future changes in requirements), flexibility (measuring the target software's adaptability to changes in business rules, user interface, data format, etc., and the richness of configuration options), user-friendliness (evaluating the target software's ease of use, user experience (UX / UI design), and the effectiveness of help documentation and training materials), and portability (assessing the ease of migrating the target software between different environments or platforms, including the conversion costs of code, data, and configuration files).
[0032] The quality measurement dimensions include individual evaluation dimensions, which include: functional testing (designing and executing test cases to verify the correctness and completeness of the target software's various functions), performance testing (simulating high concurrency, large data volume, and other stress scenarios to measure response time and resource consumption), stability testing (running the target software for a long time to monitor for issues such as memory leaks and resource exhaustion, as well as system recovery capabilities), security testing (using penetration testing, vulnerability scanning, and other methods to check for security vulnerabilities in the target software and verify the effectiveness of access control, data encryption, and other mechanisms), and compatibility testing (running the target software under various operating systems, browsers, and device combinations to ensure normal cross-platform operation).
[0033] The quality measurement dimensions include the interaction evaluation dimension, which refers to the operational response status of the target software when it collaborates with other software. The interaction evaluation dimension includes: identification of collaborative software (identifying other software systems that closely collaborate with the target software, such as database management systems, middleware, third-party APIs, etc.), interaction feature analysis (analyzing key interaction characteristics such as interface protocols, data exchange formats, message passing mechanisms, and synchronous / asynchronous processing modes between the target software and the multiple software connected to it), and interaction performance indicators (developing evaluation indicators for collaborative operations, such as interface response time, data transmission rate, message loss rate, synchronization latency, etc.).
[0034] The evaluation scenarios for the target software are determined, and valid scenario operation data is collected. Specifically, at least two random evaluation scenarios are created based on the actual application scenarios of the target software, covering various situations such as normal operation, peak load, fault recovery, and boundary conditions. Corresponding input data, configuration parameters, and user behavior scripts are prepared for each scenario to ensure the representativeness and validity of the data. The software is run in a controlled environment to execute predefined scenario operations, while recording system logs, performance indicators, user feedback, and other operation data. Using monitoring tools, log analysis tools, user surveys, and other methods, the operation data under each scenario is systematically collected, and the collected operation data under each scenario is recorded as valid scenario operation data to ensure the comprehensiveness of the data. The above operations are used to evaluate the target software using rigid and flexible quality standards, covering single evaluation dimensions and interactive evaluation dimensions, and a comprehensive and accurate evaluation result is obtained based on the valid scenario operation data.
[0035] S3: Connect to the software database, and based on the quality standard, perform an evaluation analysis based on the evaluation scenario in the intelligent evaluation module, including obtaining the homography matrix and evaluating the index interval based on the scenario interval, and determine the first evaluation result;
[0036] S4: Combine the software database to evaluate the target software’s adaptive adjustment capability and risk resistance capability, and determine the second evaluation result, which is marked with an elastic relaxation coefficient.
[0037] Connect to the software database. Specifically, obtain the necessary information such as the database access address, port, username, password, and database name, and configure the database connection parameters. Use a suitable database driver (such as JDBC, ODBC, ADO.NET, etc.) to establish a stable connection with the software database based on the configuration information. Confirm that the connected user has the appropriate permissions to query the required data tables, views, stored procedures, and other resources to ensure that data access is unrestricted during the evaluation process.
[0038] The established rigid and flexible quality standards are imported into the intelligent evaluation module as reference benchmarks for the evaluation algorithm. The designed evaluation scenario information, including scenario description, expected results, and relevant data sources, is input into the intelligent evaluation module. Based on the evaluation requirements, parameters such as the working mode, data sampling frequency, and result output format of the intelligent evaluation module are set. Furthermore, performance data, log records, and user behavior data related to the evaluation scenario are extracted from the software database. Statistical analysis is performed on the data from different evaluation scenarios to calculate the homography matrix between key indicators, reflecting the mapping relationship and changing patterns between indicators. Based on the homography matrix, the raw data is converted into interval forms (such as quantiles, interquartile ranges, standardized scores, etc.) to facilitate cross-scenario comparison and trend analysis. In actual operation, machine learning, data mining, and other technologies are applied, combined with quality standards, to conduct in-depth analysis of the intervalized indicators, identifying strengths, weaknesses, outliers, and areas for improvement. The homography matrix analysis results are summarized to form preliminary evaluation conclusions, including comparative analysis between scenarios, key indicator performance, and overall quality score, serving as the first evaluation result.
[0039] Extract data related to software runtime state adjustments from the software database, such as configuration change records, automatic optimization logs, and performance tuning events. Analyze this data to extract evidence of how the software self-adjusts resource allocation, algorithm parameters, and caching strategies under varying loads, user behaviors, and system environments. Further, using clustering, regression, and time series analysis, construct a mathematical model describing the software's adaptive adjustment capabilities, quantifying its response speed, adjustment magnitude, and effect persistence. Compare the model's predictions with actual results to evaluate the strength, stability, and effectiveness of the software's adaptive adjustment capabilities in various scenarios, and assign a score or rating.
[0040] Data related to security incidents, fault records, and anomaly detection is extracted from the software database to identify potential risks and existing risk events. The impact of each risk event on software operation is assessed, considering factors such as downtime, data loss, and decreased user satisfaction. The defensive, recovery, and compensatory measures taken by the software after a risk event, and their corresponding effectiveness, are examined. Furthermore, by integrating multiple dimensions such as risk event frequency, impact, and the effectiveness of countermeasures, a mathematical model is used to calculate the software's elasticity relaxation coefficient, reflecting its risk resistance capability. The assessment results of adaptive adjustment capability and risk resistance capability are integrated to form a second assessment result, which is marked with the elasticity relaxation coefficient. Through these steps, homography matrix acquisition, index interval evaluation, and assessment of operational adaptive adjustment capability and risk resistance capability are performed, generating a second assessment result with an elasticity relaxation coefficient to support subsequent analysis.
[0041] S5: Verify the first evaluation result and the second evaluation result, including independent verification and mutual verification, and determine the verification result;
[0042] S6: If the verification result is qualified, the average of the first evaluation result and the second evaluation result is calculated and used as the target evaluation result.
[0043] The first assessment results should be independently verified. Specifically, the first assessment results should cover all predetermined assessment dimensions and indicators without omissions. The reasonableness and credibility of the assessment results should be verified by comparing them with historical data, industry standards, or theoretical expectations. The consistency of the homography matrix and the interval assessment results should be analyzed. For example, the comparison results between scenarios should be consistent with the overall assessment conclusion.
[0044] The second assessment results were independently verified. Specifically, it was confirmed that the data used to assess the adaptive adjustment capability and risk resistance capability originated from the software database and that the collection process was error-free; the calculation process of the elastic relaxation coefficient and the data used were reviewed to ensure that the calculation process was correct; and the results of assessing the adaptive adjustment capability and risk resistance capability should be logically consistent with the software characteristics and the usage environment.
[0045] The first and second evaluation results are cross-verified. Specifically, cross-comparison verification is performed: the scenario performance in the first evaluation result is compared with the adaptive adjustment and risk mitigation performance of the software in the corresponding scenario in the second evaluation result to determine whether there is an inherent relationship between the two; if the weights of each part in the evaluation system are known, the weight allocation of the first and second evaluation results is checked to see if it conforms to the preset rules and whether the weight calculation is accurate; deviation fluctuation range verification is performed: based on the evaluation standard or experience value, the maximum allowable deviation range between the first and second evaluation results is set; the numerical differences between the first and second evaluation results on each common evaluation dimension are calculated; the calculated difference measure is compared with the allowable deviation range to confirm that the differences in all dimensions are within the allowable range.
[0046] Summarize the problems, anomalies, or confirmed correctness information found in independent and mutual verification; based on the severity and number of problems, determine whether the first and second assessment results meet the verification qualification requirements. If all inspection items pass and there are no major problems, it is considered qualified; then, according to the assessment system regulations or expert opinions, assign weights to the first and second assessment results (e.g., 50% each or other proportions); multiply the first and second assessment results by their corresponding weights, then sum them to obtain the weighted average, which is the target assessment result;
[0047] In practice, this also includes recording in detail the process, methods, problems found and solutions of independent verification and mutual verification in the evaluation report; writing the calculated target evaluation results and their calculation process into the evaluation report; adjusting the evaluation conclusions based on the target evaluation results; and proposing targeted improvement suggestions or next action plans.
[0048] Through the above steps, the first and second evaluation results are independently and mutually verified to ensure data accuracy. When the results are qualified, calculations are performed to obtain the target evaluation result and update the evaluation report. This achieves a comprehensive analysis of diversified dimensions, adaptability, and scale randomness, ensuring the completeness and accuracy of the evaluation.
[0049] Furthermore, after obtaining the target assessment results, the method of this application includes:
[0050] A homology search is performed on the target software to determine the software class set;
[0051] Perform inter-class functional ranking and overall calculation of the target software and the class software set to determine the class evaluation result, wherein the overall calculation principle is a weighted calculation based on the primary and secondary functions;
[0052] Based on the target evaluation results and the class evaluation results, the software evaluation results are determined.
[0053] The criteria for defining source software corresponding to the target software are determined, including development language, core algorithm, application field, user group, and similar functional modules. Detailed search keywords, Boolean operators, and filtering conditions are formulated using resources such as professional databases, open-source project platforms, industry reports, academic papers, and market analysis reports. Search engines and professional software indexing tools are used to find a set of software with source characteristics similar to the target software, based on the search strategy. The search results are initially screened, eliminating software that clearly does not meet the source criteria, and retaining a list of suspected source software. Additionally, if necessary, software documentation, source code, or trials are conducted to verify their source similarity. The final list of source software is compiled and confirmed, forming a software class set.
[0054] For each software in the target software and its category sets, list its main functional modules or features to form a detailed function list. For each function, assign it a primary or secondary weight based on factors such as user needs, industry standards, and technological development trends. Weights can be determined using methods such as expert scoring, questionnaires, and literature reviews. Within the category sets, conduct horizontal comparisons of identical or similar functions and rank them according to their importance weight within their respective software. The ranking of the same function across different software should reflect its relative value within that software ecosystem. For each software in the category sets, calculate a weighted total based on the primary and secondary functions according to its function ranking.
[0055] Based on the target evaluation results, the scores and analysis conclusions of each indicator are summarized to form an overall evaluation overview of the target software. For each software in the software category set, a score is calculated based on its overall performance to form a category evaluation overview, including average score, highest score, lowest score, functional advantages and disadvantages, etc. The target software evaluation overview is compared with the category evaluation overview to analyze the target software's position, advantages and disadvantages, innovation points, and unique value in the same software category. Based on the comparative analysis results, combined with the target evaluation results and category evaluation results, a comprehensive evaluation of the target software is conducted, including the target software's performance relative to other software in the software category set and the target software's applicability in specific application scenarios. Through the above steps, the target software is subjected to source retrieval, inter-category functional ranking, and overall performance calculation. After comparative analysis, a comprehensive software evaluation result is obtained, providing a scientific basis for software optimization and competitive strategies.
[0056] Furthermore, the quality measurement dimensions include an interaction evaluation dimension, and the method of this application includes:
[0057] The collaborative software of the target software is identified, and the collaborative characteristics are determined based on the software interaction functions;
[0058] The first weight distribution is based on the dominance of interaction, and the second weight distribution is based on the relevance of interaction, where the relative dominance based on software interaction is indicated.
[0059] Determine the multivariate evaluation indicators for the interaction evaluation dimension, and perform hierarchical weighting based on the first weight distribution and the second weight distribution.
[0060] By analyzing the target software's functional descriptions, interface documents, configuration files, or actual usage, identify other software that works directly or indirectly with the target software. Collaborating software includes API service providers, database management systems, third-party plugins, hardware drivers, client applications, etc. For each collaborative software, list in detail its interaction interfaces with the target software, including API interfaces, message queues, database query statements, file transfer protocols, etc., and clarify key information such as interaction methods, data formats, and access permissions.
[0061] Based on the interaction interface, the collaborative characteristics between the target software and each collaborative software are summarized, including: interaction frequency (the number of times data is exchanged or requests and responses are made between the software), data dependency (whether one software depends on the data provided by the other software to function properly), synchronous and asynchronous mode (whether real-time response is required during the interaction process, or whether asynchronous processing is allowed), fault tolerance mechanism (how the target software ensures service continuity when collaborative software fails or is delayed), and security requirements (whether security measures such as data encryption, authentication, and access control are complete).
[0062] The first weight distribution is based on interaction dominance. Specifically, interaction dominance refers to the influence or control of the target software over other software in collaborative work. This can be manifested in aspects such as the direction of data flow, initiative in task scheduling, and allocation of responsibility for error handling. Furthermore, for each collaborative software, its degree of dominance in the interaction process is evaluated based on its interaction relationship with the target software. The evaluation can use quantitative scoring (e.g., 1 to 5 points). Then, based on the evaluation results, a first weight value is assigned to each collaborative software. Generally, the stronger the dominance of the software, the larger the corresponding first weight should be, indicating that it has a greater impact on the interaction quality of the target software.
[0063] The second weight distribution is based on the degree of interaction correlation. Specifically, the degree of interaction correlation measures the degree of functional dependence and data coupling between the target software and the collaborative software, thereby reflecting the tightness of the interaction between the software and the risk of fault propagation. For each collaborative software, the criticality of its function in the business process of the target software, the depth and breadth of data exchange, and the scope of impact caused by faults are analyzed to evaluate its degree of interaction correlation with the target software. Furthermore, based on the evaluation results, a second weight value is assigned to each collaborative software. Generally, the higher the degree of correlation, the larger the corresponding second weight should be, indicating that it has a greater impact on the overall function and stability of the target software.
[0064] Select specific metrics suitable for evaluating software interaction performance, such as response time, data consistency, concurrent processing capability, fault recovery speed, interface compatibility, and user experience. Based on the first and second weight distributions, hierarchically assign weights to the multi-dimensional evaluation metrics. Further, rank all collaborative software according to their first weight to form a dominant weight sequence. For each software, adjust its position in the sequence according to its second weight to form a comprehensive weight sequence. Then, assign corresponding weights to the multi-dimensional evaluation metrics based on the position of the corresponding software in the comprehensive weight sequence, giving greater attention to the evaluation metrics corresponding to software with higher weights. Through these steps, the logical decomposition and implementation process of the target software interaction evaluation dimensions are refined, ensuring the comprehensiveness, accuracy, and operability of the interaction performance evaluation.
[0065] Furthermore, the intelligent evaluation module is communicatively connected to the software database and includes a homography processing unit and an adaptive evaluation unit. The adaptive evaluation unit includes a static evaluation branch and a dynamic evaluation branch, and performs evaluation analysis based on the evaluation scenario within the intelligent evaluation module. The method of this application includes:
[0066] Based on the software database, extract scenario operation information based on the evaluation scenario;
[0067] The scene operation information is traversed, and the indicator feature values based on the quality standard are extracted and mutually mapped to determine the homography matrix. The homography matrix is a mapping between at least two scene matrices.
[0068] Based on the homography matrix, an index interval evaluation is performed to determine the first evaluation result.
[0069] Ensure that the communication interface between the intelligent evaluation module and the software database is standardized and can transmit bidirectionally, including API calls, database connections, message queue subscriptions, and other methods; clarify the data structure, field meanings, and encoding rules of the scenario operation information stored in the software database so that the intelligent evaluation module can correctly parse and process it; furthermore, configure the necessary access permissions to ensure that the intelligent evaluation module can securely access the required data, while preventing data leakage or malicious tampering.
[0070] Based on the requirements, at least two different software operation scenarios are randomly selected, such as high-concurrency access, network environment fluctuations, data anomaly handling, and specific function usage. Furthermore, the intelligent evaluation module sends a query request to the software database, specifying the evaluation scenario to be extracted and its corresponding operation time period. The software database returns the corresponding scenario operation information according to the request, and performs data cleaning on the obtained raw scenario operation information to remove duplicate or abnormal data and perform necessary data transformations (such as timestamp conversion, unit standardization, etc.) to ensure the accuracy of subsequent analysis.
[0071] Based on pre-defined quality standards (including rigid and flexible quality standards), relevant performance indicator feature values are extracted from scenario operation information. For example, rigid quality standards may include system response time, resource utilization, and error rate; flexible quality standards may involve user experience, adaptability, and maintainability. Furthermore, for each scenario, the extracted indicator feature values are organized into a scenario feature matrix. By calculating the similarity, correlation, or transformation relationship between different scenario matrices, a homography matrix is constructed. The homography matrix reflects the correspondence and change patterns of software performance indicators under different evaluation scenarios.
[0072] Based on the eigenvalue distribution and trends revealed by the homography matrix, each performance indicator is reasonably divided into intervals. The intervals can be determined based on statistical methods (such as quartiles, standard deviation, etc.). Furthermore, for each evaluation scenario, based on its mapping relationship in the homography matrix, the indicator eigenvalues in that scenario are mapped to the corresponding indicator intervals. By comparing the proportion or average value of the same indicator falling into different intervals in different scenarios, the adaptability of the software in various scenarios is evaluated. Then, the adaptability evaluation results of each scenario are integrated to form the first evaluation result. The first evaluation result includes the adaptability score of each scenario, the overall adaptability level, and the scenario sensitivity analysis of key performance indicators.
[0073] Furthermore, in practice, the randomness of the scenario and the random distribution of characteristic states help simulate a wider range of operating conditions, thereby enhancing the generalization ability and adaptability of the evaluation results. Specifically, probabilistic models or Monte Carlo simulations are used to generate various representative random evaluation scenarios that conform to actual application environments. Moreover, when dividing the index intervals, the actual distribution characteristics of the index feature values (such as normal distribution, skewed distribution, etc.) are considered. Based on the distribution characteristics, an appropriate interval division method is selected to ensure that the intervals, while covering most observations, reflect extreme cases of the data.
[0074] Based on the scenario operation information in the software database, homography processing and adaptive evaluation techniques are used to conduct a comprehensive and detailed analysis of the adaptability of the target software in random evaluation scenarios, generating a first evaluation result with high fitness. Thus, under the condition of random scenario and random distribution of corresponding feature states, interval analysis is performed based on random distribution to improve the fitness of the result.
[0075] Furthermore, the method of this application for performing interval-based evaluation of indicators to determine the first evaluation result includes:
[0076] Identify the homography matrix, and based on the static evaluation branch, perform a scenario differentiation metric to determine the static evaluation result;
[0077] Identify the homography matrix, and based on the dynamic evaluation branch, perform dynamic trend analysis based on the scene feature interval and the index feature value interval to determine the dynamic evaluation result of scene evolution;
[0078] Based on the dynamic evaluation results and the static evaluation results, the first evaluation result is determined.
[0079] The constructed homography matrix is interpreted in detail, and the inter-mapping relationship of various performance indicators under different evaluation scenarios and their transformation rules in different scenarios are analyzed. The important features in the homography matrix that reflect significant differences in software performance and affect adaptability are identified, such as strong correlation between indicators, outliers in specific scenarios, and inflection points of indicator changes.
[0080] Based on the distribution characteristics of each indicator in the homography matrix and the preset evaluation criteria, several intervals are divided for each performance indicator, such as excellent, good, average, poor, etc. For each evaluation scenario, the proportion, average or median of the characteristic values of each indicator in that scenario falling into the corresponding interval are calculated to form a scenario-differentiated feature vector.
[0081] Comparative analysis of the differentiated feature vectors of each scenario can be performed using methods such as clustering, distance metrics, and ranking to quantify the degree of difference between scenarios and obtain static evaluation results. Static evaluation results may include scenario classification, inter-scenario difference scores, and key scenario identifiers. Furthermore, based on the feature relationships between scenarios revealed by the homography matrix, scenario intervals with similar characteristics can be defined. These intervals may be determined based on factors such as the similarity of performance indicators and the continuity of scenario changes.
[0082] Trend analysis is performed on the characteristic values of each indicator within each scenario's characteristic interval, including linear regression, time series analysis, and autoregressive moving average models, to reveal the dynamic trends of the indicators as the scenario changes. This generates dynamic trend descriptions (such as growth, decline, fluctuation range, periodicity, etc.), key turning points, and future predictions for the indicators within each scenario's characteristic interval, forming a dynamic evaluation result. The dynamic evaluation result may include scenario evolution paths, trend charts of key indicators, and predictive models.
[0083] By organically integrating static and dynamic evaluation results, static evaluation results provide an immediate comparison of differences between scenarios, while dynamic evaluation results reveal the evolutionary trend of scenarios over time or under changing conditions. Combining the two allows for a comprehensive understanding of the software's adaptability and its development in different scenarios. Furthermore, based on the scenario difference scores in the static evaluation results and the trend analysis in the dynamic evaluation results, an adaptability score is assigned to each scenario. The adaptability score can be a sub-score of multiple sub-dimensions (such as stability, response speed, resource utilization efficiency, etc.). Then, the adaptability scores of all scenarios are summarized, and combined with the importance weights of the scenarios (such as business criticality, frequency of occurrence, etc.), the overall adaptability score of the target software is calculated as the first evaluation result.
[0084] Through the above steps, based on the homography matrix and combining static and dynamic evaluation branches, an in-depth analysis of the adaptability of the target software in different scenarios is conducted, generating a first evaluation result that reflects both the current state and predicts future trends, providing a scientific basis for software performance optimization, scenario adaptability improvement, and risk warning.
[0085] Furthermore, the method for determining the second evaluation result in this application includes:
[0086] Read the extreme values of operation regulation and risk countermeasure, where the extreme value range is determined based on the upper and lower limits of the capability;
[0087] Based on the software database, a predetermined number of risk operation records are randomly extracted. The second evaluation result is determined by using the extreme values of operation control and risk resistance as a baseline. The elastic relaxation coefficient includes a coefficient plural based on diversified elasticity indicators, with response efficiency and response accuracy as evaluation elements of elasticity indicators.
[0088] Based on the software system's design specifications, industry standards, past experience, or performance test results, set clear upper and lower limits for the software's operational control and risk mitigation capabilities. These represent the target software's best performance under ideal conditions and its minimum tolerance under extreme pressure. Furthermore, for each capability indicator (such as response time, throughput, error rate, etc.), calculate its extreme value range based on the set upper and lower limits. The extreme value range represents the reasonable range of performance that the software should maintain during normal operation. Exceeding this range indicates the existence of performance bottlenecks or potential risks.
[0089] Define clear criteria for risk operation records, including situations such as abnormal event triggering, performance indicators approaching or exceeding extreme ranges, and the occurrence of specific failure modes. Determine the quantity and method for randomly selecting risk operation records to ensure sample representativeness. Consider factors such as time span, event type, and system load to guarantee sample diversity and balance. Furthermore, utilize the software database query interface to randomly select a specified number of risk operation records from the database according to preset screening conditions and sampling strategies. Records should include key information such as timestamps, event descriptions, and relevant performance indicator data.
[0090] Each risk operation record is checked one by one to verify the actual values of each performance indicator. The values are compared with the corresponding extreme value ranges to determine whether the actual values exceed the set upper and lower limits. For each record, two elastic indicators, response efficiency and response accuracy, are calculated. Response efficiency involves the speed of request processing and resource utilization, while response accuracy is related to the correctness of the system's request processing and the absence of false alarms or missed alarms.
[0091] Based on the calculated response efficiency and accuracy, a comprehensive evaluation is performed using a multivariate set of coefficients for various elasticity indices. This multivariate set includes the weights and adjustment factors for each elasticity index, used to measure the relative importance of each index in the overall elasticity assessment and its impact on system performance. Through weighted summation or other composite operations, the elasticity relaxation coefficient for each record is obtained.
[0092] Based on the elastic relaxation coefficient of each risk operation record, they are divided into different risk levels (such as low risk, medium risk, and high risk). The classification criteria should match the business risk tolerance, compliance requirements, and emergency plans. Furthermore, the records of different risk levels are statistically analyzed to determine the frequency, duration, and scope of impact of risk events, and to clarify the overall performance in responding to risks.
[0093] Then, taking into account the risk level distribution, risk event statistics, and the system's response efficiency and accuracy during risk mitigation, a second assessment result is formed, including a risk overview summary, analysis of major risk factors, distribution of elastic relaxation coefficients, and improvement suggestions. Through the above steps, the target software's operational control capabilities and risk mitigation capabilities in the face of risk events are determined, i.e., the second assessment result, which provides strong support for the software system's risk prevention and control, emergency response mechanism optimization, and performance bottleneck identification.
[0094] Furthermore, such as Figure 2 As shown, the method of this application includes independent verification and mutual verification, and includes:
[0095] Interactive indicator threshold standards are used to perform attribution verification on the first evaluation result and the second evaluation result to determine the independent verification result.
[0096] The first evaluation result and the second evaluation result are mapped and checked to determine whether the deviation fluctuation range is met and to determine the mutual verification result.
[0097] If the independent verification result is qualified, and the mutual verification result is qualified, then the verification result is qualified.
[0098] The process, which includes independent and mutual verification, aims to validate the rationality and consistency of the first and second evaluation results. According to industry standards, threshold standards are set for each interaction evaluation indicator. These threshold standards reflect the ideal range of various interaction performance indicators under normal operating conditions. Furthermore, all evaluation data related to interaction performance, including response time, interaction success rate, and user experience score, are extracted from the first evaluation results.
[0099] The interaction performance data in the first evaluation result is compared with the set threshold standard. Furthermore, if all data points fall within their respective threshold ranges, the evaluation result is considered to meet the interaction index threshold standard, the attribution verification is passed, and the independent verification result is recorded as qualified; otherwise, the index that does not meet the standard and its deviation value are recorded, and the independent verification result is unqualified.
[0100] Analyze the related indicators in the first and second assessment results to clarify the logical relationship or causal effect between them. For example, response efficiency (an indicator in the first assessment result) directly affects risk resistance capability (an indicator in the second assessment result). In this way, establish the mapping relationship between various indicators and set a reasonable deviation fluctuation range for the relevant indicators after mapping. The deviation fluctuation range should allow a certain degree of natural fluctuation or measurement error, while being able to identify significant inconsistencies or anomalies. The size of the range can be determined based on the statistical characteristics of historical data, business sensitivity, and assessment accuracy requirements.
[0101] The relevant indicator data in the first evaluation result are mapped to the indicator system corresponding to the second evaluation result. The difference between the mapped data and the actual second evaluation result is compared. Furthermore, if the difference between all mapped data points and the second evaluation result is within the set deviation fluctuation range, the mutual verification result is deemed qualified; otherwise, the indicators that exceed the fluctuation range and their deviation values are recorded, and the mutual verification result is deemed unqualified.
[0102] Collect independent and mutual verification results, record their respective pass / fail status and potential problem indicators. Furthermore, if both independent and mutual verification results are pass, the overall assessment result is pass, indicating that there is no significant contradiction or abnormality between the first and second assessment results, and the assessment process is reliable. If any verification result is fail, the overall assessment result is fail, and further investigation is needed to determine the cause. In this case, some or all of the assessment work needs to be redone, or the assessment methods and parameters need to be adjusted.
[0103] In practice, for cases where verification fails, a thorough analysis should be conducted to determine the specific reasons for the failure. These reasons could include limitations in the evaluation method, biases in data collection, or unreasonable threshold settings. Based on the analysis results, improvement measures should be proposed, such as optimizing the evaluation model, revising the data collection process, and adjusting the threshold standards. Furthermore, regardless of whether the verification result is satisfactory, a detailed verification report should be prepared, clearly describing the verification process, results, and conclusions, providing improvement suggestions for non-compliance cases, and communicating with the project team to support subsequent decision-making. Rigorous independent and cross-verification of the first and second evaluation results should be conducted to ensure the rationality and consistency of the evaluation data, providing a reliable basis for the continuous improvement and optimization of the software system.
[0104] In summary, the beneficial effects of the embodiments of this application are:
[0105] 1. By introducing diversified dimensions, adaptability of methods, and randomness of scale, a comprehensive and multi-level evaluation of software operation is achieved, covering rigid and flexible quality standards, single and interactive evaluation dimensions, ensuring the completeness and accuracy of the evaluation.
[0106] 2. By comparing the assessment scenario with the software database for intelligent analysis, the assessment is made close to the real operating environment, making full use of big data resources and improving the relevance and authenticity of the assessment.
[0107] 3. By obtaining homography matrix and evaluating index intervals, and combining static and dynamic evaluation branches, the system specifically assesses the software's adaptive adjustment capability and risk resistance capability. It introduces concepts such as elastic relaxation coefficient, strengthens the evaluation's consideration of the software's ability to cope with complex environmental changes and risk events, realizes dynamic tracking and trend analysis of the software's operating status, and improves the timeliness and predictability of the evaluation.
[0108] 4. Due to the adoption of interactive indicator threshold standards, the first and second evaluation results are subjected to attribution verification to determine independent verification results; the first and second evaluation results are then mapped and checked to determine whether they meet the deviation fluctuation range, thus determining the mutual verification results; if both the independent and mutual verification results are qualified, the overall verification result is qualified. Rigorous independent and mutual verification of the first and second evaluation results ensures the rationality and consistency of the evaluation data, providing a reliable basis for the continuous improvement and optimization of the software system.
[0109] Example 2
[0110] Based on the same inventive concept as the artificial intelligence-based software operation information evaluation method in the foregoing embodiments, such as Figure 3 As shown in the figure, this application provides an artificial intelligence-based software operation information evaluation system, wherein the system includes:
[0111] The standard determination module M100 is used to determine the quality standards of the target software. The quality standards include rigid quality standards and flexible quality standards, and the quality measurement dimensions include individual evaluation dimensions and interactive evaluation dimensions.
[0112] The runtime data determination module M200 is used to determine the evaluation scenario based on the target software and determine the runtime data of the valid scenario, wherein the evaluation scenario is at least two random scenarios;
[0113] The evaluation and analysis module M300 is used to connect to the software database and, based on the quality standard, perform evaluation and analysis based on the evaluation scenario in the intelligent evaluation module, including obtaining the homography matrix and evaluating the index interval based on the scenario interval, to determine the first evaluation result.
[0114] The capability assessment module M400 is used to combine the software database to assess the target software's adaptive adjustment capability and risk resistance capability, and determine the second assessment result, which is marked with an elastic relaxation coefficient.
[0115] The verification module M500 is used to verify the first evaluation result and the second evaluation result, including independent verification and mutual verification, and to determine the verification result.
[0116] The mean calculation module M600 is used to calculate the mean of the first evaluation result and the second evaluation result if the verification result is qualified, and use it as the target evaluation result.
[0117] Furthermore, the mean calculation module M600 is also used to perform the following method:
[0118] A homology search is performed on the target software to determine the software class set;
[0119] Perform inter-class functional ranking and overall calculation of the target software and the class software set to determine the class evaluation result, wherein the overall calculation principle is a weighted calculation based on the primary and secondary functions;
[0120] Based on the target evaluation results and the class evaluation results, the software evaluation results are determined.
[0121] Furthermore, the standard determination module M100 is used to perform the following method:
[0122] The collaborative software of the target software is identified, and the collaborative characteristics are determined based on the software interaction functions;
[0123] The first weight distribution is based on the dominance of interaction, and the second weight distribution is based on the relevance of interaction, where the relative dominance based on software interaction is indicated.
[0124] Determine the multivariate evaluation indicators for the interaction evaluation dimension, and perform hierarchical weighting based on the first weight distribution and the second weight distribution.
[0125] Furthermore, the intelligent evaluation module is communicatively connected to the software database and includes a homography processing unit and an adaptive evaluation unit. The adaptive evaluation unit includes a static evaluation branch and a dynamic evaluation branch, and performs evaluation analysis based on the evaluation scenario within the intelligent evaluation module. The AI-based software operation information evaluation system is also used to execute the following methods:
[0126] Based on the software database, extract scenario operation information based on the evaluation scenario;
[0127] The scene operation information is traversed, and the indicator feature values based on the quality standard are extracted and mutually mapped to determine the homography matrix. The homography matrix is a mapping between at least two scene matrices.
[0128] Based on the homography matrix, an index interval evaluation is performed to determine the first evaluation result.
[0129] Furthermore, in the process of performing interval-based evaluation of indicators to determine the first evaluation result, the AI-based software operation information evaluation system is also used to execute the following method:
[0130] Identify the homography matrix, and based on the static evaluation branch, perform a scenario differentiation metric to determine the static evaluation result;
[0131] Identify the homography matrix, and based on the dynamic evaluation branch, perform dynamic trend analysis based on the scene feature interval and the index feature value interval to determine the dynamic evaluation result of scene evolution;
[0132] Based on the dynamic evaluation results and the static evaluation results, the first evaluation result is determined.
[0133] Furthermore, the capability assessment module M400 is used to perform the following method:
[0134] Read the extreme values of operation regulation and risk countermeasure, where the extreme value range is determined based on the upper and lower limits of the capability;
[0135] Based on the software database, a predetermined number of risk operation records are randomly extracted. The second evaluation result is determined by using the extreme values of operation control and risk resistance as a baseline. The elastic relaxation coefficient includes a coefficient plural based on diversified elasticity indicators, with response efficiency and response accuracy as evaluation elements of elasticity indicators.
[0136] Furthermore, the verification module M500 is used to perform the following method:
[0137] Interactive indicator threshold standards are used to perform attribution verification on the first evaluation result and the second evaluation result to determine the independent verification result.
[0138] The first evaluation result and the second evaluation result are mapped and checked to determine whether the deviation fluctuation range is met and to determine the mutual verification result.
[0139] If the independent verification result is qualified, and the mutual verification result is qualified, then the verification result is qualified.
[0140] In summary, any step can be stored as a computer instruction or program in an unrestricted computer memory and can be called and recognized by an unrestricted computer processor; no further restrictions are imposed here.
[0141] Furthermore, the above technical solutions only embody the preferred technical solutions of the embodiments of this application. Any changes that those skilled in the art may make to certain parts of these solutions embody the novel principles of the embodiments of this application. Obviously, those skilled in the art can make various modifications and variations to this application without departing from the scope of this application.
Claims
1. A software operation information evaluation method based on artificial intelligence, characterized in that, The method includes: The quality standards for the target software are determined, including rigid quality standards and flexible quality standards, and the quality measurement dimensions include individual evaluation dimensions and interactive evaluation dimensions. Determine the evaluation scenarios based on the target software and determine the valid scenario operation data, wherein the evaluation scenarios are at least two random scenarios; Connect to the software database, and based on the quality standard, perform an evaluation analysis based on the evaluation scenario in the intelligent evaluation module, including obtaining the homography matrix and evaluating the index interval based on the scenario interval, to determine the first evaluation result; Based on the software database, the target software is evaluated for its adaptive adjustment capability and risk resistance capability, and a second evaluation result is determined. The second evaluation result is marked with an elastic relaxation coefficient. The first evaluation result and the second evaluation result are verified, including independent verification and mutual verification, and the verification result is determined. If the verification result is qualified, the average of the first evaluation result and the second evaluation result is calculated and used as the target evaluation result. The quality measurement dimensions include interaction evaluation dimensions, including: The collaborative software of the target software is identified, and the collaborative characteristics are determined based on the software interaction functions; The first weight distribution is based on the dominance of interaction, and the second weight distribution is based on the relevance of interaction, where the relative dominance based on software interaction is indicated. Determine the multivariate evaluation indicators for the interaction evaluation dimension, and perform hierarchical weighting based on the first weight distribution and the second weight distribution; The intelligent evaluation module is communicatively connected to the software database and includes a homography processing unit and an adaptive evaluation unit. The adaptive evaluation unit includes a static evaluation branch and a dynamic evaluation branch, and performs evaluation analysis based on the evaluation scenario within the intelligent evaluation module, including: Based on the software database, extract scenario operation information based on the evaluation scenario; The scene operation information is traversed, and the indicator feature values based on the quality standard are extracted and mutually mapped to determine the homography matrix. The homography matrix is a mapping between at least two scene matrices. Based on the homography matrix, an index interval evaluation is performed to determine the first evaluation result; The step of performing interval-based evaluation of indicators to determine the first evaluation result includes: Identify the homography matrix, and based on the static evaluation branch, perform a scenario differentiation metric to determine the static evaluation result; Identify the homography matrix, and based on the dynamic evaluation branch, perform dynamic trend analysis based on the scene feature interval and the index feature value interval to determine the dynamic evaluation result of scene evolution; Based on the dynamic evaluation results and the static evaluation results, the first evaluation result is determined; The determination of the second evaluation result includes: Read the extreme values of operation regulation and risk countermeasure, where the extreme value range is determined based on the upper and lower limits of the capability; Based on the software database, a predetermined number of risk operation records are randomly extracted. The second evaluation result is determined by using the extreme values of operation control and risk resistance as a baseline. The elastic relaxation coefficient includes a coefficient plural based on diversified elasticity indicators, with response efficiency and response accuracy as evaluation elements of elasticity indicators.
2. The method as described in claim 1, characterized in that, After obtaining the target assessment results, the following are included: A homology search is performed on the target software to determine the software class set; Perform inter-class functional ranking and overall calculation of the target software and the class software set to determine the class evaluation result, wherein the overall calculation principle is a weighted calculation based on the primary and secondary functions; Based on the target evaluation results and the class evaluation results, the software evaluation results are determined.
3. The method as described in claim 1, characterized in that, The method includes independent verification and mutual verification, including: Interactive indicator threshold standards are used to perform attribution verification on the first evaluation result and the second evaluation result to determine the independent verification result. The first evaluation result and the second evaluation result are mapped and checked to determine whether the deviation fluctuation range is met and to determine the mutual verification result. If the independent verification result is qualified, and the mutual verification result is qualified, then the verification result is qualified.
4. An information-based evaluation system for software operation based on artificial intelligence, characterized in that, The system is used to implement the artificial intelligence-based software operation information evaluation method according to any one of claims 1-3, the system comprising: The standard determination module is used to determine the quality standards of the target software, wherein the quality standards include rigid quality standards and flexible quality standards, and the quality measurement dimensions include individual evaluation dimensions and interactive evaluation dimensions. The running data determination module is used to determine the evaluation scenario based on the target software and determine the valid scenario running data, wherein the evaluation scenario is at least two random scenarios; The evaluation and analysis module is used to connect to the software database and, based on the quality standard, perform evaluation and analysis based on the evaluation scenario in the intelligent evaluation module, including obtaining the homography matrix and evaluating the index intervals based on the scenario intervals, to determine the first evaluation result. The capability assessment module is used to combine the software database to assess the target software's adaptive adjustment capability and risk resistance capability, and determine the second assessment result, which is marked with an elastic relaxation coefficient. The verification module is used to verify the first evaluation result and the second evaluation result, including independent verification and mutual verification, and to determine the verification result. The mean calculation module is used to calculate the mean of the first evaluation result and the second evaluation result if the verification result is qualified, and use it as the target evaluation result.
Citation Information
Patent Citations
Artificial intelligence multimode imaging analysis device
CN111128382A
Software evaluation method and system
CN115495348A