DevSecOps software quality entrance guard admission analysis method and system based on POMDP model
Through the DevSecOps software quality access control analysis method based on the POMDP model, the traditional method has solved the shortcomings in adaptability, real-timeness and resource efficiency, and achieved intelligent optimization of verification strategies, reducing resource consumption and defect leakage rates, which are suitable for high concurrency verification under the microservice architecture.
Patent Information
- Application Number
- CN202510616198.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-15
AI Technical Summary
The traditional DevSecOps software quality access control analysis method has shortcomings in adaptability, real-time and resource efficiency, and it is impossible to achieve real-time matching of verification strength and business needs in a dynamic environment.
The DevSecOps software quality access control access analysis method based on the POMDP model is adopted to generate the optimal strategy through multi-source data acquisition, state space construction, iterative computing update and dynamic decision-making mechanisms to achieve intelligent optimization of verification strategies.
It reduces verification resource consumption, reduces defect leakage rate, and shortens decision response time. It is suitable for high concurrency verification scenarios under the microservice architecture.
Smart Images

Figure CN120492308A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of software quality assurance, continuous delivery, and security verification, and in particular to a DevSecOps software quality access control analysis method and system based on a POMDP model. Background Art
[0002] With the popularization of DevSecOps concepts, the software delivery process requires real-time integration of quality and security verification. Traditional rule-based access control has the following core technical bottlenecks:
[0003] First, the verification strategy is static; second, it is difficult to integrate multi-dimensional indicators; third, there is a contradiction between real-time performance and resource consumption; fourth, there is insufficient adaptation to environmental dynamics; and fifth, the security verification depth is insufficient.
[0004] In response to the above technical bottlenecks, the existing technology has tried various ways to make improvements. The first is dynamic threshold adjustment. This method considers calculating the threshold range based on historical data, but there are response delays and cold start problems. Since it is essentially a threshold relaxation in a statistical sense, it cannot solve the problem of changes in the correlation between indicators in a dynamic environment. The second is the machine learning model. This method mainly uses algorithms such as random forest to integrate multiple types of indicators, but the model complexity is high (the average number of parameters reaches billions) and it is difficult to deploy lightweight. In particular, the overfitting risk and computing resource consumption in high-dimensional feature space limit its application in edge nodes. The third is elastic resource scheduling. This method dynamically allocates computing resources based on queue length, but does not solve the coordinated optimization of verification strategy and resource consumption. The decoupling of resource allocation and verification strength results in the overall efficiency of the system cannot be optimized. The fundamental defect of the above-mentioned improvement method is mainly the lack of a unified mathematical model to describe the verification decision-making process, and it is impossible to achieve real-time matching of verification strength and business needs in a dynamic environment.
[0005] Therefore, how to invent a DevSecOps software quality access control analysis method to address the shortcomings of traditional methods in adaptability, real-time performance, and resource efficiency has become an urgent problem to be solved. Summary of the Invention
[0006] To this end, this paper provides a DevSecOps software quality access control analysis method and system based on the POMDP model. By building a dynamic decision-making model to intelligently optimize verification strategies, this method addresses the shortcomings of traditional methods in terms of adaptability, real-time performance, and resource efficiency. This method can reduce verification resource consumption, lower defect leakage rates, and shorten decision response time, making it suitable for high-concurrency verification scenarios within microservices architectures.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a DevSecOps software quality access control analysis method based on the POMDP model, comprising:
[0008] Real-time data collection from multiple sources, including code quality metrics, security scan results, test coverage, and environmental health during the DevSecOps software development process, is performed through CI / CD pipeline interfaces, security scanning tool APIs, and cloud-native monitoring systems to obtain observation vectors for software quality.
[0009] Based on the observation vector, a state space of software quality is constructed; and a multi-dimensional description of the software delivery state is performed through the state space;
[0010] The state space is iteratively calculated and updated by using an exponentially weighted moving average algorithm and a particle filter algorithm to obtain a posterior probability distribution of software quality state transitions during the R&D process;
[0011] By defining the quality access control action space, a decision action set is obtained; the actions in the decision action set are associated with verification resource consumption parameters and decision delay parameters;
[0012] Based on the dynamic decision-making mechanism and the posterior probability distribution, the Bellman equation based on the POMDP model is solved to obtain the optimal strategy;
[0013] generating an admission decision according to the optimal strategy; mobilizing corresponding actions in the decision action set according to the admission decision;
[0014] After completing the corresponding actions, the parameters of the POMDP model and the reward function of the Bellman equation are adjusted and optimized.
[0015] As the preferred solution for DevSecOps software quality access control analysis based on the POMDP model, it implements three-level verification through the edge layer, control layer, and policy layer;
[0016] The edge layer deploys a lightweight inference engine on CI / CD nodes to achieve a decision response time of <200ms;
[0017] The control layer dynamically adjusts the scale of the sandbox environment based on Kubernetes scheduling and verification resources;
[0018] The strategy layer synchronizes global strategies through federated learning and supports 50+ model updates per week.
[0019] As a preferred solution for the DevSecOps software quality access control analysis method based on the POMDP model, the observation vectors include: code quality indicators, security scanning results, test coverage and environmental health.
[0020] As a preferred solution of the DevSecOps software quality access control analysis method based on the POMDP model, the expression of the state space is:
[0021] S={Q,E,P,T}
[0022] In the formula, Q is the quality confidence, which can be divided into three levels: high, medium, and low; E is the environmental stability, which is divided into three states: stable, fluctuating, and abnormal; P is the verification priority, which is divided into three levels: emergency, routine, and delayed; T is the verification stage, including rapid verification, deep verification, and manual review.
[0023] As a preferred solution of the DevSecOps software quality access control analysis method based on the POMDP model, the observation confidence is calculated by the exponentially weighted moving average algorithm; the calculation formula of the observation confidence is:
[0024] μ t =λ·o t +(1-λ)·μ t-1
[0025] Where μ t is the observation confidence; λ is the weighting coefficient, ranging from 0 to 1; o t is the observed value.
[0026] As a preferred solution of the DevSecOps software quality access control analysis method based on the POMDP model, the expression of the decision action set is:
[0027] A={A1,A2,A3,A4}
[0028] Where A1 is immediate release; A2 is enhanced verification; A3 is delayed admission; and A4 is denied admission.
[0029] As a preferred solution of the DevSecOps software quality access control analysis method based on the POMDP model, the expression of the Bellman equation is:
[0030] V * (s)=max a∈A [R(s,a)+γΣ s' P(s'|s,a)V * (s')]
[0031] Where V * (s) is the optimal value function under state s; a is an action in the action space; R(s,a) is the reward function; s is the current state; s' is the next state; P(s'|s,a) is the state transition probability matrix; γ is the discount factor.
[0032] The present invention also provides a DevSecOps software quality access control and admission analysis system based on the POMDP model, which includes:
[0033] A multi-source data collection module is used to collect multi-source data in real time through CI / CD pipeline interfaces, security scanning tool APIs, and cloud-native monitoring systems. The multi-source data includes code quality indicators, security scanning results, test coverage, and environmental health during the DevSecOps software development process, thereby obtaining observation vectors for software quality.
[0034] A state space construction module is used to construct a state space of software quality based on the observation vector; and to perform a multi-dimensional description of the software delivery state through the state space;
[0035] A posterior probability distribution acquisition module is used to iteratively calculate and update the state space through an exponentially weighted moving average algorithm and a particle filter algorithm to obtain a posterior probability distribution of software quality state transitions during the R&D process;
[0036] A decision action set definition module is used to obtain a decision action set by defining a quality access control action space; the actions in the decision action set are associated with verification resource consumption parameters and decision delay parameters;
[0037] An optimal strategy acquisition module, configured to solve the Bellman equation based on the POMDP model based on the dynamic decision-making mechanism and the posterior probability distribution to obtain the optimal strategy;
[0038] An admission decision execution module, configured to generate an admission decision according to the optimal strategy; and to mobilize corresponding actions in the decision action set according to the admission decision;
[0039] The parameter adjustment and optimization module is used to adjust and optimize the parameters of the POMDP model and the reward function of the Bellman equation after completing the corresponding action.
[0040] As the preferred solution for the DevSecOps software quality access control and analysis system based on the POMDP model, it implements three-level verification through the edge layer, control layer, and policy layer;
[0041] The edge layer deploys a lightweight inference engine on CI / CD nodes to achieve a decision response time of <200ms;
[0042] The control layer dynamically adjusts the scale of the sandbox environment based on Kubernetes scheduling and verification resources;
[0043] The strategy layer synchronizes global strategies through federated learning and supports 50+ model updates per week.
[0044] As a preferred solution for the DevSecOps software quality access control analysis system based on the POMDP model, in the multi-source data acquisition module, the observation vectors include: code quality indicators, security scanning results, test coverage and environmental health.
[0045] As a preferred solution of the DevSecOps software quality access control and admission analysis system based on the POMDP model, in the state space construction module, the expression of the state space is:
[0046] S={Q,E,P,T}
[0047] In the formula, Q is the quality confidence, which can be divided into three levels: high, medium, and low; E is the environmental stability, which is divided into three states: stable, fluctuating, and abnormal; P is the verification priority, which is divided into three levels: emergency, routine, and delayed; T is the verification stage, including rapid verification, deep verification, and manual review.
[0048] As a preferred solution of the DevSecOps software quality access control analysis system based on the POMDP model, in the posterior probability distribution acquisition module, the observation confidence is calculated by the exponentially weighted moving average algorithm; the calculation formula of the observation confidence is:
[0049] μ t =λ·o t +(1-λ)·μ t-1
[0050] Where μ t is the observation confidence; λ is the weighting coefficient, ranging from 0 to 1; o t is the observed value.
[0051] As a preferred solution of the DevSecOps software quality access control and admission analysis system based on the POMDP model, in the decision action set definition module, the expression of the decision action set is:
[0052] A={A1,A2,A3,A4}
[0053] Where A1 is immediate release; A2 is enhanced verification; A3 is delayed admission; and A4 is denied admission.
[0054] As a preferred solution of the DevSecOps software quality access control analysis system based on the POMDP model, in the optimal strategy acquisition module, the expression of the Bellman equation is:
[0055] V *(s)=max a∈A [R(s,a)+γΣ s' P(s'|s,a)V * (s')]
[0056] Where V * (s) is the optimal value function under state s; a is an action in the action space; R(s,a) is the reward function; s is the current state; s' is the next state; P(s'|s,a) is the state transition probability matrix; γ is the discount factor.
[0057] The present invention has the following advantages: the present invention collects multi-source data such as code quality indicators, security scanning results, test coverage and environmental health in the DevSecOps software development process in real time through the CI / CD pipeline interface, security scanning tool API and cloud native monitoring system to obtain observation vectors of software quality; constructs a state space of software quality based on the observation vectors; uses the state space to describe the software delivery status in multiple dimensions; it iteratively calculates and updates the state space through the exponentially weighted moving average algorithm and the particle filter algorithm to obtain the posterior probability distribution of the software quality state transition in the development process; obtains a decision action set by defining the quality access action space; the actions in the decision action set are associated with verification resource consumption parameters and decision delay parameters; solves the Bellman equation based on the POMDP model based on the dynamic decision mechanism and the posterior probability distribution to obtain the optimal strategy; generates an access decision according to the optimal strategy; mobilizes the corresponding actions in the decision action set according to the access decision; and after completing the corresponding actions, adjusts and optimizes the parameters of the POMDP model and the reward function of the Bellman equation. This paper intelligently optimizes verification strategies by building a dynamic decision-making model, addressing the shortcomings of traditional approaches in adaptability, real-time performance, and resource efficiency. This paper can reduce verification resource consumption, lower defect leakage rates, and shorten decision response times, making it suitable for high-concurrency verification scenarios within a microservices architecture. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can, without inventive effort, derive other implementation drawings based on the provided drawings.
[0059] The structures, proportions, sizes, etc. illustrated in this specification are intended solely to complement the contents disclosed herein and to facilitate understanding and reading by persons skilled in the art. They are not intended to limit the conditions under which the present invention may be implemented and therefore have no substantive technical significance. Any structural modifications, changes in proportions, or adjustments in sizes, without affecting the efficacy and objectives of the present invention, shall remain within the scope of the technical contents disclosed herein.
[0060] Figure 1 This is a flow chart of the DevSecOps software quality access control analysis method based on the POMDP model provided in Example 1 of the present invention;
[0061] Figure 2 This is a schematic diagram of a specific implementation process of the DevSecOps software quality access control analysis method based on the POMDP model provided in Example 1 of the present invention;
[0062] Figure 3 This is a schematic diagram of the state transition of the POMDP model in the DevSecOps software quality access control analysis method based on the POMDP model provided in Example 1 of the present invention;
[0063] Figure 4 This is a schematic diagram of the verification architecture in the DevSecOps software quality access control analysis method based on the POMDP model provided in Example 1 of the present invention;
[0064] Figure 5 This is a schematic diagram of verification decision in the DevSecOps software quality access control analysis method based on the POMDP model provided in Example 1 of the present invention;
[0065] Figure 6 Schematic diagram of the reward function optimization process in the DevSecOps software quality access control analysis method based on the POMDP model provided in Example 1 of the present invention;
[0066] Figure 7 This is a schematic diagram of the DevSecOps software quality access control analysis system architecture based on the POMDP model provided in Example 2 of the present invention. DETAILED DESCRIPTION
[0067] The following describes the implementation of the present invention using specific embodiments. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. Obviously, the embodiments described are only a portion of the present invention, not all of it. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.
[0068] Example 1
[0069] See also Figure 1 and Figure 2 Embodiment 1 of the present invention provides a DevSecOps software quality access control analysis method based on the POMDP model, comprising the following steps:
[0070] S1. Real-time collection of multi-source data, including code quality metrics, security scan results, test coverage, and environmental health during the DevSecOps software development process, through CI / CD pipeline interfaces, security scanning tool APIs, and cloud-native monitoring systems, to obtain observation vectors of software quality.
[0071] S2. Constructing a state space of software quality based on the observation vector; and performing a multi-dimensional description of the software delivery status through the state space;
[0072] S3. Iteratively calculating and updating the state space using an exponentially weighted moving average algorithm and a particle filter algorithm to obtain a posterior probability distribution of software quality state transitions during the R&D process;
[0073] S4. Obtain a decision action set by defining a quality access control action space; the actions in the decision action set are associated with verification resource consumption parameters and decision delay parameters;
[0074] S5. Solve the Bellman equation based on the POMDP model based on the dynamic decision-making mechanism and the posterior probability distribution to obtain the optimal strategy;
[0075] S6. Generate an admission decision according to the optimal strategy; mobilize corresponding actions in the decision action set according to the admission decision;
[0076] S7. After completing the corresponding action, the parameters of the POMDP model and the reward function of the Bellman equation are adjusted and optimized.
[0077] In this embodiment, in step S1, multi-source data such as code quality indicators, security scanning results, test coverage, and environmental health in the DevSecOps software development process are collected in real time through the CI / CD pipeline interface, security scanning tool API, and cloud native monitoring system to obtain observation vectors of software quality;
[0078] Specifically, the observation vector O is acquired through multi-source data collection, including code quality metrics (SonarQube scores), security scan results (vulnerability levels), test coverage (Jacoco metrics), and environmental health (Kubernetes monitoring data). The SonarQube score, part of the code quality metric, is a comprehensive assessment of code quality, taking into account factors such as code complexity, duplication, and potential defects. A high SonarQube score indicates good code quality and a low likelihood of issues; a low score indicates a high likelihood of issues and requires improvement. The vulnerability level in the security scan results assesses the severity of security vulnerabilities discovered in the software. These levels are generally categorized as high, medium, and low. High-level vulnerabilities pose a serious threat to software security and require immediate remediation. Medium-level vulnerabilities require attention and remediation at the appropriate time. Low-level vulnerabilities, while less impactful on security, also warrant attention and action. The Jacoco metric in the test coverage reflects the extent to which the code is covered by test cases. High test coverage means the code has been thoroughly tested and is less likely to contain defects. Low test coverage indicates that there may be untested code sections, requiring additional test cases to improve coverage. Kubernetes monitoring data in the environmental health section primarily includes resource usage and network connectivity of Kubernetes cluster nodes. By analyzing this data, we can understand the stability and availability of the environment and determine its suitability for software verification and deployment.
[0079] In this embodiment, in step S2, a state space of software quality is constructed based on the observation vector; a multi-dimensional description of the software delivery state is performed through the state space;
[0080] The expression of the state space is:
[0081] S={Q,E,P,T}
[0082] In the formula, Q is the quality confidence, which can be divided into three levels: high, medium, and low; E is the environmental stability, which is divided into three states: stable, fluctuating, and abnormal; P is the verification priority, which is divided into three levels: emergency, routine, and delayed; T is the verification stage, including rapid verification, deep verification, and manual review.
[0083] Specifically, the assessment of quality confidence takes into account multiple factors, with code quality being a key factor, measured through metrics such as code complexity and code duplication. High code complexity means the code is difficult to maintain and understand, increasing the risk of defects. High code duplication reduces maintainability and makes it more likely that errors will be missed during modifications. Test coverage is also a key metric for assessing quality confidence. High test coverage indicates that the code has been thoroughly tested, making the likelihood of defects relatively low. Furthermore, the results of code reviews, such as the number and severity of issues discovered during the review, must be considered.
[0084] The assessment of environmental stability is primarily based on the operation of the infrastructure and service layers. At the infrastructure layer, pay attention to the resource usage of the Kubernetes cluster nodes, such as CPU utilization and memory utilization. If the usage of these resources remains stable for a long time and the fluctuation range is small, the environment can be considered to be stable. If there are large fluctuations in resource usage, such as a sudden sharp increase or decrease in CPU utilization, it indicates that the environment is fluctuating. If serious problems such as node failures and network interruptions occur, the environment can be determined to be abnormal. At the service layer, observe indicators such as the response time and throughput of microservice instances. If these indicators are stable and meet expectations, the environment is stable. If the indicators fluctuate significantly or the service is unavailable, the environment may be fluctuating or abnormal.
[0085] Verification priorities require a comprehensive consideration of both business needs and project progress. For projects with a significant business impact and tight deadlines, such as those involving updates to critical business functions or fixing security vulnerabilities, the verification priority should be set to Urgent. For general feature updates or optimizations, the priority can be set to Regular. For projects with a smaller business impact and less urgent needs, such as minor adjustments to interface styles, the priority can be set to Delayed.
[0086] Quick verification primarily performs basic checks, such as code formatting and basic syntax errors, with the goal of quickly identifying obvious issues. Deep verification provides a more comprehensive and detailed check, including code logic correctness, performance testing, and security vulnerability scanning. Manual review involves further review and judgment when automated verification fails to confirm the results or uncovers complex issues.
[0087] In this embodiment, in step S3, the state space is iteratively calculated and updated by using an exponentially weighted moving average algorithm and a particle filter algorithm to obtain a posterior probability distribution of software quality state transitions during the development process;
[0088] Specifically, the observation confidence is calculated by the exponentially weighted moving average algorithm; the calculation formula of the observation confidence is:
[0089] μ t =λ·o t +(1-λ)·μ t-1
[0090] Where μ t is the observation confidence; λ is the weighting coefficient, ranging from 0 to 1; o t is the observed value.
[0091] The state posterior probability distribution is updated based on the particle filter algorithm. The exponentially weighted moving average (EWMA) method can smooth the observed data and highlight the influence of recent observations.
[0092] Here, λ is a weighting coefficient, ranging from 0 to 1. A larger λ value places more emphasis on recent observations, allowing for faster reflection of changes in observed data; a smaller λ value places more emphasis on historical observations, resulting in smoother changes in observation confidence. By adjusting the value of λ, the influence of recent and historical observations can be flexibly balanced according to different application scenarios and requirements.
[0093] In this embodiment, the particle filter algorithm is an effective method for estimating system states. It uses a set of random samples (particles) to represent the probability distribution of the state. During the state update process, the particles are weighted and resampled based on the observed data and the state transition model to update the posterior probability distribution of the state. The particle filter algorithm can handle nonlinear and non-Gaussian systems and has good real-time and adaptability, effectively coping with dynamic environmental changes.
[0094] In this embodiment, in step S4, a decision action set is obtained by defining the quality access control action space; the actions in the decision action set are associated with verification resource consumption parameters and decision delay parameters;
[0095] The expression of the decision action set is:
[0096] A={A1,A2,A3,A4}
[0097] Where A1 is immediate release; A2 is enhanced verification; A3 is delayed admission; and A4 is denied admission.
[0098] Specifically, each action is associated with a verification resource consumption parameter C and a decision delay parameter D. The immediate release action means the software has passed the current verification phase and can proceed directly to the next phase. This action consumes the least verification resources and has the shortest decision delay because no additional verification operations are required. The enhanced verification action increases the depth and breadth of existing verification. For example, this may include adding more test cases or performing more rigorous security vulnerability scans. This consumes more verification resources, such as CPU time and memory, and also increases decision delay because it takes longer to complete the additional verification tasks. The delayed admission action indicates that the software temporarily does not meet the admission requirements, but is not completely unfeasible. Re-verification is required after a period of waiting. During this waiting period, the software can be further modified and optimized. This action consumes relatively less verification resources, but has a longer decision delay due to the required waiting time. The denied admission action indicates that the software has serious issues, does not meet the admission requirements, and cannot proceed to the next phase. While this action does not consume additional verification resources, it also has a relatively long decision delay because the developer must be notified to make modifications and then re-verify after the modifications are made.
[0099] In this embodiment, in step S5, based on the dynamic decision-making mechanism and the posterior probability distribution, the Bellman equation based on the POMDP model is solved to obtain the optimal strategy;
[0100] The Bellman equation is expressed as follows:
[0101] V * (s)=max a∈A [R(s,a)+γΣ s' P(s'|s,a)V * (s')]
[0102] Where V * (s) is the optimal value function under state s; a is an action in the action space; R(s,a) is the reward function; s is the current state; s' is the next state; P(s'|s,a) is the state transition probability matrix; γ is the discount factor.
[0103] In this embodiment, the reward function comprehensively considers verification cost, defect risk, and business impact; the discount factor γ is dynamically adjusted through online learning.
[0104] The design of the reward function R(s,a) is the core of the dynamic decision-making mechanism. Verification costs include computing resources and time consumed during the verification process. Choosing the enhanced verification action results in a relatively high verification cost, while choosing the immediate release action results in a lower verification cost. Defect risk refers to the likelihood of a defect existing in the software and can be assessed based on observed code quality metrics and security scan results. Software with high-level security vulnerabilities or poor code quality has a higher defect risk. Business impact considers the impact of the software's verification results on the business. For mission-critical software, verification failure may cause business interruption and have a significant impact; for non-critical software, the impact is relatively small. The discount factor γ balances the importance of current and future rewards. Dynamically adjusting γ through online learning allows decisions to be more adaptable to different environments and business needs. In rapidly changing environments and with more urgent business needs, γ can be appropriately increased to prioritize current rewards. In relatively stable environments and with less urgent business needs, γ can be appropriately decreased to consider longer-term rewards.
[0105] Among them, the expression of the reward function R(s,a) is:
[0106] R(s,a)=α·Q(s)-β·C(a)-γ·D(a)
[0107] Where α, β, and γ are business weight parameters, which are automatically optimized through reinforcement learning; Q(s) is the quality score in the current state, reflecting the quality level of the software; C(a) is the cost of executing action a, including verification resource consumption and time cost; D(a) represents the decision delay of executing action a.
[0108] Business weight parameters α, β, and γ are used to adjust the importance of different factors in the reward function. Through reinforcement learning algorithms, the system can automatically adjust these weight parameters based on different business needs and environmental conditions, making the reward function more realistic and enabling better decisions.
[0109] In this embodiment, the state transition probability matrix P(s'|s,a) is obtained through training with historical validation data, and a long short-term memory (LSTM) network is used to capture the temporal dependencies of state transitions. Long short-term memory networks are a special type of recurrent neural network that can effectively process long-term dependencies in sequential data. When training the state transition probability matrix, the LSTM network can learn the transition patterns between different states and the relationship between state transitions and time. By learning from a large amount of historical validation data, the LSTM network can accurately predict the transition probabilities of system states under different actions, providing a more reliable basis for dynamic decision-making.
[0110] In this embodiment, the observation probability matrix O(s,a) is calculated after extracting features from multimodal data using a convolutional neural network (CNN), supporting the fusion analysis of more than ten verification indicators. A convolutional neural network is a deep learning model specifically designed for processing image and sequence data, with powerful feature extraction capabilities. When calculating the observation probability matrix, the CNN network can perform feature extraction on multimodal verification indicators and convert different types of data into a unified feature representation. By analyzing and processing these features, the CNN network can accurately calculate the probability of observing different results under different conditions, thereby realizing the fusion analysis of multi-dimensional verification indicators.
[0111] In this embodiment, the TensorFlow Lite framework is used to achieve lightweight deployment, the model size is compressed to 45MB, and it supports running on ARM architecture devices. The verification strategy is synchronized to each verification node in real time via the WebSocket protocol. TensorFlow Lite is a lightweight machine learning framework specifically designed for running models on mobile and embedded devices. By adopting this framework, it can run efficiently on devices with limited resources. The WebSocket protocol is a two-way communication protocol that enables real-time data transmission. Through this protocol, the verification strategy can be synchronized to each verification node in a timely manner to ensure the consistency and accuracy of the verification.
[0112] In this embodiment, in step S6, an admission decision is generated according to the optimal strategy; and corresponding actions in the decision action set are mobilized according to the admission decision;
[0113] Specifically, such as Figure 5 As shown, an admission decision is generated according to the optimal strategy, and corresponding actions in the decision action set are mobilized according to the admission decision, such as immediate release, enhanced verification, delayed admission, and denied admission.
[0114] Specifically, in actual operation, the multi-source data acquisition module regularly obtains verification metrics from various data sources. The posterior probability distribution acquisition module uses these metrics to update the state probability distribution in real time. The optimal strategy acquisition module calculates the expected reward of each action based on the state probability distribution, selects the action with the highest expected reward as the optimal action, and triggers the corresponding verification process. The parameter adjustment and optimization module collects the execution results of the verification process and updates the model parameters based on the results to improve system performance and accuracy.
[0115] In this embodiment, verification access decisions support a manual intervention mode. When the confidence level of an automated decision falls below a threshold, a manual review process is triggered to ensure security in high-risk scenarios. In certain complex or high-risk scenarios, automated decisions may be subject to uncertainty. By setting up a manual intervention mode, when the confidence level of an automated decision falls below a preset threshold, the system automatically triggers a manual review process, allowing for further review and judgment of the decision results. This approach effectively ensures security in high-risk scenarios and avoids the serious consequences of automated decision errors.
[0116] In this embodiment, in step S7, after completing the corresponding action, the parameters of the POMDP model and the reward function of the Bellman equation are adjusted and optimized.
[0117] Specifically, such as Figure 3 As shown in , after completing the corresponding action, the system state transition relationship; Figure 6 As shown in Figure 1, the results and feedback of decision execution are collected, and the model parameters and reward function are adjusted and optimized based on this information. Through online learning, the system can continuously adapt to changes in the environment and adjustments to business needs, improving the accuracy and efficiency of verification.
[0118] In this embodiment, a Prioritized Experience Replay (PER) mechanism is used to store historical decision data, enabling asynchronous updates of model parameters via a distributed training platform. The Prioritized Experience Replay mechanism prioritizes historical decision data based on the importance and accuracy of the decisions. When training the model, high-importance data is prioritized for learning, improving learning efficiency and model accuracy. The distributed training platform can distribute training tasks across multiple nodes for parallel execution, enabling asynchronous updates of model parameters. This approach can significantly shorten training time and improve the system's adaptability and responsiveness.
[0119] In this embodiment, Figure 4 As shown, three-level verification is achieved through the edge layer, control layer and policy layer;
[0120] The edge layer deploys a lightweight inference engine on CI / CD nodes, achieving a decision response time of <200ms. This lightweight inference engine utilizes optimized algorithms and data structures, reducing computational complexity and memory usage. Parallel computing and caching technologies further improve inference speed. The edge layer directly processes verification requests from the CI / CD pipeline, providing preliminary decision results in a short period of time, significantly shortening the verification cycle. The control layer schedules verification resources based on Kubernetes and dynamically adjusts the sandbox environment size. Kubernetes is an open-source container orchestration platform with powerful resource management and scheduling capabilities. The control layer automatically allocates and adjusts verification resources based on verification task requirements and resource usage, ensuring efficient execution. Dynamically adjusting the sandbox environment size allows for flexible provision of sandbox environments of varying sizes based on the complexity and importance of verification tasks, improving resource utilization. The policy layer synchronizes global policies through federated learning, supporting over 50 model updates per week. Federated learning is a distributed machine learning technology that allows nodes to collaboratively train models without sharing raw data. The strategy layer leverages federated learning technology to aggregate local models from edge nodes and the control layer to arrive at a globally optimal strategy. Frequent model updates enable the system to promptly adapt to environmental changes and business needs, improving verification accuracy and efficiency.
[0121] In this embodiment, the present invention adopts microservice architecture deployment, and the verification module is horizontally expanded through the gRPC protocol to support 500+ concurrent verification requests. The microservice architecture splits the system into multiple independent services, each of which can be independently developed, deployed and expanded. This architecture has the advantages of high scalability, high fault tolerance and easy maintenance. The verification module communicates through the gRPC protocol, which is a high-performance, open source remote procedure call (RPC) protocol with efficient data transmission and serialization capabilities. By horizontally expanding the verification module, the number of verification nodes can be dynamically increased or decreased according to actual verification requirements, thereby supporting a large number of concurrent verification requests.
[0122] In this embodiment, three deployment modes are supported: the single-node mode is suitable for small teams; the cluster mode supports horizontal expansion; and the cloud-edge collaboration mode achieves low-latency verification. The single-node mode deploys the system on a single node, with a simple structure and easy management, suitable for small teams with limited resources. The cluster mode achieves horizontal expansion by forming a cluster with multiple nodes, and can handle a large number of concurrent verification requests, which is suitable for large-scale development and verification scenarios. The cloud-edge collaboration mode combines the advantages of cloud computing and edge computing, performs model training and global policy management in the cloud, and performs real-time verification and decision-making at the edge nodes, achieving low-latency verification and improving the response speed and performance of the system.
[0123] In this embodiment, the present invention provides a visual monitoring interface that displays key metrics such as verification node load, decision latency, and defect rate in real time, and supports threshold alarms and policy tuning suggestions. This visual monitoring interface intuitively displays the system's operating status and performance metrics, facilitating monitoring and management. By setting threshold alarms, when key metrics exceed preset thresholds, the system promptly issues an alert, prompting management to take action. Furthermore, the system provides policy tuning suggestions based on historical data and real-time status, helping managers optimize verification strategies and improve system efficiency and accuracy.
[0124] In a possible embodiment, a specific example is provided as follows:
[0125] The NASA POMDP benchmark problem parameters are used for verification. The input parameters are as follows:
[0126] Code complexity: 5 (minimum); Security vulnerabilities: None; Test coverage: 95% (conditional coverage); Environment load: 35% CPU utilization.
[0127] The state reasoning process is as follows:
[0128] Quality confidence Q = 0.92; Environmental stability E = stable state; Verification priority P = regular; Verification stage T = rapid verification.
[0129] The decision-making process is as follows (some data is quoted from the TensorFlow Lite performance report):
[0130] Reward function calculation: 0.6*0.92-0.2*50-0.2*50=-19.448 (in this case, α=0.6, β=γ=0.2, and the cost of executing the action and the decision delay are both 50);
[0131] Action selection: Immediate release (A1), verification takes only 30ms.
[0132] The final verification results are as follows (some data is quoted from the CNCF Cloud Native White Paper):
[0133] Resource consumption: The traditional method requires 240 core hours, while the present invention requires 240*35%=84 CPU core hours;
[0134] Decision delay: 150ms for the traditional method and 30ms for the present invention.
[0135] In summary, the present invention collects multi-source data such as code quality indicators, security scanning results, test coverage and environmental health in the DevSecOps software development process in real time through the CI / CD pipeline interface, security scanning tool API and cloud native monitoring system to obtain an observation vector of software quality; based on the observation vector, constructs a state space of software quality; uses the state space to describe the software delivery status in multiple dimensions; it iteratively calculates and updates the state space through the exponentially weighted moving average algorithm and the particle filter algorithm to obtain the posterior probability distribution of the software quality state transition in the development process; by defining the quality access action space, a decision action set is obtained; the action association in the decision action set verifies resource consumption parameters and decision delay parameters; based on the dynamic decision mechanism and the posterior probability distribution, the Bellman equation based on the POMDP model is solved to obtain the optimal strategy; an access decision is generated according to the optimal strategy; the corresponding action in the decision action set is mobilized according to the access decision; after the corresponding action is completed, the parameters of the POMDP model and the reward function of the Bellman equation are adjusted and optimized. This paper intelligently optimizes verification strategies by building a dynamic decision-making model, addressing the shortcomings of traditional approaches in adaptability, real-time performance, and resource efficiency. This paper can reduce verification resource consumption, lower defect leakage rates, and shorten decision response times, making it suitable for high-concurrency verification scenarios within a microservices architecture.
[0136] It should be noted that the method of the embodiments of the present disclosure can be performed by a single device, such as a computer or server. The method of the embodiments of the present disclosure can also be applied in a distributed scenario, where multiple devices cooperate to perform the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiments of the present disclosure, and the multiple devices will interact with each other to complete the method.
[0137] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0138] Example 2
[0139] See also Figure 7 Embodiment 2 of the present invention further provides a DevSecOps software quality access control analysis system based on the POMDP model, including:
[0140] Multi-source data collection module 001 is used to collect multi-source data in real time through CI / CD pipeline interfaces, security scanning tool APIs, and cloud native monitoring systems. The multi-source data includes code quality indicators, security scanning results, test coverage, and environmental health in the DevSecOps software development process to obtain observation vectors for software quality;
[0141] A state space construction module 002 is used to construct a state space of software quality based on the observation vector; and to describe the software delivery status in multiple dimensions through the state space;
[0142] The posterior probability distribution acquisition module 003 is used to iteratively calculate and update the state space through the exponentially weighted moving average algorithm and the particle filter algorithm to obtain the posterior probability distribution of the software quality state transition during the development process;
[0143] The decision action set definition module 004 is used to obtain a decision action set by defining the quality access control action space; the actions in the decision action set are associated with verification resource consumption parameters and decision delay parameters;
[0144] The optimal strategy acquisition module 005 is used to solve the Bellman equation based on the POMDP model based on the dynamic decision-making mechanism and the posterior probability distribution to obtain the optimal strategy;
[0145] Admission decision execution module 006, used to generate an admission decision according to the optimal strategy; and mobilize corresponding actions in the decision action set according to the admission decision;
[0146] The parameter adjustment and optimization module 007 is used to adjust and optimize the parameters of the POMDP model and the reward function of the Bellman equation after completing the corresponding action.
[0147] In this embodiment, three-level verification is achieved through the edge layer, control layer, and policy layer;
[0148] The edge layer deploys a lightweight inference engine on CI / CD nodes to achieve a decision response time of <200ms;
[0149] The control layer dynamically adjusts the scale of the sandbox environment based on Kubernetes scheduling and verification resources;
[0150] The strategy layer synchronizes global strategies through federated learning and supports 50+ model updates per week.
[0151] In this embodiment, in the multi-source data acquisition module 001, the observation vectors include: code quality indicators, security scanning results, test coverage, and environmental health.
[0152] In this embodiment, in the state space construction module 002, the expression of the state space is:
[0153] S={Q,E,P,T}
[0154] In the formula, Q is the quality confidence, which can be divided into three levels: high, medium, and low; E is the environmental stability, which is divided into three states: stable, fluctuating, and abnormal; P is the verification priority, which is divided into three levels: emergency, routine, and delayed; T is the verification stage, including rapid verification, deep verification, and manual review.
[0155] In this embodiment, in the posterior probability distribution acquisition module 003, the observation confidence is calculated by the exponentially weighted moving average algorithm; the calculation formula of the observation confidence is:
[0156] μ t =λ·o t +(1-λ)·μ t-1
[0157] Where μ t is the observation confidence; λ is the weighting coefficient, ranging from 0 to 1; o t is the observed value.
[0158] In this embodiment, in the decision action set definition module 004, the expression of the decision action set is:
[0159] A={A1,A2,A3,A4}
[0160] Wherein, A1 is immediate release; A2 is enhanced verification; A3 is delayed admission; A4 is denied admission. In this embodiment, in the optimal strategy acquisition module 005, the expression of the Bellman equation is:
[0161] V * (s)=max a∈A [R(s,a)+γΣ s' P(s'|s,a)V * (s')]
[0162] Where V * (s) is the optimal value function under state s; a is an action in the action space; R(s,a) is the reward function; s is the current state; s' is the next state; P(s'|s,a) is the state transition probability matrix; γ is the discount factor.
[0163] It should be noted that the information interaction, execution process, etc. between the modules of the above-mentioned system are based on the same concept as the method embodiment in Example 1 of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and no further details will be given here.
[0164] Example 3
[0165] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which the program code of the DevSecOps software quality access control analysis method based on the POMDP model is stored. The program code includes instructions for executing the DevSecOps software quality access control analysis method based on the POMDP model of embodiment 1 or any possible implementation thereof.
[0166] Computer-readable storage media can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).
[0167] Example 4
[0168] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;
[0169] The processor and the memory communicate with each other through a bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute the DevSecOps software quality access control analysis method based on the POMDP model of Example 1 or any possible implementation thereof.
[0170] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software code stored in a memory. The memory can be integrated into the processor or located outside the processor and exist independently.
[0171] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode.
[0172] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing system. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Alternatively, they can be implemented using program code executable by a computing system, and thus, they can be stored in a storage system and executed by the computing system. In some cases, the steps shown or described herein can be performed in a different order than that shown, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0173] Although the present invention has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications and improvements may be made thereto. Therefore, such modifications and improvements, without departing from the spirit of the present invention, are intended to be within the scope of protection claimed herein.
Claims
1. DevSecOps software quality access control analysis method based on POMDP model, characterized by: include: Real-time data collection from multiple sources, including code quality metrics, security scan results, test coverage, and environmental health during the DevSecOps software development process, is performed through CI / CD pipeline interfaces, security scanning tool APIs, and cloud-native monitoring systems to obtain observation vectors for software quality. Based on the observation vector, a state space of software quality is constructed; and a multi-dimensional description of the software delivery state is performed through the state space; The state space is iteratively calculated and updated by using an exponentially weighted moving average algorithm and a particle filter algorithm to obtain a posterior probability distribution of software quality state transitions during the R&D process; By defining the quality access control action space, a decision action set is obtained; the actions in the decision action set are associated with verification resource consumption parameters and decision delay parameters; Based on the dynamic decision-making mechanism and the posterior probability distribution, the Bellman equation based on the POMDP model is solved to obtain the optimal strategy; generating an admission decision according to the optimal strategy; mobilizing corresponding actions in the decision action set according to the admission decision; After completing the corresponding actions, the parameters of the POMDP model and the reward function of the Bellman equation are adjusted and optimized.
2. The DevSecOps software quality access control analysis method based on the POMDP model according to claim 1 is characterized in that: Achieve three-level verification through edge layer, control layer and policy layer; The edge layer deploys a lightweight inference engine on CI / CD nodes to achieve a decision response time of <200ms; The control layer dynamically adjusts the scale of the sandbox environment based on Kubernetes scheduling and verification resources; The strategy layer synchronizes global strategies through federated learning and supports 50+ model updates per week.
3. The DevSecOps software quality access control analysis method based on the POMDP model according to claim 2 is characterized in that: The observation vectors include: code quality indicators, security scan results, test coverage and environment health.
4. The DevSecOps software quality access control analysis method based on the POMDP model according to claim 3 is characterized in that: The expression of the state space is: S={Q,E,P,T} In the formula, Q is the quality confidence, which can be divided into three levels: high, medium, and low; E is the environmental stability, which is divided into three states: stable, fluctuating, and abnormal; P is the verification priority, which is divided into three levels: emergency, routine, and delayed; T is the verification stage, including rapid verification, deep verification, and manual review.
5. The DevSecOps software quality access control analysis method based on the POMDP model according to claim 4 is characterized in that: The observation confidence is calculated by the exponentially weighted moving average algorithm; the calculation formula of the observation confidence is: m t =l·o t +(1-λ)·μ t-1 Where μ t is the observation confidence; λ is the weighting coefficient, ranging from 0 to 1; o t is the observed value.
6. The DevSecOps software quality access control analysis method based on the POMDP model according to claim 5 is characterized in that: The expression of the decision action set is: A={A1,A2,A3,A4} Where A1 is immediate release; A2 is enhanced verification; A3 is delayed admission; and A4 is denied admission.
7. The DevSecOps software quality access control analysis method based on the POMDP model according to claim 6 is characterized in that: The expression of the Bellman equation is: V * (s)=max a∈A [R(s,a)+γΣ s' P(s'|s,a)V * (s')] Where V * (s) is the optimal value function under state s; a is an action in the action space; R(s,a) is the reward function; s is the current state; s' is the next state; P(s'|s,a) is the state transition probability matrix; γ is the discount factor.
8. A DevSecOps software quality access control and admission analysis system based on the POMDP model, which adopts the DevSecOps software quality access control and admission analysis method based on the POMDP model according to any one of claims 1 to 7, characterized in that: include: A multi-source data collection module is used to collect multi-source data in real time through CI / CD pipeline interfaces, security scanning tool APIs, and cloud-native monitoring systems. The multi-source data includes code quality indicators, security scanning results, test coverage, and environmental health during the DevSecOps software development process, thereby obtaining observation vectors for software quality. A state space construction module is used to construct a state space of software quality based on the observation vector; and to perform a multi-dimensional description of the software delivery state through the state space; A posterior probability distribution acquisition module is used to iteratively calculate and update the state space through an exponentially weighted moving average algorithm and a particle filter algorithm to obtain a posterior probability distribution of software quality state transitions during the R&D process; A decision action set definition module is used to obtain a decision action set by defining a quality access control action space; the actions in the decision action set are associated with verification resource consumption parameters and decision delay parameters; An optimal strategy acquisition module, configured to solve the Bellman equation based on the POMDP model based on a dynamic decision-making mechanism and the posterior probability distribution to obtain an optimal strategy; An admission decision execution module, configured to generate an admission decision according to the optimal strategy; and to mobilize corresponding actions in the decision action set according to the admission decision; The parameter adjustment and optimization module is used to adjust and optimize the parameters of the POMDP model and the reward function of the Bellman equation after completing the corresponding action.
9. The DevSecOps software quality access control and admission analysis system based on the POMDP model according to claim 8 is characterized in that: Achieve three-level verification through edge layer, control layer and policy layer; The edge layer deploys a lightweight inference engine on CI / CD nodes to achieve a decision response time of <200ms; The control layer dynamically adjusts the scale of the sandbox environment based on Kubernetes scheduling and verification resources; The strategy layer synchronizes global strategies through federated learning and supports 50+ model updates per week.
10. The DevSecOps software quality access control and admission analysis system based on the POMDP model according to claim 9 is characterized in that: In the multi-source data acquisition module, the observation vectors include: code quality indicators, security scanning results, test coverage and environmental health.