Cloud Resource Selection for A/B Testing Software Deployments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional cloud service software deployments often lead to significant issues due to undetected problems in early validation stages, causing impact blasts in production, and traditional health monitoring provides only a top-level understanding of application stability, lacking detailed metrics.
Innovation Solution
A method for dynamically selecting a sample set of cloud computing resources for testing new software releases based on telemetry data and customer support data, allowing for controlled distribution of potential negative effects and more effective testing, using a selection manager to identify suitable resources and perform A/B testing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional sequential deployment across sub-regions is used, then deployment coverage is improved, but system reliability deteriorates due to undetected issues causing large impact blasts
Solution Approach 1:
The patent segments the deployment population into multiple geographic regions and further into individual customers or resource groups. By deploying to segmented subsets simultaneously rather than sequentially, the system achieves both broad coverage and isolated risk containment. Each segment acts as an independent test bed, allowing parallel validation without cascading failures across the entire system.
Solution Approach 2:
The patent performs preliminary actions by deploying software updates to a small subset of customers or resources in each region before full rollout. This staged preliminary deployment allows validation of software health in production environments with real customer data, detecting issues before they affect the entire customer base. The preliminary action includes monitoring telemetry metrics and customer support data to assess software stability.
2Ease of operation
If traditional health monitoring is used, then monitoring simplicity is improved, but measurement precision deteriorates due to lack of detailed product metrics
Solution Approach 1:
The patent introduces an intermediary layer between traditional health monitoring and detailed product metrics. This intermediary automatically collects, aggregates, and analyzes telemetry data from multiple sources (application performance, system metrics, customer support interactions) to produce synthesized health indicators. This maintains monitoring simplicity for operators while achieving precise measurement of software health through automated data processing and analysis.
Solution Approach 2:
The patent implements feedback loops that continuously monitor software performance metrics, customer support data, and telemetry information during deployment. When anomalies or degradation patterns are detected, the system automatically triggers alerts, rollback procedures, or further investigation. This feedback mechanism enables precise detection of software health issues while maintaining simple operational workflows through automated responses.
3Reliability
If insider validation programs are used for early testing, then software validation quality is improved, but device complexity and cost increase due to overhead of validation program infrastructure
Solution Approach 1:
The patent enables self-service validation by allowing the cloud service itself to participate in testing its own software updates through automated telemetry collection and performance monitoring. Customer resources serve as their own test beds, with the system automatically gathering validation data from production environments without requiring separate validation infrastructure. This eliminates the need for complex insider validation programs while maintaining high validation quality through real-world customer workloads.
Solution Approach 2:
The patent makes customer cloud resources serve multiple functions: they are both production environments for customer workloads and test beds for software validation. The same infrastructure handles both customer operations and software testing, eliminating the need for dedicated validation environments. This multi-functionality reduces infrastructure complexity while enabling comprehensive software validation across diverse real-world usage scenarios.
Data Source
AI summary
A sample set of cloud computing resources is dynamically selected for testing a software deployment. Telemetry data associated with the resources and customer support data associated with customers that utilize the resources are obtained. A subset of the customers is selected based on the customer support data, and a candidate subset of the resources is selected based on the selected subset of customers and the telemetry data. Criteria for the selection is based on usage patterns and is configurable. Resources of customers with special support agreements, and customers previously selected, may be excluded from the candidate subset. The sample set of cloud computing resources may be randomly selected from the candidate subset. Software is deployed to the sample set as a B resource group and tested for issues with an A resource group to determine whether to proceed to full deployment, roll back the deployment, and/or retest the software.


