Intelligent service access management method and system based on Python service management platform

By collecting fine-grained telemetry data to generate real-time status images and combining machine learning to diagnose silent faults, the problems of insufficient service access decision-making and difficult to identify hidden faults in the existing technology are solved, and intelligent risk perception and regulation of service access is realized, and system stability and operation and maintenance efficiency are improved.

CN120343077AActive Publication Date: 2025-07-18SOUTHERN POWER GRID DIGITAL GRID RESEARCH INSTITUTE CO LTD

Patent Information

Application Number
CN202510544086.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-18
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

The existing business management platform lacks in-depth assessment of the actual performance and potential risks of the instance handling specific types of service loads in service access management, making it difficult to identify hidden faults and evaluate their impact on the business, resulting in improper traffic allocation and inefficient operation and maintenance.

Method used

By collecting fine-grained telemetry data, generating real-time status images, using machine learning to diagnose silent faults and combining transaction context diagrams to locate fault-affected domains, realizing closed-loop regulation of intelligent decision-making and risk perception.

Benefits of technology

It improves the accuracy of service access decisions, enhances the early identification of hidden faults, realizes the automated correlation and evaluation of technical faults and business impacts, and improves system stability and operation and maintenance intelligence level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343077A_ABST
    Figure CN120343077A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent service access management method and system based on a Python service management platform, and relates to the technical field of industrial internet, and the method comprises the steps: collecting fine-grained telemetry data associated with a service context; generating a real-time state portrait which contains a predictive future risk score and faces a specific business request type for a downstream service instance based on the data; automatically diagnosing whether the service instance has a silent fault or performance slow descent or not by applying a machine learning algorithm; the business influence domain of the diagnosed fault on the specific business process or the key performance indicator is automatically determined in combination with the transaction context graph and the tracking data; and when a service access request is received, intelligent decision making and execution are carried out. According to the invention, through providing refined and foresight evaluation of the service state and automatic association analysis capability of hidden faults and business influences thereof, intelligent risk perception and regulation of service access are realized, and improvement of stability, reliability and operation and maintenance efficiency of a business platform is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial Internet, and particularly to a service access intelligent management method and system based on a Python business management platform. Background Art

[0002] The existing Business Management Platform (BMP) has limitations in service access management and health status monitoring. First, service access decisions are usually based on macroscopic resource metrics of target service instances or basic availability detection, lacking in-depth evaluation of the true performance and potential risks of instances when processing specific types of business loads, and it is also difficult to effectively perceive the conduction risks caused by downstream dependency anomalies, which easily leads to improper traffic allocation and affects service quality. Second, for hidden problems such as non-crash "silent failures" or "performance degradation", the detection ability of traditional threshold-based monitoring systems is limited. Moreover, in the existing technology, there is still a lack of effective means to automatically and accurately associate and quantitatively evaluate the identified technical-level anomalies with their actual impacts on specific business processes or key performance indicators (KPIs), which restricts the efficiency of fault handling and risk-based operation and maintenance decisions. Summary of the Invention

[0003] The present invention provides a service access intelligent management method and system based on a Python business management platform to solve the problems in the prior art such as insufficient service access decision-making information, lagging risk perception, difficulty in early detection of hidden failures, and difficulty in evaluating business impacts.

[0004] In a first aspect, the present invention provides a service access intelligent management method based on a Python business management platform, and the method includes:

[0005] Collect fine-grained telemetry data of service instances within the platform, and parse business requests to obtain business context information;

[0006] Based on the fine-grained telemetry data and the business context information, generate a real-time state portrait for downstream service instances for a specific business request type, and the real-time state portrait includes a predictive future risk score for the service instance to process the specific business request type;

[0007] Automatically diagnose whether the service instance has non-crash silent failures or performance degradation, and automatically determine its business impact domain on specific business processes or metrics when identified;

[0008] When receiving a service access request, based on the business context information of the request, the real-time state portrait of candidate downstream service instances for the request type, the predictive future risk score, the diagnosed silent failure information, and its business impact domain, intelligently decide on service access actions;

[0009] Execute the service access action.

[0010] In a second aspect, the present invention also provides a system for intelligent management of service access based on a Python business management platform. The system includes:

[0011] A data collection and context association unit, configured to collect fine-grained telemetry data, parse business requests to obtain business context information and inject tracing information, and construct and maintain a transaction context graph rich in business semantics;

[0012] A status portrait and risk prediction unit, configured to generate a real-time status portrait including a predictive future risk score for handling a specific type of business request;

[0013] A silent fault diagnosis and impact domain localization unit, configured to automatically diagnose whether there is a silent fault or performance degradation in a service instance and automatically determine its business impact domain when identified;

[0014] An intelligent decision-making unit, configured to make a decision on the service access action according to the request context, real-time status portrait, diagnosis information and business impact domain;

[0015] A policy execution unit, configured to execute the service access action instruction;

[0016] A federated learning support component, configured to support model training using a federated learning framework;

[0017] A digital twin simulation verification environment interface, used to connect or integrate a digital twin environment for simulation verification of intervention strategies.

[0018] The technical solution provided by this application has at least the following technical effects or advantages:

[0019] Improve the accuracy of service access decisions: By generating a real-time status portrait for a specific type of business request and integrating a predictive future risk score, it is possible to more accurately evaluate the ability and risk of downstream service instances to handle specific loads, thereby reducing unreasonable traffic distribution.

[0020] Enhance the ability to identify hidden faults at an early stage: Using machine learning diagnosis techniques helps to identify "silent faults" or "performance degradation" at an early stage that are difficult to detect by conventional monitoring, creating conditions for timely intervention.

[0021] Achieve automated association and evaluation of technical faults and business impacts: By combining the transaction context graph and tracing data for business impact domain localization, it is possible to associate and quantify the underlying technical problems with the impact on specific business processes or KPIs, providing a basis for prioritizing fault handling and risk control.

[0022] Improve the overall stability of the system and the level of intelligent management: Integrate in-depth state perception, risk prediction, precise diagnosis, and business impact assessment into the closed-loop control of service access, which helps enhance the operational resilience of the system and improve the intelligence level of operation and maintenance management. Brief Description of the Drawings

[0023] Figure 1 It is a flowchart of the intelligent management method for service access based on the Python business management platform of the present invention;

[0024] Figure 2 It is an architecture diagram of the intelligent management system for service access based on the Python business management platform. Detailed Embodiments

[0025] The present invention relates to an intelligent management method and system for service access based on a Python business management platform to solve the technical problems of insufficient service access decision-making information, lagging risk perception, difficulty in early detection of hidden faults, and difficulty in assessing business impacts in the prior art.

[0026] The above technical solutions will be described in detail below in combination with the drawings in the specification and specific embodiments to better understand the above technical solutions. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited to the exemplary embodiments used to explain the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings rather than all.

[0027] Embodiment 1

[0028] As Figure 1 - Figure 2 shown, the intelligent management method and system for service access based on the Python business management platform of the present invention can be deployed as a core support system in the BMP environment. Its logical architecture usually includes the following core functional units that work together:

[0029] 1. Data collection and context association unit: Responsible for building the data foundation required for analysis and decision-making. The key lies in the comprehensiveness, fine granularity, and business semantic relevance of the data.

[0030] Fine-grained Telemetry Data Collection: Metric Collection: Configure a standard metric collection client (such as integrating the OpenTelemetry SDK) in the BMP microservice (which can be built based on Python web frameworks like FastAPI). In addition to collecting system resource metrics, key performance indicators must be collected for each service instance for different business request types (uniquely identified by API endpoints, etc.) through application layer monitoring means (such as middleware). The collected metrics must at least cover: multiple statistical quantiles of processing latency (including P95 or P99 values to reflect the long tail), error rate information (distinguishing important error categories), service internal state metrics (such as the real-time length and waiting time of critical queues), and interaction performance metrics for downstream dependencies (such as database interaction time, connection pool usage rate). All metrics are strictly associated with the service instance ID and business request type ID through tags and transmitted to the central time series data storage system.

[0031] Structured Logging: Implement structured logging (preferably in JSON format). Log records contain standard metadata and TraceID and SpanID for link association. Important context information during business execution (such as user ID, order ID) is recorded as structured fields. Log streams are aggregated into the central log management system.

[0032] Distributed Tracing and Context Injection: Implement distributed tracing based on OpenTelemetry across the system. Use automated instrumentation and manual code injection. At the service entry point, extract and inject business context information. The parsed context information (including attributes such as request type, business ID, priority, etc.) is attached to the current tracing span as propagable tracing attributes (Baggage or SpanTags) using the OpenTelemetry API and ensure its correct transmission across service boundaries. Tracing data is exported to the backend system.

[0033] Business Events: Capture business state change events by subscribing to the message bus. Event messages contain the associated business ID and TraceID.

[0034] Transaction Context Graph (RBTCG) Construction and Maintenance: The backend data processing service (which can be based on technologies such as Apache Flink or Python Faust) consumes the integrated telemetry data.

[0035] Based on the call chain structure and Span attributes (including injected context) in the trace data, and combined with the business process mapping information provided externally or dynamically discovered, construct and incrementally update in real time a transaction context graph with business semantics (RBTCG). Nodes in the graph represent service instances (with additional status information), and edges represent call relationships (with additional interaction statistics and context), clearly mapping technical interactions and business processes. The graph data is stored in a suitable system such as a graph database.

[0036] 2. The status portrait and risk prediction unit extracts structured, decision-oriented deep service status information from massive and heterogeneous raw telemetry data, and provides forward-looking risk assessment in combination with machine learning techniques.

[0037] Generation of real-time status portraits for specific business request types:

[0038] Data processing flow: This function is executed through a real-time or near-real-time data processing pipeline (the implementation of which can be based on a stream processing framework such as Apache Flink or use a stream processing library in Python). The pipeline continuously consumes the fine-grained telemetry data stream with associated business context provided by the data collection and context association unit.

[0039] Core aggregation and portrait construction logic: For each service instance within the system monitoring scope and each type of business request defined as critical processed by the instance, the processing pipeline performs the following operations:

[0040] Within a preset time window (e.g., the last 5 minutes), filter and aggregate all fine-grained performance metric data points related to the specific service instance and the specific business request type.

[0041] Based on the aggregated data, independently calculate and generate a multi-dimensional performance portrait describing the current state of the service instance when processing the specific business request type. The key statistical dimensions that must be included in the performance portrait are: key statistics of the processing delay distribution (covering high percentiles such as P95 or P99), error rate statistical information (distinguishing important error categories), state quantification metrics of internal key queues (such as queue length or waiting time), and performance characteristics of downstream dependent calls when processing this request type (such as interaction delay). When appropriate, it can also include characteristics of associated resource consumption patterns.

[0042] When generating this performance portrait, dynamically query and integrate two other important analysis results: First, any identified silent fault status identifiers and their severity levels for the service instance that may affect the current business request type from the silent fault diagnosis and impact domain localization unit; Second, the predictive future risk score for the service instance to process this business request type generated by the internal risk prediction module of this unit.

[0043] Image Structure and Storage: The finally generated real-time status image is a structured data object that encapsulates comprehensive status information about a specific service instance. The key lies in its ability to distinguish and store independent performance profiles, risk scores, and fault markers when the instance processes different types of business requests internally. This image data needs to be stored in a data storage system that supports low-latency and high-concurrency access (e.g., using in-memory databases such as Redis) so that the intelligent decision-making unit can quickly and real-time obtain the status input required for decision-making.

[0044] Predictive Future Risk Score Generation: Utilize machine learning techniques to predict the probability of predefined "negative events" such as significant performance degradation or failure when a specific service instance processes a certain type of business request within a predefined short time window in the future (e.g., the next 5 to 15 minutes). Specifically, it includes:

[0045] Feature Engineering: Construct a feature set for input into the prediction model. The core input is historical multivariate time series data related to the target service instance and the target business request type (mainly extracted from the performance metric sequences in the historical real-time status image records). To improve prediction accuracy, auxiliary features usually need to be introduced, which may include the global resource utilization time series of the service instance, the health status or risk score time series of its important downstream dependent services, recent environmental change event markers (e.g., deployments, configuration modifications), and metric sequences reflecting the overall system load or the status of specific business activities, etc.

[0046] Model Training and Application: Select a machine learning model suitable for handling complex time series dependencies (e.g., deep learning models based on recurrent neural networks or Transformer architectures, or ensemble models such as gradient boosting trees applicable after good feature engineering) for training. The training process is a supervised learning task that requires a large amount of labeled historical data. Strict time series cross-validation methods need to be adopted during training and potential data imbalance problems need to be addressed. For scenarios with data privacy restrictions, it is recommended to use the federated learning framework to implement distributed training, and build a global prediction model by training local model copies and aggregating model parameter updates by a central coordinator.

[0047] Online Inference and Integration: Deploy the optimal model as a stable and low-latency online prediction service. This service receives the latest input feature vectors in real-time, performs inference, and calculates the predictive future risk score value representing the risk probability. This score value is immediately updated to the real-time status image record corresponding to the service instance and the business request type, becoming an indispensable part of the image.

[0048] 3. The Silent Fault Diagnosis and Impact Area Location Unit uses data analysis techniques to automatically detect hidden system problems that are difficult to capture by conventional monitoring means and establish their association with real business impacts.

[0049] Automated Silent Fault Diagnosis: Aims to automatically and early detect non-crash and potentially long-latent abnormal behavior patterns occurring in service instances. Implemented by a background analysis service, continuously processing fine-grained telemetry data. To ensure high accuracy and low false alarm rate, a multi-model fusion strategy is preferably adopted, that is, comprehensively processing the output results of at least two anomaly detection algorithms based on different technical principles. The technical categories of these algorithms may include:

[0050] (a) Method based on deep learning reconstruction model: Applying methods such as deep autoencoders to analyze the reconstruction error of multi-dimensional time series metric data. Abnormal data usually causes the reconstruction error to deviate significantly from the normal range.

[0051] (b) Method based on statistical process control (SPC) theory: Applying methods such as multivariate SPC control charts (e.g., Hotelling T2 chart) or univariate SPC rules (e.g., CUSUM chart rule) to detect small but statistically significant persistent shift patterns in key metric sequences.

[0052] (c) Method based on log data pattern mining: Applying an online log clustering algorithm (e.g., Drain algorithm) to extract log templates, and then monitoring the abnormal patterns of the log template sequence (e.g., using hidden Markov models or RNN to analyze transition probabilities) or the distribution drift of numerical parameters inside the templates.

[0053] Fusion Decision: Comprehensively judge the anomaly signals from different algorithms through a set fusion logic (e.g., weighted voting or meta-learning model), and output a high-confidence silent fault diagnosis conclusion. This conclusion (including fault type, confidence level, evidence, etc.) is recorded and updated to the real-time status portrait of the affected instance.

[0054] Business Impact Area Location:

[0055] Trigger Condition: Automatically starts after the system confirms a high-confidence silent fault diagnosis result.

[0056] Furthermore, the analysis process: Based on the fault instance identifier and the occurrence time window, all relevant trace links are filtered out from the distributed tracing data store, especially those links that may exhibit anomalies (such as high latency or errors) when flowing through the fault instance; Utilize the pre-built transaction context graph containing business semantics. First, locate the node representing the fault instance on this graph. Then, execute the graph path analysis algorithm, usually combined with the information of the filtered abnormal trace links, for backward tracing. This tracing process aims to identify the list of upstream business transaction types potentially affected by this fault and the end-to-end business processes to which they belong.

[0057] Business key performance indicator (KPI) impact assessment: Obtain the historical time series data of predefined business KPIs related to the affected business processes identified in the previous step. Adopt strict statistical correlation analysis or causal inference methods to evaluate whether there is a statistically significant association between the fault event and the observable changes in these KPIs during the time period when the fault occurs, and quantify the degree of this impact as much as possible (for example, estimate the percentage of KPI decline).

[0058] Report generation: Integrate the results of the entire analysis process - including fault details, technical impact scope, list of affected business transactions / processes, and (if there is a significant association) the quantified business KPI impact assessment - to generate a structured business impact domain report.

[0059] 4. The intelligent decision-making unit is responsible for receiving and understanding all the in-depth insight information provided by the aforementioned units, and accordingly formulating and outputting the optimal service access routing decision or proactive risk intervention strategy: Responding to service access requests or internally triggered intervention requirements in real time. It queries and obtains the latest and complete real-time status portraits (including predicted risks and diagnostic information) of relevant service instances, understands the business context of the current request, and combines the relevant business impact domain report, and then generates control instructions through the internal decision-making logic.

[0060] The process of formulating intelligent routing decisions:

[0061] Input acquisition: When it is necessary to select a target from a group of candidate downstream service instances for an incoming service request (whose business context is known and the request type is determined), the decision-making unit first concurrently queries and obtains the latest real-time status portraits (RTORSP) of these instances, focusing on extracting their predictive future risk scores (T-PHRS) for the current request type and any active silent fault diagnosis markers.

[0062] Evaluation and selection logic: Construct and apply an internal evaluation function or decision-making strategy to determine the routing priorities of each candidate instance. This evaluation process must comprehensively rely on the following key input factors:

[0063] (a) Predictive risk: Taking the predictive future risk score as the core consideration, generally, the lower the risk score, the more preferred.

[0064] (b) Risk tolerance related to business context: Judging the acceptance degree of potential risks based on the business attributes of the request itself (for example, is it a critical transaction or a background task?). High-priority or critical services usually require routing to instances with extremely low risks.

[0065] (c) Real-time performance and load: Referring to the current performance metrics (such as lower latency and lower error rate are more preferred) of the instance in handling similar requests reflected in the real-time status profile and the current overall load status (such as CPU and queue length, and tend to select relatively idle instances to achieve balance).

[0066] (d) Avoidance of the impact of silent failures: If a silent failure exists in an instance, it is necessary to judge whether the failure is related to the business process of the current request according to its business impact domain report. If it is related, significantly reduce the priority of the instance or apply a penalty factor; if it is not related, the impact is smaller.

[0067] Decision generation: Based on the results of the above comprehensive evaluation (for example, by calculating scores or cost ranking), combined with pre-set routing rules including risk thresholds (for example, if the predictive future risk score exceeds a certain value, it is not selected), or using a trained reinforcement learning model (which can learn the optimal decision-making strategy according to real-time status and historical feedback) to dynamically adjust routing decisions, and finally determine which specific instance to route the request to, or determine a plan including multiple instances and their corresponding traffic allocation ratios.

[0068] Trigger active intervention decisions based on high-severity silent failure diagnosis reports and their business impact domain evaluation results.

[0069] Policy selection: According to the predefined policy library and combined with the current system state, select appropriate intervention measures (such as alarm, isolation, traffic switching, service degradation, triggering self-healing, etc.).

[0070] Security verification: For automated intervention strategies with potentially wide-ranging impacts to be executed, it is recommended or required to first conduct simulation verification in a digital twin environment. This environment should be able to simulate the behavior of the real system and execute the strategy in the simulation to evaluate its expected effects and potential side effects (for example, whether it conforms to the "minimum blast radius" principle). The simulation results serve as the key basis to guide the final decision on whether to execute the intervention in the production environment.

[0071] 5. The policy execution unit is responsible for accurately converting decision instructions into operations on the underlying system.

[0072] Operations such as route updates, instance status changes, configuration modifications, or alarm notifications are performed by calling the standardized management interfaces (APIs) of each infrastructure component (API gateway, service mesh, container platform, configuration center, alarm system, etc.). It is necessary to ensure the reliability, idempotency, and security of the execution process and record detailed operation audit logs.

[0073] The system further includes a closed-loop feedback mechanism. By continuously measuring the actual effects of decision execution (such as performance, changes in business metrics, etc.) and feeding this information back to relevant intelligent analysis units (prediction, diagnosis) and decision-making units, it is used to drive the retraining of the model, the adaptive adjustment of parameters, or the iterative optimization of strategies, ensuring the effectiveness and intelligence level of the long-term operation of the system.

[0074] Through the foregoing detailed description of the system architecture, key functional units, and their interactions, the present invention provides a set of practical service access intelligent management methods and systems. Its core technical contribution lies in achieving deep state awareness and forward-looking risk prediction for specific business requests, automated silent fault diagnosis and accurate assessment of business impacts, and intelligent, risk-aware closed-loop control based on these in-depth insights. This specification has fully disclosed the technology with a Python-based technology stack and a cloud-native environment as an example, and its principles and methods can be extended and applied to other technology platforms. Implementing the present invention is expected to significantly improve the stability, performance, and operation and maintenance intelligence level of the distributed business management platform.

[0075] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0076] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalent technologies, the present application is also intended to include these changes and variations.

Claims

1. A service access intelligent management method based on a Python business management platform, characterized in that Including: Collecting fine-grained telemetry data of service instances within the platform and parsing business requests to obtain business context information; Based on the fine-grained telemetry data and the business context information, generating a real-time state portrait for downstream service instances for a specific business request type, the real-time state portrait including a predictive future risk score for the service instance to process the specific business request type; Automatically diagnosing whether there are non-crashing silent faults or performance degradation in the service instance, and automatically determining its business impact scope on specific business processes or metrics when identified; When receiving a service access request, intelligently making a decision on the service access action according to the business context information of the request, the real-time state portrait of candidate downstream service instances for the request type, the predictive future risk score, the diagnosed silent fault information and its business impact scope; Executing the service access action.

2. The intelligent management method for service access of the Python-based business management platform according to claim 1, characterized in that, In the step of collecting fine-grained telemetry data, collecting and correlating P95 or P99 latency, error rate, internal queue length, database interaction time, cache interaction time, connection pool usage when the same service instance processes different request types, and a log stream containing a trace ID injected with the business context information.

3. The intelligent management method for service access based on the Python business management platform according to claim 2, wherein In the step of parsing business context information, injecting the parsed business context information into tags of distributed tracing information, and constructing and maintaining a transaction context graph containing business semantics based on the tracing information and business process information. The transaction context graph containing business semantics includes nodes representing service instances and edges representing call relationships, and the nodes and the edges are attached with the business context and tracing information.

4. The intelligent management method for service access of the Python-based business management platform according to claim 3, wherein, In the step of generating a real-time state portrait, using the association information of the transaction context graph containing business semantics, aggregating the fine-grained telemetry data related to the specific request type, independently calculating and storing the current performance portrait and resource consumption pattern for each key request type of each service instance, and marking the diagnosed silent fault state and the predicted type-specific risk score affecting the request type.

5. The intelligent management method for service access based on the Python business management platform according to claim 4, characterized in that In the step of automatically diagnosing silent faults or performance degradation, adopting an integrated learning or multi-model fusion strategy, combining at least two algorithms among reconstruction error analysis based on autoencoders, control chart rule analysis based on statistical process control, and log template vectorization and clustering drift detection, to jointly analyze the fine-grained time series metrics and the log stream for the request type.

6. The intelligent management method for service access based on the Python business management platform according to claim 5, characterized in that, In the step of automatically determining the business impact scope, after identifying the silent fault, using the fault instance identifier and time window, performing a graph traversal or influence propagation algorithm on the transaction context graph containing business semantics, combining distributed tracing data for backward tracing, identifying the affected upstream service call chain and list of business transaction types, and quantifying or evaluating the impact degree of the fault on business key performance indicators by associating changes in business event data, generating a business impact scope report.

7. The intelligent management method for service access of the Python-based business management platform according to claim 6, characterized in that, For the steps of the intelligent decision-making service access action, when processing a service access request, query the real-time status portraits and the predictive future risk scores of all candidate downstream instances for this request type, and combine the business context of the request and the preset routing policy, or dynamically adjust the routing weights using a reinforcement learning model to decide the routing target or the traffic allocation ratio.

8. The intelligent management method for service access of the Python-based business management platform according to claim 7, wherein, For the machine learning model training in the predictive future risk score generation or the silent fault diagnosis, adopt a federated learning framework to train model replicas in the local environments of multiple service clusters or tenants, and aggregate the model parameters to construct a global model.

9. The intelligent management method for service access based on the Python business management platform according to claim 8, characterized in that, The method further includes a digital twin environment, which simulates the behavior of the business management platform based on the real-time transaction context graph containing business semantics, the status portrait, and the prediction or diagnosis model; after generating an intervention strategy for an identified silent fault, the intervention strategy is first executed and evaluated in the digital twin environment, and the simulation results are used to decide whether to execute the intervention in the production environment.

10. A system for intelligent management of service access based on a Python business management platform, the system being configured to execute the method for intelligent management of service access based on a Python business management platform as claimed in claim 9, characterized in that, The system includes: A data collection and context association unit, configured to collect fine-grained telemetry data, parse business requests to obtain business context information and inject tracing information, and construct and maintain a transaction context graph rich in business semantics; A status portrait and risk prediction unit, configured to generate a real-time status portrait containing a predictive future risk score for processing a specific business request type; A silent fault diagnosis and impact domain localization unit, configured to automatically diagnose whether a service instance has a silent fault or performance degradation, and automatically determine its business impact domain when identified; An intelligent decision-making unit, configured to decide the service access action according to the request context, the real-time status portrait, the diagnosis information, and the business impact domain; A policy execution unit, configured to execute the service access action instruction; A federated learning support component, configured to support model training using a federated learning framework; A digital twin simulation verification environment interface, used to connect or integrate a digital twin environment for the simulation verification of the intervention strategy.

Citation Information

Patent Citations

  • Fault analysis, power equipment fault analysis and fault analysis model training method

    CN119762290A

  • Monitoring device, monitoring method of monitoring object host, monitoring program, and recording medium

    JP2014120001A

  • Application performance monitoring (APM) detectors for flagging application performance alerts

    US11516269B1

  • Root cause location method, system and device

    WO2025036003A1

Cited By

  • Process decision and execution method and device, storage medium and electronic equipment

    CN121032172A

  • Proxy endpoint credible monitoring and data protection method and system for large model compatible interface

    CN122293308A