Network intelligent operation and maintenance method based on deep reinforcement learning

By generating session configuration packages and constructing state tensors and root cause graphs, the problem of inconsistent sampling granularity and clock reference in telemetry data during intelligent network operation and maintenance is solved. This enables unified acquisition and processing of telemetry data, ensures consistency between action strategies and execution trajectory records, and supports continuous decision-making and iterative updates in intelligent network operation and maintenance.

CN121940296APending Publication Date: 2026-04-28GLADIOLUS TECH (CHONGQING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GLADIOLUS TECH (CHONGQING) CO LTD
Filing Date
2026-01-21
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

In existing intelligent network operation and maintenance methods, telemetry data is difficult to align with sampling granularity and a unified clock reference. Time synchronization deviation, missing measurement rate, and noise jitter lack consistent quality labels and evidence credibility measures. Network changes lack a verification link of action constraint library consistent with state tensor and root cause correlation graph, resulting in inconsistent execution of action strategies and difficulty in forming stable execution trajectory records and model update closed loop.

Method used

By generating a session configuration package, collecting telemetry data and calculating quality labels, constructing a state tensor and root cause association graph, generating a candidate action set, inputting it into a hierarchical deep reinforcement learning decision model, outputting action policies and executing network changes, and generating execution trajectory records to update the model.

Benefits of technology

It enables unified collection and processing of telemetry data, ensuring consistency between action strategies and execution trajectory records. It is suitable for unified operation and maintenance scenarios across cell identifiers and slice identifiers, and supports continuous decision-making and iterative updates for intelligent network operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940296A_ABST
    Figure CN121940296A_ABST
Patent Text Reader

Abstract

The invention relates to the field of network intelligent operation and maintenance, in particular to a network intelligent operation and maintenance method based on deep reinforcement learning. The method comprises the following steps: performing session-level configuration on network topology, a network element list and a probe list, and combining sampling granularity, a field mapping table and a unified clock reference to form a consistent session configuration packet; collecting telemetry data, generating a quality label, calculating evidence credibility, and obtaining an admission evidence set through admission gating; constructing a state tensor and root cause association graph, forming a candidate action set in combination with the action constraint library, and performing rehearsal through a shadow network state to generate an effect prediction vector; and finally, inputting related information into the hierarchical deep reinforcement learning decision model, outputting an action strategy, executing network change, and recording an execution track for model updating. According to the method, closed-loop association of network operation decision and execution is realized through unifying the data aperture and the decision link, and the traceability and continuous evolution capability of the operation process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent network operation and maintenance, and in particular to a method for intelligent network operation and maintenance based on deep reinforcement learning. Background Technology

[0002] In the field of intelligent network operation and maintenance, existing solutions based on network topology, network element lists, and probe lists typically rely on telemetry data for status monitoring and network change orchestration. These solutions suffer from several limitations, including inconsistent session configuration packet definitions leading to difficulties in aligning telemetry data at the same sampling granularity and with a unified clock reference; a lack of consistent quality labels and evidence credibility metrics for telemetry data time synchronization deviations, missing detection rates, noise jitter, and field drift; and a lack of action constraint libraries that validate the network change process in accordance with state tensors and root cause graphs. Existing methods often lack access control and access evidence set solidification between the data acquisition and decision-making links. Furthermore, the sources of dependency and conflict markers between root cause graphs and candidate action sets are difficult to trace. Under action constraint libraries, insufficient executability of candidate action sets and inconsistent action policy dependency ordering can easily occur, making it difficult to achieve stable implementation of network changes based on action policies and the formation of execution trajectory records. For the joint processing of session configuration packages, admission evidence sets, state tensors, root cause association graphs, candidate action sets, shadow network states, effect prediction vectors, and hierarchical deep reinforcement learning decision models, existing technologies generally struggle to continuously correlate pre-executive trajectories with execution protection conditions, as well as execution trajectory records with training samples, under the same operational caliber. This makes it difficult to form a consistent process of collection—judgment—construction—constraint—pre-executive—decision—execution—recording in network intelligent operation and maintenance application scenarios. Consequently, the correspondence between policy versions and execution trajectory records is unclear, and the updates of hierarchical deep reinforcement learning decision models are difficult to align with operational data in a closed loop. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention provides a network intelligent operation and maintenance method based on deep reinforcement learning, comprising: S100: Obtain network topology, network element list, probe list, sampling granularity, field mapping table, and unified clock reference; generate session configuration package. S200: Collect telemetry data based on session configuration package, calculate time synchronization deviation, missing rate, noise jitter, and field drift, generate quality labels and calculate evidence credibility, and execute access control to generate access evidence set; S300: Based on the admission evidence set, construct a state tensor according to cell identifier, slice identifier, and time window index, and generate a root cause association graph; S400. Obtain the action constraint library, and generate a set of candidate actions carrying dependency and conflict markers based on the root cause relationship graph and the action constraint library; S500: Construct the shadow network state and perform a pre-play based on the candidate action set, generate the effect prediction vector, and filter the action set that passes the pre-play and the execution protection conditions; S600, based on pre-simulation, inputs action sets, state tensors, and effect prediction vectors into a hierarchical deep reinforcement learning decision model, and outputs action policies; S700: Based on the action policy, perform network changes and generate execution trajectory records, generate training samples based on the execution trajectory records, update the hierarchical deep reinforcement learning decision model and register the policy version.

[0004] Furthermore, the session configuration package in S100 includes a primary key binding table for probes to cell identifiers, slice identifiers, and network element identifiers, a field unit dictionary, sampling alignment rules, anomaly indicator calculation template version, and hash verification fields.

[0005] Furthermore, the quality labels in S200 include out-of-order rate, repetition rate, proportion of sudden spikes, and cross-domain time drift trend, and the credibility of the evidence is obtained by mapping the quality labels.

[0006] Furthermore, in S200, the entry control adopts a first threshold and a second threshold for grading. Data fragments with evidence credibility lower than the first threshold are written into the supplementary collection set, and data fragments with evidence credibility between the first threshold and the second threshold are written into the downgrade processing set and removed from the training sample construction process.

[0007] Furthermore, the state tensor in S300 includes radio-side load characteristics, interference characteristics, slice-level throughput characteristics, latency characteristics, core network session establishment characteristics, packet loss characteristics, and transmission-side congestion characteristics, and is aligned by time window index.

[0008] Furthermore, generating the root cause association graph in S300 includes mapping the set of anomaly candidates to a set of nodes, generating edge weights according to the temporal order and correlation strength of cross-domain indicators, and forming the root cause association graph.

[0009] Furthermore, the action constraint library in S400 includes slice service level constraints, capacity constraints, coverage constraints, change frequency constraints, parameter mutual exclusion tables, and dependency priority tables. The candidate action set is generated by matching the root cause nodes of the root cause association graph with the action templates, combined with the parameter mutual exclusion tables and dependency priority tables.

[0010] Furthermore, the set of change segments in S400 includes radio side parameter threshold change segments, slice resource weight change segments, core network traffic offloading change segments, core network migration change segments, and cross-domain rollback change segments.

[0011] Furthermore, the shadow network state in S500 includes a subset of key KPIs, a subset of key KQIs, cross-cell interference coupling terms, and congestion propagation approximation terms. The effect prediction vector includes service level default probability, oscillation risk indicators, and rollback cost prediction.

[0012] Furthermore, the hierarchical deep reinforcement learning decision model in S600 includes a cell-level decision sub-model, a slice-level decision sub-model, and a coordination sub-model. The coordination sub-model performs mutual exclusion checks and dependency ranking on the action policy. The reward function of the hierarchical deep reinforcement learning decision model includes a service level benefit term, a change stability penalty term, and a rollback cost term. The change stability penalty term is related to the number of parameter changes within a unit time window.

[0013] The key innovations of this invention include: (1) A session configuration package is generated based on the network topology, network element list, probe list, sampling granularity, field mapping table and unified clock reference. The acquisition and alignment entry of telemetry data is uniformly anchored under the session level of the session configuration package, so that the subsequent telemetry data processing link is connected around the same session configuration package.

[0014] (2) After collecting telemetry data based on the session configuration package, quality labels are generated around time synchronization deviation, missing rate, noise jitter and field drift, and the credibility of the evidence is calculated. An access evidence set is formed through access control, so that the access evidence set becomes the only evidence entry point for state tensor construction and root cause association graph generation.

[0015] (3) Based on the admission evidence set, construct the state tensor according to the cell identifier, slice identifier and time window index and generate the root cause association graph. Further obtain the action constraint library and generate a candidate action set carrying dependency and conflict labels under the constraints of the root cause association graph and the action constraint library. Combine the candidate action set to construct the shadow network state and generate the effect prediction vector through pre-drilling. Then, input the pre-drilling into the hierarchical deep reinforcement learning decision model through the action set, state tensor and effect prediction vector to output the action policy. In the network change execution driven by the action policy, generate the execution trajectory record to generate training samples, update the hierarchical deep reinforcement learning decision model and register the policy version.

[0016] The following are its main beneficial effects: (1) In response to the problems in the existing schemes, such as the scattered management scope of network topology, network element list and probe list, lack of unified binding of sampling granularity and field mapping table, and difficulty in unifying the unified clock reference through the telemetry data link, resulting in inconsistent entry points of telemetry data in subsequent processing, the network topology, network element list, probe list, sampling granularity, field mapping table and unified clock reference are solidified under the same session level scope through the session configuration package. This makes the acquisition, alignment and subsequent calculation of telemetry data refer to the same configuration baseline of the session configuration package, so that the data scope on which the subsequent access control and state tensor construction depend has a verifiable and consistent source. This is suitable for unified operation and maintenance scenarios across cell identifiers and cross slice identifiers.

[0017] (2) In response to the lack of a unified quality measurement standard for telemetry data in the existing scheme, which focuses on time synchronization deviation, missing measurement rate, noise jitter and field drift, and the lack of access control between the acquisition link and the decision link, which leads to abnormal data and low quality data directly entering the root cause analysis and decision input, a traceable evidence judgment record is formed by quality label and evidence credibility. The access control solidifies the telemetry data into access evidence set by segmenting it, so that the evidence source of the state tensor and the root cause association graph has a consistent access control entry. This allows the construction of nodes and edge weights of the root cause association graph to refer to the evidence segments under the same access evidence set standard. This is suitable for operation scenarios where telemetry data has jitter, missing measurement and field drift.

[0018] (3) In the existing solution, the root cause localization is disconnected from the network change orchestration, and the action constraint library is lacking and root cause association is not properly addressed. Figure 1 This paper addresses the problem of inconsistent action strategies across the operational chain due to the lack of structured constraints on dependencies and conflicts in the mapping links and candidate actions. It organizes the admission evidence set into a structured input usable for decision-making using state tensors and root cause association graphs. Under the constraints of the action constraint library, it generates dependency and conflict labels for the candidate action set. Furthermore, it uses shadow network states and pre-exercise to generate effect prediction vectors and filters the action set that passes the pre-exercise and execution protection conditions. The structured constraints are then injected into the hierarchical deep reinforcement learning decision model to output the action strategy. Simultaneously, during network change execution, it generates execution trajectory records and uses these records to generate training samples to update the hierarchical deep reinforcement learning decision model and register the strategy version. This allows the action strategy and execution trajectory records to form a verifiable closed-loop association under the same link, making it suitable for intelligent network operation and maintenance scenarios requiring continuous operation and maintenance decisions and iterative updates based on cell identifiers, slice identifiers, and time window indexes. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating a network intelligent operation and maintenance method based on deep reinforcement learning, provided as an embodiment of this application. Detailed Implementation

[0020] Example 1: Refer to Figure 1 This is a flowchart illustrating a network intelligent operation and maintenance method based on deep reinforcement learning provided in an embodiment of the present invention. The process may include at least steps S100-S700: S100: Obtain network topology, network element list, probe list, sampling granularity, field mapping table, and unified clock reference; generate session configuration package. S200: Collect telemetry data based on session configuration package, calculate time synchronization deviation, missing rate, noise jitter, and field drift, generate quality labels and calculate evidence credibility, and execute access control to generate access evidence set; S300: Based on the admission evidence set, construct a state tensor according to cell identifier, slice identifier, and time window index, and generate a root cause association graph; S400. Obtain the action constraint library, and generate a set of candidate actions carrying dependency and conflict markers based on the root cause relationship graph and the action constraint library; S500: Construct the shadow network state and perform a pre-play based on the candidate action set, generate the effect prediction vector, and filter the action set that passes the pre-play and the execution protection conditions; S600, based on pre-simulation, inputs action sets, state tensors, and effect prediction vectors into a hierarchical deep reinforcement learning decision model, and outputs action policies; S700: Based on the action policy, perform network changes and generate execution trajectory records, generate training samples based on the execution trajectory records, update the hierarchical deep reinforcement learning decision model and register the policy version.

[0021] To address the problems in existing technologies where network topology, network element lists, and probe lists are scattered across multiple systems and domains, and sampling granularity, field definitions, and clock references lack session-level unified solidification, leading to inconsistent primary key definitions and alignment of cross-domain telemetry data, and difficulty in tracing configuration sources, this invention obtains network topology, network element lists, probe lists, sampling granularity, field mapping tables, and unified clock references in step S100, generating a session configuration package. From the session initiation phase, it uniformly registers and solidifies the primary key definitions and alignment of cross-domain telemetry data. Specifically, this includes incorporating the primary key binding relationships of probes to cell identifiers, slice identifiers, and network element identifiers, along with field unit dictionaries, sampling alignment rules, and anomaly indicator calculation template versions, into the session configuration package, and simultaneously writing them into hash verification fields and session anomaly records. This ensures that the session configuration package can be called by subsequent steps using the same version entry point, thereby providing consistent configuration input for telemetry acquisition, quality label generation, and access control in S200. Specifically, this includes: S100: Obtain network topology, network element list, probe list, sampling granularity, field mapping table, and unified clock reference; generate session configuration package. Specifically, this step is executed during the network intelligent operation and maintenance session startup phase. Input sources include the network topology and network element list output by the network management system, the probe list output by the operation and maintenance orchestration system, and the sampling granularity, field mapping table, and unified clock reference bound to the operation and maintenance task. The network topology represents the structural connection relationship of the target network, including link associations, bearer hierarchy relationships, and coverage information between network element identifiers. The network element list represents the set of network elements participating in this operation and maintenance session, including network element identifier, network element type, network element address, domain identifier, and health status flag, where the domain identifier is used to distinguish between the wireless access network domain, core network domain, and transmission domain. The probe list represents the deployment and access configuration of the telemetry acquisition terminal, including probe identifier, probe type, attached network element identifier, acquisition interface identifier, acquisition item set identifier, and data reporting channel identifier. The probe type includes... The network element embeds telemetry probes, mirror port traffic probes, and service-side perception probes. The sampling granularity represents the sampling period and aggregation window of different acquisition items, including the sampling period value, aggregation window length, and window sliding step size. The field mapping table represents the unified caliber of cross-domain telemetry fields, including source field name, target field name, unit marker, dimensional marker, missing measurement marker rules, and discrete enumeration mapping rules. The unified clock reference represents the session-level time reference, including the reference clock source identifier, timestamp format, time synchronization mode marker, and allowable deviation threshold, used to drive the alignment and quality evaluation of subsequent telemetry data. This step uses the unified clock reference as the time anchor point to establish a session-level unified time axis and register the session start timestamp. The session start timestamp and session identifier together constitute the primary key caliber of this session, used for cross-step association.

[0022] Further, for the network topology and the network element list, this step performs network element capability registration and interface constraint registration processing. Specifically, the configuration management unit reads the telemetry capability template corresponding to the network element type from the network element management interface. The telemetry capability template includes a set identifier for collectable indicators, an upper limit for indicator sampling, an indicator reporting protocol identifier, and an identifier for the domain to which the indicator belongs. The configuration management unit then binds the telemetry capability template to the network element identifier in the network element list, generating a network element capability registration record. Simultaneously, the configuration management unit performs topology consistency verification on the network topology, verifying link bidirectional consistency, network element identifier uniqueness, and domain identifier closure. When verification fails, the reason for failure is written to the session exception record, and the corresponding network element identifier is written to the disabled network element set field of the session configuration packet. The disabled network element set field is fixed before the end of this step and does not participate in subsequent probe binding. For the probe list, this step performs probe number binding and deployment constraint registration processing. The probe number binding includes binding the probe identifier to the attached network element identifier and binding the collection interface identifier to the collection item set identifier, resulting in a probe binding table. The deployment constraint registration includes recording the deployment location marker between the probe and the network element, the upper limit of the access link bandwidth, the buffer queue length, and the upper limit of the reporting code rate. The deployment location marker is used to unify the regional aggregation caliber when performing cross-probe data association in subsequent steps. For conflicting probe bindings, the conflict type includes duplicate use of the same collection interface identifier and the same collection item set identifier exceeding the reporting code rate limit. In this step, the conflict entry is written to the session exception record and the conflict entry is written to the conflict flag field of the probe list. The conflict flag field is used for filtering in the subsequent session configuration packet distribution stage.

[0023] Further, this step links the sampling granularity with the field mapping table to form sampling alignment rules and unified field caliber rules. Specifically, the timing configuration unit generates time window index rules based on the sampling period value and the convergence window length. These time window index rules include the window sequence number generation method, the window start and end timestamp calculation method, and the window sliding step size. The time window index rules are then bound to the unified time axis to generate sampling alignment rules. The field configuration unit performs field validity checks on the field mapping table. The checks include the uniqueness of the target field name, the consistency of the unit and dimension labels, and the closure of the discrete enumeration mapping. When a check fails, the field configuration unit writes the failed entry to the session exception record and writes the corresponding source field name to the disabled field set field of the field mapping table. The disabled field set field is fixed before the end of this step. Subsequently, the field configuration unit generates a field standardization rule set based on the missing data marking rules and discrete enumeration mapping rules. This set includes missing data filling value marking, outlier value truncation rules, unit conversion rules, and enumeration normalization rules. This set, along with the field mapping table, is written into the session configuration package for use in subsequent telemetry data acquisition and quality label generation stages. Understandably, the field standardization rule set is one of the minimum sets supporting subsequent quality label calculation, and the time window index rules are one of the minimum sets supporting subsequent state tensor construction.

[0024] In an engineering implementation, the network management system issues a session initiation command when a specified maintenance task is triggered. This command includes a target area identifier and a target slice identifier. Upon receiving the session initiation command, the configuration management unit obtains telemetry capability templates from network elements in the wireless access network domain, session plane and user plane telemetry capability templates from network elements in the core network domain, and link congestion and packet loss-related telemetry capability templates from network elements in the transmission domain, and binds these templates to the network element list. The maintenance orchestration system outputs a probe list, specifying the network element identifier and acquisition interface identifier for each probe. The clock synchronization service outputs a unified clock reference and completes the registration of the reference clock source identifier during the session initiation phase. Subsequently, this step aggregates the above inputs after topology consistency verification, probe number binding, sampling alignment rule generation, and field standardization rule set generation to obtain a session configuration package. The session configuration package, in its data structure, includes a session identifier, session start timestamp, network topology summary, network element list, probe list, probe binding table, sampling granularity, sampling alignment rules, field mapping table, field standardization rule set, and unified clock reference, and includes a hash verification field for registering the integrity verification result of the session configuration package. The session configuration package is written into the configuration storage as the output of this step and is distributed to the acquisition nodes where each probe is located by the session scheduling unit. At the same time, the session configuration package serves as the input of S200, which is used by S200 to execute the processing link call for acquiring telemetry data based on the session configuration package and calculating the time synchronization deviation, missing rate, noise jitter, and field drift.

[0025] Summary of the technical effects of this step: This step completes the session-level unified registration of network topology, network element list, and probe list, and solidifies the sampling granularity, field mapping table, and unified clock reference into a session configuration package, fixing the primary key and alignment of cross-domain telemetry data from the source; This step incorporates sampling alignment rules and field standardization rule sets into the session configuration package, forming the input baseline for subsequent quality label and evidence credibility calculation; This step achieves traceable registration of the session configuration package through hash verification fields and session anomaly records, providing consistent configuration input for S200 access control.

[0026] To address the issue that while existing telemetry data acquisition technologies cover multiple probes and network elements, the lack of structured measurement and hierarchical handling mechanisms for quality issues such as time synchronization deviation, missing data, noise jitter, and field drift leads to low-quality data being mixed with high-quality data in the same processing link, resulting in unstable input for subsequent state construction and model training, this invention collects telemetry data based on the session configuration package in step S200, calculates time synchronization deviation, missing data rate, noise jitter, and field drift, generates quality labels and calculates evidence credibility, and performs access control to generate an access evidence set. Specifically, this includes: aligning the multi-probe telemetry stream with time windows and binding it to a primary key according to the sampling alignment rules in the session configuration package; after alignment, forming a quality label set around the out-of-order rate, repetition rate, sudden spike ratio, cross-domain time drift trend item, and the time synchronization deviation, missing detection rate, noise jitter, and field drift; and mapping the quality labels to evidence credibility records; then combining the evidence credibility records with the first threshold and second threshold grading rules to execute access control, simultaneously forming an access evidence set, a supplementary collection set, and a downgrade processing set; and using the access evidence set as the input caliber of S300 to solidify the evidence baseline across probes and across domains. S200: Collect telemetry data based on session configuration package, calculate time synchronization deviation, missing rate, noise jitter, and field drift, generate quality labels and calculate evidence credibility, and execute access control to generate access evidence set; Specifically, this step uses the session configuration package generated in S100 as input, and is triggered by the acquisition scheduling unit after the session start timestamp arrives. Under the constraints of the unified time axis, it drives the acquisition nodes corresponding to the probe list to carry out telemetry data acquisition. The telemetry data is the raw data stream or aggregated data stream acquired by the acquisition interface identifier corresponding to each probe identifier in the probe list. Its data structure carries the session identifier, probe identifier, connected network element identifier, source field name, and sampling timestamp; wherein the sampling timestamp follows the timestamp format specified by the unified clock reference. The acquisition scheduling unit distributes the sampling alignment rules in the session configuration package to each acquisition node, and distributes the field standardization rule set and the field mapping table to the data access unit. The data access unit performs field caliber unification processing on the telemetry data returned. The processing link includes mapping from source field name to target field name, conversion between unit marker and dimension marker, normalization of discrete enumeration mapping rules, and filling annotation of missing measurement marker rules to obtain structured telemetry segments. The structured telemetry segments carry target field name, unit marker, missing measurement marker, sampling timestamp and probe identifier at the field level, and write the attached network element identifier as the association primary key into the primary key field of the structured telemetry segments.

[0027] Further, this step performs time synchronization deviation calculation and missing rate calculation on the structured telemetry segments. The time synchronization deviation calculation is performed by the timing verification unit. This unit reads the unified clock reference and the sampling alignment rules from the session configuration package, archives the sampling timestamps of the structured telemetry segments according to the unified timeline, binds each structured telemetry segment to its corresponding time window index, and calculates the deviation between the cross-probe sampling timestamp and the window reference timestamp within the same time window index, forming a time synchronization deviation record. This time synchronization deviation record carries the probe identifier, time window index, deviation amount, and a deviation statistical summary. The missing rate calculation is performed by the missing rate analysis unit. This unit generates the expected sampling points for the target field name according to the sampling alignment rules, compares the missing rate markers in the structured telemetry segments with the expected sampling points, and outputs a missing rate record. This missing rate record carries the probe identifier, target field name, time window index, and missing rate ratio. For structured telemetry segments that are out of order, the timing verification unit rearranges them according to the unified time axis and records the out-of-order rate, which is written into the subsequent quality label; for structured telemetry segments that are duplicated, the missing detection analysis unit performs deduplication under the conditions of the same probe identifier and the same sampling timestamp and records the duplication rate, which is written into the subsequent quality label.

[0028] Further, this step performs noise jitter calculation and field drift calculation on the structured telemetry slices. The noise jitter calculation is completed by the fluctuation analysis unit. Under the constraints of the same probe identifier and the same target field name, the fluctuation analysis unit differs the numerical sequences within adjacent time window indices and combines them with the statistical summary within the window to generate a fluctuation amplitude sequence, thereby obtaining a noise jitter record. The noise jitter record carries the probe identifier, target field name, time window index, and jitter amplitude summary, and also records the proportion of sudden spikes. The proportion of sudden spikes is obtained from the percentage of samples exceeding a preset truncation rule within the window. The truncation rule is taken from the field standardization rule set. Field drift calculation is performed by the calibrator consistency unit. Based on the field mapping table and the field unit dictionary, the calibrator consistency unit monitors the statistical distribution of the target field name. Under the constraint of the same attached network element identifier and the same target field name, it compares the distribution of the statistical summary in the current time window index with the statistical summary in the historical time window index, and outputs a field drift record. This field drift record carries the attached network element identifier, target field name, time window index, and a drift degree summary. A cross-domain time drift trend item is written into the quality label. This cross-domain time drift trend item is obtained from the changing trend of time synchronization deviation records under different domain identifiers on the continuous time window index. Understandably, time synchronization deviation, missing detection rate, noise jitter, and field drift constitute the minimum set of quality labels generated in this step. Out-of-order rate, repetition rate, sudden spike ratio, and cross-domain time drift trend item are quality label components and together with the minimum set form quality label entries.

[0029] During the quality tag generation phase, the quality assessment unit aggregates time synchronization deviation records, missing test rate records, noise jitter records, and field drift records according to probe identifiers, target field names, and time window indexes to generate a quality tag set. The quality tag set uses quality tag entries as the smallest granularity. Each quality tag entry carries a probe identifier, a connected network element identifier, a target field name, a time window index, a time synchronization deviation summary, a missing test rate, a jitter amplitude summary, a drift degree summary, an out-of-order rate, a duplication rate, a sudden spike rate, and a cross-domain time drift trend item. Subsequently, the evidence scoring unit calculates the evidence credibility based on the quality label set. The evidence credibility is obtained by mapping the quality label entries. The mapping process adopts a tiered normalization and weighted synthesis approach. Tiered normalization tiers the time synchronization deviation records according to the allowable deviation threshold in the unified clock reference, tiers the missing test ratio according to the missing test marking rule, and tiers the sudden peak ratio according to the abnormal value truncation rule in the field standardization rule set. Weighted synthesis generates scores in the probe identifier dimension, target field name dimension, and time window index dimension respectively, and synthesizes the evidence credibility score. The evidence credibility score is written into the evidence credibility record. The evidence credibility record carries the probe identifier, target field name, time window index, and evidence credibility score, and establishes an association index with the quality label entries.

[0030] Further, this step executes access control to generate an access evidence set. Access control is performed by a gating decision unit. This unit reads the evidence credibility record and invokes the first and second threshold grading rules to perform gating judgment on each evidence credibility score. Structured telemetry fragments with credibility scores below the first threshold are written into the supplementary acquisition set; structured telemetry fragments with credibility scores between the first and second thresholds are written into the downgrade processing set; and structured telemetry fragments with credibility scores not lower than the second threshold are written into the access evidence set. The supplementary acquisition set carries probe identifiers, target field names, time window indexes, and missing data marker summaries, and is written into the session anomaly record for subsequent supplementary acquisition scheduling. The downgrade processing set carries probe identifiers, target field names, time window indexes, and downgrade markers, and is removed during the training sample construction process. The access evidence set carries session identifiers, probe identifiers, connected network element identifiers, target field names, time window indexes, standardized numerical sequences, quality label entry references, and evidence credibility scores, and is written into the evidence storage as the output of this step. The admission evidence set is used as input by S300 in the step connection, so that S300 can construct a state tensor and generate a root cause association graph based on the admission evidence set according to cell identifier, slice identifier and time window index.

[0031] In an engineering implementation, when the session start timestamp arrives, the data collection scheduling unit sends a data collection task to the probe nodes in the wireless access network domain. The data collection item set includes target field names related to wireless load and target field names related to interference. It also sends a data collection task to the probe nodes in the core network domain. The data collection item set includes target field names related to session establishment and target field names related to packet loss. Finally, it sends a data collection task to the probe nodes in the transmission domain. The data collection item set includes target field names related to congestion. Each probe node transmits structured telemetry fragments back based on the sampling period value. The data access unit maps the source field names to the target field names based on the field mapping table and the field standardization rule set. The time sequence verification unit generates a time window index based on a unified time axis and calculates the time synchronization deviation record. The missing detection analysis unit generates a missing detection rate record and writes the repetition rate. The fluctuation analysis unit generates a noise jitter record and writes the sudden peak ratio. The caliber consistency unit generates a field drift record and writes the cross-domain time drift trend item. The quality assessment unit aggregates and generates a quality label set, and the evidence scoring unit generates an evidence credibility record. The gating judgment unit outputs the admission evidence set according to the first threshold and the second threshold, and simultaneously outputs the supplementary collection set and the downgraded processing set. Finally, the admission evidence set is written to the evidence storage and made available for S300 to call.

[0032] In summary, the technical effects of this step are as follows: This step implements the session configuration package into a unified operational link for telemetry data acquisition and caliber standardization. It generates a quality label set and establishes an evidence credibility record based on time synchronization deviation, missing measurement rate, noise jitter, and field drift. This step combines the evidence credibility record with the first and second threshold grading rules to implement access control, forming an access evidence set and simultaneously creating a set of supplementary data acquisition and downgrade processing. This step uses the access evidence set as the input caliber of the S300 to solidify the evidence baseline across probes and domains, providing a consistent data entry point for subsequent state tensor construction and root cause correlation graph generation.

[0033] To address the problems in existing technologies where fault manifestations across wireless access network domains, core network domains, and transmission domains are distributed across different indicator systems, lacking a unified state representation organized by cell identifier, slice identifier, and time window index, and where the temporal sequence and correlation strength between anomalies are difficult to express in a structured manner, leading to root cause localization relying on manual experience and difficulty in forming calculable correlation constraints, this invention constructs a state tensor and generates a root cause correlation graph based on the admission evidence set according to cell identifier, slice identifier, and time window index in step S300. Specifically, this includes: aggregating the admission evidence set by primary key according to the cell identifier and slice identifier, and aligning it according to the time window index to form a state tensor. The state tensor at least includes wireless side load characteristics, interference characteristics, slice-level throughput characteristics, latency characteristics, core network session establishment characteristics, packet loss characteristics, and transmission side congestion characteristics; extracting anomaly candidate sets from the state tensor and mapping them to node sets; generating edge weights by combining the temporal sequence and correlation strength of cross-domain indicators to form a root cause correlation graph; and using the root cause correlation graph as the structured entry point for generating dependency and conflict markers in step S400. Specifically, this includes: S300: Based on the admission evidence set, construct a state tensor according to cell identifier, slice identifier, and time window index, and generate a root cause association graph; Specifically, this step uses the admission evidence set output by S200 as input. It is triggered by the evidence orchestration unit upon receiving the evidence storage write completion event, and references the network topology, network element list, probe binding table, and sampling alignment rules in the session configuration package as co-location configuration items in the processing link. The record granularity of the admission evidence set is structured telemetry fragmentation, carrying session identifier, probe identifier, attached network element identifier, target field name, time window index, standardized numerical sequence, quality label entry reference, and evidence credibility score. This step first performs network domain normalization and identifier completion processing, specifically completed by the identifier resolution unit: the identifier resolution unit reads the probe binding table, maps the probe identifier to the attached network element identifier and writes it into the domain identifier. Simultaneously, based on the coverage relationship information and bearer hierarchical relationship information in the network topology, it parses the service unit identifier corresponding to the attached network element identifier and generates a cell identifier; and based on the service bearer mapping relationship in the network element list, it links and parses the attached network element identifier and service unit identifier to obtain the slice identifier. For structured telemetry fragments for which cell or slice identifiers cannot be parsed, the identifier parsing unit adds them to the anomaly candidate set and registers the reason for non-parsing. Simultaneously, the structured telemetry fragment is marked as unarrangeable and does not proceed to the subsequent state tensor construction process. Understandably, the cell identifier and slice identifier are the minimum set identifier fields for constructing the state tensor in this step, and the time window index is the minimum set index field for cross-domain alignment in this step.

[0034] Further, this step constructs a state tensor based on cell identifier, slice identifier, and time window index. The state tensor construction is performed by a state tensor construction unit. Under the same session identifier constraint, this unit performs window aggregation on structured telemetry slices according to the time window index, and performs secondary grouping within each time window index according to cell identifier and slice identifier, generating an index skeleton for the state tensor. Subsequently, the state tensor construction unit maps the structured telemetry slices to the feature channels of the state tensor based on the target field name and domain identifier. The mapping process is driven by a feature caliber template, which is taken from the abnormal indicator calculation template version in the session configuration package and is consistent with the field standardization rule set. In the specific implementation process, within the radio access network domain, the state tensor construction unit extracts a sequence of target field names related to radio-side load characteristics from the admission evidence set and aggregates them according to a time window index to obtain a radio-side load feature vector. Simultaneously, it extracts a sequence of target field names related to interference characteristics and aggregates them according to a time window index to obtain an interference feature vector. Within the core network domain, the state tensor construction unit extracts a sequence of target field names related to session establishment characteristics and aggregates them according to a time window index to obtain a core network session establishment feature vector. Simultaneously, it extracts a sequence of target field names related to packet loss characteristics and aggregates them according to a time window index to obtain a packet loss feature vector. Within the transport domain, the state tensor construction unit extracts a sequence of target field names related to congestion and aggregates them according to a time window index to obtain a transport-side congestion feature vector. The slice-level throughput and latency features are generated by the slice aggregation unit. The slice aggregation unit uses the slice identifier as the primary key and aggregates the target field name sequences related to throughput and latency from the radio access network domain, core network domain, and transport domain within the same time window index to obtain slice-level throughput and slice-level latency feature vectors. For duplicate features reported by multiple probes under the same cell identifier, slice identifier, and time window index, the state tensor construction unit performs weighted fusion based on the evidence credibility score and records the participating probe identifier set and weight allocation record in the fusion metadata. For feature channels with missing test markers, the state tensor construction unit generates a missing test compensation marker by referencing the missing test ratio and jitter amplitude summary in the quality label entry, and binds the missing test compensation marker to the feature channel and writes it into the quality side information field of the state tensor. Thus, the state tensor contains three types of index fields: cell identifier, slice identifier, and time window index, and includes feature channels such as radio-side load features, interference features, slice-level throughput features, latency features, core network session establishment features, packet loss features, and transmission-side congestion features, as well as evidence credibility side information and quality side information. The state tensor is written into the state storage as a stage product of this step and is used for subsequent generation of root cause association graphs in step transitions. It also serves as one of the inputs of S600 in cross-step transitions, allowing S600 to input the state tensor into the hierarchical deep reinforcement learning decision model.

[0035] Further, this step generates a root cause correlation graph. The root cause correlation graph generation is completed by the root cause graph construction unit, which performs anomaly candidate set construction and causal correlation edge weight calculation processing based on the state tensor. Specifically, the anomaly candidate set construction unit extracts the time window index sequence of feature channels from the state tensor, calls the anomaly indicator calculation template version to calculate deviation summaries and trend summaries within the window for each feature channel, and writes feature channel instances that meet the trigger conditions in the anomaly indicator calculation template version into the anomaly candidate set. The entries in the anomaly candidate set carry cell identifier, slice identifier, time window index, domain identifier, target field name reference, deviation summary, trend summary, and evidence credibility information. Subsequently, the root cause graph construction unit maps the anomaly candidate set to a node set. Nodes in the node set carry node identifier, domain identifier, attached network element identifier reference, cell identifier, slice identifier, time window index, and indicator summary reference. The node identifier is generated and registered by the session identifier in conjunction with the cell identifier, slice identifier, and time window index. In the edge weight calculation stage, the root cause graph construction unit generates edge weights and forms a root cause association graph based on the temporal order and correlation strength of cross-domain indicators. The temporal order is obtained by comparing the time window index sequences of anomaly candidate set entries under different domain identifiers, and the candidate directions of upstream and downstream domain identifiers are determined by combining the bearer hierarchy relationship information in the network topology. The correlation strength is calculated by the correlation calculation unit. Under the constraints of the same cell identifier and the same slice identifier, the correlation calculation unit aligns the deviation summary sequence and trend summary sequence of anomaly candidate set entries to obtain a correlation strength summary. Simultaneously, the coupling adjacency relationship of cross-cell interference coupling items is introduced from the coverage relationship information of the network topology to generate a correlation strength summary between cross-cell nodes. For edges with conflict directions, the root cause graph construction unit writes them into the conflict edge set and records the conflict cause. The conflict edge set serves as one of the sources of dependency and conflict markers in the subsequent action constraint library matching stage. For nodes with low evidence credibility information, the root cause graph construction unit performs weight reduction processing based on the evidence credibility score and writes the weight reduction record into the edge element information of the root cause association graph. Finally, the root cause association graph contains a set of nodes, a set of edges, and edge element information. The edge element information carries edge weights, temporal order markers, relevance strength summaries, and weight reduction records. The set of conflicting edges is written into the conflict record field of the root cause association graph.

[0036] In the engineering implementation, for a target network comprising a radio access network domain, a core network domain, and a transport domain, the evidence orchestration unit triggers a state tensor construction once after each time window index reaches the window termination timestamp. The identifier resolution unit maps probe identifiers to attached network element identifiers based on the probe binding table and generates domain identifiers. Simultaneously, it associates attached network element identifiers with cell identifiers based on the coverage relationship information of the network topology and generates slice identifiers based on the service bearer mapping relationship of the network element list. The state tensor construction unit extracts radio-side load characteristics and interference characteristics in the radio access network domain and extracts session establishment characteristics and packet loss characteristics in the core network domain. The system extracts congestion-related features in the transmission domain, and the slice aggregation unit generates slice-level throughput and latency features, forming a state tensor with an indexed skeleton and writing it into the state storage. The anomaly candidate set construction unit calculates deviation and trend summaries within the calculation window based on the anomaly indicator calculation template version, filters out anomaly candidate sets, and maps them to node sets. The root cause graph construction unit generates edge weights and forms a root cause association graph based on the time window index sequence of cross-domain indicators, the bearer hierarchy information of network topology, and the correlation strength summary output by the correlation calculation unit, and writes the conflict edge set into the conflict record field of the root cause association graph. The root cause association graph is written into the root cause graph storage as the output product of this step, and serves as the input of S400 in the cross-step connection, allowing S400 to generate a set of candidate actions carrying dependency and conflict markers based on the root cause association graph and the action constraint library.

[0037] In summary, the technical effects of this step are as follows: This step arranges the admission evidence set under the cell identifier, slice identifier, and time window index and constructs a state tensor to form a consistent input structure across the radio access network domain, core network domain, and transmission domain; this step constructs an anomaly candidate set based on the state tensor and maps it to a node set, and generates edge weights by combining the temporal order and correlation strength of cross-domain indicators to form a root cause association graph; this step uses the root cause association graph as the input caliber of S400 to solidify dependency markers and conflict records, providing structured root cause constraints for the subsequent generation of candidate action sets.

[0038] To address the problem that existing technologies typically rely on static rule bases or single-point alarm associations for selecting operational actions, lacking a traceable mapping link between "root cause nodes and action templates," and where parameter mutual exclusion and dependency order between actions often depend on manual verification, making it difficult to achieve consistency verification and pre-fixed executability constraints for candidate action sets under the same session caliber, this invention obtains an action constraint library in step S400. Based on the root cause association graph and the action constraint library, a candidate action set carrying dependency and conflict markers is generated. Specifically, this includes: loading the action constraint library and performing version verification; the action constraint library at least includes slice service level constraints, capacity constraints, coverage constraints, change frequency constraints, parameter mutual exclusion tables, and dependency order tables; matching root cause nodes in the root cause association graph with action templates to generate action candidates bound to the root cause nodes; and generating conflict and dependency markers for each action candidate under the constraints of the parameter mutual exclusion table and dependency order table, while writing sorting fields and constraint reference records to form a candidate action set. The structured markers of the candidate action set are then used as a unified input for S500 shadow network status pre-simulation and execution protection condition screening. Specifically, this includes: S400. Obtain the action constraint library, and generate a set of candidate actions carrying dependency and conflict markers based on the root cause relationship graph and the action constraint library; Specifically, this step uses the root cause correlation graph output by S300 as input. It is triggered by the policy orchestration unit after the root cause graph storage write completion event arrives, and simultaneously references the network topology and network element list in the session configuration package generated by S100 as the structured basis for action constraint parsing and action template matching. The root cause correlation graph includes a node set, an edge set, edge element information, and conflict record fields. The nodes in the node set carry domain identifiers, attached network element identifier references, cell identifiers, slice identifiers, and time window indexes. Meanwhile, the edge element information carries temporal sequence markers and correlation strength summaries. This step first obtains the action constraint library, which is read from the policy repository and loaded into session memory by the constraint library management unit. The minimum set of the action constraint library consists of slice service level constraints, capacity constraints, coverage constraints, change frequency constraints, parameter mutual exclusion tables, and dependency order tables. Among them, slice service level constraints are used to describe the mapping relationship between the service level entries and indicator calibers corresponding to slice identifiers; capacity constraints are used to describe the available resource limits and reserved resource entries of network element identifiers and cell identifiers under different time window indices; coverage constraints are used to describe the neighbor cell association and overlapping coverage boundary entries of cell identifiers in network topology coverage relationship information; change frequency constraints are used to describe the upper limit of the number of changes of attached network element identifiers and cell identifiers on continuous time window indices and cooling window entries; parameter mutual exclusion tables are used to describe the combination of parameter items that cannot appear in parallel under the same network element identifier and mutual exclusion condition entries; and dependency order tables are used to describe the sequential dependency relationship between action segments and dependency satisfaction condition entries. Furthermore, the constraint library management unit performs version verification and scope alignment processing on the action constraint library. Specifically, each constraint entry in the action constraint library carries a constraint library version number and an effective scope field. Version verification uses the hash digest registered in the hash verification field of the session configuration package for consistency comparison, and the comparison result is written to the constraint loading record. For constraint entries that fail to match, the constraint library management unit writes them to the constraint exception set and records the reason for the exception. At the same time, it marks the constraint entry as unavailable and does not enter the candidate action set generation process. Understandably, the parameter mutual exclusion table and dependency priority table in the action constraint library belong to the minimum set of constraint table entries for subsequent generation of dependency markers and conflict markers, and the slice service level constraint and change frequency constraint belong to the minimum set of constraint table entries for subsequent screening of candidate action sets.

[0039] Further, this step generates a candidate action set based on the root cause association graph and the action constraint library. The generation of the candidate action set is collaboratively completed by an action template matching unit, a conflict resolution unit, and a dependency orchestration unit. The action template matching unit establishes a mapping relationship between root cause nodes and action templates; the conflict resolution unit generates conflict markers and performs conflict pruning by combining a parameter mutual exclusion table and a conflict record field; and the dependency orchestration unit generates dependency markers and sorts action segments by combining a dependency priority table. Specifically, the action template matching unit extracts a root cause node set from the node set of the root cause association graph. The determination of the root cause node set uses the temporal sequence markers and correlation strength summaries in the edge cell information written by the root cause graph construction unit. Specifically, nodes located upstream of the temporal sequence markers under the same cell identifier and the same slice identifier, and whose correlation strength summaries meet the threshold conditions, are selected as root cause nodes. The threshold conditions and participating node identifiers of the selection process are recorded as root cause selection records. Subsequently, the action template matching unit performs domain-specific action template selection based on the domain identifier of the root cause node. Action templates are stored in the action template library and grouped by domain identifier. Each action template carries an action template identifier, applicable domain identifier, applicable network element type, parameter set, precondition entries, and output action fragment type. The action template matching unit inputs the attached network element identifier reference of the root cause node into the network element list, parses it to obtain the network element type, and matches the network element type with the domain identifier against the applicable network element type and applicable domain identifier in the action template library to obtain a candidate action template set. For root cause nodes that do not match an action template, the action template matching unit writes it into the unmatched root cause set and records the reason for the unmatch; simultaneously, the root cause node does not enter the candidate action set. For the matched candidate action template set, the action template matching unit further loads slice service level constraints and capacity constraints into the template instantiation context. Specifically, the slice service level constraint is bound to the slice identifier in the root cause node, and the capacity constraint is bound to the cell identifier and attached network element identifier reference in the root cause node. Constraint adaptation is performed on the parameter set of the action template to generate an action template instance. The action template instance carries an action template identifier, a target cell identifier, a target slice identifier, a target attached network element identifier reference, a set of parameter items, and a reference to precondition items, and records the instantiation process as a template instantiation record.

[0040] Further, this step forms a set of change fragments and assembles them into a set of candidate actions. Specifically, the action template matching unit converts the action template instance into a set of change fragments. The set of change fragments includes radio-side parameter threshold change fragments, slice resource weight change fragments, core network traffic offloading change fragments, core network migration change fragments, and cross-domain rollback change fragments. The radio-side parameter threshold change fragment is generated by a radio access network domain action template instance, and its parameter set points to the radio-side configuration parameter threshold entry and is bound to the cell identifier. The slice resource weight change fragment is generated by a slice domain action template instance, and its parameter set points to the slice resource weight entry and is bound to the slice identifier. The core network traffic offloading change fragment and the core network migration change fragment are generated by a core network domain action template instance. The parameter set of the core network traffic offloading change fragment points to the traffic offloading rule entry and is bound to the slice identifier, while the parameter set of the core network migration change fragment points to the migration path entry and is bound to the attached network element identifier reference. The cross-domain rollback change fragment is generated by a cross-domain action template instance, and its parameter set points to the rollback point entry and is bound to the conflict record field of the root cause association graph. After the change fragment set is generated, the conflict resolution unit inputs the parameter mutual exclusion table into the conflict detection process. Specifically, the conflict resolution unit extracts the parameter item set from the change fragment set and constructs a parameter item conflict index. The parameter item conflict index uses the attached network element identifier reference and the time window index as a combined primary key to map the parameter item set to candidate parameter item pairs. Then, the parameter mutual exclusion table is called to compare the mutual exclusion conditions of the candidate parameter item pairs to obtain mutual exclusion hit records, which are then written into the conflict marker field. The conflict resolution unit also reads the conflict record field of the root cause association graph, maps the set of conflict edges in the conflict record field to cross-domain conflict constraints, and merges the cross-domain conflict constraints with the mutual exclusion hit records to form a conflict list. For conflict entries in the conflict list, the conflict resolution unit performs conflict pruning. Conflict pruning is driven by priority rules, which are jointly determined by slice service level constraints and relevance strength summaries. Specifically, change fragments corresponding to high-priority slice identifiers indicated by slice service level constraints are retained, while change fragments with lower relevance strength summaries are pruned and written into the conflict pruning record. Simultaneously, the pruned change fragments are marked as conflict-removed in the candidate action set. After conflict pruning, the dependency orchestration unit reads the dependency priority table and performs dependency resolution and dependency sorting on the remaining change fragment set. Specifically, the dependency orchestration unit maps the fragment type, target attached network element identifier reference, and parameter item set of the change fragment set to the dependency satisfaction condition entries in the dependency priority table, generating dependency hit records and writing them into the dependency tag field. Subsequently, based on the priority dependencies in the dependency priority table, a dependency sorting sequence is generated for the change fragment set, and the dependency sorting sequence is written into the sorting field of the candidate action set.The orchestration unit also uses frequency change constraints to perform frequency change constraint checks on the candidate action set. Specifically, it counts the number of candidate change segments for the same cell identifier within the continuous time window index and compares them with the cooling window entries. Candidate change segments that exceed the upper limit are marked as frequency over-limit and written into the downgrade processing set reference field. At the same time, the candidate change segment is still retained in the candidate action set but carries the frequency over-limit mark for subsequent S500 rehearsal and protection condition screening.

[0041] In an engineering implementation, for intelligent operation and maintenance scenarios of multi-domain networks of operators, the policy orchestration unit triggers this step after the root cause association graph is written in each time window index. The constraint library management unit reads the action constraint library from the policy repository and completes version verification, loading the slice service level constraints, capacity constraints, coverage constraints, change frequency constraints, parameter mutual exclusion table and dependency priority table into the session memory. The action template matching unit filters root cause nodes from the node set of the root cause association graph, matches action templates from the action template library according to the domain identifier and network element type and instantiates them to generate wireless side parameter threshold change fragments and slices. The system comprises a set of change segments, including resource weight change segments, core network traffic splitting change segments, core network migration change segments, and cross-domain rollback change segments. A conflict resolution unit generates a conflict list and performs conflict trimming based on the conflict record fields of the parameter mutual exclusion table and root cause correlation diagram, writing conflict flag fields and conflict trimming records into the candidate action set. A dependency orchestration unit generates a dependency sorting sequence based on the dependency priority table and writes it into the dependency flag field and sorting field. Simultaneously, it writes a frequency exceedance flag and a degradation processing set reference field based on change frequency constraints, forming a candidate action set carrying dependency and conflict flags. This candidate action set is written into the action candidate storage and serves as input to S500 during cross-step transitions, allowing S500 to construct and rehearse the shadow network state based on the candidate action set.

[0042] In summary, the technical effects of this step are as follows: This step loads, verifies, and matches the root cause relationship graph and the action constraint library under the same session caliber, establishing a traceable mapping chain between root cause nodes and action templates; This step generates conflict markers and dependency markers under the constraints of parameter mutual exclusion tables and dependency priority tables, forming a set of candidate actions with sorting fields; This step pre-solidifies the executability constraints of the candidate action set into structured markers and records, providing a unified input for subsequent shadow network state pre-simulation and execution protection condition screening.

[0043] In response to the problems in existing technologies where operational changes are often executed directly on the real network, lacking replayable session-level state representation and auditable pre-execution trajectory records, making it difficult to obtain traceable input of "action consequences" in the decision-making process, and lacking a protection condition screening mechanism based on risk and rollback costs before execution, this invention constructs a shadow network state and performs a pre-execution based on the candidate action set in step S500, generates an effect prediction vector, and filters the action set that passes the pre-execution and the execution protection conditions. Specifically, this includes: The shadow state construction unit extracts a subset of key KPIs and a subset of key KQIs from the state tensor and combines them with cross-cell interference coupling terms and congestion propagation approximations to form a shadow network state; the candidate action set is transformed into an executable set of change fragments and pre-rendered on the shadow network state, generating a pre-rendering trajectory record bound to the action identifier; an effect prediction vector is generated on the pre-rendering trajectory record, the effect prediction vector at least including service level default probability, oscillation risk index, and rollback cost prediction, and the evidence index field is simultaneously solidified; execution protection conditions are derived based on the effect prediction vector and the action constraint library, and a set of actions that pass the pre-rendering is selected, so that the set of actions that pass the pre-rendering, the effect prediction vector, and the execution protection conditions serve as the basic input for constraint injection in S600, specifically including: S500: Construct the shadow network state and perform a pre-play based on the candidate action set, generate the effect prediction vector, and filter the action set that passes the pre-play and the execution protection conditions; Specifically, this step takes the candidate action set output by S400 and uses the state tensor output by S300 and the network topology and network element list output by S100 as input sources for shadow state construction and pre-playing orchestration. This step is executed collaboratively by the shadow state construction unit, the pre-playing orchestration unit, the effect evaluation unit, and the protection condition generation unit. The shadow state construction unit is responsible for compressing the running representation of the real network corresponding to the time window index into the shadow network state. The pre-playing orchestration unit is responsible for mapping the candidate action set into a replayable change sequence and driving the shadow state evolution. The effect evaluation unit is responsible for calculating the effect prediction vector on the pre-playing trajectory. The protection condition generation unit is responsible for extracting the execution protection conditions from the pre-playing trajectory and constraint records and writing them into the action candidate storage. The shadow network state is a session-level data structure. Its primary key includes cell identifier, slice identifier, time window index, and reference to the attached network element identifier. The field set of the shadow network state includes a subset of key performance indicators (KPIs), a subset of key service experience indicators (KSIs), a cross-cell interference coupling term, and a congestion propagation approximation term. The KPI subset is mapped from the KPI's metric definition. The full English names for KPIs are Key Performance Indicators (KPIs) and Key Service Experience Indicators (KSIs). The cross-cell interference coupling term characterizes the coupling relationship of interference propagation between neighboring cells, and the congestion propagation approximation term characterizes the approximate propagation relationship of transmission-side congestion on adjacent links and adjacent time window indices. To adapt to the wireless communication network operation and maintenance process, this step uses a joint index management for cell identifiers and slice identifiers in the shadow network state, and uses the neighbor cell relationships and bearer link relationships obtained from the network topology parsing as the basis for constructing the cross-cell interference coupling term and the congestion propagation approximation term.

[0044] During the shadow network state construction phase, the shadow state construction unit extracts radio-side load features, interference features, slice-level throughput features, latency features, core network session establishment features, packet loss features, and transmission-side congestion features aligned with the time window index from the state tensor. Based on the field mapping table and field unit dictionary in the session configuration packet, it performs caliber normalization and unit normalization on each feature to obtain the shadow state input slice. Subsequently, the shadow state construction unit performs indicator projection and coupling term assembly processing. Indicator projection maps the shadow state input slice to the field sets of key performance indicator subsets and key service experience indicator subsets through indicator caliber mapping, and writes the mapping process into the indicator projection record. Coupling term assembly extracts neighboring cell identifier pairs through the neighbor cell relationship of the network topology and generates cross-cell interference coupling terms by combining them with the interference features in the state tensor. Simultaneously, it extracts neighboring link pairs through the bearer link relationship of the network topology and generates congestion propagation approximation terms by combining them with the transmission-side congestion features in the state tensor. For missing test fields, the shadow state construction unit calls the missing test rate record corresponding to the quality label in the admission evidence set output by S200, performs missing test field annotation, and writes it into the shadow state missing test mark field; for fields with high noise jitter, the shadow state construction unit calls the noise jitter record in the admission evidence set, performs weight reduction annotation, and writes it into the shadow state weight reduction mark field, thereby forming the shadow network state carrying traceable marks, and writes the shadow network state into the shadow state storage as the initial state input of the pre-playing orchestration unit.

[0045] During the rehearsal phase, the rehearsal orchestration unit performs dependency sorting sequence parsing and conflict marker parsing on the candidate action set, and associates the parsing results with the sorting field, conflict marker field, and frequency exceedance marker field in the candidate action set to generate a rehearsal action queue. The rehearsal action queue employs a playback step-by-step mechanism within a time window index. Each playback step is bound to a cell identifier, slice identifier, and attached network element identifier reference. A rehearsal step record is generated at the start of each playback step, containing action candidate identifiers, action template identifiers, changed segment types, and a summary of parameter item sets. The rehearsal orchestration unit applies the rehearsal action queue to the shadow network state to obtain the shadow state evolution trajectory, which includes a step sequence, a shadow state snapshot before the step, a shadow state snapshot after the step, conflict resolution records, and dependency satisfaction records. For radio-side parameter threshold change segments, the pre-drill orchestration unit locates the radio-side load characteristic and interference characteristic fields of the target cell identifier in the shadow state snapshot, and updates the corresponding threshold mapping entries according to the parameter item set summary; for slice resource weight change segments, the pre-drill orchestration unit locates the slice-level throughput characteristic and latency characteristic fields of the target slice identifier in the shadow state snapshot, and updates the resource weight mapping entries according to the parameter item set summary; for core network traffic offloading change segments and core network migration change segments, the pre-drill orchestration unit locates the core network session establishment characteristic and packet loss characteristic fields referenced by the attached network element identifier in the shadow state snapshot, and updates the traffic offloading rule mapping entries or migration path mapping entries according to the parameter item set summary; for cross-domain rollback change segments, the pre-drill orchestration unit calls the conflict handling record to locate the rollback point entry, rolls back the shadow state snapshot to the shadow state snapshot version corresponding to the rollback point, and writes the rollback trigger reason into the rollback trigger record. During the rehearsal, the rehearsal arrangement unit executes a rehearsal downgrade path for candidate actions carrying frequency exceeding the limit in the candidate action set, lengthens their step interval and writes it into the downgrade step record. The downgrade step record is used for elimination or weighted reference when constructing subsequent training samples.

[0046] In the effect prediction vector generation and screening stage, the effect evaluation unit calculates the effect prediction vector on the shadow state evolution trajectory. The effect prediction vector belongs to the action candidate level data structure, and its field set includes service level default probability, oscillation risk index, and rollback cost prediction. The input for calculating the service level default probability comes from the step-after shadow state snapshot of the subset of key business experience indicators in the shadow state evolution trajectory, and is compared with the threshold of the slice service level constraints in the action constraint library to form a default event sequence and write it into the default event record. The service level default probability is obtained by aggregating the default event records. The input for calculating the oscillation risk index comes from the step-before and after difference sequence of the subset of key performance indicators in the shadow state evolution trajectory, and is combined with the cooling window entries of the change segment type and change frequency constraints in the candidate action set to generate oscillation candidate events, and the oscillation candidate events are written into the oscillation event record. The oscillation risk index is obtained by aggregating the oscillation event records. The input for calculating the rollback cost prediction comes from the rollback trigger record, rollback point entry, and conflict handling record in the shadow state evolution trajectory, and is combined with the bearer link relationship of the network topology to generate cost decomposition records. The rollback cost prediction is obtained by aggregating the cost decomposition records. Subsequently, the effect evaluation unit binds the effect prediction vector with the action candidate identifier and writes it into the effect vector storage. Simultaneously, it writes the index of key evidence fragments in the shadow state evolution trajectory into the vector evidence index field, which is used for subsequent policy version audits. The protection condition generation unit generates the execution protection conditions based on default event records, oscillation event records, rollback trigger records, and conflict handling records. The field set of the execution protection conditions includes a protection condition identifier, a trigger condition summary, a protection action summary, and an effective scope. The trigger condition summary is associated with the cell identifier, slice identifier, and time window index, and the protection action summary is associated with the action candidate identifier in the candidate action set. The pre-playing arrangement unit performs pre-playing screening on the candidate action set based on the effect prediction vector and the execution protection conditions, and outputs the pre-playing passed action set. The pre-playing passed action set includes action candidate identifier, dependency marker field, conflict marker field, sorting field, frequency exceeding limit marker field, and vector evidence index field. The pre-playing passed action set and the effect prediction vector are written into the action candidate storage as input to S600. At the same time, the execution protection conditions are written into the protection condition storage for S600 to inject policy constraints during the action policy generation stage.

[0047] In an engineering implementation, for a multi-cell, multi-slice wireless communication network operation and maintenance scenario, the shadow state construction unit reads the candidate action set written by S400 when each time window index arrives, and retrieves the shadow state input slice that matches the time window index from the shadow state storage to complete the shadow network state construction; the pre-playing orchestration unit generates a pre-playing action queue according to the dependency sorting sequence, drives the wireless side parameter threshold change segment, slice resource weight change segment, core network traffic diversion change segment, and core network migration change segment to perform replay steps on the shadow state, and writes the shadow state snapshots before and after each replay step into the shadow state evolution trajectory; the effect evaluation unit generates the effect prediction vector on the shadow state evolution trajectory, and binds the service level default probability, oscillation risk index, and rollback cost prediction to the action candidate identifier; the protection condition generation unit forms the execution protection condition and writes it into the protection condition storage; the pre-playing orchestration unit completes the pre-playing pass screening at the action candidate level, outputs the pre-playing pass action set, and writes the pre-playing pass action set, the effect prediction vector, and the vector evidence index field into the action candidate storage as input to S600.

[0048] Summary of the technical effects of this step: This step compresses the multi-domain state tensor into a replayable session-level representation through the shadow network state, and transforms the candidate action set into an auditable pre-simulation trajectory record; This step generates effect prediction vectors on the pre-simulation trajectory and simultaneously solidifies the evidence index field, enabling traceable input for subsequent decision-making processes; This step outputs the pre-simulation, which, through the action set and execution protection conditions, provides a constraint injection basis for the generation of action strategies in the hierarchical deep reinforcement learning decision model.

[0049] To address the problems in existing technologies where reinforcement learning decision-making in network operation and maintenance scenarios often directly faces high-dimensional states and large-scale action spaces, making it difficult to simultaneously handle multi-granularity decision collaboration at the cell and slice levels, and where action mutual exclusion and dependency constraints are mostly processed post-processed outside the model, resulting in a disconnect between policy output and executability constraints, and difficulty in forming auditable input batches and policy version alignment entry points, this invention, through step S600, inputs the action set, the state tensor, and the effect prediction vector into a hierarchical deep reinforcement learning decision-making model based on the pre-exercise, and outputs an action policy. Specifically, this includes: on the time-window index-aligned input assembly link, the pre-playing is uniformly encapsulated into a model input batch through the action set, the state tensor, and the effect prediction vector, and the input hash record is registered; the hierarchical deep reinforcement learning decision model includes at least a cell-level decision sub-model, a slice-level decision sub-model, and a coordination sub-model. The coordination sub-model performs mutual exclusion checks and dependency ordering on the action policy and writes the conflict resolution record and dependency ordering record into the action policy reference field; simultaneously, the reward function of the hierarchical deep reinforcement learning decision model includes a service level benefit term, a change stability penalty term, and a rollback cost term, wherein the change stability penalty term is related to the number of parameter changes within a unit time window, thereby outputting an action policy carrying a policy version record, providing a version alignment entry for the S700 network change execution and training sample generation, specifically including: S600, based on pre-simulation, inputs action sets, state tensors, and effect prediction vectors into a hierarchical deep reinforcement learning decision model, and outputs action policies; Specifically, this step takes the pre-playback action set, the effect prediction vector, and the execution protection conditions output by S500, and uses the state tensor output by S300 as the decision input. It also uses the sampling alignment rules, anomaly indicator calculation template version, and hash verification field from the session configuration package output by S100 as the input consistency verification basis. This step is completed collaboratively by a hierarchical deep reinforcement learning decision model running unit, an input assembly and normalization unit, a policy constraint injection unit, a mutual exclusion check and dependency sorting unit, and a policy version registration unit. The hierarchical deep reinforcement learning decision model running unit includes a cell-level decision sub-model, a slice-level decision sub-model, and a coordination sub-model. The cell-level decision sub-model generates candidate action scores and local action selections at the cell identifier granularity. The slice-level decision sub-model generates candidate action scores and local action selections at the slice identifier granularity. The coordination sub-model performs action merging, mutual exclusion checks, and dependency sorting across cell identifiers and slice identifiers to form an executable action policy. The hierarchical deep reinforcement learning decision model is a reusable model component for online inference and offline training. Its input port receives the time window index alignment feature of the state tensor, the action candidate identifier and dependency label field, conflict label field, ranking field, frequency over-limit label field of the action set, and the service level default probability, oscillation risk index, rollback cost prediction, and other fields of the effect prediction vector. Its output port outputs the action policy and carries an auditable decision evidence index.

[0050] During the input assembly phase, the input assembly and normalization unit performs primary key alignment and time alignment processing on the pre-simulation action set, the state tensor, and the effect prediction vector. Specifically, firstly, using the time window index as the alignment anchor point, the radio-side load characteristics, interference characteristics, slice-level throughput characteristics, latency characteristics, core network session establishment characteristics, packet loss characteristics, and transmission-side congestion characteristics corresponding to the current time window index are extracted from the state tensor, and the above features are formed into state fragment records according to cell identifier and slice identifier; secondly, using the action candidate identifier in the pre-simulation action set as the primary key, the service level default probability, oscillation risk index, and rollback cost prediction bound to the action candidate identifier are read from the effect prediction vector, and written into the action evaluation record; thirdly, the state fragment records and the action evaluation records are associated and assembled according to cell identifier, slice identifier, and action candidate identifier to generate a model input batch. The model input batch contains state input fragments and action input fragments, wherein the state input fragments are obtained by mapping the state tensor, and the action input fragments are obtained by mapping the pre-simulation action set and the effect prediction vector. Furthermore, the input assembly and normalization unit performs unit normalization processing on the state input fragments according to the field unit dictionary in the session configuration package, resamples and aligns cross-domain sampling granularity differences according to the sampling alignment rules, and calls the hash verification field to generate input hash records for the model input batch. These input hash records are written into the decision audit record for subsequent S700 training sample construction and policy version backtracking. For the action candidate identifiers carrying the frequency exceedance flag field in the action set used in the pre-playing, the input assembly and normalization unit writes a frequency reduction flag into the model input batch and points the reduction source to the vector evidence index field of the S500's degradation step record to maintain cross-step evidence chain consistency.

[0051] During the model inference phase, the hierarchical deep reinforcement learning decision-making model execution unit groups and feeds the model input batches according to cell identifiers and slice identifiers. Specifically, the cell-level decision sub-model reads state input segments belonging to the same cell identifier and combines them with the service level default probability, oscillation risk index, and rollback cost prediction in the corresponding action input segments to generate a cell-level action preference record. The cell-level action preference record includes action candidate identifiers, preference scores, and confidence marker fields. The slice-level decision sub-model reads state input segments belonging to the same slice identifier and combines them with the corresponding action input segments to generate a slice-level action preference record. The slice-level action preference record includes action candidate identifiers, preference scores, and confidence marker fields. The preference score is derived from the reward function constraints within the hierarchical deep reinforcement learning decision-making model. The reward function includes a service level benefit term, a change stability penalty term, and a rollback cost term. The service level benefit term corresponds to the inverse mapping of the service level default probability, the change stability penalty term is related to the number of parameter changes within the unit time window index, and the rollback cost term corresponds to the mapping of the rollback cost prediction. To ensure that the change stability penalty item has a feasible input, the model running unit reads the sorting field and frequency exceedance flag field from the action set in the pre-exercise, and calls the change count template in conjunction with the abnormal indicator calculation template version in the session configuration package. It generates a parameter change count record within the current time window index. The parameter change count record is written into the decision audit record and simultaneously serves as the input for the change stability penalty item.

[0052] During the policy constraint injection and coordination phase, the policy constraint injection unit injects the execution protection conditions output by S500 into the policy filtering link of the coordination sub-model. Specifically, the policy constraint injection unit maps the trigger condition summary in the execution protection conditions to protection condition determination rules, and binds the protection condition determination rules with action candidate identifiers to form protection condition binding records. When the coordination sub-model receives cell-level action preference records and slice-level action preference records, it first performs a protection condition pre-check on the action candidate identifiers based on the protection condition binding records. If the pre-check is successful, a protection interception record is generated and the corresponding action candidate identifier is written into the protection interception set. At the same time, an alternative candidate query tag is generated for the action candidate identifier. Furthermore, the coordination sub-model performs mutual exclusion checks and dependency sorting on the uninterrupted action candidate identifiers. The mutual exclusion check and dependency sorting unit reads the conflict marker field and dependency marker field from the pre-rehearsed action set, and references the rule definitions of the parameter mutual exclusion table and the dependency priority table in S400, mapping the conflict marker field to mutual exclusion conflict pair records and the dependency marker field to dependency precedence pair records. For mutual exclusion conflict pair records, the coordination sub-model resolves conflicts according to preference scores and generates conflict resolution records. For dependency precedence pair records, the coordination sub-model performs topological sorting according to the dependency priority table and generates dependency sorting records, wherein the dependency sorting records contain action candidate identifier sequences and sorting evidence index fields. Understandably, the coordination sub-model also checks the cross-domain consistency between cell identifiers and slice identifiers during the action merging process. Specifically, the check is based on the cross-domain feature correlation strength in the state tensor. If a cross-domain feature mutation is found in the same time window index and the corresponding action candidate identifier involves a core network migration change segment, the action candidate identifier is written into the risk review set and the risk review reason field is recorded. The risk review set is not included in the output of this action strategy, but its record is used as a source of negative samples when the S700 execution trajectory record and training sample generation are generated.

[0053] During the action policy output phase, the hierarchical deep reinforcement learning decision model execution unit encapsulates the action candidate identifier sequence output by the coordination sub-model into the action policy. This action policy belongs to an executable policy data structure, and its field set includes at least a policy identifier, a time window index, an action sequence, action candidate identifiers within the action sequence, a dependency sorting record reference, a conflict resolution record reference, and a protection interception record reference. Further, to support the traceability of subsequent network change execution, the model execution unit associates each action candidate identifier in the action policy with its source cell-level action preference record, slice-level action preference record, and the vector evidence index field of the effect prediction vector, generating a policy evidence chain record and writing it into the decision audit record. The policy version registration unit generates a policy version number when the action policy is written to disk and registers the policy version number, along with the input hash record, the anomaly indicator calculation template version, and the hash verification field of the session configuration package, as a policy version record. This policy version record is used as a version alignment basis when the S700 performs network changes and generates execution trajectory records.

[0054] In an engineering embodiment, for online operation and maintenance scenarios of the same wireless communication network across multiple consecutive time window indices, the input assembly and normalization unit reads the pre-exercise action set and the effect prediction vector from the action candidate storage when each time window index is triggered, and reads the state tensor from the state storage to complete the model input batch construction and write it into the input hash record; the hierarchical deep reinforcement learning decision model running unit runs the cell-level decision sub-model and the slice-level decision sub-model in parallel according to the cell identifier and slice identifier to generate local preference records, and then the coordination sub-model introduces the execution protection conditions to complete protection interception, and combines the parameter mutual exclusion table and the dependency priority table to complete mutual exclusion checks and dependency sorting, and outputs the action policy; the policy version registration unit writes the action policy into the policy storage and registers the policy version number and policy evidence chain record, so that S700 can execute the action policy according to the policy version number and return the execution trajectory record.

[0055] In summary, the technical effects of this step are as follows: In the time-window index-aligned input assembly chain, this step unifies the pre-training action set, state tensor, and effect prediction vector into an auditable model input batch and registers the input hash record; this step, through the sub-model collaboration and coordination of the hierarchical deep reinforcement learning decision model, forms an action policy with conflict resolution records and dependency ranking records; this step outputs the action policy and simultaneously registers the policy version record, ensuring a stable version alignment entry point for subsequent network change execution and training sample generation.

[0056] To address the problems in existing technologies where the network change execution link, decision link, and evidence link are disconnected, change issuance lacks traceable constraints for dependency ordering and conflict resolution, and the records of protection condition monitoring, effectiveness verification, and rollback during execution are incomplete, making it difficult to trace the execution results and form highly consistent training samples for continuous iteration, this invention, through step S700, executes network changes based on the action strategy and generates execution trajectory records, generates training samples based on the execution trajectory records, updates the hierarchical deep reinforcement learning decision model, and registers the strategy version. Specifically, this includes: parsing the action policy into executable change tasks and completing network change distribution under the constraints of dependency ordering records and conflict resolution records, generating an execution trajectory record that includes the policy version number and the input hash record; during execution, performing online monitoring and trigger determination based on the execution protection conditions to form auditable protection trigger records and rollback records, and associating and archiving telemetry evidence during execution with the execution trajectory record under the time window index; in the training sample construction stage, extracting state-action-result segments based on the execution trajectory record and writing them into the training samples, completing the incremental update and policy version registration of the hierarchical deep reinforcement learning decision model, so that subsequent sessions have a continuously iterative call chain under the same policy version record entry, specifically including: S700: Based on the action policy, perform network changes and generate execution trajectory records, generate training samples based on the execution trajectory records, update the hierarchical deep reinforcement learning decision model and register the policy version; Specifically, this step takes the action strategy output by S600 and uses the strategy version record registered by S600, the execution protection condition output by S500, the effect prediction vector output by S500, the state tensor output by S300, and the session configuration package output by S100 as the input basis for execution and learning. The action strategy includes the time window index, the action sequence, the action candidate identifier, the dependency sorting record reference, the conflict resolution record reference, and the protection interception record reference; the strategy version record includes the strategy version number, the input hash record, the anomaly indicator calculation template version, and a hash verification field associated with the session configuration package. This step is completed collaboratively by the change orchestration and distribution unit, the protection condition monitoring and interception unit, the effectiveness verification and rollback unit, the execution trajectory record generation unit, the training sample construction unit, the model update unit, and the policy version registration unit. The change orchestration and distribution unit is deployed on the network operation and maintenance control plane and connected to the network element management interface. The protection condition monitoring and interception unit is connected to the telemetry subscription channel and reads the sampling alignment rule execution time window index alignment of the session configuration package. The model update unit shares model parameter storage with the hierarchical deep reinforcement learning decision model running unit and performs versioned writing.

[0057] During the network change execution phase, the change orchestration and distribution unit parses the action policy into change tasks to be executed, generates change task numbers, and writes them into the execution context. Specifically, the change orchestration and distribution unit reads the action candidate identifier sequence from the action sequence, determines the distribution order based on the dependency sorting record reference, and determines the suppressed action candidate identifiers based on the conflict resolution record reference. Subsequently, each action candidate identifier is mapped to a target change segment in the change segment set formed in S400. The change segment set includes radio side parameter threshold change segments, slice resource weight change segments, core network traffic offloading change segments, core network migration change segments, and cross-domain rollback change segments. Further, the change orchestration and distribution unit determines the scope of each change segment based on the primary key binding table of probe to cell identifier, slice identifier, and network element identifier in S100. The scope at least includes the network element identifier and can be extended to include the cell identifier or slice identifier. The binding relationship between the action candidate identifier and the scope is solidified in the execution context. Subsequently, the change orchestration and distribution unit encapsulates the change fragments into configuration change requests according to the network element identifier, and calls the configuration management function unit of the network element to complete the distribution. Among them, the radio side parameter threshold change fragment is received by the base station side parameter management function unit and written into the parameter pending effect area; the slice resource weight change fragment is received by the slice orchestration function unit and written into the slice resource scheduling table; the core network traffic offloading change fragment is received by the user plane traffic offloading function unit and written into the traffic allocation rule table; and the core network migration change fragment is received by the session management function unit and written into the migration plan table. Each distribution generates a distribution timestamp and records the request digest hash. The request digest hash is associated with the input hash record and written into the execution trajectory index.

[0058] During the protection condition monitoring, effectiveness verification, and rollback phases, the protection condition monitoring and interception unit establishes a protection monitoring session based on the executed protection conditions and continuously reads and aligns telemetry data during the execution of the change task. Specifically, the protection condition monitoring and interception unit performs time window indexing and merging of the timestamps of the telemetry data according to the sampling alignment rules, and calls the abnormal indicator calculation template version to generate online monitoring indicator records. The online monitoring indicator records contain at least a subset of Key Performance Indicators (KPIs) and a subset of Key Quality Indicators (KQIs), and can be expanded to include online estimates of cross-cell interference coupling terms and congestion propagation approximations. Further, the protection condition monitoring and interception unit aligns the online monitoring indicator records with the service level default probability, oscillation risk index, and rollback cost prediction in the effect prediction vector using time window indexing, and performs trigger determination according to the executed protection conditions; when the trigger determination is successful, a protection trigger flag is generated and written to the protection trigger record, and a rollback request is sent to the effectiveness verification and rollback unit. After each change segment is issued, the effectiveness verification and rollback unit reads the network element receipt and generates an effectiveness confirmation timestamp. If the receipt is missing or abnormal, an effectiveness failure flag is generated, and the corresponding action candidate identifier is written into the execution failure set. Upon receiving a rollback request or detecting an effectiveness failure flag, the effectiveness verification and rollback unit calls the cross-domain rollback change segment to generate a rollback task and issues the rollback request in reverse order of the dependency sorting record references. Simultaneously, a rollback flag is generated, and the rollback time is recorded. Understandably, if a protection trigger flag appears accompanied by an increasing trend in the oscillation risk indicator, the effectiveness verification and rollback unit writes the change task number into the oscillation risk handling record and suspends the execution of subsequent action candidate identifiers within the same scope. The suspended state is written into the execution trajectory index for the training sample construction unit to read.

[0059] During the execution trajectory recording and training sample generation phase, the execution trajectory recording generation unit writes the entire process of the change task into the execution trajectory record. The execution trajectory record is a structured record set and includes cross-step traceable fields. Specifically, the execution trajectory record includes at least the policy version number, change task number, time window index, network element identifier, cell identifier, slice identifier, action candidate identifier, change segment type, distribution timestamp, effectiveness confirmation timestamp, protection trigger flag, effectiveness failure flag, rollback flag, rollback time record, online monitoring indicator record reference, and request digest hash. Among them, the request digest hash is associated with the input hash record, and the online monitoring indicator record reference points to the telemetry data segment index within the same time window index. The training sample construction unit extracts state-action-result triples from the execution trajectory records and generates training samples. Specifically, the state input segment corresponding to the time window index in the state tensor of S300 is used as the state field of the training sample, the action candidate identifier and its changed segment type are used as the action field of the training sample, the differential summary of online monitoring indicator records before and after the effective confirmation timestamp, the protection trigger flag, the effective failure flag, and the rollback flag are used as the result field of the training sample, and the service level default probability, oscillation risk index, and rollback cost prediction bound to the action candidate identifier in the effect prediction vector are written into the prediction field of the training sample, forming a sample alignment structure that can be used for error attribution. Further, the training sample construction unit filters the training sample source according to the admission control classification logic of S200: when the telemetry data segment corresponding to the execution trajectory is written into the pending acquisition set or downgrade processing set in S200, the training sample construction unit writes the sample into the sample removal record and registers the removal reason field in the sample index, so that the model update unit only receives samples from the admission evidence set link.

[0060] During the model update and policy version registration phase, the model update unit reads training samples and performs parameter updates on the hierarchical deep reinforcement learning decision model, while maintaining version isolation between online inference and offline updates. Specifically, the model update unit bins the training samples according to cell identifiers and slice identifiers, and feeds them to the training entry points of the cell-level decision sub-model and the slice-level decision sub-model, respectively. Simultaneously, it feeds the sample summaries related to cross-domain conflict resolution and dependency ranking to the training entry point of the coordination sub-model. During the model update process, the model update unit generates reward label records based on the result fields of the training samples. These reward label records correspond one-to-one with the reward function terms of the hierarchical deep reinforcement learning decision model and are written into the training audit record. Further, after completing the parameter update, the model update unit generates a model version number and associates it with the policy version number, the anomaly indicator calculation template version, and the input hash record, registering them as a policy version registration record. The policy version registration record is written into the policy version library for subsequent S600 calls. The policy version registration record also serves as the associated object of the hash verification field of the S100 session configuration package, supporting version backtracking across time windows. For the automated operation scenario of the engineering implementation, after each time window index ends or the protection trigger flag is hit, the model update unit triggers an incremental update process and registers the trigger source field and sample batch number in the policy version registration record, so that the evolution trajectory of the hierarchical deep reinforcement learning decision model has a continuously searchable link in the policy version library.

[0061] In summary, the technical effects of this step are as follows: This step transforms the action policy into an executable change task and completes the network change distribution under the constraints of dependency ordering and conflict resolution. Simultaneously, it generates an execution trajectory record that associates the policy version number with the input hash record. This step creates auditable protection trigger records and rollback records on the execution protection condition monitoring, effectiveness verification, and rollback chain, ensuring the traceability of the execution process and telemetry evidence under the time window index. This step constructs training samples based on the execution trajectory record and completes the incremental update and policy version registration of the hierarchical deep reinforcement learning decision model, providing a continuous iteration entry point for subsequent decision steps.

Claims

1. A network intelligent operation and maintenance method based on deep reinforcement learning, characterized in that, include: S100: Obtain network topology, network element list, probe list, sampling granularity, field mapping table, and unified clock reference; generate session configuration package. S200: Collect telemetry data based on session configuration package, calculate time synchronization deviation, missing rate, noise jitter, and field drift, generate quality labels and calculate evidence credibility, and execute access control to generate access evidence set; S300: Based on the admission evidence set, construct a state tensor according to cell identifier, slice identifier, and time window index, and generate a root cause association graph; S400. Obtain the action constraint library and generate a set of candidate actions carrying dependency and conflict markers based on the root cause relationship graph and the action constraint library. S500: Construct the shadow network state and perform a pre-play based on the candidate action set, generate the effect prediction vector, and filter the action set that passes the pre-play and the execution protection conditions; S600, based on pre-simulation, inputs action sets, state tensors, and effect prediction vectors into a hierarchical deep reinforcement learning decision model, and outputs action policies; S700: Based on the action policy, perform network changes and generate execution trajectory records, generate training samples based on the execution trajectory records, update the hierarchical deep reinforcement learning decision model and register the policy version.

2. The method according to claim 1, characterized in that, The session configuration package in S100 includes a primary key binding table for probes to cell identifiers, slice identifiers, and network element identifiers, a dictionary of field units, sampling alignment rules, anomaly indicator calculation template version, and hash verification fields.

3. The method according to claim 1, characterized in that, The quality labels in S200 include out-of-order rate, repetition rate, proportion of sudden spikes, and cross-domain time drift trend. The credibility of the evidence is obtained by mapping the quality labels.

4. The method according to claim 1, characterized in that, In S200, the entry control adopts a first threshold and a second threshold for classification. Data fragments with evidence credibility lower than the first threshold are written into the supplementary collection set, and data fragments with evidence credibility between the first threshold and the second threshold are written into the downgrade processing set and removed from the training sample construction process.

5. The method according to claim 1, characterized in that, The state tensor in S300 includes radio-side load characteristics, interference characteristics, slice-level throughput characteristics, latency characteristics, core network session establishment characteristics, packet loss characteristics, and transmission-side congestion characteristics, and is aligned by time window index.

6. The method according to claim 1, characterized in that, In S300, generating a root cause association graph involves mapping the set of anomaly candidates to a set of nodes, generating edge weights based on the temporal order and correlation strength of cross-domain indicators, and forming a root cause association graph.

7. The method according to claim 1, characterized in that, The action constraint library in S400 includes slice service level constraints, capacity constraints, coverage constraints, change frequency constraints, parameter mutual exclusion tables and dependency priority tables. The candidate action set is generated by matching the root cause nodes of the root cause association graph with the action templates and combining the parameter mutual exclusion tables and dependency priority tables.

8. The method according to claim 1, characterized in that, The S400 change fragment set includes radio side parameter threshold change fragments, slice resource weight change fragments, core network traffic offloading change fragments, core network migration change fragments, and cross-domain rollback change fragments.

9. The method according to claim 1, characterized in that, The shadow network state in S500 includes a subset of key KPIs, a subset of key KQIs, cross-cell interference coupling terms, and congestion propagation approximation terms. The effect prediction vector includes service level default probability, oscillation risk indicators, and rollback cost prediction.

10. The method according to claim 1, characterized in that, The hierarchical deep reinforcement learning decision model in S600 includes a cell-level decision sub-model, a slice-level decision sub-model, and a coordination sub-model. The coordination sub-model performs mutual exclusion checks and dependency ordering on action policies. The reward function of the hierarchical deep reinforcement learning decision model includes a service level benefit term, a change stability penalty term, and a rollback cost term. The change stability penalty term is related to the number of parameter changes within a unit time window.