Industrial private network fusion deployment method, device, equipment and storage medium

By deploying a converged core network in an open-source container orchestration platform cluster, constructing digital twin environment data, and using a reinforcement learning model for dynamic scheduling, the problems of resource silos and traffic fluctuations in industrial private networks are solved, achieving efficient resource utilization and rapid response.

CN121887644APending Publication Date: 2026-04-17E SURFING IOT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
E SURFING IOT CO LTD
Filing Date
2026-01-22
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing industrial private networks, the independent deployment of core networks of various standards forms resource silos, resulting in low resource utilization, high costs, inability to adapt to complex and random traffic fluctuations, and delayed response of traditional resource scheduling strategies, which cannot meet the agility requirements of industrial production.

Method used

Deploy a converged core network in an open-source container orchestration platform cluster to form a converged resource pool. Construct digital twin environment data by collecting data on terminal type, traffic matrix, and service level agreement requirements. Train a reinforcement learning model using a dual-latency deep deterministic policy gradient. Collect resource and business metrics in real time and dynamically schedule resources.

Benefits of technology

It enables the sharing and reuse of computing resources, improves resource utilization, reduces capital expenditure, enhances resource efficiency and deployment flexibility, and ensures business continuity and rapid response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121887644A_ABST
    Figure CN121887644A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial private network fusion deployment method, device and equipment and a storage medium, and the method comprises the steps: deploying a fusion core network in an open source container arrangement platform cluster, and forming a fusion resource pool; acquiring terminal type, traffic matrix and service level protocol requirement data, and constructing digital twin environment data; based on the digital twinning environment data, using a double-delay depth deterministic strategy gradient to train the reinforcement learning model; acquiring a resource index, a business index and a service grade protocol index of the fusion resource pool in real time to form state data; when the state data meets any one of preset triggering conditions, inputting the state data into a trained reinforcement learning scheduling model to obtain a resource scheduling action; and executing the resource scheduling action through an open source container arrangement platform interface to generate an execution result. According to the method, resource islands are eliminated, and sharing and multiplexing of computing resources are realized, so that the resource utilization rate is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet of Things (IoT) communication technology, and in particular to a method, apparatus, equipment, and storage medium for the integrated deployment of industrial private networks. Background Technology

[0002] In the field of industrial IoT, traditional industrial private networks typically adopt a separate deployment scheme, that is, the 4G evolved packet core network (EPC), 5G core network (5GC), narrowband IoT (NB-IoT) core network, and IP multimedia subsystem (IMS) are all independently built on dedicated hardware equipment.

[0003] This solution has significant drawbacks: First, the independent deployment of core networks for each standard creates resource silos, with overall resource utilization generally below 40%. Moreover, for small and medium-sized scenarios, the capital expenditure (CAPEX) often exceeds one million, resulting in high costs. Second, the IMS voice system is built separately and cannot share underlying computing, storage, and network resources with 4GEPC, ​​5GC, and NB-IoT data services, further exacerbating resource idleness and waste. Third, when service traffic changes require expansion, the reliance on manual adjustments to physical equipment configurations often results in response delays exceeding 48 hours, failing to meet the agility requirements of industrial production.

[0004] With the application of cloud-native technologies, existing technologies attempt to deploy 5G core networks using Kubernetes container orchestration platforms (such as the free5GC open-source project), achieving containerization of network elements. However, significant shortcomings remain: Firstly, such solutions typically do not support a converged architecture for multi-standard core network functions including 4G EPC, 5GC, NB-IoT, and IMS, failing to fundamentally solve the resource pool sharing problem. Secondly, their resource scheduling engines rely on static rules, making it difficult to accurately predict and respond quickly to the complex and ever-changing random traffic surges in industrial scenarios, easily leading to service level agreement (SLA) breaches. Furthermore, while some industry resource allocation algorithms have introduced automation, they are still based on preset rule engines with fixed strategies, lacking self-learning and dynamic optimization capabilities, and similarly unable to effectively adapt to the complex and random traffic fluctuations in industrial settings. Summary of the Invention

[0005] The purpose of this invention is to provide a method, apparatus, equipment and storage medium for the integrated deployment of industrial private networks, which aims to solve the problem that existing technologies cannot effectively adapt to the complex and random traffic fluctuations in industrial sites.

[0006] In a first aspect, embodiments of the present invention provide a method for the integrated deployment of industrial private networks, including: A converged core network is deployed in an open-source container orchestration platform cluster to form a converged resource pool; the converged core network includes 4G evolved packet core network components, 5G core network components, narrowband IoT technology core components, and IP multimedia subsystem components. Collect data on terminal types, traffic matrices, and service level agreement requirements to construct digital twin environment data; Based on the digital twin environment data, a reinforcement learning model is trained using a dual-delay deep deterministic policy gradient. The system collects resource metrics, business metrics, and service level agreement metrics of the integrated resource pool in real time to form status data. When the state data meets any of the preset triggering conditions, the state data is input into the trained reinforcement learning scheduling model to obtain resource scheduling actions. The resource scheduling action is executed through the interface of the open-source container orchestration platform, and the execution result is generated.

[0007] Secondly, embodiments of the present invention provide an industrial private network converged deployment device, comprising: The deployment unit is used to deploy the converged core network in the open-source container orchestration platform cluster to form a converged resource pool; the converged core network includes 4G evolved packet core network components, 5G core network components, narrowband IoT technology core components, and IP multimedia subsystem components; The building unit is used to collect data on terminal type, traffic matrix, and service level agreement requirements to build digital twin environment data; The training unit is used to train the reinforcement learning model using a dual-delay deep deterministic policy gradient based on the digital twin environment data. The data acquisition unit is used to collect resource indicators, business indicators, and service level agreement indicators of the fusion resource pool in real time to form status data. The input unit is used to input the state data into the trained reinforcement learning scheduling model when the state data meets any one of the preset trigger conditions, so as to obtain the resource scheduling action. The execution unit is used to execute the resource scheduling action through the open-source container orchestration platform interface and generate the execution result.

[0008] Thirdly, embodiments of the present invention provide a computer device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the industrial private network converged deployment method described in the first aspect.

[0009] Fourthly, embodiments of the present invention also provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, which, when executed by a processor, implements the industrial private network converged deployment method described in the first aspect.

[0010] This invention discloses a method, apparatus, device, and storage medium for the converged deployment of an industrial private network. The method includes: deploying a converged core network in an open-source container orchestration platform cluster to form a converged resource pool; the converged core network includes 4G evolved packet core network components, 5G core network components, narrowband IoT technology core components, and IP multimedia subsystem components; collecting terminal type, traffic matrix, and service level agreement (SLA) requirement data to construct digital twin environment data; training a reinforcement learning model using a dual-delay deep deterministic policy gradient based on the digital twin environment data; collecting resource indicators, service indicators, and SLA indicators of the converged resource pool in real time to form state data; when the state data meets any of the preset trigger conditions, inputting the state data into the trained reinforcement learning scheduling model to obtain a resource scheduling action; and executing the resource scheduling action through the open-source container orchestration platform interface to generate an execution result. This invention constructs a cloud-native core network architecture that integrates multiple standards, including 4G evolved packet core network components, 5G core network components, narrowband IoT technology core components, and IP multimedia subsystem components. Each network element is uniformly deployed as a containerized microservice within an open-source container orchestration platform cluster, fundamentally eliminating resource silos and enabling the sharing and reuse of computing resources. This significantly improves resource utilization and reduces capital expenditure. Simultaneously, it introduces a reinforcement learning model based on dual-latency deep deterministic policy gradient (TD3) for dynamic resource scheduling. This reinforcement learning model can automatically make scaling decisions based on real-time monitoring data and continuously optimize through an online learning mechanism, replacing the traditional rule engine that relies on static thresholds. This significantly improves the resource efficiency, deployment flexibility, and service continuity of industrial private networks. The embodiments of this invention also provide an industrial private network converged deployment device, a computer-readable storage medium, and a computer device, all possessing the aforementioned beneficial effects, which will not be elaborated further here. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating the deployment method for integrated industrial private networks; Figure 2This is a schematic diagram of the first sub-process of the industrial private network integrated deployment method. Figure 3 This is a schematic diagram of the second sub-process of the industrial private network converged deployment method; Figure 4 A schematic diagram of the third sub-process of the industrial private network converged deployment method; Figure 5 A schematic block diagram of an industrial private network converged deployment device. Detailed Implementation

[0013] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] It should be understood that, when used in this specification and the appended claims, the terms “comprising” and “including” indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more of its features, integrals, steps, operations, elements, components and / or collections thereof.

[0015] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0016] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the relevant listed items and all possible combinations, and includes such combinations.

[0017] Please see Figures 1-4 This embodiment provides a method for the converged deployment of industrial private networks, including: S101: Deploy a converged core network in an open-source container orchestration platform cluster to form a converged resource pool; the converged core network includes 4G evolved packet core network components, 5G core network components, narrowband IoT technology core components, and IP multimedia subsystem components. In this embodiment, deploying a converged core network in an open-source container orchestration platform cluster includes: The mobility management entity, service gateway, and packet data network gateway of the 4G evolved packet core network components are decomposed into microservices for deployment; The access and mobility management functions, session management functions, and user plane functions of the 5G core network components are deployed in the form of independent minimum deployment units; Microservice-based cellular IoT service gateway nodes, a core component of narrowband IoT technology; The proxy call session control function of the IP Multimedia Subsystem component and the data plane of the core control node share the user plane function.

[0018] This embodiment breaks away from the siloed architecture of traditional private networks by deploying various network elements (4GEPC, ​​5GC, NB-IoT, and IMS) as microservices or independent minimum deployment units (Pods) on the same open-source container orchestration platform (such as Kubernetes). This allows computing, storage, and network resources to be flexibly allocated and reused across different services, significantly improving overall resource utilization. Furthermore, the fine-grained deployment based on microservices enables each network element's function to scale and update independently. Combined with cloud-native orchestration capabilities, it achieves seamless elastic scaling and agile operation and maintenance, providing a standardized and automated execution foundation for upper-layer intelligent scheduling (such as reinforcement learning models). Ultimately, this supports industrial private networks in achieving the goals of simplified deployment and efficient resource utilization while ensuring high reliability and low latency services.

[0019] Specifically, independent namespaces and resource configuration declarations are created for the deployment of 4G evolved packet core network components within the Kubernetes cluster. According to the technical solution, the three traditional monolithic network elements—Mobility Management Entity (MME), Service Gateway (SGW), and Packet Data Network Gateway (PGW)—are each decomposed into independent, lightweight microservices. In practice, a dedicated container image is built for each component, and a corresponding Kubernetes deployment manifest is written. The MME, SGW, and PGW are deployed as three independent applications, each running as one or more Pods on cluster nodes. Through this microservice decomposition, each 4G core network functional component gains independent lifecycles, resource quotas, and scaling capabilities. Subsequently, a pre-developed, customized Kubernetes Operator is used to uniformly manage the entire lifecycle of these microservice network elements. This Operator continuously monitors the cluster status and automatically performs maintenance operations such as fault recovery and version rolling updates based on the declared expected state, thereby transferring the complexity of network element management and maintenance from manual processes to an automated framework. This deployment laid the foundation for the subsequent construction of a converged resource pool: the microservice-ized 4GEPC components, along with the simultaneously deployed 5GC, NB-IoT, and IMS microservices, are registered to the K8s control plane, sharing the same physical resource pool. Their resource usage is uniformly monitored and used as raw data input, providing a unified management object and execution interface for subsequent reinforcement learning models to perform cross-system and cross-business global resource dynamic scheduling.

[0020] In the industrial private network deployment embodiment, the 5G core network components are deployed using an independent minimum deployment unit architecture. First, a dedicated namespace is created in the Kubernetes cluster, establishing an independent deployment unit for each core network element. The Access and Mobility Management (AMF) component is encapsulated as a stateless Pod, configured with corresponding service discovery mechanisms and health check policies. The Session Management Function (SMF) component is deployed as a stateful Pod, persistently storing session data and configuring network policies. The User Plane Function (UPF) component, as a key unit of the data plane, adopts a multi-replica deployment mode to achieve network acceleration.

[0021] During deployment, configuration files for each network element are managed uniformly using ConfigMaps (Kubernetes objects used to store non-confidential configuration data, saving configuration files in key-value pairs), and security credentials are handled using Secret objects. Resource request and limit parameters are set for each Pod to ensure reasonable allocation of CPU and memory resources. Service mesh technology enables communication between network elements, with sidecar proxies handling service discovery and load balancing. After deployment, connectivity tests are performed to verify normal communication between the interfaces between AMF and SMF, and between SMF and UPF.

[0022] Next, the KubernetesOperator is used to monitor the operational status of each network element, enabling automatic fault recovery and rolling upgrades. When an AMF instance anomaly is detected, the replica restart mechanism is automatically triggered; when the UPF load reaches a threshold, the HPA controller elastically scales up or down according to a preset strategy. This independent minimum deployment unit architecture ensures functional isolation between network elements and enables dynamic resource scheduling, laying the foundation for subsequent reinforcement learning optimization.

[0023] Then, a dedicated namespace is created in the Kubernetes cluster to split the C-SGN functionality into two microservice modules: a control plane and a user plane. The control plane microservice is responsible for connection management, mobility management, and session management, while the user plane microservice handles packet forwarding and routing. Each microservice is encapsulated as an independent Pod instance, configured with corresponding resource requests and limit parameters.

[0024] Subsequently, C-SGN network parameters and business policies are uniformly configured via ConfigMap, and sensitive information such as security credentials is handled using Secret objects. During deployment, readiness probes and liveness probes are set to ensure that microservices only receive traffic after full initialization and to automatically restart in case of anomalies. Service mesh technology enables communication between microservices, with sidecar proxies handling service discovery and load balancing.

[0025] The microservice-based C-SGN component interconnects with the 5GC's AMF / SMF components via a standard interface, supporting NB-IoT terminal access. The monitoring system collects real-time business metrics such as connection count and message throughput, providing data input for subsequent reinforcement learning scheduling. When a large-scale concurrent terminal access is detected, the HPA controller automatically adjusts the number of microservice replicas according to a preset strategy to ensure service quality.

[0026] Next, the Proxy Call Session Control Function (P-CSCF) and the Core Control Node (I-CSCF) are built as independent container images and deployed as independent Pods in the Kubernetes cluster. Each Pod is configured with explicit resource requests and service quality levels. To achieve the core design of sharing with the data plane, when deploying the User Plane Function (UPF) Pod, not only are rules and resources for processing 5G user plane data streams allocated to it, but its network interface and data plane capabilities are also exposed to the IMS control plane components through additional network policies and configurations. Specifically, a dedicated network service is created within the Kubernetes cluster, which publishes the data plane network endpoints of the UPFPod (e.g., specific IP addresses and port ranges) through internal service discovery. Subsequently, in the Pod configurations deploying P-CSCF and I-CSCF, the service discovery address of the UPF data plane is injected through environment variables or configuration files. When the IMS components need to establish or forward voice or video media streams, P-CSCF no longer seeks to establish independent data channels, but instead, based on the injected configuration, directly points the media stream forwarding request to the shared UPF data plane service endpoint. After receiving media stream processing requests from IMS, UPF forwards these media streams efficiently and with low latency, just like it would with ordinary 5G data packets, according to its embedded forwarding rules. Through this architecture, IMS media plane functions are converged into a unified, high-performance UPF data plane, while session establishment and control logic are still handled independently by P-CSCF and I-CSCFPod. In subsequent operational monitoring, IMS service metrics (such as the number of concurrent sessions) and shared UPF performance metrics (such as network latency and throughput) are collected together to form the system state vector. When the reinforcement learning model decides to expand or shrink the data plane, its actions (such as adjusting the number of UPFPod replicas or resources) will simultaneously affect the carrying capacity of both 5G data services and IMS media services, achieving unified elastic scheduling and optimization of cross-service data plane resources.

[0027] S102: Collect terminal type, traffic matrix and service level agreement requirements data to construct digital twin environment data; Specifically, the system collects terminal type data in real time through network monitoring probes, including automated guided vehicles, various industrial sensors, and augmented reality glasses. Simultaneously, it utilizes a traffic analysis engine to collect traffic matrix data, covering hourly traffic peaks, type distribution (such as video streams, control commands, and data acquisition), and bandwidth demand patterns. Furthermore, it extracts Service Level Agreement requirements from the SLA management interface, including latency thresholds, bandwidth guarantees, and availability metrics. The system integrates the collected terminal type, traffic matrix, and SLA requirement data into a digital twin engine. This engine constructs an environmental data model based on time series analysis and network topology, generating digital twin environmental data that includes terminal distribution density, dynamic traffic patterns, and SLA constraints, providing a real-time input foundation for subsequent dynamic resource scheduling.

[0028] The process involves integrating collected terminal type, traffic matrix, and SLA requirement data into a digital twin engine. This engine constructs an environmental data model based on time series analysis and network topology, generating digital twin environmental data that includes terminal distribution density, dynamic traffic patterns, and SLA constraints. Collect historical access data, service traffic data, and predefined service level agreement requirements data of various terminals in the industrial private network; The collected terminal access data, business traffic data, and service level agreement requirement data are cleaned and time-aligned to form a multi-source time-series data set; The processed multi-source time-series data set is input into the digital twin engine; The digital twin engine is based on a multi-source time-series data set. It extracts the access patterns and traffic fluctuation characteristics of different terminal groups through time series analysis algorithms, and establishes a network state evolution model based on the logical topology of the industrial private network. The digital twin engine runs a network state evolution model to simulate and generate digital twin environment data that includes the patterns of terminal distribution density changes over time, dynamic business traffic fluctuation patterns, and constraints bound to service level agreement requirements.

[0029] The system collects historical access data, service traffic data, and predefined service level agreement (SLA) requirements from various terminals in the industrial private network, covering multi-dimensional information such as terminal type, service mode, and SLA constraints. The collected data is cleaned and time-aligned to eliminate data source heterogeneity and time-series bias, forming a unified multi-source time-series data set. The processed data is then input into a digital twin engine, which uses time-series analysis algorithms to extract access patterns and traffic fluctuation characteristics of different terminal groups and establishes a network state evolution model based on the logical topology of the industrial private network. The digital twin environment data generated by running this model includes the patterns of terminal distribution density changes over time, dynamic service traffic fluctuation patterns, and constraints bound to service level agreement requirements, providing accurate and structured input data for training reinforcement learning models.

[0030] Specifically, during the system initialization phase, when building the digital twin training environment, a data acquisition agent is first deployed in the running converged core network Kubernetes cluster. By monitoring Pod logs, capturing network node metering information, and querying network element management interfaces, historical access data of various terminals within the industrial private network is continuously collected. This includes connection and disconnection timestamps, identification, and access standards of devices such as automated guided vehicles, industrial sensors, and augmented reality glasses. Simultaneously, service traffic data is collected, covering uplink and downlink byte counts, packet rates, and time windows of traffic peaks. At the same time, predefined service level agreement (SLA) requirements for each service type are extracted from the network management policy library, such as latency limits for voice services, packet loss rate thresholds for data services, and availability targets for critical control signaling. The collected raw data streams are transmitted to the data processing module for cleaning and time alignment. This module removes abnormal records caused by network jitter, uniformly calibrates timestamps from different data sources to a standard timeline, and correlates and aggregates terminal access events, traffic statistics sequences, and SLA constraint labels according to time windows, ultimately generating a well-structured, time-synchronized multi-source time-series data set. This dataset is fully input into the data loading layer of the digital twin engine through a standard data interface. After receiving this dataset, the engine uses it as the core data foundation for building the virtual simulation environment.

[0031] In some embodiments, based on a multi-source time-series data set, the access patterns and traffic fluctuation characteristics of different terminal groups are extracted using time-series analysis algorithms, and a network state evolution model is established according to the logical topology of the industrial private network, including: Acquire a multi-source time-series data set, which includes terminal access data, business traffic data, and service level agreement requirement data, and has completed time alignment processing to generate a standardized time-series dataset; Next, based on the standardized time series dataset, time series analysis algorithms are applied to extract multidimensional features, including using an autoregressive integral moving average model to analyze the access patterns of terminal groups, using spectrum analysis methods to identify traffic fluctuation characteristics, and generating a feature set containing periodic patterns, trend components, and residual sequences. Then, combining the logical topology of the industrial private network, which includes the connection relationship of core network elements, the layout of the access network, and the service flow path, The feature set is then mapped to network topology nodes, a state transition probability matrix and a load propagation model are established, and a network state evolution framework is generated. Furthermore, stochastic process simulation is introduced into the network state evolution framework. The network state transition is modeled through Markov decision process, and a network state evolution model is generated by combining traffic fluctuation characteristics, which includes dynamic changes in terminal distribution density, spatiotemporal distribution of service traffic, and service quality constraints.

[0032] This embodiment employs time series analysis algorithms for multidimensional feature extraction, including using an autoregressive integral moving average model to analyze terminal access patterns and spectral analysis to identify traffic fluctuation characteristics. This allows for in-depth mining of periodic patterns, trend components, and residual sequences within the data, revealing the inherent patterns of terminal group behavior and dynamic traffic characteristics, thus enhancing the scientific rigor and explanatory power of feature engineering. Simultaneously, a network state evolution framework is established based on the logical topology of the industrial private network, mapping the feature set to specific network nodes. Through state transition probability matrices and load propagation models, a precise characterization of network dynamics is achieved, enabling the model to reflect the connection relationships and service flow paths in the actual network, enhancing the realism of the simulation. The introduction of stochastic process simulation and Markov decision process modeling of network state transitions further improves the model's adaptability to uncertainty, generating an evolutionary model that includes dynamic changes in terminal distribution density, spatiotemporal distribution of service traffic, and quality of service constraints, providing strong theoretical support for resource scheduling and optimization decisions. The overall process is progressively layered, with each step's output serving as subsequent input, forming a closed-loop optimization that significantly improves the intelligence level of industrial private network management.

[0033] Specifically, a multi-source time-series dataset that has been cleaned and time-aligned is loaded. This dataset integrates three types of data: terminal access events, service traffic statistics, and service level agreement requirements, and has been transformed into a standardized time-series dataset with unified timestamps and standard dimensions. Subsequently, a time-series analysis pipeline is deployed on this standardized dataset to perform multi-dimensional feature extraction: For different categories of terminal access sequences, an autoregressive integral moving average model is applied for modeling. This model identifies the autocorrelation and differential stationarity of the sequences, analyzes the access patterns of each terminal group within historical periods, and separates the trend components representing the long-term direction of change, the seasonal patterns reflecting fixed periods, and the random residual sequences that cannot be explained by the model. At the same time, a spectral analysis method is used on the parallel service traffic time-series data. The time-domain traffic signal is converted to the frequency domain through a fast Fourier transform to identify the key frequency components and their amplitudes that dominate service traffic fluctuations, thereby characterizing their periodic fluctuation characteristics. Finally, the decomposition results from the ARIMA model and the frequency domain features from the spectrum analysis are merged and encoded to generate a structured feature set. This set quantitatively describes the regularity of terminal access behavior, the periodic oscillation pattern of business traffic, and their respective trend and random components in industrial scenarios.

[0034] Then, combining the logical topology of the industrial private network, the feature set is mapped to the network topology nodes, a state transition probability matrix and a load propagation model are established, and a network state evolution framework is generated, including: Obtain the logical topology of the industrial private network. The logical topology is predefined and includes the microservice call connection relationship between core network elements, the physical and logical layout of the wireless access network and the core network, and the end-to-end transmission path of typical business data flow. Obtain a feature set extracted from a standardized multi-source time series dataset using a time series analysis algorithm. The feature set includes periodic patterns and trend components of terminal group access behavior, as well as frequency domain features and residual sequences identified from business traffic fluctuations. Each pattern and component in the feature set is assigned and associated with the corresponding network topology nodes and communication links in the logical topology structure according to its corresponding business type and origin terminal category, forming node feature association mapping data; Based on node feature association mapping data, we analyze the statistical correlation between load levels of different nodes and link performance indicators in historical network state sequences, construct a state transition probability matrix with nodes and links as state dimensions, and establish a mathematical propagation model of load propagation between topology nodes based on business flow paths and load diffusion principles. By integrating the state transition probability matrix and the load propagation model, a network state evolution framework is constructed to simulate the dynamic evolution of network state with time and external feature input.

[0035] Specifically, in constructing the network state evolution framework, the logical topology of the industrial private network is first obtained from the network design document or configuration management database. This structure defines, in a machine-readable form, the calling dependencies between core network elements (such as microservices like AMF, SMF, UPF, and MME), the connection layout between wireless access network base stations and core network gateways, and the end-to-end transmission paths of typical services (such as AGV control commands, sensor data reporting, and IMS voice streams) from the terminal to the application server. Simultaneously, a structured feature set extracted from standardized multi-source time-series datasets using time-series analysis algorithms is received. This set includes periodic patterns such as the concentrated access of AGV groups every 2 hours, the trend component of the slow increase in the number of sensors, the frequency domain characteristics of the main fluctuation periods of video traffic, and the random residual sequences of various services. Subsequently, the mapping from features to topology is performed: based on the service type (e.g., control signaling, periodic reporting, voice stream) and originating terminal category (e.g., AGV, sensor, AR glasses) corresponding to the features, they are assigned and associated with relevant network topology nodes and communication links in the logical topology structure. For example, the periodic access pattern of AGVs is associated with the specific access base station serving it and the corresponding AMF node, and the traffic trend reported by sensors is associated with the C-SGN node and its data link with the UPF, thereby generating a node feature association mapping data that details which behavioral features each topology element is associated with. Based on this mapping data, the time series of network status in historical monitoring data is further analyzed, and the statistical correlation between the CPU load level of different nodes (e.g., UPFPod) and the latency performance index of related links (e.g., N6 interface) is calculated. Using these statistical laws, a state transition probability matrix is ​​constructed with the resource utilization rate of each node and the link performance index as the state dimension. At the same time, based on the principle of service flow path and network load diffusion (e.g., the overload of a UPF node will affect the downstream nodes on all service flow paths it processes), a mathematical propagation model describing the propagation and superposition effect of load between topology nodes is established. Finally, the state transition probability matrix describing random state changes is integrated with the mathematical propagation model describing deterministic load propagation to construct a network state evolution framework that can simulate the dynamic evolution of the overall network state (including all nodes and links) over time and with the changes in terminal behavior characteristics and traffic fluctuation characteristics from external inputs.

[0036] In some embodiments, stochastic process simulation is introduced into the network state evolution framework. Network state transitions are modeled using Markov decision processes, and a network state evolution model is generated by combining traffic fluctuation characteristics, including dynamic changes in terminal distribution density, spatiotemporal distribution of service traffic, and service quality constraints. In the network state evolution framework, the state space is defined by the resource indicators of topology nodes and the performance indicators of links, and the action space is defined by the resource configuration operations on nodes and links, thus constructing the basic structure of the Markov decision process. Based on the state transition probability matrix and load propagation model in the network state evolution framework, a state transition function of the Markov decision process is established. The function describes the probability distribution of the system transitioning from one state to another after performing a specific action. The traffic fluctuation characteristics extracted from historical business traffic data are obtained and injected into the state transition function as external input variables to modulate the dynamic process of state transition. The reward function is designed based on predefined service quality constraints. The reward function is calculated based on the performance indicators in the state space and the operation costs in the action space. By integrating the defined state space and action space, the established state transition function, and the designed reward function, a complete network state evolution model is generated that can simulate the dynamic changes in terminal distribution density, the spatiotemporal evolution of service traffic, and embed service quality constraints.

[0037] Specifically, in the final stage of constructing a complete network state evolution model, based on the generated network state evolution framework, the key resource indicators (including CPU utilization and memory utilization) of all topology nodes (such as microservice Pods) and the core performance indicators of communication links (including network latency and packet loss rate) are first enumerated and vectorized to define the state space of the Markov decision process. Simultaneously, a series of resource configuration operations that can be performed on the aforementioned nodes and links (such as adjusting Pod CPU limits, increasing or decreasing the number of replicas, and modifying link bandwidth policies) are encoded to define the action space of the Markov decision process, thus completing its basic structure construction. Subsequently, using the established state transition probability matrix (describing the statistical probability of state transitions) and load propagation model (describing the deterministic laws of load changes) within the framework, mathematical integration and formal expression are performed to construct the state transition function of the Markov decision process. This function precisely characterizes the probability distribution of the system evolving to various possible subsequent states after performing a specific action in any given state. Next, traffic fluctuation characteristics (such as periodic peaks and burst intensity) are extracted from historical business traffic data using methods such as spectrum analysis. These characteristics are quantified into time-varying parameters and dynamically injected into the aforementioned state transition function as key external input variables, thereby modulating the dynamic process of state transition to respond to simulated traffic changes. Simultaneously, based on predefined service quality constraints for each business (such as latency limits and availability targets), a reward function is designed. This function takes real-time performance indicators in the state space (used to measure SLA satisfaction) and resource costs and stability overhead caused by operations in the action space as inputs to calculate the immediate reward for each decision step. Finally, the defined state space and action space, the established state transition function modulated by external characteristics, and the designed reward function are systematically integrated to generate a fully functional network state evolution model. During runtime, this model can drive state transitions based on the input terminal activity patterns, thereby simulating the dynamic changes in terminal distribution density and the evolutionary distribution of business traffic in the spatiotemporal dimensions with high fidelity. Furthermore, its reward mechanism inherently incorporates considerations and optimization guidance regarding service quality constraints.

[0038] After constructing the network state evolution model, the simulation and generation phase of the digital twin environment data begins. The digital twin engine loads the network state evolution model and simultaneously receives a set of terminal behavior and traffic fluctuation characteristics generated by time series analysis as the core driving input. The engine starts the simulation, iteratively deducing the dynamic changes of the network state over simulation time based on the state transition function defined within the model and the externally input characteristic data. During the deduction process, the engine calculates and records in real time the number of terminal accesses in different geographical areas or logical groups in the network at each simulation moment, thereby generating data on the continuous change of terminal distribution density over time. Simultaneously, the model simulates the spatiotemporal dynamic fluctuation pattern of service traffic on various nodes and links in the network based on the input traffic fluctuation characteristics and the propagation effect of load in the topology. At the same time, the service quality constraints (such as latency and packet loss rate thresholds) preset in the model reward function are converted into monitoring logic, continuously evaluating whether the network state at each moment meets the constraints during the simulation process, and binding the evaluation result as a constraint label to the aforementioned dynamic data. Finally, the engine outputs a structured digital twin environment dataset, which fully includes the simulation's time series, the evolution trajectory of terminal distribution density, dynamic service traffic fluctuation patterns, and the corresponding service level agreement constraint satisfaction status at each moment. This dataset is then packaged and transmitted as the core data source for subsequent offline training and validation of reinforcement learning scheduling models.

[0039] S103: Based on the digital twin environment data, train the reinforcement learning model using a dual-delay deep deterministic policy gradient; In this embodiment, training the reinforcement learning model using a dual-delay deep deterministic policy gradient based on digital twin environment data includes: Encode historical operational data in the digital twin environment data into a state vector; The action space is defined based on the system observation dimension represented by the state vector; The reward value is calculated using a reward function based on the resource utilization rate and service level agreement compliance rate of the state vector and the frequency of action changes in the action space. Select actions from the action space and interact with the digital twin environment data to generate interaction data, and store the interaction data in the experience playback buffer pool; Batch data is sampled from the experience replay buffer, and the model parameters of the reinforcement learning model are iteratively optimized using a dual-network structure and a policy delay update mechanism based on the reward function, batch data, and reward value, until the model converges, thus obtaining the trained reinforcement learning model.

[0040] The training process encodes historical operational data from the digital twin environment into state vectors and defines an action space to represent resource allocation operations. Based on the resource utilization rate, service level agreement compliance rate, and action change frequency of the state vectors, a reward function is designed to calculate the reward value. Data generated by action-environment interaction is selected and stored in an experience replay buffer to achieve efficient reuse of historical interaction data. Data is sampled from the buffer, and the model parameters are iteratively optimized using a dual-network structure (online network and target network) and a policy delay update mechanism. The training stability is improved by dynamically modulating the policy update frequency, ensuring that the reinforcement learning model converges efficiently in scenarios with random traffic fluctuations.

[0041] Specifically, structured environment data generated by the digital twin engine is loaded, containing historical operation records throughout the complete simulation cycle. Data fields directly corresponding to predefined system observation dimensions are extracted from these records. For CPU utilization and memory utilization, resource usage snapshots of each microservice Pod at each simulation time point are obtained from the data and aggregated to obtain the overall average utilization percentage of the system. For UPF network latency, end-to-end or node-to-node latency measurements at corresponding time points are extracted from the simulated business flow performance logs. For NB-IoT connection count and IMS session count, the total number of active connections and sessions at each moment is counted from the simulated session management event logs. For SLA compliance rate, the comprehensive compliance percentage within each time window is calculated by comparing the business flow latency and packet loss rate recorded in the simulated monitoring logs with preset thresholds. Subsequently, the above six scalar indicators (CPU utilization) under the same simulation time stamp are... Memory usage UPF network latency NB-IoT connection count IMS session count SLA compliance rate The elements are combined in a fixed order to form a six-dimensional real vector. This vector is the state vector S corresponding to that specific simulation moment. t .

[0042] For example, This extraction and combination process is repeated for all historical simulation time points to generate a state vector sequence arranged in chronological order. This sequence fully encodes the historical evolution trajectory of the system state in the digital twin environment, providing well-formatted and semantically clear input state data for subsequent reinforcement learning training.

[0043] In defining the action space, the six system observation dimensions encoded by the state vector are first analyzed. For the CPU utilization and memory utilization dimensions in the state vector, vertical scaling operations are defined for specific microservice Pods running in the converged core network. This operation dynamically adjusts the CPU core quota and memory limit values ​​of the target Pod by calling the Kubernetes API to achieve real-time increases or decreases in resource allocation. For the NB-IoT connection count and IMS session count dimensions in the state vector, horizontal scaling operations are defined. This involves automatically adjusting the number of Pod replicas in the corresponding microservice workload controller based on the changing trends and preset thresholds of the connection count or session count. For example, when the number of NB-IoT connections surges, the number of Pod instances responsible for access management (AMF service) is increased. Furthermore, considering the periodic characteristics of the business load reflected in the state vector, especially when the number of IMS sessions is consistently low during business downtime (such as late at night), start / stop operations are defined for specified non-core IMS component Pods to release their occupied computing resources. Finally, the aforementioned set of vertical scaling operations, set of horizontal scaling operations, and set of microservice start / stop operations are integrated and defined as a discretized action space that is directly related to the system's observation dimensions and can be precisely executed through an orchestration platform.

[0044] In some embodiments, the reward value is calculated using a reward function based on the resource utilization rate and service level agreement compliance rate of the state vector and the frequency of action changes in the action space, including: The reward value is calculated using the following formula: in, Indicates the reward value; This represents the resource consumption weight, with a value ranging from 0 to 1. It is usually set to 0.3 to optimize resource conservation. This indicates the SLA compliance weight, with a value ranging from 0 to 1. It is usually set to 0.6 to prioritize performance. This represents the policy stability weight, with a value ranging from 0 to 1. It is usually set to 0.1 to ensure scheduling smoothness. Indicates the rate of resource waste; Indicates the service level agreement compliance rate; Indicates the frequency of action changes.

[0045] Furthermore, the constraint is: α+β+γ=1, with priority given to ensuring β=0.6 in the SLA.

[0046] During the interactive training of the reinforcement learning model and the digital twin environment, the current system state is first obtained from the digital twin environment, which has been encoded as a state vector S. tSubsequently, based on a predefined exploration strategy (e.g., adding noise to the output of a deterministic strategy), a specific action 'a' is selected from the defined action space. t This action might involve instructing the number of Pod replicas for the AMF service to increase to 5. Then, this action a... t As a control input, it is submitted to the digital twin environment. The network state evolution model within the environment extrapolates the state based on this action and the built-in state transition rules, simulating the network system's response after executing the action, and outputting the new state vector S for the next time step. t+1 At the same time, the environment, based on the state vector S t Action a to be performed t And the new state S reached t+1 The pre-defined reward function is called to calculate the immediate reward r obtained in this interaction step. t+1 Thus, a complete interaction has produced a four-tuple of data, namely (current state S) t Execute action a t Receive reward r t+1 Next state S t+1 Finally, this quadruple is added as a separate data entry to the experience replay buffer. The buffer, acting as a first-in-first-out queue or priority storage structure, continuously accumulates this type of interactive data, providing a training sample set for subsequent batch updates of model parameters.

[0047] In some embodiments, batch data is sampled from the experience replay buffer, and the model parameters of the reinforcement learning model are iteratively optimized based on the reward function, batch data, and reward value, using a dual-network structure and a policy delay update mechanism, until the model converges, resulting in the trained reinforcement learning model, including: A batch of historical interaction data is randomly sampled from the experience replay buffer. The historical interaction data contains multiple sets of state vectors, executed actions, immediate rewards, and subsequent state vectors. Based on a predefined reward function, and combined with the state vector and execution action of each group of data in the batch, the corresponding reward prediction value is calculated. The state vectors, actions, immediate rewards, subsequent state vectors, and calculated reward predictions in the batch are all input into a dual-network structure, which includes an actor network for generating action policies and a critic network containing two independent Q networks for evaluating the value of actions. In the commentator network, two independent Q networks are used to calculate the state-action value, and the smaller value is taken as the target value estimate. Based on the immediate reward, subsequent state and target value estimate in the batch data, the loss of the current Q network is calculated through temporal difference error. The model parameters of the critic network are updated using the calculated loss through backpropagation algorithm. After updating the critic network parameters a predetermined number of times, the actor network model parameters are then updated using the policy gradient method, thereby achieving policy delayed update. Repeat the iterative process of sampling data from the experience replay buffer, calculating reward predictions, updating commentator network parameters, and delaying actor network parameter updates until the model's output policy evaluation metric stabilizes and the reward predictions converge, thus obtaining a trained reinforcement learning scheduling model.

[0048] Specifically, a fixed-size batch of data is randomly selected from the experience replay buffer pool that has accumulated a large number of interaction records. This batch contains multiple independent sets of historical interaction quadruples. Each set of data consists of a historical state vector, the action performed in that state, the immediate reward from the environmental feedback, and the subsequent state vector after the action is performed.

[0049] Based on the predefined reward function, the historical state vector and execution action of each group of data in the batch are used as input to calculate a corresponding reward prediction value. This value represents the model's expected evaluation of the state-action pair under the current policy.

[0050] Subsequently, all state vectors, executed actions, actual immediate rewards, subsequent state vectors, and calculated reward predictions from this batch of data are input into a dual-network structure employing a dual-delay deep deterministic policy gradient algorithm architecture. This structure includes an actor network responsible for outputting specific action policies based on the input state vectors; and a critic network containing two structurally identical but parameter-independent Q-networks used to evaluate the value of a given state-action pair.

[0051] During training, within the critic network, two independent Q-networks are used to calculate the value of each state-action pair in the batch. The smaller of the two calculation results is then selected as a more conservative and stable estimate of the target value to reduce overestimation bias.

[0052] Next, based on the actual immediate reward from the environmental feedback in the batch data, the subsequent state vector transitioned to, and the calculated target value estimate, the loss function value of the currently updated Q-network is calculated using the temporal difference error formula. Using the backpropagation algorithm, the model parameters of the critic network are updated with a gradient based on the calculated loss value. This process is repeated, i.e., sampling from the buffer pool and updating the critic network parameters multiple times consecutively. After reaching a preset fixed number of iterations (e.g., updating the critic network twice), the model parameters of the actor network are calculated and updated using the policy gradient method based on the action policy output by the actor network and the value evaluated by the critic network, thus achieving delayed updates to the policy network to enhance training stability. This complete iterative process of sampling from the buffer pool, calculating reward predictions, updating the critic network, and delayed updating the actor network is repeated. Model performance is evaluated after each iteration, for example, by monitoring the trend of the average reward value. When the model's output policy evaluation metrics (such as average reward) remain stable and the reward prediction value tends to converge over a continuous training period, training is stopped, resulting in a trained and parameter-stable reinforcement learning scheduling model.

[0053] In some embodiments, it also includes: Collect current operating data of the converged core network, including real-time resource indicators, service connection count indicators, and service level agreement indicators; Based on the collected current operating data, a time series prediction model is used to calculate the predicted values ​​of business traffic trends and terminal connection trends within a future preset time window. The time series prediction model is built based on a long short-term memory network or time series transformer architecture. The current operating data is fused with the calculated predicted values ​​to construct an enhanced state vector, which includes both instantaneous indicators reflecting the current system state and predicted indicators reflecting future short-term changes. The enhanced state vector is input into a pre-trained reinforcement learning scheduling model; The reinforcement learning scheduling model outputs a resource scheduling action based on the input reinforcement state vector. This action is used to pre-adjust the resource configuration of the converged core network before the predicted peak of service load arrives.

[0054] This embodiment constructs a complete view of the system's current state by collecting current operating data from the converged core network, including real-time resource indicators, service connection count indicators, and service level agreement (SLA) indicators. A time series prediction module based on a long short-term memory (LSTM) network or time series transformer architecture calculates the service traffic trend and terminal connection count trend within a preset future time window, generating high-precision predicted values. The current operating data and predicted values ​​are fused to form an enhanced state vector, which simultaneously contains indicators reflecting the system's instantaneous state and predicted indicators reflecting short-term future changes. After inputting the enhanced state vector into a reinforcement learning scheduling model, the model dynamically generates resource scheduling actions based on historical states and future expectations. This enables proactive pre-adjustment of resource allocation before the predicted load peak arrives, avoiding resource allocation delays and SLA default risks caused by relying on real-time responses in traditional solutions, and significantly improving the foresight of scheduling decisions and the system's robustness.

[0055] Specifically, during the real-time operation phase of the system, to achieve accurate perception of the converged core network, the system initiates a periodic data collection process. Monitoring agents deployed on each worker node of the Kubernetes cluster, and sidecar containers deployed within core network microservice Pods, are woken up at fixed time intervals (e.g., every 10 seconds) to execute collection tasks. The monitoring agents obtain real-time resource consumption data for all Pods on their respective nodes through the container runtime interface, including the CPU utilization percentage and memory usage percentage for each Pod, while simultaneously collecting node-level load metrics from the host operating system. For service connection metrics, the collection module obtains the total number of currently active 5G sessions, the number of established NB-IoT connections, and the number of concurrent voice and video sessions in the IP multimedia subsystem by calling the northbound management interface of core network control plane elements (such as AMF and SMF) or listening to their internal event bus. For service level agreement (SLA) metrics, the end-to-end transmission latency of critical service data packets is measured using passive probes deployed on data plane paths (such as N3 and N6 interfaces) or by utilizing performance counters provided by the user plane function elements themselves, and the packet loss rate of user plane function data packet processing is statistically analyzed. All collected raw indicator data are appended with precise timestamps and source tags, and are reported in real time to the central time-series database through a lightweight message channel for aggregation and storage, forming a complete dataset that reflects the current operating status of the converged core network.

[0056] Next, the current operational data of the converged core network is acquired. This dataset includes real-time resource metrics, service connection count metrics, and service level agreement (SLA) metrics from the most recent collection period. This current operational data sequence is then input into a pre-deployed time series prediction model. This model is built on a Long Short-Term Memory (LSTM) network or time series transformer architecture and has been trained using historical data. The model processes the input time series data to calculate predicted trends in service traffic and terminal connection counts for a future preset time window (e.g., the next 5 minutes).

[0057] Next, the operational data reflecting the current instantaneous state of the system is fused with the calculated future trend predictions, and encoded into a higher-dimensional enhanced state vector. This vector not only includes instantaneous indicators such as CPU utilization, memory utilization, current connection count, and real-time latency, but also adds predicted traffic load and connection count change trend indicators.

[0058] Then, this enhanced state vector is input into a pre-trained reinforcement learning scheduling model. The reinforcement learning model infers based on this panoramic state view containing future information and outputs an optimal resource scheduling action. For example, 1-2 scheduling cycles before the predicted peak load actually arrives, the instruction expands the number of replicas of user plane function instances from 3 to 5 and pre-allocates additional CPU computing resources to them. This pre-adjustment action aims to improve the system's service capacity in advance to smoothly cope with upcoming business peaks, thereby achieving better resource utilization efficiency than passive reactive scaling while ensuring service level agreements are maintained.

[0059] S104: Collect resource indicators, business indicators, and service level agreement indicators of the fusion resource pool in real time to form status data; Specifically, during the real-time operation phase, monitoring agents deployed on each node and key network element of the Kubernetes cluster initiate data collection tasks at fixed intervals of 10 seconds. Regarding resource metrics, the agents obtain the real-time CPU utilization and memory usage percentages of each microservice Pod through the container runtime interface, and collect overall host load data through the node resource manager. For business metrics, the agents collect the total number of currently active 5G sessions by querying session management statistics of 5G core network elements; they also count the number of NB-IoT uplink messages received within a period by monitoring the message processing queue of the narrowband IoT core; and simultaneously obtain the total number of concurrent voice and video sessions through the session control layer interface of the IP multimedia subsystem. Regarding service level agreement (SLA) metrics, the agents measure the end-to-end round-trip latency of selected service flows using timestamp probes deployed on the data plane path; and they count the packet loss rate of data packets processed by the counters of user plane function network elements. All collected raw metric data is encapsulated by the agents and synchronously reported to the central monitoring service. After receiving the data, the monitoring service aligns, verifies, and aggregates all metrics according to a unified time window. For example, it calculates the average CPU utilization of all UPFPods and counts the total number of NB-IoT messages across the entire network. The processed, time-synchronized set of multi-dimensional metrics is organized into a structured snapshot of state data. This snapshot fully contains quantitative information on the three observation planes of resources, services, and SLA, providing real-time input for subsequent system state coding and scheduling decision triggering.

[0060] S105: When the state data meets any of the preset triggering conditions, the state data is input into the trained reinforcement learning scheduling model to obtain resource scheduling actions; In this embodiment, the preset triggering conditions include: CPU utilization exceeding a first threshold and the duration reaching a first time, end-to-end latency continuously exceeding a second time, or the number of newly accessed terminals exceeding a second threshold within a unit time.

[0061] During the real-time operation and dynamic decision-making phase of the system, the monitoring module continuously receives status data snapshots generated every 10 seconds. This module has built-in condition judgment logic that analyzes these snapshot data in real time to detect whether any preset trigger condition is met. For the CPU utilization condition, the module calculates the moving average of CPU utilization for a specified Pod or node group. If this average remains above a first threshold of 70% for 60 seconds, the condition is considered met. For the end-to-end latency condition, the module checks the latency sequence in the SLA metrics. If the latency measured for three consecutive collection cycles (i.e., 30 seconds) is greater than 15 milliseconds, the trigger condition is considered met. For the new terminal access condition, the module analyzes the rate of change of the number of 5G sessions and NB-IoT connections in the service metrics, calculates the number of new accesses per minute, and if this number exceeds a second threshold of 100 units, it is also considered a trigger event. Once any of the above conditions is triggered, the monitoring module immediately formats and encapsulates the complete status data snapshot (including all resource, service, and SLA metrics) corresponding to the trigger time. Subsequently, the system invokes the trained reinforcement learning scheduling model deployed in the inference service, submitting the formatted state data as input for forward computation. Based on its learned optimal policy, the model performs inference analysis on the input state and ultimately outputs a specific, executable resource scheduling action instruction, such as expanding the number of replicas deployed in UPF to 4 or adding a resource limit of 2 CPU cores to AMFPod.

[0062] S106: Execute the resource scheduling action through the open-source container orchestration platform interface and generate the execution result.

[0063] Specifically, the executor parses resource scheduling actions and converts them into one or more operation requests conforming to the Kubernetes API specification of the open-source container orchestration platform. For example, if the instruction is to expand the number of Pod replicas deployed in AMF to 5, the executor will generate a sub-resource update request for the corresponding deployment resources, with the target replica number field set to 5. Subsequently, the executor uses authentication credentials with appropriate permissions to send the operation request by calling the APIServer endpoint of the Kubernetes cluster. After receiving the request, the APIServer performs authentication and verification, and finally sends it to the corresponding controller manager to execute the actual scaling operation. The executor synchronously starts a monitoring coroutine to continuously query the status of relevant Pods or deployments to track the execution progress of the scheduling action in real time, such as confirming whether the new Pod instance has been successfully scheduled and entered the running state. When the action is completed or times out, the executor collects the final execution status information, such as whether the new number of replicas is ready, the specified resource limits have been updated, or errors such as insufficient resources were encountered during execution. This execution status, along with the timestamps of the start and end of the execution, is encapsulated into a structured execution result report.

[0064] In this embodiment, after executing resource scheduling actions through the open-source container orchestration platform interface and generating the execution result, the following is included: Monitor the service level agreement indicators of the converged core network; When a service level agreement (SLA) metric is violated, the configuration of the converged core network will be rolled back to the secure configuration prior to the execution of resource scheduling actions.

[0065] Specifically, after executing resource scheduling actions and generating execution results through the open-source container orchestration platform interface, the system immediately initiates post-execution monitoring and security assurance processes. The system continuously monitors key service level agreement (SLA) metrics of the converged core network, focusing on service performance directly or indirectly affected by the resource scheduling actions. For example, for UPF expansion actions, it continuously tracks the end-to-end latency of related data streams; for AMF downsizing actions, it monitors the success rate of control plane signaling processing. The monitoring module compares the collected real-time SLA metrics with preset assurance thresholds, such as determining whether latency has decreased to the expected range or whether the signaling success rate remains above the requirements. When monitoring detects a violation of any key SLA metric, such as latency not decreasing to below 10 milliseconds within 3 minutes after UPF expansion, the system determines that the scheduling action failed to achieve the expected optimization effect and may have introduced service risks. At this point, the security control module is activated. This module has persistently saved all relevant configuration snapshots of the affected network elements as a security baseline before each scheduling action. Based on the type of violation event triggered, the module locates the corresponding security configuration snapshot and issues a rollback command to the cluster via the Kubernetes API. This could involve restoring the number of previously expanded UPFDeployment replicas to their original value, or reverting adjusted Pod resource limits to their previous values. The cluster controller executes these rollback operations, restoring the configuration state of the relevant microservices to a known safe state before the scheduling action, thereby quickly terminating the SLA violation and ensuring the stability of network services.

[0066] In this embodiment, after executing resource scheduling actions through the open-source container orchestration platform interface and generating the execution result, the following is included: The actual reward value is calculated based on the actual running data in the execution results, and then compared with the predicted reward value of the reinforcement learning scheduling model. When the actual reward value is lower than the preset proportion of the predicted reward value, the model parameters of the reinforcement learning scheduling model are incrementally updated.

[0067] This mechanism calculates the reward value of the actual running data in the execution results in real time and dynamically compares it with the reward value predicted by the reinforcement learning scheduling model. When the actual reward value is lower than the predicted value by a preset proportion, it automatically triggers an incremental update of the model parameters. This design enables the model to continuously absorb the latest network running data, dynamically correct prediction biases, and avoid scheduling strategy inaccuracies caused by sudden environmental changes. At the same time, the incremental update mechanism significantly reduces model iteration overhead, ensuring the real-time performance and accuracy of scheduling decisions in high-concurrency traffic fluctuation scenarios in industrial private networks, and effectively preventing the risk of service level agreement (SLA) defaults.

[0068] Specifically, the system extracts actual operational data from the monitoring data storage within a complete observation window after the scheduling action is executed. This data includes: average CPU and memory utilization of each relevant Pod, end-to-end latency compliance statistics, and the number of Pod changes caused by the scheduling action during this period. Using the same reward function formula as during model training, the system substitutes this actual data into the calculation: resource consumption is calculated based on actual utilization to determine resource waste rate; SLA compliance is calculated based on the compliance of indicators such as latency; and policy stability is calculated based on the action execution frequency to determine operational cost. Combining these three factors and their weight coefficients, an actual reward value R_actual, representing the overall effect of this scheduling action, is calculated. Simultaneously, the system retrieves the expected reward output by the reinforcement learning model based on the input state at the time the scheduling was triggered from the decision log, i.e., the predicted reward value R_predicted. Subsequently, the system compares the actual reward value R_actual with the predicted reward value R_predicted, calculating their relative ratio (R_actual / R_predicted). This comparison result is recorded and used as a key evaluation indicator to determine the model's prediction accuracy in the current real-world environment.

[0069] When the actual reward value is lower than a preset percentage (e.g., 80%) of the predicted reward value, a batch of historical interaction data is randomly sampled from the experience replay buffer. This data includes state vectors, actions, immediate rewards, and subsequent states. The predicted reward value is calculated based on the reward function and the batch data. The batch data is then input into a dual-network structure, where the commentator network contains two independent Q-networks. The state-action values ​​are calculated separately, and the smaller value is taken as the target value estimate. The Q-network loss is calculated using temporal difference error, and the commentator network model parameters are updated using the backpropagation algorithm. After the commentator network parameters have been updated a preset fixed number of times, the actor network model parameters are updated using the policy gradient method. The sampling, calculation, and update iteration process is repeated until the policy evaluation index output by the model stabilizes, thus achieving incremental updates of the reinforcement learning scheduling model parameters.

[0070] In the simplified private network deployment example for an automotive factory, the initialization phase is performed first. A cloud-native converged core network, integrating 4G EPC, 5GC, NB-IoT core network, and IMS components, is deployed in a Kubernetes cluster consisting of three physical nodes. The initial shared computing resources total 32 CPU cores and 64GB of memory. Simultaneously, based on the collected historical AGV movement trajectory data and periodic reporting feature data from the NB-IoT sensors, a digital twin environment is constructed and a reinforcement learning scheduling model is trained, completing the system preparation.

[0071] After entering the operational phase, the system encountered a typical industrial scenario: at the start of the morning shift, 500 AGV devices simultaneously started and connected to the network, causing a 40% surge in user plane traffic within a short period. The monitoring layer detected this anomaly in real time, and status data showed that UPF-related indicators had triggered preset conditions. The system then invoked its pre-trained reinforcement learning model for decision-making. Based on the current system state and historical learning experience, the model output the optimal resource scheduling action: horizontally expanding the number of Pod instances for the UPF microservice from 2 to 4, and simultaneously vertically increasing the reserved CPU resources for these Pods, increasing the total allocation from 8 cores to 16 cores. After executing these actions via the Kubernetes API within seconds, the network user plane processing capacity was instantly enhanced. Performance verification showed that the end-to-end latency of the business flow rapidly decreased from 15 milliseconds at the time of the event to 8 milliseconds, meeting the SLA requirement of less than 10 milliseconds, while the overall infrastructure resource consumption increased by this elastic scaling was only 22%. Compared to traditional solutions that typically require reserving and scaling resources by more than 50% to cope with similar peak traffic, this solution significantly improves resource efficiency while ensuring performance.

[0072] In an embodiment of NB-IoT large-scale access optimization, the system handled a sudden event during operation. At 3:00 AM, due to task scheduling leading to a compressed reporting cycle, 1900 NB-IoT sensors in the factory simultaneously accessed and reported data. Business metrics showed that the number of NB-IoT sessions surged to 1900, causing the CPU utilization of the AMF network element responsible for access control to spike to 95%, posing an overload risk. At this time, the user plane latency remained at a normal level of 5 milliseconds. The monitoring and acquisition layer continuously reported these abnormal metrics.

[0073] When the CPU utilization of the AMF exceeds the 90% threshold and remains above it for 30 seconds, the system meets the preset decision trigger conditions. The reinforcement learning scheduling model is activated to make real-time decisions. The model not only analyzes the current real-time state but also incorporates learning from historical data patterns, executing a series of compound actions: First, it adopts a predictive strategy, referring to historical patterns to preemptively expand AMF instances to cope with the expected sustained load; second, it performs intelligent resource reclamation, recognizing that the IMS voice service is currently in a low-intensity period, and decides to scale down non-core IMS components. Specifically, it expands the number of AMF deployment replicas to 5 via the Kubernetes command line, while reducing the number of IMS deployment replicas to 1.

[0074] The results were validated after execution: AMF's CPU utilization significantly decreased from a high load of 95% to 62%, avoiding service overload and potential signaling loss. By scaling down idle IMS components, computing resources from four CPU cores were freed up. This dynamic scheduling significantly reduced the system's resource waste rate from 61% in the traditional static deployment mode to 19% when dealing with sudden large-scale IoT access, resulting in an overall resource consumption reduction of 18%. It achieved efficient and precise elastic scheduling of the resource pool while ensuring the stability of core business operations.

[0075] This embodiment constructs a cloud-native core network architecture that integrates multiple standards, including 4G evolved packet core network components, 5G core network components, narrowband IoT technology core components, and IP multimedia subsystem components. Each network element is deployed uniformly as a containerized microservice within an open-source container orchestration platform cluster, fundamentally eliminating resource silos and enabling the sharing and reuse of computing resources. This significantly improves resource utilization and reduces capital expenditure. Simultaneously, dynamic resource scheduling is achieved by introducing a reinforcement learning model based on dual-latency deep deterministic policy gradient (TD3). This reinforcement learning model can automatically make scaling decisions based on real-time monitoring data and continuously optimize through an online learning mechanism, replacing the traditional rule engine that relies on static thresholds. This significantly improves the resource efficiency, deployment flexibility, and service continuity of the industrial private network.

[0076] Please see Figure 5 This embodiment provides an industrial private network converged deployment device 200, including: Deployment unit 201 is used to deploy a converged core network in an open-source container orchestration platform cluster to form a converged resource pool; the converged core network includes 4G evolved packet core network components, 5G core network components, narrowband IoT technology core components, and IP multimedia subsystem components; Construction unit 202 is used to collect terminal type, traffic matrix and service level agreement requirements data to construct digital twin environment data; Training unit 203 is used to train a reinforcement learning model using a dual-delay deep deterministic policy gradient based on the digital twin environment data. The acquisition unit 204 is used to collect the resource indicators, business indicators and service level agreement indicators of the fusion resource pool in real time to form status data; The input unit 205 is used to input the state data into the trained reinforcement learning scheduling model when the state data meets any one of the preset triggering conditions, so as to obtain the resource scheduling action. The execution unit 206 is used to execute the resource scheduling action through the open-source container orchestration platform interface and generate the execution result.

[0077] Furthermore, the deployment unit 201 includes: The sub-unit is used to split the mobility management entity, service gateway, and packet data network gateway of the 4G evolved packet core network component into microservices for deployment; The unit deployment subunit is used to deploy the access and mobility management functions, session management functions, and user plane functions of the 5G core network components in the form of independent minimum deployment units; The microservice subunit is used to microservice the cellular IoT service gateway node, a core component of the narrowband IoT technology. The shared subunit is used to share the data plane of the user plane function with the proxy call session control function and the core control node of the IP multimedia subsystem component.

[0078] Furthermore, the training unit 203 includes: The data encoding subunit is used to encode the historical operation data in the digital twin environment data into a state vector; A spatial definition subunit is used to define the action space based on the system observation dimension represented by the state vector; The function calculation subunit is used to calculate the reward value through the reward function based on the resource utilization rate and service level agreement compliance rate of the state vector and the action change frequency of the action space. The interaction subunit is used to select actions from the action space and interact with the digital twin environment data to generate interaction data, and store the interaction data in the experience playback buffer pool. The iterative optimization subunit is used to sample batch data from the experience replay buffer pool, and based on the reward function, the batch data, and the reward value, iteratively optimize the model parameters of the reinforcement learning model using a dual network structure and a policy delay update mechanism until the model converges, thus obtaining the trained reinforcement learning model.

[0079] Furthermore, the function computation subunit includes: The reward value calculation subunit is used to calculate the reward value according to the following formula: in, Indicates the reward value; Indicates the weight of resource consumption; Indicates the SLA compliance weight; Indicates the strategy stability weight; Indicates the rate of resource waste; Indicates the service level agreement compliance rate; Indicates the frequency of action changes.

[0080] Furthermore, the preset triggering conditions include: CPU utilization exceeding a first threshold for a duration of a first time, end-to-end latency continuously exceeding a second time, or the number of newly connected terminals exceeding a second threshold within a unit time.

[0081] Furthermore, the execution unit 206 includes: The indicator monitoring subunit is used to monitor the service level protocol indicators of the converged core network. The rollback subunit is used to roll back the configuration of the converged core network to the secure configuration before the resource scheduling action was performed when the service level agreement indicator is violated.

[0082] Furthermore, the execution unit 206 includes: The comparison subunit is used to calculate the actual reward value based on the actual running data in the execution result, and compare the actual reward value with the predicted reward value of the reinforcement learning scheduling model. The incremental update subunit is used to incrementally update the model parameters of the reinforcement learning scheduling model when the actual reward value is lower than a preset proportion of the predicted reward value.

[0083] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described apparatus and unit can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0084] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed, can implement the methods provided in the above embodiments. The storage medium may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0085] The present invention also provides a computer device, which may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, it can implement the methods provided in the above embodiments. Of course, the computer device may also include various network interfaces, power supplies, and other components.

[0086] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this invention.

[0087] It should also be noted that, in this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusivity.

[0088] The term "comprises" implies that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

Claims

1. An industrial private network convergence deployment method, characterized in that, include: Deploy a converged core network in an open-source container orchestration platform cluster to form a converged resource pool; The converged core network includes 4G evolved packet core network components, 5G core network components, narrowband Internet of Things technology core components, and IP multimedia subsystem components. Collect data on terminal types, traffic matrices, and service level agreement requirements to construct digital twin environment data; Based on the digital twin environment data, a reinforcement learning model is trained using a dual-delay deep deterministic policy gradient. The system collects resource metrics, business metrics, and service level agreement metrics of the integrated resource pool in real time to form status data. When the state data meets any of the preset triggering conditions, the state data is input into the trained reinforcement learning scheduling model to obtain resource scheduling actions. The resource scheduling action is executed through the interface of the open-source container orchestration platform, and the execution result is generated.

2. The industrial private network converged deployment method according to claim 1, characterized in that, The deployment of the converged core network in the open-source container orchestration platform cluster includes: The mobility management entity, service gateway, and packet data network gateway of the 4G evolved packet core network components are decomposed into microservices and deployed. The access and mobility management functions, session management functions, and user plane functions of the 5G core network components are deployed in the form of independent minimum deployment units; Microservice-based cellular IoT service gateway nodes, which are core components of the narrowband IoT technology; The proxy call session control function and core control node of the IP multimedia subsystem components share the data plane of the user plane function.

3. The industrial private network converged deployment method according to claim 1, characterized in that, The step of training the reinforcement learning model using a dual-delay deep deterministic policy gradient based on the digital twin environment data includes: The historical operational data in the digital twin environment data is encoded into a state vector; The action space is defined based on the system observation dimension represented by the state vector; The reward value is calculated using the reward function based on the resource utilization rate and service level agreement compliance rate of the state vector and the action change frequency of the action space. Actions are selected from the action space and interacted with the digital twin environment data to generate interaction data, which is then stored in the experience playback buffer pool. Batch data is sampled from the experience replay buffer, and the model parameters of the reinforcement learning model are iteratively optimized using a dual-network structure and a policy delay update mechanism based on the reward function, the batch data, and the reward value, until the model converges, thus obtaining the trained reinforcement learning model.

4. The industrial private network converged deployment method according to claim 3, characterized in that, The reward value calculated using the reward function based on the resource utilization rate and service level agreement compliance rate of the state vector and the action change frequency of the action space includes: The reward value is calculated using the following formula: in, Indicates the reward value; Indicates the weight of resource consumption; Indicates the SLA compliance weight; Indicates the strategy stability weight; Indicates the rate of resource waste; Indicates the service level agreement compliance rate; Indicates the frequency of action changes.

5. The industrial private network converged deployment method according to claim 1, characterized in that, The preset triggering conditions include: CPU utilization exceeding a first threshold and the duration reaching a first time, end-to-end latency continuously exceeding a second time, or the number of newly connected terminals exceeding a second threshold within a unit time.

6. The industrial private network converged deployment method according to claim 1, characterized in that, The step of executing the resource scheduling action through the open-source container orchestration platform interface and generating the execution result includes: Monitor the service level agreement (SLP) metrics of the converged core network; When the service level agreement indicator is violated, the configuration of the converged core network is rolled back to the secure configuration before the resource scheduling action was performed.

7. The industrial private network converged deployment method according to claim 1, characterized in that, The step of executing the resource scheduling action through the open-source container orchestration platform interface and generating the execution result includes: The actual reward value is calculated based on the actual running data in the execution result, and the actual reward value is compared with the predicted reward value of the reinforcement learning scheduling model. When the actual reward value is lower than a preset proportion of the predicted reward value, the model parameters of the reinforcement learning scheduling model are incrementally updated.

8. An industrial private network converged deployment device, characterized in that, include: Deployment units are used to deploy the converged core network in an open-source container orchestration platform cluster to form a converged resource pool; The converged core network includes 4G evolved packet core network components, 5G core network components, narrowband Internet of Things technology core components, and IP multimedia subsystem components. The building unit is used to collect data on terminal type, traffic matrix, and service level agreement requirements to build digital twin environment data; The training unit is used to train the reinforcement learning model using a dual-delay deep deterministic policy gradient based on the digital twin environment data. The data acquisition unit is used to collect resource indicators, business indicators, and service level agreement indicators of the fusion resource pool in real time to form status data. The input unit is used to input the state data into the trained reinforcement learning scheduling model when the state data meets any one of the preset trigger conditions, so as to obtain the resource scheduling action. The execution unit is used to execute the resource scheduling action through the open-source container orchestration platform interface and generate the execution result.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the industrial private network converged deployment method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to perform the industrial private network converged deployment method as described in any one of claims 1 to 7.