Intelligent service starting arrangement system based on container life cycle

By using an intelligent service startup orchestration system, the problems of fixed startup order, lack of pre-detection, single fault recovery, and rudimentary dependency management in container orchestration systems are solved. It realizes dynamic startup, multi-dimensional detection, intelligent recovery, and full lifecycle management, thereby improving startup success rate, resource utilization, and fault diagnosis efficiency.

CN121116486APending Publication Date: 2025-12-12CHINA IND INTERNET RES INST
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511174630.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing container orchestration systems cannot dynamically adjust the service startup order, lack a robust pre-detection mechanism, have simplistic fault recovery strategies, poor dependency management, and lack lifecycle awareness, leading to startup failures and resource waste.

Method used

An intelligent service launch orchestration system based on container lifecycle is adopted, including an orchestration control layer, a pre-detection layer, and a lifecycle management layer. It uses the DAG algorithm to manage complex dependencies, and features multi-dimensional pre-detection, intelligent fault recovery, and full-stage lifecycle monitoring.

Benefits of technology

Significantly improves startup success rate, optimizes resource utilization, enhances fault recovery efficiency, increases lifecycle visibility and automation, shortens fault diagnosis time, and improves scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121116486A_ABST
    Figure CN121116486A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent service starting arrangement system based on container life cycles, and particularly relates to the field of containerized starting service management. Comprising an arrangement control layer, a pre-detection layer and a life cycle management layer. The arrangement control layer comprises a starting manager module, a dependency relationship analysis sub-module, an arrangement strategy control sub-module and a starting sequence scheduling sub-module; the pre-detection layer comprises a multi-dimensional pre-detection engine, a port availability detection sub-module, a database connection verification sub-module, a configuration integrity verification sub-module and a resource ready state check sub-module. The life cycle management layer comprises a container life cycle monitoring module, a construction stage management sub-module, an operation state tracking sub-module and a stop process control sub-module. A multi-dimensional pre-detection engine, DAG-based complex dependency relationship management, intelligent fault classification and adaptive recovery, container life cycle full-stage perception and dynamic starting of an arrangement strategy are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of containerized startup service management, and more specifically to an intelligent service startup orchestration system based on the container lifecycle. Background Technology

[0002] Existing container orchestration systems such as Kubernetes, Docker Swarm, and Docker Compose employ a declarative configuration approach, including modules for container definition, resource scheduling, service discovery, and health checks. Their working principle is as follows: users define container runtime specifications and dependencies; the orchestration engine creates container instances based on the configuration file; it monitors container status through heartbeat detection; and it executes a restart strategy when a container encounters an anomaly.

[0003] A simple service startup script, a service startup system based on shell scripts or batch files, includes a service startup module, a dependency checking module, a logging module, and an error handling module. Its working principle is: start each service process in a predefined order, check the process startup status, record startup logs, and output error messages and terminate execution when startup fails.

[0004] Application frameworks such as Spring Boot and Node.js have built-in service management mechanisms, including application startup modules, configuration loading modules, component initialization modules, and lifecycle management modules. The working principle is as follows: when the application starts, it loads various components according to the configuration, executes initialization methods, establishes dependencies between components, and provides a unified shutdown interface.

[0005] The shortcomings of existing technology Fixed startup order: Existing systems typically use static startup order configurations, which cannot dynamically adjust the service startup order according to the actual operating environment. When a dependent service starts slowly, it can cause the entire startup process to be blocked.

[0006] Lack of pre-detection mechanism: Most container orchestration systems lack a robust pre-startup detection mechanism, which cannot verify prerequisites such as port availability, database connection status, and configuration file integrity before service startup, leading to immediate service failure after startup.

[0007] Limited fault recovery strategies: Existing systems typically rely on simple container restarts for fault recovery, lacking intelligent fault analysis and differentiated recovery strategies, and failing to adopt different recovery solutions based on fault type.

[0008] Crude dependency management: The dependency management between services is too simple, only able to indicate the basic startup order, and unable to describe complex dependency conditions, such as "service A can only start when the database connection is available and port 8080 is idle".

[0009] Lack of lifecycle awareness: Existing systems cannot perceive the complete lifecycle state of containers and cannot distinguish between the container's build, startup, running, and stopping phases, leading to the execution of certain operations at inappropriate times. Summary of the Invention

[0010] The purpose of this invention is to address the above-mentioned shortcomings by proposing an intelligent service startup orchestration system based on the container lifecycle, which enables dynamic startup orchestration based on real-time status, provides a comprehensive multi-dimensional pre-detection mechanism, establishes intelligent fault analysis and adaptive recovery strategies, supports accurate expression and management of complex dependencies, and realizes full-stage perception and management of the container lifecycle.

[0011] The present invention specifically adopts the following technical solution: An intelligent service launch orchestration system based on container lifecycle, including an orchestration control layer, a pre-detection layer, and a lifecycle management layer; The orchestration control layer includes a startup manager module, a dependency resolution submodule, an orchestration strategy control submodule, and a startup order scheduling submodule; The startup manager module is responsible for resolving service dependencies, formulating startup strategies, and controlling the startup process. The dependency resolution submodule uses the Directed Acyclic Graph (DAG) algorithm to represent and manage complex dependencies between services; The orchestration strategy control submodule dynamically adjusts the startup strategy based on the current system state; The startup sequence scheduling submodule implements an intelligent scheduling algorithm based on priority and dependency. The pre-detection layer includes a multi-dimensional pre-detection engine, a port availability detection submodule, a database connection verification submodule, a configuration integrity verification submodule, and a resource readiness status check submodule; The multi-dimensional pre-detection engine performs a series of pre-detection operations before the service actually starts, ensuring that the startup environment meets all prerequisites. The port availability detection submodule is responsible for checking the availability status of the ports required by the service. It tests port occupancy by creating temporary socket connections and supports detection of TCP and UDP ports. The database connection verification submodule establishes test connections to various database services to verify the correctness of connection parameters and the availability of database services. The configuration integrity verification submodule verifies the integrity of the configuration files and environment variables required for service startup; The resource readiness status check submodule monitors the availability status of system resources; The lifecycle management layer includes a container lifecycle monitoring module, a build phase management submodule, a runtime status tracking submodule, and a stop process control submodule. The container lifecycle monitoring module enables the tracking and management of containers from creation to destruction, providing accurate status information for intelligent orchestration; The build phase management submodule monitors the container image build process; The startup phase monitoring submodule tracks each stage of the container startup process in a fine-grained manner; The runtime status tracking submodule continuously monitors the runtime status of the container; The stop process control submodule manages the graceful shutdown process of containers, ensuring data integrity and consistency. This invention has the following beneficial effects: The startup success rate is significantly improved. Through a multi-dimensional pre-detection mechanism, the system can identify and resolve more than 90% of common startup problems before the service starts, including port conflicts, configuration errors, and unavailability of dependent services. The startup success rate has increased from 75% in the traditional solution to more than 95%.

[0012] The predictability of startup time is enhanced. Based on historical startup data and the current system status, the dynamic scheduling algorithm can accurately predict service startup time with a startup time prediction accuracy of over 90%, providing strong support for capacity planning and user experience optimization.

[0013] Improved fault recovery efficiency: The intelligent fault recovery engine can automatically select the most suitable recovery strategy based on the fault type, reducing the average fault recovery time from 5 minutes to 1 minute, and achieving an automatic recovery success rate of 85%.

[0014] Precise dependency management: Based on the representation of complex dependencies using directed acyclic graphs, it can accurately describe and manage complex dependency conditions between services, avoid startup sequence errors caused by unclear dependencies, and achieve 99% system startup consistency.

[0015] Resource utilization is optimized. Through resource readiness checks and dynamic scheduling strategies, the system can avoid startup failures caused by resource contention, while improving resource utilization. Peak CPU utilization is reduced by 30%, and memory usage is more balanced.

[0016] Improved lifecycle visibility: Full-stage container lifecycle monitoring provides complete service status visibility, allowing operations personnel to understand the detailed status of each service in real time, reducing fault location time by 60%.

[0017] With enhanced automation, the system's automatic pre-detection, intelligent recovery, and adaptive scheduling mechanisms significantly reduce the need for manual intervention, reducing maintenance workload by 50% while improving the consistency of maintenance quality.

[0018] The fault diagnosis capability has been enhanced. Detailed fault classification and recovery strategies provide strong support for fault diagnosis. The diagnosis time for new faults has been shortened from an average of 30 minutes to 10 minutes, and the diagnosis accuracy has been improved to over 90%.

[0019] The scalability is significantly improved. The modular architecture design allows the system to easily adapt to new service types and deployment environments. The access time for new services is reduced from 2 hours to 30 minutes, and the configuration complexity is reduced by 70%. Attached Figure Description

[0020] Figure 1 The overall architecture diagram of the intelligent service launch orchestration system is shown, illustrating the five-layer architecture and the components of each layer's modules; Figure 2 The intelligent startup orchestration flowchart illustrates the detailed process from dependency resolution to service startup completion. Figure 3 This is a flowchart illustrating the multidimensional pre-detection process, showing the parallel pre-detection and result analysis steps. Figure 4 A flowchart for selecting a fault recovery strategy is provided, illustrating the logic for fault classification and the selection of recovery strategies. Detailed Implementation

[0021] The specific embodiments of the present invention will be further described below with reference to the accompanying drawings and specific examples: Combination Figures 1-4 A container-based intelligent service launch orchestration system, comprising an orchestration control layer, a pre-detection layer, and a lifecycle management layer.

[0022] The orchestration control layer includes a startup manager module, a dependency resolution submodule, an orchestration strategy control submodule, and a startup order scheduling submodule.

[0023] The startup manager module is responsible for resolving service dependencies, formulating startup strategies, and controlling the startup process. It uses Apache Kafka as the event bus to implement asynchronous message passing and event distribution mechanisms. The state machine uses a Redis cluster as the state storage backend, leveraging Redis's atomic operation characteristics to ensure consistency in state transitions. The state machine model adopts a Finite State Automaton (FSA) design, defining eight core states: PENDING, INITIALIZING, STARTING, RUNNING, STOPPING, STOPPED, ERROR, and ROLLBACK. State transition rules are stored in a Redis hash data structure with the key-value format "state_machine:{service_id}:{timestamp}".

[0024] The dependency resolution submodule uses a Directed Acyclic Graph (DAG) algorithm to represent and manage complex dependencies between services. Each service node includes attributes such as service identifier, startup conditions, resource requirements, and timeout settings. Dependencies not only include simple startup order but also support conditional dependencies, such as "Service A depends on a database connection being available and the gateway service being in a healthy state." The dependency resolution submodule uses a graph database (e.g., Neo4j) as the storage engine for dependencies, leveraging its native graph traversal capabilities and the Cypher query language to model and query complex dependencies. The node model of the dependency graph includes service nodes and condition nodes, and the edge model defines relationship types such as DEPENDS_ON, CONFLICTS_WITH, REQUIRES_RESOURCE, and PROVIDES_SERVICE.

[0025] The orchestration strategy control submodule dynamically adjusts startup strategies based on the current system state. This module maintains a real-time system state table, recording the current state, resource usage, and historical startup time of each service. Based on this information, the module can dynamically adjust the concurrency, timeout, and retry strategies for service startup. The orchestration strategy control submodule is implemented based on the Drools rule engine. Rule definitions use DRL (Drools Rule Language) syntax, and rule files are stored in the distributed file system HDFS, supporting hot reloading and version management. The rule engine's working memory uses the Hazelcast distributed memory grid to ensure state synchronization between multiple nodes.

[0026] The startup order scheduling submodule implements an intelligent scheduling algorithm based on priority and dependency. This algorithm first identifies all services without prerequisite dependencies as startup candidates, and then sorts them according to factors such as service priority, estimated startup time, and resource requirements to select the optimal startup sequence. The fitness function is designed as a multi-objective optimization problem, including the optimization objectives of minimizing total startup time and maximizing resource utilization.

[0027] The pre-detection layer includes a multi-dimensional pre-detection engine, a port availability detection submodule, a database connection verification submodule, a configuration integrity verification submodule, and a resource readiness status check submodule.

[0028] The multi-dimensional pre-detection engine performs a series of pre-detection operations before the service actually starts, ensuring that the startup environment meets all prerequisites. The intelligent detection strategy is implemented based on a Bayesian network, with the network structure defining the impact of factors such as service type, historical success rate, and system load on the detection strategy. Network parameters are stored in an entity table and inferred and calculated using a custom strategy algorithm.

[0029] The port availability detection submodule is responsible for checking the availability status of ports required by the service. It tests port occupancy by creating temporary socket connections, supporting both TCP and UDP port detection. In case of port conflicts, the module can provide port alternatives or wait for the port to be released. It performs actual connection tests based on Java NIO's SocketChannel and DatagramChannel, with a connection timeout set to 3 seconds, and supports IPv4 and IPv6 dual-stack detection. The detection results include port status (AVAILABLE, OCCUPIED, FILTERED, UNREACHABLE), response time, and error messages.

[0030] The database connection validation submodule establishes test connections to various database services, verifying the correctness of connection parameters and the availability of the database services. This module supports multiple database types, including MySQL, MongoDB, and Redis. For connection failures, the module provides corresponding repair suggestions based on the cause of the failure. Based on a design combining adapter and factory patterns, it supports multiple database types such as MySQL, PostgreSQL, MongoDB, Redis, and Elasticsearch. Each database type corresponds to a specific validation adapter, implementing the IDatabaseValidator interface.

[0031] The configuration integrity verification submodule verifies the integrity of the configuration files and environment variables required for service startup. This module maintains a checklist of configuration items for each service, including required, optional, and value ranges. Verification is performed across multiple dimensions, including the existence, format correctness, and value validity of configuration items. The configuration verification engine is implemented based on the everit-org / json-schema library and supports all features of the JSON Schema Draft 7 specification, including type validation, format validation, constraint validation, and condition validation. Verification results include verification status, error details, and suggested remediation solutions.

[0032] The resource readiness status check submodule monitors the availability status of system resources. It includes three layers: system-level monitoring, container-level monitoring, and application-level monitoring.

[0033] The lifecycle management layer includes a container lifecycle monitoring module, a build phase management submodule, a runtime status tracking submodule, and a stop process control submodule.

[0034] Container-level monitoring is implemented through the cgroups v2 interface. Monitoring data includes CPU usage: ` / sys / fs / cgroup / cpu / cpu.stat`, memory usage: ` / sys / fs / cgroup / memory / memory.stat`, network I / O statistics: ` / sys / fs / cgroup / net_cls / net_cls.classid`, and disk I / O statistics: ` / sys / fs / cgroup / blkio / blkio.throttle.io_service_bytes`. Process-level monitoring is implemented through the ` / proc` filesystem. Key monitoring files include: ` / proc / {pid} / stat` (basic process statistics), ` / proc / {pid} / status` (detailed process status), ` / proc / {pid} / io` (process I / O statistics), and ` / proc / {pid} / fd / ` (process file descriptor information). Application-level monitoring uses the Micrometer metrics library and supports various monitoring system backends (Prometheus, InfluxDB, CloudWatch, etc.). Custom metrics include business processing volume, response time distribution, error rate, connection pool status, etc. The main technical components for building a log stream processing pipeline are as follows: Log collector: collects Docker build logs based on Filebeat; Message queue: Kafka Topic stores the raw log stream; Stream processor: Kafka Streams is used to parse the log content; State storage: RocksDB stores the build progress status; Result output: sends the parsed results to downstream systems.

[0035] The container lifecycle monitoring module enables the tracking and management of containers from creation to destruction, providing accurate status information for intelligent orchestration.

[0036] The build phase management submodule monitors the container image build process, including steps such as base image download, dependency installation, and application packaging. This module records key time points and resource consumption during the build process, providing data support for subsequent startup time estimation.

[0037] The startup phase monitoring submodule tracks each stage of the container startup process in a fine-grained manner, including container creation, network configuration, volume mounting, and process startup. This module identifies the current startup phase by parsing the event stream during container runtime and detects anomalies.

[0038] The runtime status tracking submodule continuously monitors the runtime status of containers, including process health, resource usage, and network connectivity. This module obtains container status information through various methods such as health check probes, resource monitoring metrics, and log analysis.

[0039] The stop process control submodule manages the graceful shutdown process of the container, ensuring data integrity and consistency. This module sends a stop signal to the application process within the container, waits for the application to complete its cleanup work, and then forcibly terminates the process after a timeout. The graceful shutdown protocol employs a multi-stage shutdown strategy: Pre-stop phase: sends a custom stop signal, allowing the application to execute cleanup logic; Graceful stop phase: sends a SIGTERM signal, waiting for the application to exit voluntarily; Forced stop phase: sends a SIGKILL signal, forcibly terminating the process; Resource cleanup phase: cleans up temporary files, network connections, shared memory, etc.

[0040] The multi-dimensional pre-detection engine described in this application performs multi-dimensional detections such as port availability, database connectivity, configuration integrity, and resource readiness status before service startup, which is an important function that traditional container orchestration systems do not have.

[0041] The complex dependency management based on DAG described in this application uses a directed acyclic graph algorithm to accurately represent and manage complex dependencies between services, and supports conditional dependencies and dynamic dependency adjustments.

[0042] The intelligent fault classification and adaptive recovery described in this application automatically selects the optimal recovery strategy based on the fault type, supports strategy upgrades and adaptive adjustments, and achieves truly intelligent fault handling.

[0043] This application describes a container lifecycle awareness system covering all stages, from construction, startup, operation to shutdown, providing precise status information and control capabilities.

[0044] The dynamic startup orchestration strategy described in this application is based on a dynamic scheduling algorithm using real-time system status and historical data, and can automatically adjust the startup strategy according to environmental changes.

[0045] The pre-detection algorithm used in this application protects the execution logic of multi-dimensional pre-detection, including parallel execution of detection items, result summary and analysis, failure cause classification and other technical methods.

[0046] The dependency representation method used in this application protects the design of dependency data structures based on DAG, including technical solutions such as node attribute definition, edge weight calculation, and loop detection algorithm.

[0047] The fault recovery strategy selection algorithm used in this application is a decision-making algorithm that automatically selects a recovery strategy based on factors such as fault type, historical success rate, and system load.

[0048] The lifecycle state machine design adopted in this application protects the definition and management methods of container lifecycle state transitions, including mechanisms such as state identification, transition conditions, and exception handling.

[0049] The dynamic scheduling optimization algorithm used in this application protects the startup order optimization algorithm based on multiple factors such as priority, resource requirements, and estimated time.

[0050] The configuration management mechanism adopted in this application protects the technical implementation of service configuration verification, repair, and version management, including configuration item checklists and automatic repair rules.

[0051] The pre-detection timing control adopted in this application requires that pre-detection be performed before the service is actually started. At the same time, it is necessary to balance the comprehensiveness of detection and the efficiency of execution to avoid startup delays caused by excessive detection.

[0052] This application adopts dependency consistency, and the dependency graph must maintain consistency and accuracy. Any change to the dependency requires re-verification of the graph's validity to prevent circular dependencies.

[0053] The fault recovery boundary adopted in this application requires that the automatic recovery mechanism must have clearly defined boundary conditions, including the maximum number of retries, recovery timeout, and resource usage limits, to avoid infinite retries.

[0054] The state synchronization mechanism adopted in this application requires that the state information between multiple modules be kept synchronized, especially during fault recovery and dynamic scheduling, to ensure the consistency of state information.

[0055] The performance monitoring adopted in this application requires the system to monitor the performance indicators of the orchestration process, including pre-detection time, startup success rate, resource usage, etc., to provide data support for optimization.

[0056] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. An intelligent service launch orchestration system based on container lifecycle, characterized in that, This includes the orchestration control layer, the pre-detection layer, and the lifecycle management layer; The orchestration control layer includes a startup manager module, a dependency resolution submodule, an orchestration strategy control submodule, and a startup order scheduling submodule; The startup manager module is responsible for resolving service dependencies, formulating startup strategies, and controlling the startup process. The dependency resolution submodule uses the Directed Acyclic Graph (DAG) algorithm to represent and manage complex dependencies between services; The orchestration strategy control submodule dynamically adjusts the startup strategy based on the current system state; The startup order scheduling submodule implements an intelligent scheduling algorithm based on priority and dependency. The fitness function is designed as a multi-objective optimization problem, including the optimization objectives of minimizing the total startup time and maximizing resource utilization. The pre-detection layer includes a multi-dimensional pre-detection engine, a port availability detection submodule, a database connection verification submodule, a configuration integrity verification submodule, and a resource readiness status check submodule; The multi-dimensional pre-detection engine performs a series of pre-detection operations before the service actually starts, ensuring that the startup environment meets all prerequisites. The port availability detection submodule is responsible for checking the availability status of the ports required by the service. It tests port occupancy by creating temporary socket connections and supports detection of TCP and UDP ports. The database connection verification submodule establishes test connections to various database services to verify the correctness of connection parameters and the availability of database services. The configuration integrity verification submodule verifies the integrity of the configuration files and environment variables required for service startup; The resource readiness status check submodule monitors the availability status of system resources; The lifecycle management layer includes a container lifecycle monitoring module, a build phase management submodule, a runtime status tracking submodule, and a stop process control submodule. The container lifecycle monitoring module enables the tracking and management of containers from creation to destruction, providing accurate status information for intelligent orchestration; The build phase management submodule monitors the container image build process; The startup phase monitoring submodule tracks each stage of the container startup process in a fine-grained manner; The runtime status tracking submodule continuously monitors the runtime status of the container; The Stop Process Control submodule manages the graceful shutdown process of containers, ensuring data integrity and consistency.

Citation Information

Cited By

  • Process life cycle management method and device based on state machine and electronic equipment

    CN121433848A

  • Abnormality processing method, service platform and computing device

    CN121705083A