Digital drainage twin large model system microservice architecture, knowledge graph and high availability integration method and system
Patent Information
- Application Number
- CN202611195674.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-07
- Publication Date
- 2026-09-11
AI Technical Summary
首先,系统整体架构耦合度较高,各功能模块相互关联、绑定紧密,模块化拆分与独立迭代能力薄弱,导致系统功能拓展、版本升级与日常运维的难度大幅增加,迭代更新及后期运维的时间成本、人力成本与资源成本居高不下,系统灵活性与可拓展性严重不足
(1)鉴于数字排水孪生系统需满足低耦合与高扩展需求,故而采用中心调度层、业务服务层与终端适配层三级轻量化微服务架构,并结合轻量化剪枝与量化感知训练组合实施方法。通过去除非核心参数和通道特征映射压缩,实现服务独立部署和弹性扩展,从而将系统耦合度显著降低,提升整体维护效率与终端算力适配能力。
Smart Images

Figure CN122733480A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of smart water management digital twins, artificial intelligence large model integration, microservice architecture design, knowledge graph construction and high availability system technology, and in particular to a microservice architecture, knowledge graph and high availability integration method and system for a digital drainage twin large model system. Background Technology
[0002] Currently, digital twin technology has been gradually popularized and deeply applied in the field of intelligent management and control of urban drainage, becoming a core technical means to improve the intelligent level of urban drainage network operation and maintenance, flood control scheduling, and fault handling. Existing urban drainage digital twin systems mostly adopt traditional monolithic architecture designs. However, as urban drainage business scenarios become increasingly complex and the demand for intelligent management and control continues to upgrade, the technical shortcomings of these traditional monolithic drainage twin systems have gradually become apparent, making them unable to meet the current practical application needs of high precision, real-time monitoring, intelligence, and high reliability in urban drainage.
[0003] Specifically, existing monolithic drainage twin systems suffer from several key technical defects. First, the overall system architecture is highly coupled, with each functional module being closely interconnected and bound together. The ability to modularize and iterate independently is weak, which significantly increases the difficulty of system function expansion, version upgrades, and daily operation and maintenance. The time, manpower, and resource costs of iterative updates and subsequent operation and maintenance remain high, and the system's flexibility and scalability are severely lacking.
[0004] Secondly, the integration of existing drainage twin systems with large models is at a low level, only achieving shallow interface and simple calls to the large model API, without building a deeply integrated fusion architecture. This interface mode has obvious technical drawbacks, including high model inference response latency and low call execution efficiency, as well as a lack of standardized context storage, update, and management mechanisms. This makes it prone to problems such as context information confusion, loss, and errors, failing to meet the business requirements of real-time data analysis, dynamic judgment, and intelligent scheduling in urban drainage scenarios, and making it difficult to support high-intensity, high-timeliness real-time intelligent decision-making.
[0005] Furthermore, the urban drainage field encompasses a vast amount of professional industry knowledge, including pipeline topology, drainage equipment operating parameters, historical operation and maintenance experience, fault and defect identification rules, and operational scheduling standards. However, the existing technology system lacks standardized modeling, structured storage, and unified management mechanisms for the aforementioned multi-dimensional professional knowledge. It has not formed a systematic drainage industry knowledge graph and knowledge base, resulting in weak system knowledge retrieval, knowledge iteration, and intelligent reasoning capabilities. This makes it impossible to achieve accurate intelligent analysis and autonomous judgment based on professional knowledge, thus limiting the level of intelligent empowerment of the system.
[0006] Meanwhile, the existing high availability guarantee mechanism of the drainage digital twin system is imperfect, and a complete fault self-healing, data disaster recovery and rapid recovery system has not been established. Under extreme weather conditions such as rainstorms and urban flooding, as well as high-load operation scenarios, the system is prone to problems such as fault stagnation, data interruption and functional failure. The system lacks the ability to autonomously detect and repair faults and to quickly recover after disasters, resulting in insufficient system stability and reliability, and failing to meet the requirements for continuous and stable operation under extreme conditions.
[0007] In addition, the existing system's computing resource scheduling mode is rigid, generally adopting a static resource allocation method, which cannot dynamically allocate computing power according to the real-time load of drainage business and the intensity of scenario operation. During peak business periods, computing resources are prone to overload and system lag, while during off-peak business periods, computing resources are idle and wasted, resulting in extremely low computing resource utilization and poor resource allocation rationality and adaptability.
[0008] Finally, existing technologies lack unified and standardized cross-platform integration specifications and interface systems, resulting in insufficient system openness. The integration processes with various third-party monitoring systems, operation and maintenance management platforms, and dispatch and command platforms are cumbersome and difficult to adapt, with poor module and data reusability and weak cross-platform collaboration capabilities. Furthermore, existing intelligent models do not employ lightweight optimization techniques such as lightweight pruning, model quantization, and parameter compression, resulting in large model sizes and high computational consumption. This makes them unsuitable for deployment and operation on low-computing-power devices such as edge terminals and mobile terminals, hindering scenario adaptability and deployment flexibility, and limiting the large-scale, end-to-end application of drainage twin systems. Summary of the Invention
[0009] This invention aims to provide a microservice architecture, knowledge graph, and high-availability integration solution for a digital drainage twin model system. Through a combination of lightweight and quantitative models, it achieves the following objectives: 1. Build a lightweight microservice architecture to reduce system coupling and improve expansion and maintenance efficiency; 2. Achieve deep integration between the Volcano Ark large model and the twin system to optimize inference efficiency and context management; 3. Construct a knowledge graph for the drainage system to achieve unified knowledge storage, association mining, and intelligent retrieval; 4. Establish a highly available architecture and disaster recovery mechanism to ensure stable operation of the system under extreme scenarios; 5. Enable efficient collaboration and dynamic allocation of computing power among microservices, thereby improving resource utilization; 6. Establish standardized interfaces and integration methods to support rapid integration with multiple platforms; 7. Adapt to lightweight and quantization models to reduce terminal computing power consumption and expand application scenarios.
[0010] This invention provides a microservice architecture, knowledge graph, and high-availability integration method for a digital drainage twin large-scale model system, including: Establish a three-tier architecture consisting of a central scheduling layer, a business service layer, and a terminal adaptation layer; Build a context cache, incremental synchronization, and quantized call mechanism; Performs lightweight large-model knowledge extraction, quantitative storage of entity relationships, and semantic association retrieval; Deploy multi-node clusters, quantitative load balancing, off-site active-active backup, and automatic restart mechanism in case of failure; Configure a lightweight data bus, a quantized compression transmission channel, and a standardized model container; Perform real-time load monitoring, GPU / CPU elastic quota allocation, and closed-loop control of resource utilization.
[0011] In one embodiment of the present invention, the establishment of the three-tier architecture of the central scheduling layer, the business service layer, and the terminal adaptation layer includes: A lightweight service registry, configuration center, and gateway are deployed in the central scheduling layer, and the INT8 quantized load balancing algorithm is used to perform service routing and traffic scheduling. The business service layer independently deploys twin scenario services, large model integration services, knowledge graph services, computing power management services, and high availability services, and each service adopts the core logic after lightweight pruning; Configure access interfaces for PCs, mobile devices, large screens, and edge devices in the terminal adaptation layer, and adapt to different terminal computing power through dynamic LOD and quantization rendering.
[0012] In one embodiment of the present invention, the construction context cache, incremental synchronization, and quantized call mechanism include: Historical data of drainage scenarios, pipeline topology, and operation and maintenance rules are quantitatively stored through a lightweight vector database to build a persistent context cache; Synchronize new pipeline network data, real-time water level change data, and incremental business data to the Volcano Ark API; Input prompts are quantized using INT4 to compress the inference request body size.
[0013] In one embodiment of the present invention, the execution of lightweight large model knowledge extraction, entity relationship quantitative storage, and semantic association retrieval includes: Entities and relationships are extracted from drainage documents, drawings, and inspection records using pruned BERT or GPT models; Construct a knowledge graph that includes pipeline entities, equipment entities, defect entities, and operation entities; The entities and relationships in the knowledge graph are quantified and encoded, and then stored in a graph database. Semantic retrieval and association reasoning are performed based on the embedded vectors of queries and knowledge entities, as well as the graph path weights.
[0014] In one embodiment of the present invention, the deployment of multi-node clusters, quantitative load balancing, off-site active-active backup, and automatic fault restart mechanism include: Deploy core microservices on a cluster of 3 or more nodes, distribute traffic using a quantitative load balancing algorithm, and trigger traffic migration when a node's health falls below a threshold. Daily scheduled quantitative incremental backups are performed using a multi-site active-active and incremental backup strategy. The status of microservices is detected by heartbeat detection and service health monitoring. When a microservice fails, container restart and traffic splitting are triggered.
[0015] In one embodiment of the present invention, the configuration of the lightweight data bus, the quantization compression transmission channel, and the standardized model container includes: A data communication channel is built based on a message queue, and data is transmitted using quantized compression. The lightweight and quantized large models and inference models are encapsulated into standardized container images and managed uniformly through an image repository. Based on task priority and computing load, collaborative scheduling is implemented to distribute tasks among microservices.
[0016] In one embodiment of the present invention, the execution of real-time load monitoring, GPU / CPU elastic quota allocation, and resource utilization closed-loop control includes: Collect real-time CPU and GPU load data of each node in the cluster; Dynamically adjust the GPU / CPU elastic quota of microservices based on load data; The target utilization rate is maintained between 60% and 80% through closed-loop control of resource utilization.
[0017] In one embodiment of the present invention, it further includes: configuring RESTful and gRPC unified specifications, interface signatures and RBAC permission verification to achieve bidirectional data synchronization with third-party platforms.
[0018] This invention also provides a microservice architecture, knowledge graph, and highly available integrated system for a digital drainage twin large-scale model system, including: The lightweight microservice module is used to establish a three-tier architecture consisting of a central scheduling layer, a business service layer, and a terminal adaptation layer. The Volcano Ark API integration module is used to build context caching, incremental synchronization, and quantitative call mechanisms; The knowledge graph module is used to perform lightweight large model knowledge extraction, quantitative storage of entity relationships, and semantic association retrieval. The high-availability self-healing module is used to deploy multi-node clusters, quantitative load balancing, off-site active-active backup, and automatic restart mechanism in case of failure. The microservice collaboration module is used to configure the lightweight data bus, the quantized compression transmission channel, and the standardized model container; The computing power management module is used to perform real-time load monitoring, GPU / CPU elastic quota allocation, and closed-loop control of resource utilization.
[0019] In one embodiment of the present invention, it further includes: a standardized interface module, configured with RESTful and gRPC unified specifications, interface signature and RBAC permission verification components, for realizing bidirectional data synchronization with third-party platforms.
[0020] The present invention has the following beneficial effects: (1) Given that the digital drainage twin system needs to meet the requirements of low coupling and high scalability, a three-level lightweight microservice architecture consisting of a central scheduling layer, a business service layer, and a terminal adaptation layer is adopted, combined with a lightweight pruning and quantitative perception training approach. By removing non-core parameters and compressing channel feature mappings, services can be deployed independently and expanded elastically, thereby significantly reducing the system coupling and improving overall maintenance efficiency and terminal computing power adaptation capabilities.
[0021] (2) In order to solve the inference latency and context management bottleneck of large models, a context quantization cache, incremental data synchronization and Prompt quantization encoding mechanism were constructed. Thanks to synchronizing only the newly added pipeline data and real-time water level changes, the data transmission volume was reduced by ≥90% and the cache hit rate was ≥90%, thereby significantly reducing network bandwidth consumption and inference latency, and achieving a call efficiency improvement of ≥60%.
[0022] (3) To achieve the goal of unified association mining of pipeline network, equipment and operation and maintenance knowledge, an entity relationship extraction and graph database quantitative storage scheme based on lightweight large model is adopted. By quantizing and encoding entities and relationships and storing them in graph databases such as NebulaGraph, the storage cost is reduced by ≥75% and the retrieval accuracy is ≥95%. Moreover, semantic retrieval based on embedding vectors and graph path weights ensures accurate knowledge matching.
[0023] (4) Given that the system needs to ensure stable operation 24 / 7 under extreme scenarios, a multi-node cluster, quantitative load balancing and off-site active-active disaster recovery mechanism are built. When the node health is lower than the threshold, traffic migration and container restart are triggered, so that the recovery time target RTO is ≤30 minutes and the data recovery point target RPO is ≤15 minutes, thereby effectively avoiding business interruption caused by single point of failure.
[0024] (5) To avoid data silos between microservices and achieve efficient collaboration, a lightweight data bus based on Kafka and a standardized model container sharing mechanism are established. By quantizing and compressing data transmission and prioritizing cross-service task scheduling, transmission efficiency is improved by ≥3 times and bandwidth usage is reduced by ≥80%, thereby ensuring low latency and optimal task allocation for cross-service model calls.
[0025] (6) To improve the utilization rate of computing resources, a closed-loop control system for real-time load monitoring, GPU / CPU elastic quota and resource utilization is constructed. Due to the use of model fusion scheduling algorithm to weighted balance load balancing loss and inference performance loss, dynamic allocation of computing power in the system is realized, and the target utilization rate is stably maintained at 60%–80%, thereby achieving efficient flow of computing resources and cost control.
[0026] (7) To support rapid multi-platform integration and bidirectional data synchronization, standardized interfaces and integration methods were developed. By seamlessly integrating with more than three third-party platforms and supplementing with error compensation coefficients to correct for quantization accuracy loss, data interaction is ensured without loss, the compatibility range of lightweight application scenarios is expanded, and ultimately, the open interconnection of the system ecosystem is achieved. Attached Figure Description
[0027] Figure 1 The flowchart illustrating the microservice architecture, knowledge graph, and high-availability integration method of a digital drainage twin big model system according to an embodiment of the present invention is shown. Detailed Implementation
[0028] In the following description, the invention is described with reference to various embodiments. However, those skilled in the art will recognize that the embodiments may be practiced without one or more specific details or with other alternatives and / or additional methods, materials, or components. In other instances, well-known structures, materials, or operations are not shown or described in detail so as not to obscure the inventive points of the invention. Similarly, for illustrative purposes, specific quantities, materials, and configurations are set forth to provide a comprehensive understanding of embodiments of the invention. However, the invention is not limited to these specific details.
[0029] In this invention, the various embodiments are merely intended to illustrate the solutions of the invention and should not be construed as limiting.
[0030] In this specification, references to "an embodiment" or "this embodiment" mean that a particular feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment of the invention. The phrase "in one embodiment" appearing throughout this specification does not necessarily refer to the same embodiment in all instances.
[0031] Furthermore, the numbering of the steps in the methods of the present invention does not limit the execution order of the method steps. Unless otherwise specified, the method steps may be executed in different orders.
[0032] The present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0033] Figure 1 The flowchart illustrating the microservice architecture, knowledge graph, and high-availability integration method of a digital drainage twin big model system according to an embodiment of the present invention is shown.
[0034] like Figure 1 As shown in this embodiment, given the characteristics of the drainage network business scenario—large data volume, frequent concurrent requests, and high real-time requirements—a technical approach combining layered decoupling and quantitative compression is adopted. Specifically, this method first establishes a three-tier architecture: a central scheduling layer, a business service layer, and a terminal adaptation layer, achieving physical separation of computing resources and business logic through hierarchical isolation. Based on this, a context cache, incremental synchronization, and quantitative calling mechanism are constructed to reduce network bandwidth consumption while ensuring data consistency. Subsequently, lightweight large-model knowledge extraction, entity relationship quantitative storage, and semantic association retrieval are performed to transform unstructured drainage documents into a computable structured knowledge network. To improve system resilience, a multi-node cluster, quantitative load balancing, off-site multi-active backup, and automatic fault restart mechanism are deployed. Once a core service node experiences an anomaly, traffic switching and instance reconstruction are automatically executed. Furthermore, a lightweight data bus, quantitative compression transmission channel, and standardized model container are configured to enable data exchange between microservice components in a unified format. Finally, real-time load monitoring, GPU / CPU elastic quota allocation, and closed-loop resource utilization control are implemented to maintain the cluster's computing power within the optimal operating range through dynamic parameter tuning.
[0035] Specifically, the steps include the following: S1. Lightweight microservice architecture construction steps: establish a three-tier architecture consisting of a central scheduling layer, a business service layer, and a terminal adaptation layer.
[0036] The steps for building the lightweight microservice architecture include: Step 1: Deploy a lightweight service registry, configuration center, and gateway at the central scheduling layer, and use the INT8 quantized load balancing algorithm to perform service routing and traffic scheduling. Step 2: Independently deploy twin scenario services, large model integration services, knowledge graph services, computing power management services, and high availability services at the business service layer. Each service adopts the core logic after lightweight pruning. Step 3: Configure access interfaces for PCs, mobile devices, large screens, and edge devices in the terminal adaptation layer, and adapt to the computing power of different terminals through dynamic LOD and quantization rendering.
[0037] Central scheduling layer: Deploys a lightweight service registration and discovery center, configuration center, and gateway, employing an INT8 quantized load balancing algorithm. Core formula: ; in For service instances The routing weight value is used by the gateway / load balancer to allocate traffic. The larger the value, the more traffic is allocated to the service. This is the i-th backend service instance; For example Current CPU resource utilization, with a value range of [0,1], where 0 represents no load and 1 represents full CPU load; For example The current overall business load includes comprehensive indicators such as the number of concurrent connections, the length of the pending request queue, and memory usage, and is a normalized value. This is the weighting adjustment coefficient.
[0038] It is used to implement service routing and traffic scheduling, and reduce the computing power consumption of the central node.
[0039] Business service layer: It is divided into 5 core microservices: twin scenario service, large model integration service, knowledge graph service, computing power management service, and high availability service. Each service is deployed independently and adopts the core logic after lightweight pruning, with a size compression of ≥60%.
[0040] Terminal adaptation layer: Supports lightweight access for PCs, mobile devices, large screens, and edge devices. It adapts to different terminal computing power through dynamic LOD+ quantization rendering. Core formula: ; in This represents the model loading scheme matched based on the terminal's computing power T; T is the peak computing power of the terminal hardware; 16 TOPS is the computing power threshold used to distinguish between high-performance terminals and low-computing-power edge / mobile terminals; INT8 is 8-bit integer quantization; LOD0 is the highest level of detail (LevelofDetail0): the original high-precision model, complete triangle faces, high-definition textures, and full lighting effects, suitable for high-computing-power devices such as PCs and large screens; INT4 is 4-bit integer quantization: an extremely lightweight compression scheme that significantly reduces data volume with slight loss of precision, suitable for devices with limited computing power; LOD2 is the low simplified level of detail (LevelofDetail2): simplified model faces, reduced texture resolution, and disabled complex real-time rendering effects, significantly reducing rendering overhead.
[0041] S2, Volcano Ark API Deep Integration and Call Optimization Steps, Building Context Caching, Incremental Synchronization and Quantized Call Mechanism.
[0042] The steps for deep integration and call optimization of the Volcano Ark API include: Step 1: Quantize and store historical data of drainage scenarios, pipeline topology, and operation and maintenance rules through a lightweight vector database to build a persistent context cache; Step 2: Only update incremental business data such as new pipeline network data and real-time water level changes to the Volcano Ark API; Step 3: INT4 quantization encoding is performed on the input prompt words to compress the inference request body size.
[0043] Context persistent caching: Historical data, pipeline topology, and operation and maintenance rules for drainage scenarios are quantitatively stored using a lightweight vector database. The cache hit rate formula is as follows: ; Reduce redundant context transmissions to improve call efficiency. Cache hit rate is a performance metric for the vector cache set C in the drainage scenario. For vector query hit count; This represents the total number of vector query requests.
[0044] Volumetric call optimization: Only synchronize incremental business data (such as new pipeline data, real-time water level changes) to the Volcano Ark API. Incremental data volume formula: ; Data transmission volume is reduced by ≥90%, thus reducing network bandwidth consumption. For incremental data volume, This represents the latest total amount of full data for the current period. This represents the total amount of existing data from the previous statistical period.
[0045] Quantization-based inference acceleration: Input prompts are quantized using INT4 encoding, compressing the inference request size by ≥70%. Core acceleration formula: ; in, The acceleration factor is quantized and ranges from 0.6 to 0.8. For the inference latency optimized by quantization encoding, The baseline inference delay.
[0046] S3, the steps for constructing a knowledge graph for the drainage system include lightweight large-scale model knowledge extraction, quantitative storage of entity relationships, and semantic association retrieval.
[0047] The steps for constructing the knowledge graph of the drainage system include: Step 1: Extract entities and relationships from drainage documents, drawings, and inspection records using the pruned BERT or GPT model; Step 2: Construct a knowledge graph that includes pipeline entities, equipment entities, defect entities, and operation entities; Step 3: Quantify and encode the entities and relationships in the knowledge graph and store them using a graph database; Step four: Perform semantic retrieval and association reasoning based on the embedding vectors of the query and knowledge entities, as well as the graph path weights.
[0048] Knowledge Extraction and Construction: Entities and relationships are extracted from drainage documents, drawings, and inspection records using a lightweight large model (pruned BERT / GPT). Entity recognition formula: ; Construct a knowledge graph containing pipeline entities, equipment entities, defect entities, and operation entities. The overall identification confidence score for entity e. The similarity score is the embedding vector score. The model predicts the vector representation of the candidate entities. This serves as the baseline vector for the standard entity library within the knowledge graph. Score for matching rules. This is the rule weight adjustment coefficient.
[0049] Quantized storage: The entities and relationships of the knowledge graph are quantized and encoded, and stored using a graph database (such as NebulaGraph), reducing storage costs by ≥75%.
[0050] Intelligent retrieval and reasoning: Based on a large model, semantic retrieval and relational reasoning are achieved. For example, for "querying historical maintenance plans for pipeline defects in a certain area", the retrieval formula is: ; in, For the embedding vector of the query / knowledge entity, Assigning path weights to the graph to achieve accurate knowledge matching. To query the total similarity score between entity Q and entity K in the knowledge base, Calculate the query vector for cosine similarity. With entity vector Semantic similarity between them These are the weighting coefficients for the spectral paths.
[0051] S4, high availability disaster recovery and fault self-healing architecture configuration steps, deployment of multi-node clusters, quantitative load balancing, off-site active-active backup and automatic fault restart mechanism.
[0052] The configuration steps for the high-availability disaster recovery and fault self-healing architecture include: Step 1: Deploy the core microservices on a cluster with 3 or more nodes, and distribute traffic using a quantitative load balancing algorithm. When the health of a node falls below a threshold, traffic migration is triggered. Step 2: Use a multi-site active-active and incremental backup strategy to perform daily scheduled quantitative incremental backups; Step 3: Detect the status of microservices through heartbeat detection and service health monitoring. When a microservice fails, trigger container restart and traffic splitting.
[0053] The system adopts a high-availability architecture with cluster deployment, multi-active disaster recovery, and microservice circuit breaking, combined with a lightweight self-healing mechanism to ensure stable operation 24 / 7.
[0054] Cluster Deployment and Load Balancing: Core microservices are deployed in a cluster of 3 or more nodes. Traffic is distributed using a quantitative load balancing algorithm. The cluster health formula is as follows: ; When a node's health falls below a threshold, traffic migration is automatically triggered. The overall health score of cluster S is given by n, where n is the total number of nodes in the cluster. For the i-th service node instance in the cluster, For nodes Request error rate, with values in the range [0, 1]. This represents the maximum CPU resource limit for a node. This represents the current average CPU utilization of the node.
[0055] Disaster recovery: A multi-site active-active + incremental backup strategy is adopted, and quantitative incremental backups are performed daily on a regular schedule. The recovery time objective (RTO) is ≤30 minutes and the data recovery point objective (RPO) is ≤15 minutes.
[0056] Fault self-healing: Through heartbeat detection and service health monitoring, when a microservice fails, it automatically triggers container restart and traffic splitting. The self-healing formula is as follows: ; The self-healing execution strategy for microservice instance s selects different automated repair actions based on the number of failures; s can be any business microservice instance. This refers to the number of consecutive abnormal requests / service errors counted within the heartbeat detection period. Lightweight self-healing action: Performs a restart operation on the faulty microservice container, suitable for short-term, occasional failures; For scaling and traffic offloading self-healing actions: horizontally scale up the microservice by adding new instances, and at the same time split the traffic of the original faulty node to the new container, which is suitable for scenarios with continuous high-frequency errors.
[0057] S5, the steps for establishing a data communication and model sharing mechanism between microservices, including configuring a lightweight data bus, a quantized compression transmission channel, and a standardized model container.
[0058] The steps for establishing the inter-microservice data interoperability and model sharing mechanism include: Step 1: Build a data communication channel based on a message queue and use quantized compression to transmit data; Step 2: Package the lightweight and quantized large models and inference models into standardized container images, and manage them uniformly through an image repository; Step 3: Based on task priority and computing load, perform collaborative scheduling to realize the allocation of tasks among microservices.
[0059] Establish a unified data bus and model standard to achieve efficient collaboration between microservices and avoid data silos. Lightweight data bus: Data communication channels are built based on message queues (such as Kafka), and data is transmitted using quantized compression, which improves transmission efficiency by ≥3 times and reduces bandwidth usage by ≥80%.
[0060] Model sharing mechanism: Lightweight and quantized large models and inference models are encapsulated into standardized container images, which are managed uniformly through an image repository. The model call formula between services is as follows: ; Ensure low latency for cross-service model calls. Among these... This is the model shared call latency coefficient, used to measure the latency gain constraint of a containerized model m shared by multiple services and inference remotely by service s compared to a locally deployed model. This coefficient is required to be no greater than 0.5. The total time taken for service s to remotely call the shared container model m to complete one inference operation; The baseline time for a single inference run when model m is deployed locally on the service and does not share an image.
[0061] Cooperative scheduling strategy: A cooperative mechanism based on task priority and computing load, with the core formula as follows: ; Achieve optimal task allocation among microservices and improve collaboration efficiency.
[0062] in The overall scheduling priority score for task t. , , These are the weight adjustment coefficients corresponding to the three features. The urgency score for task t. The current comprehensive load index of the computing power node s to be allocated. This is the volume normalization score after model quantization.
[0063] S6, the dynamic allocation of computing power and quantitative management of resources, executes real-time load monitoring, elastic quota allocation of GPU / CPU, and closed-loop control of resource utilization.
[0064] The dynamic allocation of computing power and quantitative management of resources include the following steps: Step 1: Collect real-time CPU and GPU load data for each node in the cluster; Step 2: Dynamically adjust the GPU / CPU elastic quota for microservices based on load data; Step 3: Maintain the target utilization rate between 60% and 80% through closed-loop control of resource utilization.
[0065] Construct a management system that integrates computing power monitoring, dynamic scheduling, and quantified resource pools to achieve efficient utilization of computing resources.
[0066] Dynamic computing power allocation: Microservice computing power is dynamically adjusted based on business load (e.g., high concurrency during rainstorm warnings, low load during routine inspections). The allocation formula is as follows: ; Here, Δ is a dynamic adjustment coefficient; the higher the load, the more resources are allocated to ensure computing power for core tasks. The final GPU computing resource quota allocated to microservice s For microservices, a fixed computing power quota is allocated. This represents the real-time business load factor for the microservice s.
[0067] Resource Quantitative Management: This involves quantitative monitoring and quota management of resources such as CPU, memory, and GPU memory. Resource utilization formula: ; in Let R be the resource utilization rate. This represents the actual amount of resource R currently occupied by microservices, model inference, and simulation tasks. The upper limit of the total resource quota allocated to the business platform for nodes / clusters.
[0068] The target utilization rate should be maintained at 60%-80% to avoid resource overload or idleness.
[0069] Elastic scaling of computing power: Combining the resource consumption characteristics of lightweight large model inference, it realizes automatic scaling up and down with a scaling latency of ≤1 minute.
[0070] S7. Standardized interface integration steps across multiple platforms: Configure RESTful and gRPC unified specifications, interface signatures, and RBAC permission verification to achieve bidirectional data synchronization with third-party platforms.
[0071] Establish standardized interface specifications for RESTful+gRPC to support rapid integration with third-party systems such as campus platforms, operation and maintenance systems, and emergency command platforms.
[0072] Interface standardization: Unifying interface naming, parameter formats, and return status codes; automatically generating interface documentation; and ensuring interface compatibility. ; in For the overall compatibility rate of interface set I, This refers to the number of compatible interfaces in interface set I that meet standardized specifications and are interoperable between new and old versions. This represents the total number of interfaces within interface set I. It ensures backward compatibility and does not affect existing systems.
[0073] Multi-protocol support: Supports multiple protocols such as HTTP / HTTPS, WebSocket, and gRPC to adapt to the communication needs of different platforms.
[0074] Security Integration: Employs a security mechanism combining RBAC access control and interface signature to ensure secure cross-platform data interaction. Signature formula: ; To prevent unauthorized API calls. For the encrypted signature string of request req, This is a complete business interface request initiated by the client. It uses the MD5 hash digest algorithm. A unique application key assigned to the caller by the platform. The timestamp when the request was initiated. Request all business parameters for the interface.
[0075] This invention adopts a unified approach across the entire process, including microservice architecture, large model integration, and computing power management, combining lightweight pruning, quantization-aware training, and model fusion scheduling. The core formula is as follows: Lightweight pruning: ; in, For microservice / inference task loss, For sparse regularization coefficients, For the first Attention weights for each channel, This involves channel feature mapping. Non-core parameters are removed through pruning to compress the model size.
[0076] Quantization perception training (INT4 / INT8): ; in, For model input and output, To quantify weights, For quantitative scale, The number of quantization bits is 4 / 8. Quantization reduces model storage and computing power consumption, and supports edge deployment.
[0077] Model fusion scheduling: ; in, For load balancing losses, For inference performance loss, To facilitate synchronized service loss, Optimal resource allocation and service collaboration are achieved through weighted fusion.
[0078] Inference error compensation: ; in, For quantified data, For raw floating-point data, The error compensation coefficient (0.01-0.1) compensates for the accuracy loss caused by quantization and ensures the accuracy of inference.
[0079] Although various embodiments of the invention have been described above, it should be understood that they are presented by way of example only and not as limitations. It will be apparent to those skilled in the art that various combinations, modifications, and alterations can be made without departing from the spirit and scope of the invention. Therefore, the breadth and scope of the invention disclosed herein should not be limited by the exemplary embodiments disclosed above, but should be defined solely by the appended claims and their equivalents.
Claims
1. A microservice architecture, knowledge graph, and high-availability integration method for a digital drainage twin large-scale model system, characterized in that: include: Establish a three-tier architecture consisting of a central scheduling layer, a business service layer, and a terminal adaptation layer; Build a context cache, incremental synchronization, and quantized call mechanism; Performs lightweight large-model knowledge extraction, quantitative storage of entity relationships, and semantic association retrieval; Deploy multi-node clusters, quantitative load balancing, off-site active-active backup, and automatic restart mechanism in case of failure; Configure a lightweight data bus, a quantized compression transmission channel, and a standardized model container; Perform real-time load monitoring, GPU / CPU elastic quota allocation, and closed-loop control of resource utilization.
2. The method according to claim 1, characterized in that, The establishment of the three-tier architecture consisting of a central scheduling layer, a business service layer, and a terminal adaptation layer includes: A lightweight service registry, configuration center, and gateway are deployed in the central scheduling layer, and the INT8 quantized load balancing algorithm is used to perform service routing and traffic scheduling. The business service layer independently deploys twin scenario services, large model integration services, knowledge graph services, computing power management services, and high availability services, and each service adopts the core logic after lightweight pruning; Configure access interfaces for PCs, mobile devices, large screens, and edge devices in the terminal adaptation layer, and adapt to different terminal computing power through dynamic LOD and quantization rendering.
3. The method according to claim 1, characterized in that, The construction context cache, incremental synchronization, and quantized call mechanism include: Historical data of drainage scenarios, pipeline topology, and operation and maintenance rules are quantitatively stored through a lightweight vector database to build a persistent context cache; Synchronize new pipeline network data, real-time water level change data, and incremental business data to the Volcano Ark API; Input prompts are quantized using INT4 to compress the inference request body size.
4. The method according to claim 1, characterized in that, The execution of lightweight large-model knowledge extraction, entity relationship quantitative storage, and semantic association retrieval includes: Entities and relationships are extracted from drainage documents, drawings, and inspection records using pruned BERT or GPT models. Construct a knowledge graph that includes pipeline entities, equipment entities, defect entities, and operation entities; The entities and relationships in the knowledge graph are quantified and encoded, and then stored in a graph database. Semantic retrieval and association reasoning are performed based on the embedded vectors of queries and knowledge entities, as well as the graph path weights.
5. The method according to claim 1, characterized in that, The deployment of multi-node clusters, quantitative load balancing, off-site active-active backup, and automatic restart mechanisms include: Deploy core microservices on a cluster of 3 or more nodes, distribute traffic using a quantitative load balancing algorithm, and trigger traffic migration when a node's health falls below a threshold. Daily scheduled quantitative incremental backups are performed using a multi-site active-active and incremental backup strategy. The status of microservices is detected by heartbeat detection and service health monitoring. When a microservice fails, container restart and traffic splitting are triggered.
6. The method according to claim 1, characterized in that, The configuration of the lightweight data bus, quantization compression transmission channel, and standardized model container includes: A data communication channel is built based on a message queue, and data is transmitted using quantized compression. The lightweight and quantized large models and inference models are encapsulated into standardized container images and managed uniformly through an image repository. Based on task priority and computing load, collaborative scheduling is implemented to distribute tasks among microservices.
7. The method according to claim 1, characterized in that, The implementation of real-time load monitoring, GPU / CPU elastic quota allocation, and closed-loop control of resource utilization includes: Collect real-time CPU and GPU load data of each node in the cluster; Dynamically adjust the GPU / CPU elastic quota of microservices based on load data; The target utilization rate is maintained between 60% and 80% through closed-loop control of resource utilization.
8. The method according to claim 1, characterized in that, Also includes: Configure RESTful and gRPC unified specifications, interface signatures, and RBAC permission verification to achieve bidirectional data synchronization with third-party platforms.
9. A microservice architecture, knowledge graph, and high-availability integrated system for a digital drainage twin large-scale model system, characterized in that: include: The lightweight microservice module is used to establish a three-tier architecture consisting of a central scheduling layer, a business service layer, and a terminal adaptation layer. The Volcano Ark API integration module is used to build context caching, incremental synchronization, and quantitative call mechanisms; The knowledge graph module is used to perform lightweight large model knowledge extraction, quantitative storage of entity relationships, and semantic association retrieval. The high-availability self-healing module is used to deploy multi-node clusters, quantitative load balancing, off-site active-active backup, and automatic restart mechanism in case of failure. The microservice collaboration module is used to configure the lightweight data bus, the quantized compression transmission channel, and the standardized model container; The computing power management module is used to perform real-time load monitoring, GPU / CPU elastic quota allocation, and closed-loop control of resource utilization.
10. The system according to claim 9, characterized in that, Also includes: The standardized interface module is configured with RESTful and gRPC unified specifications, interface signatures, and RBAC permission verification components to achieve bidirectional data synchronization with third-party platforms.