Optimization method for improving service continuity of core system
Through system architecture optimization, data management and backup, performance monitoring, emergency plans and operation and maintenance management, the problem of rising failure rates caused by high complexity of modern core systems is solved, and efficient business continuity and resource utilization are achieved.
Patent Information
- Application Number
- CN202510430641.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-22
AI Technical Summary
Modern core systems adopt multi-level and distributed architectures and integrate a variety of advanced technologies, resulting in increased system complexity and increased probability of failures, and tiny failure points may lead to system paralysis.
Through system architecture optimization, data management and backup, performance monitoring, emergency plans and drills, and operation and maintenance management optimization, redundant configuration, distributed storage, load balancing, automated operation and maintenance tools and other technical means are used to improve the system's fault tolerance and response speed.
It significantly improves the fault tolerance of the system, reduces the failure rate, ensures system stability and business continuity, improves resource utilization and team collaboration capabilities, and reduces economic losses.
Smart Images

Figure CN120353645A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to an optimization method for improving the business continuity of a core system. Background Art
[0002] With the rapid development of information technology, the core business processes of enterprises increasingly rely on complex computer systems and network architectures; whether it is the transaction settlement system of financial institutions, the order processing system of e-commerce platforms, or the production management system of manufacturing enterprises, the stable operation of the core system is directly related to the daily operations and market competitiveness of enterprises.
[0003] Modern core systems usually adopt a multi-layered and distributed architecture, integrating various advanced technologies such as cloud computing, big data, and artificial intelligence; although these technologies bring powerful functions and performance improvements to the system, they also greatly increase the complexity of the system; the increase in the number of system components and the diversification of technology stacks significantly increase the probability of failures; the intricate interdependencies between different components mean that a tiny fault point can quickly spread through a chain reaction, leading to the paralysis of the entire system; for example, in an e-commerce core system based on cloud computing, it may involve the infrastructure of multiple cloud service providers, numerous microservice modules, and complex data interaction processes; any problem such as network latency, software vulnerability, or hardware failure in any link may trigger serious issues such as order processing failures and users being unable to log in, affecting the normal operation of the business. Therefore, we propose an optimization method for improving the business continuity of the core system to solve the problems raised above. Summary of the Invention
[0004] The purpose of the present invention is to provide an optimization method for improving the business continuity of a core system to solve the problems in the above background art, where modern core systems usually adopt a multi-layered and distributed architecture, integrating various advanced technologies. Although these technologies bring powerful functions and performance improvements to the system, they also greatly increase the complexity of the system, and the increase in the number of system components and the diversification of technology stacks significantly increase the probability of failures. At the same time, the intricate interdependencies between different components mean that a tiny fault point can quickly spread through a chain reaction, leading to the paralysis of the entire system.
[0005] To achieve the above object, the present invention provides the following technical solutions: An optimization method for improving the business continuity of the core system, including the following steps: S1. System architecture optimization, redundant configuration of key components, and splitting the core system into multiple microservices; S2. Data management and backup, adopting database master-slave replication technology or distributed data storage solution to achieve real-time synchronization of data between multiple nodes, and formulating a strict backup plan; S3. Deploying performance monitoring tools to monitor the usage of resources such as CPU, memory, disk I / O, and network bandwidth of the server in real time, and setting reasonable warning thresholds; S4. Emergency plan and drill, formulating a detailed emergency plan, clarifying the emergency handling process and responsibility division in different failure scenarios, regularly organizing emergency drills, simulating real failure scenarios, and testing the effectiveness and operability of the emergency plan; S5. Optimization of operation and maintenance management, adopting automated operation and maintenance tools to achieve functions such as automatic deployment, configuration management, and software upgrade of the server, regularly organizing operation and maintenance personnel to participate in technical training and learning exchange activities, and continuously improving their technical level and fault handling ability.
[0006] Preferably, in S1, multiple application servers are introduced to build a cluster, and a load balancer (such as F5 Big-IP) is used to evenly distribute user requests. Taking the order processing module as an example, when a single server processes orders, response delays or even service crashes may occur due to excessive load during peak business hours. Now, through load balancing, order requests are dispersed to 5 servers, and each server bears an average load of 20%, effectively improving the overall processing capacity and stability.
[0007] Preferably, in S1, the original centralized database is transformed into a distributed database architecture, and the TiDB distributed database is selected to reasonably shard and store commodity information, user data, etc. Different shards are distributed on multiple data nodes. For example, user data is divided into two different nodes according to the parity of the user ID. In this way, when querying specific user data, the corresponding node can be directly located, greatly reducing the query time, and at the same time, the failure of a single node will not affect the entire database service.
[0008] Preferably, in S2, based on the master-slave replication technology of the MySQL database, a one-master-two-slave architecture is built. The master database is responsible for handling all write operations, such as writing order data generated by user orders. The slave databases synchronize the data of the master database in real time, and data replication is achieved through binary logs (binlog). When the master database fails, one of the slave databases can be promoted to the master database within seconds through an automatic switching mechanism (such as MHA, Master High Availability) to continue providing services, ensuring data consistency and business continuity.
[0009] Preferably, in step S2, full - volume data backup is performed every day at dawn. The backup data is transmitted to a remote data center for storage through a dedicated network. Incremental backup is performed once a week, only backing up the data that has changed since the last full - volume backup. A data recovery test is arranged once a month to simulate the scenario of local data loss and recover the system data from the remote backup data to ensure the integrity and availability of the backup data. In one test, all the data of the core transaction system was successfully recovered within 3 hours, providing a reliable guarantee for actual disaster recovery.
[0010] Preferably, in step S3, a monitoring platform is built using Prometheus + Grafana. Prometheus is responsible for collecting various metric data of servers, databases, and application programs, such as server CPU usage rate, memory occupancy, database query time - consuming, application program interface response time, etc. Grafana displays the collected data in an intuitive chart form, and operation and maintenance personnel can view the system operation status in real - time through the monitoring panel. For example, it was found through the monitoring chart that the CPU usage rate of a certain application server often exceeded 80% during the business peak period, and resource optimization and load adjustment were carried out in a timely manner.
[0011] Preferably, in step S3, warning rules are set in the monitoring platform. For example, when the server CPU usage rate exceeds 85% for 10 consecutive minutes, the number of slow database queries exceeds 10 times per minute, and the application program interface error rate exceeds 5%, the system automatically sends warning messages to the operation and maintenance team via text messages and emails. Once, the number of slow database queries suddenly increased, and the warning system promptly notified the operation and maintenance personnel. After investigation, it was found that a complex query statement did not have a suitable index added, and potential performance problems were avoided after timely optimization.
[0012] Preferably, in step S4, detailed emergency plans are formulated for common fault scenarios such as server failures, network interruptions, and database crashes. For example, when a server hardware failure occurs, it is clearly stipulated that operation and maintenance personnel need to confirm the faulty server within 15 minutes, enable the standby server within 30 minutes, and switch the business to the standby server. For network interruptions, the plan stipulates how to quickly switch to the standby network link and the process of communicating and coordinating with the network provider to restore the main link.
[0013] Preferably, in step S4, an emergency drill is organized once a quarter to simulate different fault scenarios. In a drill simulating the failure of the main database, the operation and maintenance team completed the operation of switching the standby database to the main database within 5 minutes according to the emergency plan, and the order - processing business returned to normal within 10 minutes. Through the drill, the emergency response ability and collaborative cooperation ability of the team were improved, and at the same time, the effectiveness of the emergency plan was verified, and some detailed problems in the process were discovered and optimized.
[0014] Preferably, in S5, the Ansible automated operation and maintenance tool is used to write automated scripts to achieve batch deployment, software installation, and configuration management of servers. For example, when a new application server needs to be deployed, the server operating system installation, application deployment, and related configurations can be completed within 1 hour through Ansible scripts, saving a large amount of time compared to manual operations before, reducing human errors at the same time, and organizing internal technical training every month to invite external experts or internal technical backbones to share the latest technical knowledge and troubleshooting experience, and regularly carrying out operation and maintenance skills competitions to encourage operation and maintenance personnel to improve their technical levels. After a server performance optimization training, the operation and maintenance personnel optimized the servers of the core system with the knowledge they learned, improving the overall performance of the system by 20%.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: The optimization method for improving the business continuity of the core system can strengthen competitiveness, promote business expansion, and reduce economic losses at the business level, ensure system stability, promote technological innovation, and improve resource utilization rate at the technical level, and can improve the operation and maintenance process, enhance team collaboration, and improve risk management ability at the management level. The specific content is as follows: (1) The optimized system architecture adopts redundant design, distributed technology, etc., significantly improving the fault tolerance of the system; even if some components fail, the system can still automatically switch to standby components to maintain normal operation, avoiding system crashes caused by single-point failures; for example, in a cloud computing environment, through multi-node deployment and load balancing mechanisms, it is ensured that the application always maintains a stable response during high-concurrency access, significantly reducing the system failure rate.
[0016] (2) During the optimization process, the enterprise actively introduced new technologies such as cloud computing, big data, and artificial intelligence, not only improving the performance of the core system, but also promoting technological innovation and integration; for example, using big data analysis technology to deeply mine the system operation data can predict potential fault risks in advance and achieve preventive maintenance; applying artificial intelligence algorithms to the customer service system to realize automatic answering of intelligent customer service, improving service efficiency and quality, and also providing possibilities for the enterprise to explore more intelligent business application scenarios.
[0017] (3) With the help of automated operation and maintenance tools and intelligent resource allocation technologies, precise management and efficient utilization of hardware and software resources are achieved; during the business low season, resources are automatically reduced to lower energy consumption and costs; during the business peak season, resources are quickly allocated to meet demand, avoiding resource waste and over-configuration; taking server resources as an example, the resource utilization rate can be increased from 30% - 40% to 60% - 70% after optimization, reducing the technical operation cost of the enterprise.
[0018] (4) A comprehensive and accurate monitoring and early warning system has been established, which can capture the changes in the system operation status in real time and detect potential fault hazards in a timely manner; through intelligent analysis and dynamic adjustment of thresholds, the accuracy of monitoring and early warning has been greatly improved, reducing false alarms and missed alarms; at the same time, a perfect emergency response mechanism ensures that in the face of faults, the operation and maintenance team can take actions quickly and flexibly, efficiently handle problems according to the optimized process, and shorten the business interruption time; for example, when a network fault occurs, it can switch to the standby network link within a few minutes to ensure the normal operation of the business.
[0019] (5) The responsibilities of each technical team in the optimization and maintenance of the core system have been clarified, and an effective communication and cooperation mechanism has been established; the development team, operation and maintenance team, security team, etc. can cooperate closely and form a joint force in work such as system upgrades and fault handling; through regular cross-team meetings, information sharing platforms, and joint training, etc., the understanding and trust between teams have been enhanced, reducing problems caused by poor communication and unclear responsibilities, and improving the overall work efficiency and cooperation effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a schematic diagram of the optimization method process of the present invention; Figure 2 It is the architecture diagram of the core system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0022] Please refer to Figure 1 and Figure 2, the present invention provides a technical solution: an optimization method for improving the business continuity of the core system, including the following steps: S1. System architecture optimization, redundant configuration of key components, and splitting the core system into multiple microservices. In S1, multiple application servers are introduced to build a cluster, and a load balancer (such as F5 Big-IP) is used to evenly distribute user requests. Taking the order processing module as an example, when a single server originally processed orders, response delays or even service crashes might occur due to high load during peak business hours. Now, through load balancing, order requests are distributed to 5 servers, and each server bears an average load of 20%, effectively improving the overall processing capacity and stability. In S1, the original centralized database is transformed into a distributed database architecture, and TiDB distributed database is selected to reasonably shard and store commodity information, user data, etc. Different shards are distributed on multiple data nodes. For example, user data is divided into two different nodes according to the parity of the user ID. In this way, when querying specific user data, the corresponding node can be directly located, significantly reducing the query time, and at the same time, the failure of a single node will not affect the entire database service; S2. Data management and backup, using database master-slave replication technology or distributed data storage solutions to achieve real-time synchronization of data between multiple nodes, and formulating a strict backup plan. In S2, based on the master-slave replication technology of the MySQL database, a one-master-two-slave architecture is built. The master database is responsible for handling all write operations, such as writing order data generated by user orders. The slave databases synchronize the data of the master database in real time, and data replication is achieved through binary logs (binlog). When the master database fails, one of the slave databases can be promoted to the master database within seconds through an automatic switching mechanism (such as MHA, Master High Availability) to continue providing services, ensuring data consistency and business continuity. In S2, a full data backup is performed every day at midnight, and the backup data is transmitted to a remote data center for storage through a dedicated network. An incremental backup is performed once a week, only backing up the data that has changed since the last full backup. A data recovery test is arranged once a month to simulate the scenario of local data loss and restore the system data from the remote backup data to ensure the integrity and availability of the backup data. In one test, all the data of the core trading system was successfully restored within 3 hours, providing a reliable guarantee for actual disaster recovery;S3. Deploy a performance monitoring tool to monitor the usage of resources such as the CPU, memory, disk I / O, and network bandwidth of the server in real time, and set reasonable warning thresholds. In S3, use Prometheus + Grafana to build a monitoring platform. Prometheus is responsible for collecting various metric data of the server, database, and application program, such as the CPU usage rate of the server, memory occupancy, database query time consumption, application program interface response time, etc. Grafana displays the collected data in an intuitive chart form. The operation and maintenance personnel can view the running status of the system in real time through the monitoring panel. For example, it is found through the monitoring chart that the CPU usage rate of a certain application server often exceeds 80% during the business peak period, and resource optimization and load adjustment are carried out in a timely manner. In S3, set warning rules in the monitoring platform. For example, when the CPU usage rate of the server exceeds 85% for 10 consecutive minutes, the number of slow database queries exceeds 10 times per minute, and the error rate of the application program interface exceeds 5%, the system automatically sends warning messages to the operation and maintenance team via text message and email. Once, the number of slow database queries suddenly increased, and the warning system promptly notified the operation and maintenance personnel. After investigation, it was found that a complex query statement did not add a suitable index, and potential performance problems were avoided after timely optimization; S4. Emergency plan and drill. Develop a detailed emergency plan to clarify the emergency handling process and division of responsibilities in different failure scenarios, and regularly organize emergency drills to simulate real failure scenarios to test the effectiveness and operability of the emergency plan. In S4, for common failure scenarios such as server failure, network interruption, and database crash, develop a detailed emergency plan. For example, when a server hardware failure occurs, it is clearly stipulated that the operation and maintenance personnel need to confirm the failed server within 15 minutes, enable the standby server within 30 minutes, and switch the business to the standby server. For network interruption, the plan stipulates how to quickly switch to the standby network link and the process of communicating and coordinating with the network provider to restore the main link. In S4, organize an emergency drill every quarter to simulate different failure scenarios. In a drill simulating the failure of the main database, the operation and maintenance team completed the operation of switching the standby database to the main database within 5 minutes according to the emergency plan, and the order processing business returned to normal within 10 minutes. Through the drill, the emergency response ability and collaborative cooperation ability of the team were improved, and at the same time, the effectiveness of the emergency plan was tested, and some detailed problems in the process were discovered and optimized;S5. Optimize operation and maintenance management. Adopt automated operation and maintenance tools to achieve functions such as automatic deployment, configuration management, and software upgrade of servers. Regularly organize operation and maintenance personnel to participate in technical training and learning exchange activities to continuously improve their technical level and fault handling ability. In S5, use the Ansible automated operation and maintenance tool to write automated scripts to achieve batch deployment, software installation, and configuration management of servers. For example, when deploying a new application server, the Ansible script can complete the installation of the server operating system, application deployment, and related configurations within 1 hour, saving a large amount of time compared to manual operations before, reducing human errors at the same time, and organizing internal technical training every month. Invite external experts or internal technical backbones to share the latest technical knowledge and fault handling experience, and regularly carry out operation and maintenance skills competitions to encourage operation and maintenance personnel to improve their technical level. After a server performance optimization training, the operation and maintenance personnel optimized the servers of the core system using the knowledge learned, improving the overall performance of the system by 20%.
[0023] Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An optimization method for improving the business continuity of the core system, characterized in that, It includes the following steps: S1. System architecture optimization, redundantly configuring key components and splitting the core system into multiple microservices; S2. Data management and backup, adopting database master-slave replication technology or distributed data storage solutions to achieve real-time synchronization of data among multiple nodes and formulating a strict backup plan; S3. Deploying performance monitoring tools to monitor the usage of resources such as the CPU, memory, disk I / O, and network bandwidth of the server in real time and setting reasonable warning thresholds; S4. Emergency plan and drill, formulating a detailed emergency plan, clarifying the emergency handling procedures and division of responsibilities in different failure scenarios, regularly organizing emergency drills, simulating real failure scenarios, and testing the effectiveness and operability of the emergency plan; S5. Optimization of operation and maintenance management, adopting automated operation and maintenance tools to achieve functions such as automatic deployment, configuration management, and software upgrade of the server, regularly organizing operation and maintenance personnel to participate in technical training and learning and communication activities, and continuously improving their technical level and fault handling ability.
2. The optimization method for improving the business continuity of the core system according to claim 1, wherein: In S1, multiple application servers are introduced to build a cluster, and a load balancer (such as F5 Big-IP) is used to evenly distribute user requests. Taking the order processing module as an example, when a single server processes orders, during peak business hours, response delays or even service crashes may occur due to excessive load. Now, through load balancing, order requests are distributed to 5 servers, and each server bears an average load of 20%, effectively improving the overall processing capacity and stability.
3. An optimization method for improving the business continuity of the core system according to claim 1, characterized in that: In S1, the original centralized database is transformed into a distributed database architecture, and the TiDB distributed database is selected to reasonably shard and store commodity information, user data, etc. Different shards are distributed on multiple data nodes. For example, user data is divided into two different nodes according to the parity of the user ID. In this way, when querying specific user data, the corresponding node can be directly located, greatly reducing the query time, and at the same time, the failure of a single node will not affect the entire database service.
4. An optimization method for improving the business continuity of the core system according to claim 1, characterized in that: In S2, based on the master-slave replication technology of the MySQL database, a one-master-two-slave architecture is built. The master database is responsible for handling all write operations, such as writing order data generated by user orders. The slave databases synchronize the data of the master database in real time, and data replication is achieved through binary logs (binlog). When the master database fails, one of the slave databases can be promoted to the master database within seconds through an automatic switching mechanism (such as MHA, Master High Availability) to continue providing services, ensuring data consistency and business continuity.
5. An optimization method for improving the business continuity of the core system according to claim 1, characterized in that: In S2, a full data backup is performed every day at midnight, and the backup data is transmitted to an off-site data center for storage through a dedicated network. An incremental backup is performed once a week, only backing up the data that has changed since the last full backup. A data recovery test is arranged once a month to simulate the scenario of local data loss and recover system data from the off-site backup data to ensure the integrity and availability of the backup data. In a test, all the data of the core trading system was successfully recovered within 3 hours, providing a reliable guarantee for actual disaster recovery.
6. An optimization method for improving the business continuity of the core system according to claim 1, characterized in that: In S3, a monitoring platform is built using Prometheus + Grafana. Prometheus is responsible for collecting various metric data of servers, databases, and application programs, such as server CPU usage, memory occupancy, database query time consumption, application program interface response time, etc. Grafana displays the collected data in an intuitive chart form. Operation and maintenance personnel can view the system operation status in real time through the monitoring panel. For example, it is found through the monitoring chart that the CPU usage of a certain application server often exceeds 80% during the business peak period, and resource optimization and load adjustment are carried out in a timely manner.
7. An optimization method for improving the business continuity of the core system according to claim 1, characterized in that: In S3, warning rules are set in the monitoring platform. For example, when the server CPU usage exceeds 85% for 10 consecutive minutes, the number of slow database queries exceeds 10 times per minute, and the application program interface error rate exceeds 5%, the system automatically sends warning messages to the operation and maintenance team via text messages and emails. Once, the number of slow database queries suddenly increased, and the warning system promptly notified the operation and maintenance personnel. After investigation, it was found that a complex query statement did not add a suitable index, and potential performance problems were avoided after timely optimization.
8. An optimization method for improving the business continuity of the core system according to claim 1, characterized in that: In S4, detailed emergency plans are formulated for common fault scenarios such as server failures, network interruptions, and database crashes. For example, when a server hardware failure occurs, it is clearly stipulated that operation and maintenance personnel need to confirm the faulty server within 15 minutes, enable the standby server within 30 minutes, and switch the business to the standby server. For network interruptions, the plan stipulates how to quickly switch to the standby network link and the process of communicating and coordinating with the network provider to restore the main link.
9. An optimization method for improving the business continuity of the core system according to claim 1, characterized in that: In S4, an emergency drill is organized once every quarter to simulate different fault scenarios. In a drill simulating the failure of the main database, the operation and maintenance team completed the operation of switching the standby database to the main database within 5 minutes according to the emergency plan, and the order processing business returned to normal within 10 minutes. Through the drill, the emergency response ability and collaborative cooperation ability of the team were improved. At the same time, the effectiveness of the emergency plan was also tested, and some detailed problems in the process were discovered and optimized.
10. An optimization method for improving the business continuity of the core system according to claim 1, characterized in that: In S5, the Ansible automated operation and maintenance tool is used to write automated scripts to achieve batch deployment, software installation, and configuration management of servers. For example, when a new application server needs to be deployed, the server operating system installation, application program deployment, and related configurations can be completed within 1 hour through the Ansible script, saving a lot of time compared with the previous manual operation, reducing human errors at the same time, and organizing internal technical training every month, inviting external experts or internal technical backbones to share the latest technical knowledge and fault handling experience, and regularly carrying out operation and maintenance skills competitions to encourage operation and maintenance personnel to improve their technical levels. After a server performance optimization training, the operation and maintenance personnel optimized the servers of the core system using the knowledge learned, improving the overall performance of the system by 20%.