A service fault operation and maintenance method, device, medium and program product

By identifying historical fault feature vectors in e-commerce platforms and utilizing Milvus and cosine similarity algorithms, business faults can be quickly located and repaired, solving the problem of long fault repair times and improving operational efficiency and profitability.

CN122195708APending Publication Date: 2026-06-12BEIJING YOUTEJIE INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING YOUTEJIE INFORMATION TECH
Filing Date
2026-03-05
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

The long repair time for business failures such as e-commerce platforms has severely impacted revenue, and the high proportion of recurring failures makes it difficult to resolve efficiently through expert experience.

Method used

By determining the feature vectors of historical business failures, fault location association data is generated and stored in Milvus. The cosine similarity algorithm is used to calculate the target fault feature vector with a similarity greater than a threshold, and the current fault repair strategy is determined.

Benefits of technology

Leveraging Milvus's millisecond-level similarity search capabilities and cosine similarity algorithm on its massive vector dataset, fault repair strategies can be quickly queried and associated, significantly reducing fault maintenance time and improving overall business benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195708A_ABST
    Figure CN122195708A_ABST
Patent Text Reader

Abstract

The application discloses a service fault operation and maintenance method, device, medium and program product, and relates to the field of data processing. The service fault operation and maintenance method comprises the following steps: determining a historical service fault feature vector; generating fault positioning associated data based on the historical service fault feature vector, and storing the fault positioning associated data to Milvus; calculating a target fault feature vector with a similarity greater than a similarity threshold value to the current service fault feature vector in Milvus according to a cosine similarity algorithm; and determining a current fault repair strategy based on the target fault feature vector. The technical scheme of the embodiment of the application can greatly reduce the service fault operation and maintenance time and improve the overall service revenue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a business fault operation and maintenance method, equipment, medium and program product. Background Technology

[0002] With the development of computer technology, the processing of massive amounts of data has become an urgent problem to be solved. For example, the business data of e-commerce platforms, financial platforms, and cloud platforms are quite large.

[0003] Taking e-commerce platforms as an example, with over 200 microservices and a daily order volume exceeding 5,000,000, the daily application performance management monitoring data volume can reach 1.2TB, with over 1,000 monitoring metrics. The average time to resolve business failures on e-commerce platforms is 85 minutes, resulting in hourly losses of up to 250,000 yuan and annual losses of nearly 86 million yuan, severely impacting overall revenue. Furthermore, recurring failures account for a high percentage. Due to the difficulty in inheriting expert experience, current methods primarily involve comparing original business time series to identify specific failures; however, the massive amounts of data pose significant challenges for manual analysis. Summary of the Invention

[0004] This invention provides a business failure operation and maintenance method, equipment, media, and program product to solve the problem of long operation and maintenance time for business failures, which seriously affects revenue.

[0005] According to one aspect of the present invention, a service fault operation and maintenance method is provided, comprising: Determine the feature vector of historical service failures; Based on historical business fault feature vectors, fault location association data is generated and stored in Milvus. Based on the cosine similarity algorithm, the target fault feature vector in Milvus with a similarity greater than the similarity threshold is calculated. Based on the target fault feature vector, the current fault repair strategy is determined.

[0006] According to another aspect of the present invention, a service failure maintenance device is provided, comprising: The first vector determination module is used to determine the historical service fault feature vector; The fault location data generation and storage module is used to generate fault location association data based on historical business fault feature vectors and store the fault location association data in Milvus. The second vector determination module is used to calculate, based on the cosine similarity algorithm, the target fault feature vector in Milvus that has a similarity greater than the similarity threshold with the current business fault feature vector. The fault repair strategy determination module is used to determine the current fault repair strategy based on the target fault feature vector.

[0007] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the business fault operation and maintenance method described in any embodiment of the present invention.

[0008] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions, the computer instructions being configured to cause a processor to execute and implement the service fault operation and maintenance method described in any embodiment of the present invention.

[0009] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the business fault operation and maintenance method described in any embodiment of the present invention.

[0010] The technical solution of this invention determines historical service fault feature vectors, generates fault location association data based on these vectors, and stores the fault location association data in Milvus. Then, using a cosine similarity algorithm, it calculates target fault feature vectors in Milvus that have a similarity greater than a similarity threshold with the current service fault feature vector. Based on these target fault feature vectors, it determines the current fault repair strategy. This solution leverages the high proportion of recurring service faults, combined with Milvus's millisecond-level similarity search capability across its massive vector dataset and the cosine similarity algorithm. With millisecond-level query speed, it retrieves target fault feature vectors from Milvus that have a similarity greater than a similarity threshold with the current service fault feature vector, and accurately associates them with the current fault repair strategy. This improves service fault operation and maintenance efficiency, solves the problem of long service fault operation and maintenance times that severely impact revenue, and significantly reduces service fault operation and maintenance time, thereby increasing overall service revenue.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A flowchart of a service fault operation and maintenance method provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart of a service fault operation and maintenance method provided in Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of a service fault operation and maintenance device provided in Embodiment 4 of the present invention; Figure 4 A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. Detailed Implementation

[0014] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0015] It should be noted that the terms "current," "target," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0016] Example 1 Figure 1 This is a flowchart of a service failure operation and maintenance method provided in Embodiment 1 of the present invention. This embodiment is applicable to the efficient operation and maintenance of service failures. The method can be executed by a service failure operation and maintenance device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1As shown, the method includes: Step 110: Determine the feature vector of historical service failures.

[0017] Among them, the historical service failure feature vector can be the feature vector of historical service failure.

[0018] In this embodiment of the invention, key indicators for business operation can be set first, that is, some key indicators can be selected from the routine monitoring indicators for business operation, and feature vectors corresponding to the key indicators for business operation can be extracted from historical business failure data as historical business failure feature vectors.

[0019] Among them, historical service failure data refers to relevant data on historical service failures.

[0020] Step 120: Generate fault location association data based on historical business fault feature vectors, and store the fault location association data in Milvus.

[0021] Among them, fault location correlation data can be used for root cause localization of business faults and operation and maintenance guidance. Fault location correlation data may include, but is not limited to, historical business fault feature vectors, root causes of historical business faults, fault resolution strategies, fault duration, and timestamps.

[0022] Specifically, historical business failure data is parsed to obtain data related to business failure location and maintenance. The parsed data is then summarized with the historical business failure feature vectors to obtain failure location-related data, which is then stored in Milvus.

[0023] Step 130: Calculate the target fault feature vector in Milvus that has a similarity greater than the similarity threshold with the current business fault feature vector, based on the cosine similarity algorithm.

[0024] The current service fault feature vector can be any feature vector of the current service fault. The similarity threshold can be a pre-set similarity evaluation threshold. The target fault feature vector can be any feature vector stored in Milvus that has a similarity greater than the similarity threshold to the current service fault feature vector.

[0025] In this embodiment of the invention, the cosine similarity algorithm can be used to calculate the cosine similarity between the historical service fault feature vector in Milvus and the current service fault feature vector, and the historical service fault feature vector in Milvus whose cosine similarity with the current service fault feature vector is greater than the similarity threshold is used as the target fault feature vector.

[0026] Step 140: Determine the current fault repair strategy based on the target fault feature vector.

[0027] The current fault repair strategy can be a strategy for repairing current business faults. For example, assuming the current business fault is that the database connection pool is exhausted, the current fault repair strategy can be to restart the service and adjust the maximum number of connections to 200 (which can be adjusted as needed).

[0028] In this embodiment of the invention, the fault resolution strategy corresponding to the target fault feature vector in the historical business fault data can be used as the current fault repair strategy.

[0029] The technical solution of this invention determines historical service fault feature vectors, generates fault location association data based on these vectors, and stores the fault location association data in Milvus. Then, using a cosine similarity algorithm, it calculates target fault feature vectors in Milvus that have a similarity greater than a similarity threshold with the current service fault feature vector. Based on these target fault feature vectors, it determines the current fault repair strategy. This solution leverages the high proportion of recurring service faults, combined with Milvus's millisecond-level similarity search capability across its massive vector dataset and the cosine similarity algorithm. With millisecond-level query speed, it retrieves target fault feature vectors from Milvus that have a similarity greater than a similarity threshold with the current service fault feature vector, and accurately associates them with the current fault repair strategy. This improves service fault operation and maintenance efficiency, solves the problem of long service fault operation and maintenance times that severely impact revenue, and significantly reduces service fault operation and maintenance time, thereby increasing overall service revenue.

[0030] Example 2 Figure 2 This is a flowchart of a service fault operation and maintenance method provided in Embodiment 2 of the present invention. This embodiment is based on the above embodiment and provides specific optional implementation methods for determining historical service fault feature vectors. Figure 2 As shown, the method includes: Step 210: Calculate the mean and standard deviation of key business operation indicators based on historical business failure sampling data.

[0031] Among them, historical service failure sampling data can be downsampled results from historical service failure data. The mean of key service operation indicators can be the average value of all key service operation indicator values. The standard deviation of key service operation indicators can be the standard deviation of all key service operation indicator values.

[0032] In this embodiment of the invention, historical service failure data can be downsampled to obtain historical service failure sampling data. Based on the historical service failure sampling data, the indicator values ​​of each service operation key indicator can be determined. Based on the indicator values ​​of each service operation key indicator, the mean value and standard deviation of the corresponding service operation key indicator can be calculated.

[0033] In an optional embodiment of the present invention, before calculating the mean and standard deviation of key business operation indicators based on historical business failure sampling data, the method may further include: determining the key business operation indicators of historical business failure data and the sampling period of the key business operation indicators; and downsampling the historical business failure data according to the sampling period of the key business operation indicators to obtain historical business failure sampling data.

[0034] In this embodiment of the invention, key indicators for business operation and the sampling period for key indicators for business operation can be determined for different businesses in historical business failure data. Then, according to the sampling period of key indicators for business operation, the relevant data of key indicators for each business operation in historical business failure data are downsampled to obtain historical business failure sampling data.

[0035] In an optional embodiment of the present invention, determining the key indicators of business operation of historical business failure data and the sampling period of the key indicators of business operation may include: determining the key indicators of business operation of historical business failure data according to the historical business type, and determining the sampling strategy of the key indicators of business operation; generating the sampling period of the key indicators of business operation according to the sampling strategy of the key indicators of business operation; wherein, the sampling strategy may include aggregate sampling, balanced sampling or hierarchical sampling.

[0036] Among these, historical business types can be the business types of historical businesses that generated historical business failure data. Aggregate sampling can be used to take the average of multiple samples of key business operation indicators as a single sampling result. Balanced sampling can be used to collect data on different key business operation indicators according to the same sampling period. Hierarchical sampling can be used to collect data on different key business operation indicators according to at least one sampling period.

[0037] In this embodiment of the invention, a pre-defined mapping relationship between historical business types and key indicators of business operation can be obtained. Based on the business types of historical business failure data and the mapping relationship between historical business types and key indicators of business operation, the key indicators of business operation for each historical business in the historical business failure data can be determined. Then, the sampling strategy of key indicators of business operation set by technical personnel can be obtained. Based on the sampling strategy of key indicators of business operation, sampling periods can be dynamically allocated for different key indicators of business operation.

[0038] Step 220: Determine the historical business failure feature vector based on the indicator values, the mean of the key business operation indicators, and the standard deviation of the key business operation indicators.

[0039] In this embodiment of the invention, the difference between the index value of each key business operation indicator and the mean of the corresponding key business operation indicator can be calculated. The difference between the index value of each key business operation indicator and the mean of the corresponding key business operation indicator can be further calculated, and the quotient of the difference between the index value of each key business operation indicator and the mean of the corresponding key business operation indicator can be calculated. The calculated quotient constitutes the historical business fault feature.

[0040] Step 230: Based on historical business fault feature vectors, generate fault location association data and store the fault location association data in Milvus.

[0041] Step 240: Calculate the target fault feature vector in Milvus that has a similarity greater than the similarity threshold with the current business fault feature vector, based on the cosine similarity algorithm.

[0042] In an optional embodiment of the present invention, calculating a target fault feature vector in Milvus with a similarity greater than a similarity threshold based on a cosine similarity algorithm may include: obtaining fault feature weights; calculating the similarity between historical service fault feature vectors and current service fault feature vectors in Milvus based on the cosine similarity algorithm and the fault feature weights; and searching for a target fault feature vector in Milvus based on the similarity between historical service fault feature vectors and current service fault feature vectors in Milvus and the similarity threshold.

[0043] The fault feature weight can be a weighting coefficient set for key business operation indicators. Each key business operation indicator corresponds to a fault feature weight. The key business operation indicator corresponding to the current business fault feature vector can be some or all of the key business operation indicators corresponding to the historical business fault feature vectors.

[0044] In this embodiment of the invention, fault feature weights corresponding to key indicators of business operation can be obtained. Then, based on the cosine similarity algorithm and the fault feature weights corresponding to each key indicator of business operation, the similarity between historical business fault feature vectors and current business fault feature vectors in Milvus can be calculated. The similarity between historical business fault feature vectors and current business fault feature vectors in Milvus can be compared with a similarity threshold. The historical business fault feature vectors in Milvus with a similarity greater than the similarity threshold are used as target fault feature vectors.

[0045] Optionally, based on the cosine similarity algorithm and fault feature weights, the similarity between historical service fault feature vectors and current service fault feature vectors in Milvus is calculated, including calculation based on the following formula: .in, This represents the fault feature weight corresponding to the i-th key business operation indicator. This represents the i-th element in the historical service failure feature vector. This represents the i-th element in the current business fault feature vector.

[0046] Step 250: Determine the current fault repair strategy based on the target fault feature vector.

[0047] In an optional embodiment of the present invention, determining the current fault repair strategy based on the target fault feature vector may include: querying fault location association data corresponding to the target fault feature vector in Milvus; and determining the current fault repair strategy based on the fault location association data corresponding to the target fault feature vector and the similarity between the target fault feature vector and the current business fault feature vector.

[0048] In this embodiment of the invention, the target fault feature vector can be used as a keyword to query fault location association data with the target fault feature vector in Milvus, and the similarity between the target fault feature vector and the current business fault feature vector can be sorted. Then, the fault location association data corresponding to the top-ranked (such as the first or the top K) target fault feature vectors can be parsed to obtain the current fault repair strategy.

[0049] In an optional embodiment of the present invention, after determining the current fault repair strategy based on the target fault feature vector, the method may further include: collecting the fault repair results of the current fault repair strategy; and updating the historical service fault feature vector and fault feature weights according to the fault repair results of the current fault repair strategy.

[0050] The fault repair results can be used to characterize the repair outcome of the current fault repair strategy. These results may include, but are not limited to, the correctness of the fault diagnosis, the effectiveness of the repair strategy, the actual fault maintenance time, and expert repair experience.

[0051] In this embodiment of the invention, after executing the current fault repair strategy, the fault repair results of the current fault repair strategy can be collected. Then, when the current fault repair strategy successfully repairs the fault, the current business fault feature vector of the current fault repair strategy is added to the historical business fault feature vector, and higher fault feature weights are assigned to the key business operation indicators with greater impact in the current fault repair strategy.

[0052] Optionally, the key business operation indicators with a greater impact in the current fault repair strategy, as well as the fault characteristic weights of the key business operation indicators with a greater impact, can be set by technical personnel.

[0053] The technical solution of this invention calculates the mean and standard deviation of key business operation indicators based on historical business failure sampling data. Then, based on these key indicators, their mean, and standard deviation, historical business failure feature vectors are determined. Subsequently, fault location association data is generated based on these historical business failure feature vectors and stored in Milvus. Furthermore, using a cosine similarity algorithm, target fault feature vectors in Milvus with a similarity greater than a similarity threshold to the current business failure feature vector are calculated. Based on these target fault feature vectors, the current fault repair strategy is determined. This solution leverages the high proportion of recurring business failures, combined with Milvus's millisecond-level similarity search capability across its massive vector dataset and the cosine similarity algorithm. With millisecond-level query speeds, it retrieves target fault feature vectors from Milvus with a similarity greater than a similarity threshold to the current business failure feature vector and accurately associates them with the current fault repair strategy. This improves business failure maintenance efficiency, solves the problem of long maintenance times that severely impact revenue, and significantly reduces overall business failure maintenance time, thereby increasing overall business revenue.

[0054] Example 3 Embodiment 3 of the present invention provides an optional embodiment of a service fault operation and maintenance method, the specific implementation of which can be found in the following embodiments. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.

[0055] Based on the business failure operation and maintenance method, historical business failure data can be collected and historical business failure features can be extracted to obtain historical business failure feature vectors. Real-time monitoring of business operation status and anomaly detection are also performed. When a business failure is detected, the current business failure feature vector can be extracted, and a search can be conducted in Milvus for historical business failure feature vectors similar to the current business failure feature vector, i.e., the target failure feature vector. Then, the fault repair strategy corresponding to the target failure feature vector, which has been manually confirmed, is executed, and the fault repair result is obtained. The fault repair result can be used to update the fault feature vector and the fault feature weights.

[0056] First, historical business failure data (such as data on 243 business failures that occurred in the past 12 months) is collected and stored in a failure knowledge base. For each business failure in the historical failure data, key operational metrics (CPU utilization, memory utilization, database connection count, latency response time, and error rate) are collected. Furthermore, the fault location-related data for each business failure can be stored in a table. This table can include the historical business failure ID (identifier), failure occurrence time, root cause of the historical business failure, fault resolution strategy, and failure duration.

[0057] For example, the failure corresponding to historical business failure ID F001 occurred at 02:30 on March 15, 2023, and lasted for 45 minutes. The root cause of this failure was database connection pool exhaustion, and the solution was to restart the service and adjust the maximum number of connections to 200. The failure corresponding to historical business failure ID F002 occurred at 14:15 on April 22, 2023, and lasted for 120 minutes. The root cause was cache avalanche, and the solution was to preheat the cache and add a circuit breaker. The failure corresponding to historical business failure ID F003 occurred at 09:45 on May 10, 2023, and lasted for 75 minutes. The root cause was memory leak, and the solution was to restart the service and fix the code.

[0058] Assume that the key performance indicators (KPIs) corresponding to historical business failure ID F001 are CPU utilization, memory utilization, database connection count, latency, and error rate. For example, CPU utilization is 58.2%, memory utilization is 96.0%, database connection count is 382, ​​latency is 241ms, and error rate is 0.183%. The historical business failure feature vector for failure ID F001 is [0.867, 2.236, 0.040, 2.379, 3.062]. The first element in the historical business failure feature vector is calculated based on (58.2-45.2) / 15.0, where 45.2 represents the mean CPU utilization and 15 represents the standard deviation of CPU utilization. The second element in the historical business failure feature vector is calculated based on (96.0-68.5) / 12.3, where 68.5 represents the mean memory utilization and 12.3 represents the standard deviation of memory utilization. The third element in the historical service failure feature vector is calculated based on (382-380.2) / 45.2, with a mean database connection rate of 380.2 and a standard deviation of 45.2. The fourth element is calculated based on (241-156.8) / 35.4, with a mean latency response rate of 156.8 and a standard deviation of 35.4. The fifth element is calculated based on (0.183-0.085) / 0.032, with a mean error rate of 0.085 and a standard deviation of 0.032.

[0059] In Milvus, fault location-related data can be stored as a single record: {"id":"F001", "vector":[0.867,2.236,0.040,2.379,3.062],"metadata":{"root_cause":"Database connection pool exhausted","solution":"Restart service, adjust maximum connection count to 200","duration":45,"timestamp":1678854600}}. Here, duration is the fault duration, and timestamp is the timestamp.

[0060] When a new business failure occurs, the current business failure feature vector is quickly extracted. For example, if the monitoring system detects an anomaly (such as CPU utilization >80% for 5 minutes), the key business operation indicators of the most recent 30 minutes are collected. That is, the 5 key business operation indicators that are the same as those of historical business failures are extracted. According to the calculation logic of the feature vector of historical business failures, the current business failure feature vector is calculated as [1.087, 2.089, -0.049, 2.011, 2.562].

[0061] Optionally, an actionable solution can be provided based on the matching results between the current business fault feature vector and historical business fault feature vectors. This involves searching for the target fault feature vector in Milvus, sorting it by similarity, and extracting the solution for the top-ranked business fault. The matching results are recorded according to the matching ranking, historical business fault ID, similarity, root cause of the historical business fault, and fault resolution strategy. The matching results may look like the following: {1; F001; 0.995; Database connection pool exhausted; Restart service and adjust maximum number of connections to 200.}

[0062] 2; F045; 0.872; Cache avalanche; Warm up the cache, add circuit breaker.

[0063] 3; F128; 0.821; Memory leak; Restart the service and fix the code. Optionally, the similarity calculated based on cosine similarity can be normalized, resulting in a similarity range of 0-1, with 0 to 1 indicating increasing similarity. For example, 0.7-0.9 indicates relatively similar, greater than 0.9 indicates highly similar, and less than 0.7 indicates not very similar.

[0064] The fault diagnosis report for the current service fault corresponding to the current service fault feature vector can include: Fault time: 2024-01-15 14:30; An anomaly detected: CPU usage increased from 30% to 96% within 30 minutes; Matching result: The most similar historical fault ID is F001, and the time is 2023-03-15 02:30; Similarity: 99.5%; The root cause of the business failure: database connection pool exhaustion; Troubleshooting strategy: Restart the service and adjust the maximum number of connections to 200; Historical service failure resolution time: 12 minutes.

[0065] The troubleshooting strategy is to immediately check the number of database connections. If the number of connections is greater than 250, restart the order service and modify the connection pool configuration to set the maximum number of connections to 200 and the maximum duration to 5000ms.

[0066] The execution log of the troubleshooting strategy is shown below: 14:30: The system detected an anomaly and generated a fault diagnosis report; 14:32: The engineer received the alarm and reviewed the report; 14:33: Performed a database connection count check, found 285 connections, with the longest duration being 300ms; 14:35: Order service restarted; 14:38: Service restart complete; 14:40: CPU utilization dropped from 95% to 45%; 14:45: All indicators have returned to normal; Actual resolution time: 15 minutes (70 minutes faster than the traditional method).

[0067] After troubleshooting based on the current troubleshooting strategy, collect the troubleshooting results: Was the diagnosis correct? Was the solution effective? How long did the actual troubleshooting take? If the diagnosis is correct, strengthen the importance of that failure mode; if the diagnosis is incorrect, adjust the failure feature weights (e.g., the initial failure feature weights for key business operation indicators are [0.2, 0.2, 0.2, 0.2, 0.2], and after learning, it is found that CPU utilization and error rate are more important features, so the optimized weights are [0.3, 0.1, 0.2, 0.2, 0.2] or add new modes, update the mean and standard deviation of key business operation indicators.

[0068] Based on the mean and standard deviation of key business operation indicators, the historical business fault feature vector is calculated to solve the problem of differences in the dimensions, units and numerical ranges of different monitoring indicators. The amount of data is reduced by 99% and the computational complexity is reduced by 90%. Furthermore, fault feature weights are introduced on the basis of cosine similarity to dynamically adjust the importance of each feature according to different fault types.

[0069] Example 4 Figure 3 This is a schematic diagram of a service fault operation and maintenance device provided in Embodiment 4 of the present invention. Figure 3 As shown, the device includes: The first vector determination module 310 is used to determine the historical service fault feature vector; The fault location data generation and storage module 320 is used to generate fault location association data based on historical business fault feature vectors and store the fault location association data in Milvus. The second vector determination module 330 is used to calculate the target fault feature vector in Milvus with a similarity greater than the similarity threshold based on the cosine similarity algorithm. The fault repair strategy determination module 340 is used to determine the current fault repair strategy based on the target fault feature vector.

[0070] The technical solution of this invention determines historical service fault feature vectors, generates fault location association data based on these vectors, and stores the fault location association data in Milvus. Then, using a cosine similarity algorithm, it calculates target fault feature vectors in Milvus that have a similarity greater than a similarity threshold with the current service fault feature vector. Based on these target fault feature vectors, it determines the current fault repair strategy. This solution leverages the high proportion of recurring service faults, combined with Milvus's millisecond-level similarity search capability across its massive vector dataset and the cosine similarity algorithm. With millisecond-level query speed, it retrieves target fault feature vectors from Milvus that have a similarity greater than a similarity threshold with the current service fault feature vector, and accurately associates them with the current fault repair strategy. This improves service fault operation and maintenance efficiency, solves the problem of long service fault operation and maintenance times that severely impact revenue, and significantly reduces service fault operation and maintenance time, thereby increasing overall service revenue.

[0071] Optionally, the first vector determination module 310 is used to calculate the mean and standard deviation of key business operation indicators based on historical business failure sampling data; and to determine the historical business failure feature vector based on the indicator values, the mean and standard deviation of the key business operation indicators.

[0072] Optionally, the service failure operation and maintenance device further includes a data downsampling module, used to determine the key service operation indicators of historical service failure data, and the sampling period of the key service operation indicators; and to perform data downsampling on the historical service failure data according to the sampling period of the key service operation indicators to obtain the historical service failure sampled data.

[0073] Optionally, the data downsampling module is further configured to determine key indicators of business operation based on historical business types, and determine the sampling strategy for the key indicators of business operation; and generate the sampling period for the key indicators of business operation based on the sampling strategy for the key indicators of business operation; wherein the sampling strategy includes aggregate sampling, balanced sampling, or hierarchical sampling.

[0074] Optionally, the second vector determination module 330 is used to obtain fault feature weights; calculate the similarity between historical service fault feature vectors and current service fault feature vectors in Milvus based on the cosine similarity algorithm and the fault feature weights; and search for target fault feature vectors in Milvus based on the similarity between historical service fault feature vectors and current service fault feature vectors in Milvus and the similarity threshold.

[0075] Optionally, the fault repair strategy determination module 340 is used to query the fault location association data corresponding to the target fault feature vector in Milvus; and determine the current fault repair strategy based on the fault location association data corresponding to the target fault feature vector and the similarity between the target fault feature vector and the current business fault feature vector.

[0076] Optionally, the service fault operation and maintenance device further includes a data update module, used to collect the fault repair results of the current fault repair strategy; and update the historical service fault feature vector and the fault feature weights according to the fault repair results of the current fault repair strategy.

[0077] The service fault operation and maintenance device provided in the embodiments of the present invention can execute the service fault operation and maintenance method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.

[0078] Example 5 Figure 4 A schematic diagram of an electronic device that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0079] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as ROM 12, RAM 13, etc., communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from the storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An I / O interface 15 is also connected to the bus 14. The ROM 12 is a read-only memory, the RAM 13 is a random access memory, and the I / O interface 15 is an input / output interface.

[0080] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0081] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as business fault operation and maintenance methods.

[0082] In some embodiments, the service failure operation and maintenance method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the service failure operation and maintenance method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the service failure operation and maintenance method by any other suitable means (e.g., by means of firmware).

[0083] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0084] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0085] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0086] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0087] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0088] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS servers, such as high management difficulty and weak business scalability.

[0089] This application also discloses a computer program product, which includes a computer program that, when executed by a processor, implements the business fault operation and maintenance method provided in any embodiment of this application. This program product shares the same inventive concept as the business fault operation and maintenance methods disclosed in the embodiments of this application, and therefore will not be described in detail here.

[0090] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0091] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for operational maintenance during business failures, characterized in that, include: Determine the feature vector of historical service failures; Based on the historical business fault feature vectors, fault location association data is generated and stored in the cloud-native vector database Milvus. Based on the cosine similarity algorithm, the target fault feature vector in Milvus with a similarity greater than the similarity threshold with the current business fault feature vector is calculated. Based on the target fault feature vector, the current fault repair strategy is determined.

2. The method according to claim 1, characterized in that, Determine the feature vector of historical service failures, including: Based on historical business failure sampling data, calculate the mean and standard deviation of key business operation indicators; The historical business failure feature vector is determined based on the indicator values, the mean values, and the standard deviations of the key business operation indicators.

3. The method according to claim 2, characterized in that, Before calculating the mean and standard deviation of key business operation indicators based on historical business failure sampling data, the following steps are also included: Determine the key operational indicators of historical service failure data, and the sampling period for the key operational indicators; Based on the sampling period of the key indicators of business operation, the historical business failure data is downsampled to obtain the historical business failure sampling data.

4. The method according to claim 3, characterized in that, Determine the key operational indicators of historical service failure data, and the sampling period for these key operational indicators, including: Based on the historical service types, determine the key service operation indicators of historical service failure data, and determine the sampling strategy for the key service operation indicators; Based on the sampling strategy of the key business operation indicators, the sampling period of the key business operation indicators is generated; The sampling strategies include aggregate sampling, balanced sampling, or hierarchical sampling.

5. The method according to claim 1, characterized in that, Based on the cosine similarity algorithm, target fault feature vectors in Milvus with a similarity greater than a similarity threshold are calculated, including: Obtain fault feature weights; Based on the cosine similarity algorithm and the fault feature weights, the similarity between the historical service fault feature vector and the current service fault feature vector in Milvus is calculated. Based on the similarity between historical service fault feature vectors and current service fault feature vectors in Milvus and the similarity threshold, the target fault feature vector in Milvus is searched.

6. The method according to claim 1, characterized in that, Based on the target fault feature vector, the current fault repair strategy is determined, including: Retrieve fault location association data corresponding to the target fault feature vector from Milvus; The current fault repair strategy is determined based on the fault location association data corresponding to the target fault feature vector and the similarity between the target fault feature vector and the current business fault feature vector.

7. The method according to claim 5, characterized in that, After determining the current fault repair strategy based on the target fault feature vector, the method further includes: Collect the fault repair results of the current fault repair strategy; Based on the fault repair results of the current fault repair strategy, update the historical service fault feature vector and the fault feature weights.

8. An electronic device, characterized in that, The electronic device includes: At least one processor, and a memory communicatively connected to said at least one processor; The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the business fault operation and maintenance method according to any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the service fault operation and maintenance method according to any one of claims 1-7.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the business fault operation and maintenance method according to any one of claims 1-7.