LOF algorithm-based work order abnormal information detection method
By considering the priority of work orders in the LOF algorithm, filtering and updating the reachable distance of data points, and calculating the local reachable density and LOF values, the problem of inaccurate detection of work orders of different priority levels in the prior art is solved, and the accuracy of work order abnormal detection is improved.
Patent Information
- Application Number
- CN202510559011.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The existing LOF algorithms are difficult to adapt to work orders of different priority levels in work order abnormal detection, resulting in the abnormal points in high-density areas being judged as normal points, or the normal points in low-density areas being misjudged as abnormal points, affecting the accuracy of the detection.
By obtaining the processing time and priority of the work order, filtering the relevant data points in the data point collection, determining the key data points based on the priority, and updating the reachable distance based on the distance between the key data points and the current data points, calculating the local reachable density and LOF values to determine the abnormal work order.
By considering the priority differences between data points, the work ticket abnormality detection is more accurately judged, which improves the accuracy of abnormality detection.
Smart Images

Figure CN120067844A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and particularly relates to a method for detecting abnormal information of work orders based on the LOF algorithm. Background Art
[0002] A work order is a written or electronic carrier of work instructions, which details information such as work tasks to be completed, relevant requirements, performers, time limits, etc., in order to arrange, track, and manage the work process to ensure that each task can be completed accurately and in a timely manner.
[0003] The LOF algorithm is a density-based unsupervised anomaly detection method that identifies anomaly points by calculating the local density deviation between data points and their neighbors, and is widely used in the field of anomaly detection. When using the LOF algorithm for work order anomaly detection, due to the possible uneven density distribution of data points, a fixed k value is difficult to adapt to work orders with different priorities; and the calculation of the reachable distance in the algorithm is also based on a fixed neighborhood, without considering the priority differences between data points, which may lead to misjudgment of data points, resulting in anomaly points in high-density areas being misjudged as normal points, or normal points in low-density areas being misjudged as anomaly points; and the importance and potential impacts of different work orders are different, which leads to errors in the judgment of abnormal work orders and affects the accuracy of work order abnormal information detection. Therefore, it is necessary to adjust the relevant parameters in the algorithm according to the priorities of different data points. Summary of the Invention
[0004] The present invention provides a method for detecting abnormal information of work orders based on the LOF algorithm to overcome the deficiencies and defects of the above prior art.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is: A method for detecting abnormal information of work orders based on the LOF algorithm, the method includes the following steps: Obtain a number of work orders, and the data of each work order includes processing duration and priority; Use the processing duration in each work order as basic data to obtain a data point set of all work orders; Screen the data points within the k-distance neighborhood range of the data point set to obtain relevant data points; Screen the relevant data points according to the priority to obtain key data points; Screen the data points within the k-distance neighborhood range of the key data points to obtain target data points, and determine the reachable distance of each key data point in combination with the distances among the key data points, the current data points; Determine the priority level importance coefficient of key data points according to the priority of key data points, and update and adjust the reachable distance of each key data point according to the priority level importance coefficient; Obtain the local reachable density of the current data point according to the updated reachable distances of all key data points; Calculate the LOF value of the current data point according to the local reachable density of the current data point, and determine the abnormal work order based on this.
[0006] Preferably, screen the data points within the k-distance neighborhood range of the data point set to obtain relevant data points, including: Confirm the initial k-nearest neighbor distance of each data point in the data point set; Take the range with each data point as the center and the k-nearest neighbor distance as the radius as the k-distance neighborhood of each data point; Screen according to the k-distance neighborhood of each data point and determine all data points included in the k-distance neighborhood; Record all data points included in the k-distance neighborhood as relevant data points.
[0007] Preferably, screen the relevant data points according to the priority to obtain key data points, including: According to the actual situation of the business to which the work order belongs, determine a priority difference range with reference to the priority of the current data point; Traverse all relevant data points of the current data point and calculate the priority difference between each relevant data point and the current data point; Screen out the relevant data points within the priority difference range and record them as key data points.
[0008] Preferably, screen the data points within the k-distance neighborhood range of the key data points to obtain target data points, including: For each key data point, take the range with each key data point as the center and the k-nearest neighbor distance as the radius as the k-distance neighborhood of each key data point according to the determined k-nearest neighbor distance; Screen according to the k-distance neighborhood of each data point and determine all data points included in the k-distance neighborhood, and record them as target data points.
[0009] Furthermore, determine the reachable distance of each key data point in combination with the distances among the key data point, the current data point, including Calculate the distance between each key data point and each target data point; Take the average value of the distances between each key data point and all its target data points as the new k-nearest neighbor distance of this key data point; Compare the new k-nearest neighbor distance of each key data point with the distance between each key data point and the current data, and take the larger value of the two as the reachable distance of each key data point to the current data point; thus, the reachable distance of each key data point is obtained.
[0010] Preferably, determining the importance coefficient of the priority level of the key data point according to the priority of the key data point includes: Judge whether the priorities of each key data point and other key data points are the same; when the priority levels of two key data points are the same, assign the similarity coefficient of the priority level of each key data point as the first priority level similarity coefficient; when the priority levels of two key data points are different, assign the similarity coefficient of the priority level of each key data point as the second priority level similarity coefficient; the first priority level similarity coefficient is greater than the second priority level similarity coefficient; According to the similarity coefficient of the priority level between each key data point and other key data points, calculate the mean value of the similarity coefficient of the priority level between each key data point and all key data points, and use it as the importance coefficient of the priority level of each key data point.
[0011] Furthermore, update and adjust the reachable distance of each key data point according to the importance coefficient of the priority level, including: Determine the priority formation reason sequence according to the reason for the priority of the work order corresponding to each key data point; Determine the importance coefficient of the priority formation reason of each key data point according to the priority formation reason sequence of each key data point; Calculate the importance coefficient of each key data point according to the importance coefficient of the priority formation reason and the importance coefficient of the priority level of each key data point; Update and adjust the reachable distance of the key data point according to the importance coefficient of each key data point to obtain the updated reachable distance.
[0012] Still further, determining the importance coefficient of the priority level of the key data point according to the priority of the key data point also includes: For another key data point with the same priority level as each key data point, determine the number of formation reasons included in the intersection and union of the priority formation reason sequences of the two, so as to adjust the importance coefficient of the priority level of each key data point.
[0013] Preferably, determining the importance coefficient of the priority formation reason of each key data point according to the priority formation reason sequence of each key data point also includes: Classify the reasons for the priority of all key data points according to the reasons, obtain the importance sequence of the reasons for the formation of each key data point, and count the number of all reasons in the importance sequence of the reasons for the formation of each key data point, as well as the number of important reasons. According to the creation time of each work order, sort the importance sequences of the reasons for the formation of all work orders in chronological order. Centering on a certain work order, judge the reasons and their quantities that exist simultaneously in all the importance sequences of the reasons for the formation within a preset number of time points before and after it, sort the quantities of the reasons that exist simultaneously from largest to smallest, and record the top two reasons with the largest quantities as high-frequency reasons, and determine the quantity of high-frequency reasons. Determine the number of all reasons in the reason sequence for the formation of each key data point. Calculate the importance coefficient of the reasons for the formation of the priority of each key data point based on the number of all reasons in the importance sequence of the reasons for the formation, the number of important reasons in the importance sequence of the reasons for the formation, the number of high-frequency reasons, and the number of all reasons in the reason sequence for the formation.
[0014] Preferably, calculate the LOF value of the current data point according to the local reachability density of the current data point, and determine the abnormal work order based on this, including Calculate the LOF value of the current data point according to the local reachability density of the current data point. Compare the LOF value of the current data point with a preset first threshold. If the LOF value of the current data point is greater than the preset first threshold, then regard the current data point as an abnormal data point, and the corresponding work order is an abnormal work order; otherwise, the current data point is a normal data point.
[0015] Compared with the prior art, the beneficial effects of the present invention are: The present invention screens the data points within the k-distance neighborhood range of the data point set to obtain relevant data points; screens the relevant data points according to the priority to obtain key data points; screens the data points within the k-distance neighborhood range of the key data points to obtain target data points, and determines the reachable distance of each key data point by combining the distances among the key data points, the current data point; determines the importance coefficient of the priority level of the key data point according to the priority of the key data point, and updates and adjusts the reachable distance of each key data point according to the importance coefficient of the priority level. According to the updated reachable distances of all key data points, obtain the local reachability density of the current data point; calculate the LOF value of the current data point according to the local reachability density of the current data point, and determine the abnormal work order based on this.
[0016] The present invention screens the data points within the k-neighborhood range according to the priorities between different data points, which helps to focus more on the work orders that are similar to the current work order in terms of importance. At the same time, the reachable distance of each key data point is updated and adjusted according to the importance coefficient of the priority level, further refining the judgment of work order similarity, which helps to more accurately identify the work orders with similar potential problems, thereby improving the accuracy of anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a flowchart of the steps of the method for detecting work order anomaly information based on the LOF algorithm according to the present invention; Figure 2 is an attached drawing of a work order example provided by the present invention; Figure 3 is an example of the data point set of all work orders of the present invention; Figure 4 is an example of obtaining key data points according to the present invention; Figure 5 is a schematic block diagram of a computer device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. The present invention will be described in detail below in conjunction with the drawings and specific embodiments.
[0019] It should be understood that when used in this specification, the terms "including" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0020] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in this specification of the present invention, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.
[0021] It should be further understood that the term " / and" as used in this specification of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0022] Such as Figure 1As shown in the figure, a method for detecting abnormal information in work orders based on the LOF algorithm, the method includes the following steps: Obtain a number of work orders, and the data of each work order includes processing duration and priority; Take the processing duration in each work order as basic data to obtain a set of data points for all work orders; Screen the data points within the k-distance neighborhood range of the data point set to obtain relevant data points; Screen the relevant data points according to the priority to obtain key data points; Screen the data points within the k-distance neighborhood range of the key data points to obtain target data points, and determine the reachable distance of each key data point by combining the distances among the key data points, the current data point; Determine the importance coefficient of the priority level of the key data points according to the priority of the key data points, and update and adjust the reachable distance of each key data point according to the importance coefficient of the priority level; Obtain the local reachable density of the current data point according to the updated reachable distances of all key data points; Calculate the LOF value of the current data point according to the local reachable density of the current data point, and determine the abnormal work order based on this.
[0023] According to the priorities among different data points, the present invention screens the data points included within the k-neighborhood range, which helps to focus more on the work orders that are similar in importance to the current work order. At the same time, the reachable distance of each key data point is updated and adjusted according to the importance coefficient of the priority level, further refining the judgment of work order similarity, which helps to more accurately identify the work orders with similar potential problems, thereby improving the accuracy of abnormal detection.
[0024] In a specific embodiment, a number of work orders are obtained, and the data of each work order includes processing duration and priority; Specifically, the work orders are collected from the enterprise's internal management systems, such as customer support systems, operation and maintenance management systems, or project management systems; the detailed information of the work orders is recorded in these systems, and the data of each work order includes processing duration, priority, and also includes the creation time, processing status, completion time, problem description, processing information, current status, etc. of the work order.
[0025] Most enterprises have their own work order management systems, and the work order-related data can be directly exported from the system; for example, by writing scripts or using the API interfaces provided by the system, the work order data is extracted to the local area or stored in the data warehouse regularly (such as daily, weekly), or the log analysis tool is used to parse and extract the log files to obtain the required work order data; the work orders listed in this embodiment are as Figure 2 shown.
[0026] The collected work order information shall include, but not be limited to, the following key information: Work order number: The unique identifier for each work order, used to track and manage the work order; Creation time: Records the specific moment when the work order is created; Completion time: The time when the work order is processed. Combining with the creation time, the processing duration of the work order can be calculated; Priority: Usually divided into different levels such as high, medium, and low, reflecting the importance and urgency of the work order, which helps to distinguish normal and abnormal situations of work orders with different priorities during detection; Problem description: A detailed description of the problem involved in the work order; Processing information: Includes information about the personnel handling the work order, the solutions taken for the problems in the work order, etc.; Current status: Such as pending, in progress, completed, closed, etc. The change of the status can reflect whether the work order processing process is normal.
[0027] In this embodiment, taking the processing duration of each work order as the basic data, the data point set of all collected work orders is obtained. As Figure 3 shown in the example, the processing durations of each work order are not exactly the same. Horizontally represents the processing durations of different work orders, and vertically represents that there are multiple work orders with the same processing duration under this processing duration.
[0028] In a specific embodiment, the data points within the k-distance neighborhood range of the data point set are screened to obtain relevant data points, including: Confirm the initial k-nearest neighbor distance of each data point in the data point set; In this embodiment, taking point 1 in the data point set as an example, the initial k-nearest neighbor distance of point 1 is determined to be 5.
[0029] Taking the range with each data point as the center and the k-nearest neighbor distance as the radius as the k-distance neighborhood of each data point; Taking point 1 as the center and the k-nearest neighbor distance as the radius, the range is the k-distance neighborhood of point 1. The k-nearest neighbor distance is used to determine the neighborhood range of each data point. The k-distance neighborhood range can reflect the local density situation around the data point, and a suitable k value can accurately reflect the local environmental characteristics where the data point is located.
[0030] Screen and determine all data points included in the k-distance neighborhood according to the k-distance neighborhood of each data point; In this embodiment, according to the k-distance neighborhood of point 1, all data points included in the k-distance neighborhood can be determined. When the k-nearest neighbor distance is 5, data points 2 - 11 are all within the k-distance neighborhood of point 1, and the data points included in the k-distance neighborhood are recorded as relevant data points.
[0031] All the data points included within the k-distance neighborhood are denoted as relevant data points.
[0032] When determining relevant data points through the above steps, all data points are treated equally, ignoring the differences in importance among different data points. The importance, processing urgency, and potential impact of different work order data vary. In work order anomaly detection, high-priority work orders may be more susceptible to anomalies, or rather, the anomalies of high-priority work orders are more worthy of attention. Therefore, based on the priorities of all relevant data points determined by the above steps, the relevant data points are screened to exclude the interference of irrelevant or low-correlation information.
[0033] In a specific embodiment, the relevant data points are screened according to priorities to obtain key data points, including: According to the actual situation of the business to which the work order belongs, with the priority of the current data point as a reference, a priority difference range is determined. For example, the priority difference range is set such that the difference does not exceed one level. The difference between "low" and "high" is two levels, and the differences between "low" and "medium", and between "medium" and "high" are one level.
[0034] Traverse all relevant data points of the current data point, and calculate the priority difference between each relevant data point and the current data point; Screen out the relevant data points within the priority difference range and denote them as key data points.
[0035] For example, represent the three priorities of "low", "medium", and "high" as A, B, and C respectively. When taking the 1st data point as the center, if its priority is A, then after screening its relevant data points, the obtained key data points are the 2nd, 3rd, 4th, 6th, 9th, 10th, and 11th data points. The priority differences between these data points and the 1st data point are within the set difference range, while the priority differences between other data points and the 1st data point are relatively large, so they are excluded, as Figure 4 shown in the example.
[0036] Through the above steps, in order to exclude the interference of irrelevant or low-correlation data points, the relevant data points included within the k-neighborhood range of the current data point are screened, which is equivalent to optimizing the k-distance. Also, since the drawback of the reachable distance of each key data point is related to the k-neighborhood distance, the k-neighborhood distance of each key data point needs to be adjusted, specifically as follows: Screen the data points within the k-distance neighborhood range of the key data points to obtain target data points, including: For each key data point, according to the determined k-neighborhood distance, the range centered on each key data point with the k-neighborhood distance as the radius is used as the k-distance neighborhood of each key data point; Filter according to the k-distance neighborhood of each data point, determine all data points included in the k-distance neighborhood, and denote them as target data points.
[0037] In this embodiment, after determining the target data points, the reachable distance of each key data point is determined by combining the distances among the key data points, the current data point, including: Calculate the distance between each key data point and each target data point; Take the average value of the distances between each key data point and all its target data points as the new k-nearest distance of this key data point; Compare the new k-nearest distance of each key data point with the distance between each key data point and the current data point, and take the larger value of the two as the reachable distance from each key data point to the current data point; thus, the reachable distance of each key data point is obtained, denoted as d2.
[0038] In this embodiment, in the work order data, work orders with different priorities have different business importance and different degrees of impact on the business. High-priority work orders usually involve key business processes, major customer requirements, etc. Once an anomaly occurs, it may cause relatively serious consequences. The reasons for generating work orders with the same priority may also vary, and the importance of different generation reasons is also different; for example, work orders under different business modules may have the same priority due to the same problem.
[0039] When calculating the local reachability density of the current data point using the reachable distance, only the spatial distance between data points is considered, while the relevance of priorities between data points is ignored, and work orders with different priorities may be distributed in different density regions.
[0040] Therefore, this embodiment combines the similarity of priorities between data points, assigns different weights to the reachable distances of different data points, so that when performing the lof algorithm to detect work order anomalies, it can pay more attention to the potential associations between data points, adapt to the different density distributions of work orders with different priorities, and improve the accuracy of anomaly detection.
[0041] In a specific embodiment, determine the priority level importance coefficient of the key data point according to the priority of the key data point, including: Judge whether the priorities of each key data point and other key data points are the same; when the priority levels of two key data points are the same, assign the priority level similarity coefficient of each key data point as the first priority level similarity coefficient, for example, 1; when the priority levels of two key data points are different, then assign the priority level similarity coefficient of each key data point as the second priority level similarity coefficient, for example, 0.5; the first priority level similarity coefficient is greater than the second priority level similarity coefficient; According to the priority level similarity coefficient between each key data point and other key data points, calculate the mean value of the priority level similarity coefficients of each key data point and all key data points, and use it as the priority level importance coefficient of each key data point.
[0042] In this embodiment, the priority level importance coefficient of the i-th key data point is calculated The calculation formula is as follows: ; In the formula, represents the priority level similarity coefficient between the i-th key data point and the q-th key data point; represents the mean value of the priority level similarity coefficients of the i-th key data point and all key data points, where the larger the value of, the more similar the level of the i-th key data point is to other key data points, and a higher importance should be assigned to it; z represents the number of all key data.
[0043] In a specific embodiment, the reachable distance of each key data point is updated and adjusted according to the priority level importance coefficient, including: Determine the priority formation reason sequence according to the reason for the priority of the work order corresponding to each key data point.
[0044] Since work orders with different priorities have different formation reasons; for example, high-priority work orders may be due to network failures, system failures, privilege management vulnerabilities, data leaks, time-limited service requests, special time requirements, key customer needs, etc.; medium-priority work orders may be due to performance degradation, local blockage of business processes, partial function failures, etc.; low-priority work orders may be due to interface display problems, minor error reports, etc.
[0045] For two data points with the same priority level, the reasons for their priorities are not necessarily the same, and there may not be only one reason for the formation of each work order priority.
[0046] Therefore, according to the reason for the priority of the work order corresponding to each data point, determine its formation reason sequence, and the formation reason sequences of data points with the same priority are not exactly the same.
[0047] For example, the priority of data point 1 is A, high priority, and its formation reason sequence may be {network failure, system failure, data leak, key customer need}.
[0048] The priority of data point 2 is also high priority, and its formation reason sequence may be {special time requirement, local blockage of business process}.
[0049] In this embodiment, it also includes: For another key data point with the same priority level as each of the key data points, determine the number of contributing factors contained in the intersection and union of the sequences of contributing factors for the two, denoted as m1 and m2 respectively; and adjust the importance coefficient of the priority level of each of the key data points accordingly.
[0050] Among them, The larger the value of, the more similar the contributing factors of the two are. Based on this, adjust the importance coefficient of the priority level of the i-th key data point.
[0051] The adjusted importance coefficient of the priority level can be expressed as: ; Among them, represents the number of key data points with the same priority level as the i-th key data point; and respectively represent the number of intersections and unions of the factors within the sequence of contributing factors of the i-th key data point and the t-th key data point. The more data points with the same priority as the i-th data point, and the larger the intersection-union ratio for each, the more similar they are, indicating that the importance coefficient of this data point is larger, that is, the larger the value.
[0052] According to the sequence of contributing factors for the priority of each key data point, determine the importance coefficient of the contributing factors for the priority of each key data point. In this embodiment, there may be more than one contributing factor for the priority of each key data point, and the importance of different contributing factors may also vary; data points with the same priority may contain contributing factors of different importance. Based on this, determine the importance coefficient of the contributing factors for the priority of each key data point.
[0053] Calculate the importance coefficient of each key data point based on the importance coefficient of the contributing factors for the priority of each key data point and the importance coefficient of the priority level; Update and adjust the reachable distance of each key data point according to the importance coefficient of each key data point to obtain the updated reachable distance.
[0054] In this embodiment, determining the importance coefficient of the contributing factors for the priority of each key data point according to the sequence of contributing factors for the priority of each key data point further includes: Classify the contributing factors for the priority of all key data points according to the factors, obtain the sequence of importance of the contributing factors for each key data point, and count the number of all factors and the number of important factors in the sequence of importance of the contributing factors for each key data point.
[0055] In this embodiment, since there are multiple reasons for the formation of all key data points, but they can be divided into several major categories. For example, network failures and system failures can be classified as reasons for the interruption of critical business processes; permission management vulnerabilities and data leaks can be classified as reasons for major security issues; time-limited service requests and special time requirements are for urgent customer needs; key customer needs are for important customers; performance degradation, partial blockage of business processes, and partial function failures are for the impact on general business functions; interface display problems and minor error messages are for minor faults.
[0056] For different types of reasons, their importance is also different. The types such as interruption of critical business processes, major security issues, urgent customer needs, and important customers usually have a relatively high importance; the impact on general business functions is of medium importance; and the importance of minor faults is relatively low.
[0057] Classify the reasons for the formation of the priority of all key data points according to the reason categories to obtain the importance sequence of the reasons for the formation of each key data point, expressed in terms of reason types. For example: The importance sequence of the reasons for the formation of data point 1 can be expressed as {important reason, important reason, important reason, important reason}; the importance sequence of the reasons for the formation of data point 2 can be expressed as {important reason, medium important reason}.
[0058] Count the number of all reasons and the number of important reasons in the importance sequence of the reasons for the formation of each key data point; thereby determine the number of times the reasons of different importance appear in the importance sequence of the reasons for the formation of each key data point. Among them, the more times the important reasons appear and the larger the proportion, the greater the importance coefficient of the reasons for the formation of this data point.
[0059] Since the importance of different reasons is also affected by their occurrence frequencies, some reasons of medium importance or low importance may appear continuously within a certain period of time, or there may be a large number of work orders caused by medium or low importance reasons appearing simultaneously at a certain time. At this time, the importance of these reasons is also relatively large.
[0060] According to the creation time of each work order, sort the importance sequences of the reasons for the formation of all work orders in chronological order; Taking a certain work order as the center, judge the reasons and their quantities that exist simultaneously in all the importance sequences of the reasons for the formation within a preset number of time points before and after it, sort the quantities of the reasons that exist simultaneously from large to small, and record the top two reasons with the largest quantities as high-frequency reasons, and determine the quantities of the high-frequency reasons; In this embodiment, the preset quantity is 5. Taking the i-th work order, that is, the i-th sequence of importance of formation reasons as the center, judge the reasons and their quantities that exist simultaneously in the sequences of importance of formation reasons included in the first 5 and the last 5 time points before and after it and the i-th sequence of importance of formation reasons. Sort the quantities of the reasons that exist simultaneously in descending order, and record the first two reasons with the largest quantities as high-frequency reasons.
[0061] For example, for the key data points 2, 3, 4, 6, 9, 10, 11 after screening, when taking the 9th key data point as the center, all data points are included within its first 5 and last 5 time points before and after, that is, the sequences of importance of formation reasons are included. Among these sequences, the reasons that exist simultaneously include performance degradation, interface display problem, partial function failure, data leakage, and customer importance. However, the interface display problem appears in 10 of these sequences, and the reason of performance degradation appears in 8 sequences of importance of formation reasons. Therefore, the two reasons of performance degradation and interface display problem are recorded as high-frequency reasons.
[0062] Determine the quantity of all reasons in the sequence of formation reasons for each key data point; Calculate the importance coefficient of the formation reasons with priority for each key data point according to the quantity of all reasons in the sequence of importance of formation reasons, the quantity of important reasons in the sequence of importance of formation reasons, the quantity of high-frequency reasons, and the quantity of all reasons in the sequence of formation reasons.
[0063] In this embodiment, the importance coefficient of the formation reasons with priority calculated for each key data point can be expressed as: ; wherein, represents the importance coefficient of the formation reasons with priority for the i-th key data point; represents the length of the sequence of importance of formation reasons for the i-th key data point, that is, the quantity of all reasons included; represents the quantity of important reasons in the sequence of importance of formation reasons; represents the quantity of all reasons in the sequence of formation reasons for the i-th key data point; represents the quantity of high-frequency reasons in the sequence of importance of formation reasons; the larger the proportion of important reasons and high-frequency reasons, the larger the importance coefficient of its formation reasons.
[0064] In this embodiment, according to the importance coefficient of the formation reasons with priority and the importance coefficient of the priority level for each key data point, the importance coefficient of each key data point is calculated. Specifically as follows, the importance of each key data point is jointly determined by its priority and formation reasons. Therefore, the importance coefficient It can be expressed as: .
[0065] Update and adjust the reachable distance of each key data point according to the importance coefficient of the key data point to obtain the updated reachable distance of the i-th key data point , specifically as follows: , represents the reachable distance of the i-th key data point.
[0066] According to the updated reachable distances of all key data points (data points 2, 3, 4, 6, 9, 10, 11), obtain the local reachable density of the current data point (data point 1).
[0067] In a specific embodiment, according to the local reachable density of the current data point, calculate the LOF value of the current data point, and determine the abnormal work order based on this, including: Calculate the LOF value of the current data point according to the local reachable density of the current data point; In this embodiment, the LOF value of each data point can be calculated in the same way; the larger the LOF value, the more likely the data point is an abnormal point.
[0068] Compare the LOF value of the current data point with a preset first threshold. If the LOF value of the current data point is greater than the preset first threshold, then regard the current data point as an abnormal data point, and the corresponding work order is an abnormal work order; otherwise, the current data point is a normal data point.
[0069] In this embodiment, the first threshold can be preset to 1.5, and compare the LOF value of each data point with the preset first threshold; when the LOF value of the data point is greater than 1.5, regard it as an abnormal data point, and the corresponding work order is an abnormal work order.
[0070] Further verify the abnormal work order, determine the problems and severity thereof, and perform corresponding processing.
[0071] Please refer to Figure 5 , Figure 5 is a schematic block diagram of a computer device provided by an embodiment of the present invention. The computer device 500 is a server, and the server can be an independent server or a server cluster composed of multiple servers.
[0072] Refer to Figure 5 , the computer device 500 includes a processor 502, a memory, and a network interface 505 connected through a system bus 501. Among them, the memory may include a non-volatile storage medium 503 and an internal memory 504.
[0073] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, it can cause the processor 502 to execute a work order exception information detection method based on the LOF algorithm.
[0074] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500.
[0075] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, it can cause the processor 502 to execute a work order exception information detection method based on the LOF algorithm.
[0076] The network interface 505 is used for network communication, such as providing the transmission of data information, etc. Those skilled in the art can understand that Figure 5 the structure shown in is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the computer device 500 to which the solution of the present invention is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0077] Among them, the processor 502 is used to run the computer program 5032 stored in the memory to implement the work order exception information detection method based on the LOF algorithm disclosed in the embodiments of the present invention.
[0078] Those skilled in the art can understand that Figure 5 the embodiments of the computer device shown in do not constitute a limitation on the specific composition of the computer device. In other embodiments, the computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements. For example, in some embodiments, the computer device may only include a memory and a processor. In such an embodiment, the structures and functions of the memory and the processor are the same as those in Figure 5 the shown embodiment, and will not be described in detail here.
[0079] It should be understood that in the embodiments of the present invention, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0080] In another embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium may be a non-volatile computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the work order exception information detection method based on the LOF algorithm disclosed in the embodiments of the present invention.
[0081] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for detecting abnormal information of a work order based on a LOF algorithm, characterized in that: The method comprises the following steps: Get several work orders, each of which has data including processing time and priority; The processing time in each work order is used as the basic data to obtain a set of data points for all work orders; Screen the data points within the k-distance neighborhood of the data point set to obtain relevant data points; Filter relevant data points according to priority to obtain key data points; The data points within the k-distance neighborhood of the key data point are screened to obtain the target data point, and the reachable distance of each key data point is determined by combining the distance between the key data point and the current data point; Determine the priority level importance coefficient of the key data point according to the priority of the key data point, and update and adjust the reachable distance of each key data point according to the priority level importance coefficient; According to the updated reachable distances of all key data points, the local reachable density of the current data point is obtained; According to the local reachable density of the current data point, the LOF value of the current data point is calculated and used to determine the abnormal work order.
2. The method for detecting abnormal work order information based on the LOF algorithm according to claim 1, characterized in that: The data points within the k-distance neighborhood of the data point set are screened to obtain relevant data points, including: Determine the initial k-neighbor distances of each data point in the data point set; A range with each data point as the center and k neighboring distances as the radius is used as the k distance neighborhood of each data point; Screening is performed according to the k-distance neighborhood of each data point and determining all data points contained in the k-distance neighborhood; All data points contained in the k-distance neighborhood are recorded as relevant data points.
3. The work order abnormal information detection method based on the LOF algorithm according to claim 1 is characterized in that: Filter relevant data points according to priority to obtain key data points, including: According to the actual situation of the business to which the work order belongs, a priority difference range is determined with reference to the priority of the current data point; Traverse all related data points of the current data point and calculate the priority difference between each related data point and the current data point; Filter out relevant data points within the priority difference range and record them as key data points.
4. The work order abnormal information detection method based on the LOF algorithm according to claim 1 is characterized in that: The data points within the k-distance neighborhood of the key data point are screened to obtain the target data points, including: For each key data point, according to the determined k-neighboring distances, a range with the k-neighboring distances as the radius and the center of each key data point is used as the k-distance neighborhood of each key data point; The k-distance neighborhood of each data point is screened and all data points contained in the k-distance neighborhood are determined and recorded as target data points.
5. The method for detecting abnormal work order information based on the LOF algorithm according to claim 4, characterized in that: The reachable distance of each key data point is determined by combining the distances between the key data point and the current data point, including: Calculate the distance between each key data point and each target data point; The average value of the distances between each key data point and all its target data points is used as the new k-neighbor distance of the key data point; The new k-neighbor distance of each key data point is compared with the distance between each key data point and the current data, and the larger value of the two is used as the reachable distance from each key data point to the current data point; thereby obtaining the reachable distance of each key data point.
6. The method for detecting abnormal work order information based on the LOF algorithm according to claim 1, characterized in that: The priority level importance coefficient of the key data point is determined according to the priority of the key data point, including: Determine whether each key data point has the same priority as other key data points; when the priority levels of two key data points are the same, assign a priority level similarity coefficient of each key data point as a first priority level similarity coefficient; when the priority levels of two key data points are different, assign a priority level similarity coefficient of each key data point as a second priority level similarity coefficient; the first priority level similarity coefficient is greater than the second priority level similarity coefficient; According to the priority level similarity coefficients between each key data point and other key data points, the average value of the priority level similarity coefficients between each key data point and all key data points is calculated and used as the priority level importance coefficient of each key data point.
7. The method for detecting abnormal work order information based on the LOF algorithm according to claim 6, characterized in that: The reachable distance of each key data point is updated and adjusted according to the priority level importance coefficient, including: Determine the priority formation reason sequence based on the priority formation reason of each key data point corresponding to the work order; According to the priority formation reason sequence of each key data point, determine the importance coefficient of the priority formation reason of each key data point; The importance coefficient of each key data point is calculated based on the importance coefficient of the priority formation reason and the importance coefficient of the priority level of each key data point; The reachable distance of each key data point is updated and adjusted according to the importance coefficient of the key data point to obtain an updated reachable distance.
8. The method for detecting abnormal work order information based on the LOF algorithm according to claim 7, characterized in that: The priority level importance coefficient of the key data point is determined according to the priority of the key data point, and also includes: For another key data point with the same priority level as each key data point, the intersection and the number of causes contained in the union of the priority cause sequences of the two are determined, so as to adjust the priority level importance coefficient of each key data point.
9. The method for detecting abnormal work order information based on the LOF algorithm according to claim 1, characterized in that: According to the priority formation reason sequence of each key data point, the importance coefficient of the priority formation reason of each key data point is determined, which also includes: Classify the priority formation reasons of all key data points according to the reasons, obtain the importance sequence of the formation reasons of each key data point, and count the number of all reasons in the importance sequence of the formation reasons of each key data point, as well as the number of important reasons; According to the creation time of each work order, sort the importance sequence of the reasons for the formation of all work orders in chronological order; Taking a work order as the center, determine the causes and their number that exist simultaneously in all the cause importance sequences within a preset number of time points before and after the work order, sort the number of the causes that exist simultaneously from large to small, record the first two causes with the largest number as high-frequency causes, and determine the number of high-frequency causes; Determine the number of all causes in the cause sequence for each key data point; Based on the number of all causes in the cause importance sequence, the number of important causes in the cause importance sequence, the number of high-frequency causes, and the number of all causes in the cause sequence, the priority cause importance coefficient of each key data point is calculated.
10. The method for detecting abnormal work order information based on the LOF algorithm according to claim 1, characterized in that: According to the local reachable density of the current data point, the LOF value of the current data point is calculated and used to determine the abnormal work order, including: According to the local reachable density of the current data point, the LOF value of the current data point is calculated; The LOF value of the current data point is compared with the preset first threshold. If the LOF value of the current data point is greater than the preset first threshold, the current data point is regarded as an abnormal data point and the corresponding work order is an abnormal work order; otherwise, the current data point is a normal data point.
Citation Information
Patent Citations
LOF-based abnormal high-voltage metering point screening method and system
CN112083371A
Feeder adaptive control method of traction power supply wide-area protection measurement and control system
CN116231603A
Targeted short message reaching method based on dynamic position tracking and crowd density analysis
CN119052736A
Multi-source data quality improvement method and system and storage medium
CN119513078A
Abnormal data detection method and device, storage medium and electronic equipment
CN119760593A