Processing method and device for nearest neighbor calculation
By sampling access traffic and asynchronous real neighbor calculations, the problem of poor observability of nearest neighbor calculation results is solved, real-time monitoring of approximate nearest neighbor calculation recall is realized, and systemic risks are reduced.
Patent Information
- Application Number
- CN202111125713.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-09-24
AI Technical Summary
Nearest neighbor calculation results have poor observability and are difficult to monitor their accuracy in real time, resulting in a decrease in systemic risks and recall rates.
By sampling access traffic, asynchronous real neighbor calculations are performed in real time, and compared with the approximate neighbor calculation results, the recall rate of approximate neighbor calculations is determined.
Real-time and efficient monitoring of the approximate neighbor calculation recall rate in large-scale vector computing scenarios is realized, reducing systemic risks and improving computing efficiency.
Smart Images

Figure CN113836440B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of artificial intelligence and big data, and in particular to the fields of intelligent search, intelligent recommendation, advertising recommendation, knowledge graphs, and user understanding, and can be used to evaluate scenarios such as the recall rate of approximate neighbor calculations. Background Art
[0002] Nearest neighbor computing is often a global service, and the stability of its system is critical. At the same time, since the result of nearest neighbor computing is not ordinary numerical data, but relational data that represents the nearest neighbor relationship, it is difficult to track and monitor its accuracy through simple methods. Summary of the invention
[0003] The present disclosure provides a processing method, apparatus, device, storage medium and computer program product for neighbor computing.
[0004] According to one aspect of the present disclosure, a processing method for nearest neighbor calculation is provided, comprising: in the process of performing approximate nearest neighbor calculation, sampling access traffic to obtain at least one sampled traffic; performing asynchronous true nearest neighbor calculation on the at least one sampled traffic in real time to obtain at least one corresponding true nearest neighbor calculation result; and determining the recall rate of the approximate nearest neighbor calculation based on the at least one true nearest neighbor calculation result and the corresponding at least one approximate nearest neighbor calculation result, wherein the at least one approximate nearest neighbor calculation result is obtained by performing the approximate nearest neighbor calculation on the at least one sampled traffic.
[0005] According to another aspect of the present disclosure, a processing device for nearest neighbor calculation is provided, including: a sampling module, used to sample access traffic during approximate nearest neighbor calculation to obtain at least one sampled traffic; a real nearest neighbor calculation module, used to perform asynchronous real nearest neighbor calculation on the at least one sampled traffic in real time to obtain at least one corresponding real nearest neighbor calculation result; and a determination module, used to determine the recall rate of the approximate nearest neighbor calculation based on the at least one real nearest neighbor calculation result and the corresponding at least one approximate nearest neighbor calculation result, wherein the at least one approximate nearest neighbor calculation result is obtained by performing the approximate nearest neighbor calculation on the at least one sampled traffic.
[0006] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in an embodiment of the present disclosure.
[0007] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method according to the embodiment of the present disclosure.
[0008] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the method according to the embodiment of the present disclosure is implemented.
[0009] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0011] Figure 1 The system architecture suitable for the embodiments of the present disclosure is exemplified;
[0012] Figure 2 A flowchart of a processing method for nearest neighbor calculation according to an embodiment of the present disclosure is exemplarily shown;
[0013] Figure 3 A schematic diagram exemplarily shows a true neighbor calculation according to an embodiment of the present disclosure;
[0014] Figure 4 A schematic diagram exemplarily shows an approximate nearest neighbor calculation according to an embodiment of the present disclosure;
[0015] Figure 5 A block diagram exemplarily shows a processing device for neighbor calculation according to an embodiment of the present disclosure; and
[0016] Figure 6 The block diagram of an electronic device for implementing the method and apparatus of the embodiments of the present disclosure is exemplarily shown. DETAILED DESCRIPTION
[0017] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0018] It should be understood that the full name of ANN is approximate nearest neighbor, that is, approximate nearest neighbor calculation. The full name of KNN is k-nearest neighbors, that is, K nearest neighbor calculation, taking K nearest neighbors. In practical applications, the nearest neighbor can be solved by true nearest neighbor calculation, or the nearest neighbor can be solved by approximate nearest neighbor calculation at an accelerated speed. For example, KNN algorithm, ANN algorithm, etc. are widely used in the fields of calculating the distance between vectors and correlation applications. For example, they are widely used in the fields of Internet recommendation systems, advertising systems, search systems, etc.
[0019] It should also be understood that the issues involved in the industrially available implementation of approximate neighbor calculations are quite complex. These are likely to lead to serious deviations or even errors in the results of approximate neighbor calculations, which will bring serious systemic risks. For example, vector calculations are often a complex process, and vector expressions do not have semantic characteristics, so vector calculation anomalies are difficult to detect, and abnormal calculation results can lead to global errors. For another example, the approximate vector calculation algorithm (also known as the fast neighbor calculation algorithm or the accelerated neighbor calculation algorithm) itself will directly affect the construction of the index. Each update of the approximate vector calculation algorithm requires an update of the index. If the upgrade and update of the approximate vector calculation algorithm is not synchronized with the upgrade and update of the index, it will cause abnormal vector calculation results.
[0020] In addition, the index built to achieve vector accelerated computing needs to be calculated offline first and then synchronized to online use. The update process has risks such as update failure and update anomaly.
[0021] At the same time, the results obtained by using the ANN algorithm are not ordinary numerical results, but other objects, vectors and their spatial distance rankings, so the observability is relatively poor.
[0022] In summary, ANN algorithms are widely used, large-scale, and have a vital impact. They can accelerate vector calculations, but the ANN calculation results do not have good observability.
[0023] Therefore, the embodiments of the present disclosure provide a processing method for nearest neighbor computing, which can efficiently monitor the recall rate of approximate nearest neighbor computing used in industrial application scenarios involving large-scale vector computing in real time.
[0024] The present disclosure will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0025] The system architecture of the processing method and apparatus for neighbor computing suitable for the embodiments of the present disclosure is introduced as follows.
[0026] Figure 1 The system architecture suitable for the embodiment of the present disclosure is exemplified. It should be noted that: Figure 1The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, in order to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other environments or scenarios.
[0027] like Figure 1 As shown, the system architecture 100 in the embodiment of the present disclosure may include: a server 101 and a server 102. The server 101 may be used to provide online approximate neighbor computing services. The server 102 may provide real neighbor computing services, that is, real neighbor computing may be performed on a sample of online access traffic in an asynchronous manner. The server 102 may be a server or a distributed server cluster.
[0028] The application scenarios of the processing method and device for nearest neighbor calculation suitable for the embodiments of the present disclosure are introduced as follows.
[0029] The processing method for neighbor calculation provided by the embodiments of the present disclosure can be applied to the fields of intelligent recommendation, intelligent search, advertising recommendation, knowledge graph, user understanding, etc., and has key application value in core fields such as intelligent recommendation, intelligent search, and advertising recommendation. For example, the technical solution can be used to support neighbor applications of big data social network relationship data.
[0030] According to an embodiment of the present disclosure, the present disclosure provides a processing method for nearest neighbor calculation.
[0031] Figure 2 A flowchart of a processing method for neighbor calculation according to an embodiment of the present disclosure is exemplified.
[0032] like Figure 2 As shown, the processing method 200 for neighbor calculation may include: operations S210 to S230.
[0033] In operation S210, in the process of performing approximate neighbor calculation, access traffic is sampled to obtain at least one sampled traffic.
[0034] In operation S220, asynchronous true neighbor calculation is performed on at least one sampled flow in real time to obtain at least one corresponding true neighbor calculation result.
[0035] In operation S230, a recall rate of the approximate neighbor calculation is determined based on at least one true neighbor calculation result and at least one corresponding approximate neighbor calculation result, wherein the at least one approximate neighbor calculation result is obtained by performing an approximate neighbor calculation on at least one sampled flow.
[0036] It should be understood that neighbor calculation refers to finding the closest vector, i.e., the nearest neighbor (such as top N), by calculating the spatial distance from a vector (a numerical abstraction of an object) to other vectors (which can be called candidate vectors).
[0037] In an embodiment of the present disclosure, the real neighbor calculation may include: finding the vector closest to it, i.e., the nearest neighbor, by calculating the Euclidean distance or cosine distance from a vector to other vectors. In other words, the above-mentioned real neighbor calculation may include: neighbor calculation based on Euclidean distance calculation or cosine distance calculation.
[0038] For example, if the nearest neighbor of vector A is calculated using true nearest neighbor calculation, the Euclidean distance or cosine distance between vector A and all candidate vectors must be calculated one by one, and then sorted according to the true nearest neighbor calculation results. Figure 3 As shown, taking candidate vector B as an example, when the true nearest neighbor calculation method is used to calculate the distance between two vectors, the Euclidean distance or cosine distance from vector A to vector B is directly calculated.
[0039] Therefore, the true nearest neighbor calculation method obtains the true distance between vector A and vector B. Therefore, the result obtained by the true nearest neighbor calculation can be used as a reference value to judge the recall rate of the approximate nearest neighbor calculation.
[0040] However, the complexity of true neighbor calculation is very high. It is necessary to calculate the (Euclidean or cosine) spatial distance for each candidate object or candidate vector, which is time-consuming and costly, making it difficult to apply to large-scale industrial applications. In actual scenarios, the scale of candidate vectors or candidate objects may reach tens of millions, hundreds of millions, or even larger. In this case, if the true neighbor calculation method is used, there are problems of high computational cost and slow response speed.
[0041] Therefore, for industrial application scenarios involving large-scale and ultra-large-scale vector computing, approximate neighbor computing solutions can be used. For example, the Annoy algorithm (Approximate Nearest Neighbors Oh Yeah, referred to as the simulated nearest neighbor algorithm), the HNSW algorithm (Hierarchical Navigable Small World, referred to as the hierarchical clustering algorithm), the Faiss PQ algorithm (Product Quantize, referred to as the product quantization algorithm), etc. can be used to implement approximate neighbor computing.
[0042] It should be understood that the goal of the annoy algorithm is to establish a data structure that can find the nearest point to any query point in a relatively short time, sacrificing accuracy in exchange for a much faster speed than brute force search under the condition that accuracy allows. The HNSW algorithm is a hierarchical optimization of the NSW (Navigable Small World, referred to as the clustering algorithm) algorithm, which can improve query performance. It starts with searching on a sparse graph and gradually goes deeper into the underlying graph. The Faiss PQ algorithm is a product quantization algorithm, where the product is the Cartesian product, which means decomposing the original vector space into the Cartesian product of several low-dimensional vector spaces, and quantizing the decomposed low-dimensional vector spaces separately. In this way, each vector can be represented by a quantized combination of multiple low-dimensional vector spaces.
[0043] In the embodiments of the present disclosure, approximate neighbor calculation (also known as accelerated neighbor calculation or fast neighbor calculation) may include: calculating the spatial distance from a vector to other vectors by clustering method, and finding the vector closest to it, that is, the nearest neighbor. In other words, approximate neighbor calculation may include: neighbor calculation based on clustering method.
[0044] For example, if the nearest neighbor of vector A is calculated using approximate nearest neighbor, it is not necessary to calculate the Euclidean distance or cosine distance between vector A and all candidate vectors one by one. Instead, all clusters corresponding to all candidate vectors can be determined first, and then the spatial distance from vector A to each cluster center point can be calculated. Then, the spatial distance from each candidate vector to its cluster center point can be determined by table lookup. Then, the spatial distance from vector A to each candidate vector can be obtained by fitting the two, and finally, the candidate vectors can be sorted according to the approximate nearest neighbor calculation results. Figure 4 As shown in the figure, taking candidate vector B as an example, when using the approximate nearest neighbor calculation method to calculate the distance between two vectors, the distance L1 from vector A to the cluster center C of vector B is calculated first, and then the distance L2 between vector B and vector C is obtained by looking up the table, and L1 and L2 are fitted to obtain the distance between the two. For example, L1+L2 can be used as the distance from vector A to vector B.
[0045] Since the number of cluster centers corresponding to all candidate vectors is much smaller than the size of all candidate vectors, and the distance from the candidate vector in each vector cluster to its cluster center can be obtained by pre-calculation, the distance from the candidate vector in the vector cluster to its cluster center can be obtained by table lookup in actual use. Therefore, the computational cost of approximate neighbor calculation can be greatly reduced, and the overall computational efficiency can be greatly improved. However, the accuracy of approximate neighbor calculation will decrease. In other words, approximate neighbor calculation can simplify the calculation process and achieve high-performance computing and fast response, but it also brings another problem, that is, the recall rate will be compromised and no longer 100%.
[0046] Furthermore, the accuracy of a single distance calculation will eventually affect the overall recall rate of the nearest neighbor calculation. In the disclosed embodiment, the recall rate refers to the similarity between the nearest neighbor result obtained by the approximate nearest neighbor calculation and the nearest neighbor result obtained by the true nearest neighbor calculation.
[0047] Moreover, the approximate nearest neighbor calculations available in the industry usually involve more complex issues, which may lead to serious deviations or even serious errors in the nearest neighbor results obtained through approximate nearest neighbor calculations, which will bring serious systemic risks.
[0048] At the same time, the result of the nearest neighbor calculation is not ordinary numerical data, but the ranking of other objects, vectors and their spatial distances, which has poor observability.
[0049] Therefore, for approximate nearest neighbor computation, there is a need for robust, real-time, and industrially viable recall monitoring.
[0050] Therefore, in the embodiments of the present disclosure, for large-scale industrial application scenarios, approximate neighbor calculations can be used online to obtain corresponding nearest neighbors, thereby simplifying the vector calculation process and achieving the purpose of high-performance vector calculations and rapid response. At the same time, in the process of performing approximate neighbor calculations, the access traffic can be sampled to obtain at least one sampled traffic. For each sampled traffic, real neighbor calculations can be performed in real time in an asynchronous manner to obtain the corresponding real nearest neighbor results. Then, whether for each individual sampled traffic or for all sampled traffic as a whole, the recall rate of the approximate neighbor calculation can be determined by comparing the nearest neighbor results obtained by the real neighbor calculation with the corresponding nearest neighbor results obtained based on the approximate neighbor calculation.
[0051] It should be understood that in the embodiments of the present disclosure, the computing scale can be reduced overall by sampling. At the same time, since the scale of access traffic in industrial applications is usually large, the sampled traffic can fully express the actual situation online. In addition, by asynchronous computing, the real neighbor computing and the online high-performance approximate neighbor computing can be separated from each other, and the asynchronous computing method will not affect the response and performance of the online service.
[0052] In addition, compared with the pre-evaluation recall rate that cannot reflect the actual online situation, and the offline evaluation recall rate that has poor timeliness and cannot discover abnormal conditions of online services in the first time, in the embodiments of the present disclosure, by sampling access traffic and performing real neighbor calculations on the sampled traffic in a real-time, asynchronous manner, and then evaluating the recall rate of approximate neighbor calculations by comparing the results obtained from the two types of neighbor calculations, the above problems can be overcome, that is, it can reflect the actual online situation and improve timeliness.
[0053] As an optional embodiment, approximate neighbor calculation is performed by a first task; and true neighbor calculation is performed by a second task, wherein the first task and the second task are independent of each other.
[0054] In the disclosed embodiment, different tasks are used to implement approximate neighbor calculation and real neighbor calculation, so that the two neighbor calculations can be performed asynchronously and the real neighbor calculation does not affect the performance and response of the upward service, i.e., the online approximate neighbor calculation.
[0055] In the embodiment of the present disclosure, the online approximate nearest neighbor calculation in the first task may adopt the hnsw algorithm to simplify the calculation and improve timeliness.
[0056] In addition, during the online approximate neighbor calculation process, the online access traffic may be sampled, for example, by random sampling or any other sampling method, and the online access traffic may be sampled at a certain ratio (eg, 1%).
[0057] Furthermore, the sampled access traffic (i.e., sampled traffic) can be used to perform real neighbor calculations in an asynchronous manner in the second task. Thus, the original process of the online service is not affected. Moreover, after the online service completes the calculation for each sampled traffic, it can return the corresponding nearest neighbor result to the second task, so the online service response (i.e., the calculation time of the online service) is not affected.
[0058] In addition, in the embodiments of the present disclosure, asynchronous means that the calculation process of the real neighbor calculation is independent of the calculation process of the online high-performance approximate neighbor calculation. Among them, the real neighbor calculation in the second task can be performed asynchronously by multi-threading or multi-process. Alternatively, the real neighbor calculation can be performed asynchronously by multi-threading or independent process. The calculation process and calculation results of these methods do not directly affect the first task.
[0059] It should be noted that each thread in the multi-threaded mode can be an independent thread. Among them, the first task process (i.e., the online high-performance approximate neighbor calculation process) is used to transfer the sampled traffic and the calculation results of the first task to the asynchronous real neighbor calculation in an asynchronous manner (such as a queue). And the first task process can return directly without waiting for the result.
[0060] It should be noted that the multi-process and independent process mode refers to performing real neighbor calculations with one or more separate processes (which can run on different machines). Among them, the first task process (i.e., the online high-performance approximate neighbor calculation process) is used to pass the sampled traffic and the calculation results of the approximate neighbor calculation to the real neighbor calculation process. And the first task process can return directly without waiting for the result.
[0061] As an optional embodiment, determining the recall rate of the approximate nearest neighbor calculation includes at least one of the following.
[0062] At least one instantaneous recall of the approximate nearest neighbor computation is determined.
[0063] An average recall rate of the approximate neighbor calculations within a first preset time period is determined.
[0064] Determine the recall of approximate nearest neighbor calculations in a single scenario.
[0065] Determine the recall of approximate neighbor computation in multiple scenarios.
[0066] In one embodiment, a recall rate can be calculated for each sampled flow. The recall rate is the instantaneous recall rate. Exemplarily, the instantaneous recall rate can be calculated by the following formula 1.
[0067] Formula 1: Recall = M / N
[0068] Where Recall represents the recall rate; M represents the number of vectors in the nearest neighbors calculated by the accelerated nearest neighbors that belong to the nearest neighbors calculated by the real nearest neighbors; N represents the number of vectors in the nearest neighbors (such as Top N) calculated by the real nearest neighbors.
[0069] In another embodiment, for a plurality of sampled flows within a period of time, a recall rate may be calculated for each sampled flow, and then a weighted average of all the recall rates may be calculated to obtain a corresponding average recall rate.
[0070] In the embodiments of the present disclosure, by comparing the nearest neighbor result obtained by online high-performance approximate nearest neighbor calculation with the nearest neighbor result obtained by real nearest neighbor calculation, a variety of corresponding recall rates can be obtained. For example, the recall rate can be calculated for a single module (i.e., a single scene), and the global recall rate calculation (i.e., multiple scenes) can also be realized. For another example, the instantaneous recall rate can be obtained, and the average recall rate of different time spans can also be obtained, which is more flexible in application.
[0071] Through the embodiments of the present disclosure, the recall rate of approximate nearest neighbor calculation can be evaluated from multiple angles, and the application is flexible.
[0072] As an optional embodiment, the method further includes: recording the recall rate calculated by the approximate nearest neighbor to obtain record information about the recall rate.
[0073] As an optional embodiment, the method further includes: based on the record information, counting the recall rate that meets the preset conditions to obtain statistical data on the recall rate.
[0074] As an optional embodiment, the method further includes: in response to the statistical data of the recall rate exceeding a preset value within a second preset time period, sending corresponding prompt information.
[0075] In the embodiments of the present disclosure, each time a recall rate is calculated, a record information about the recall rate can be recorded, that is, a record log can be generated. In other words, the recall rate data can be recorded online using logs and other methods. For example, the online recall rate data within one or more time periods (such as 1 second, 1 minute) can be aggregated and calculated in combination with a single or multiple online services, and the overall recall rate can be counted and recorded.
[0076] Example 1: The average recall rate can be calculated based on all sampled traffic for all machines within 1 second, and the average recall rate can be obtained. In this example, the average can be directly accumulated or a threshold can be set. First, each recall rate is judged as [0, 1] (where 0 indicates failure and 1 indicates success), and then the average is calculated based on the number of successes and failures to obtain the average recall rate.
[0077] In Example 2, the recall rates calculated based on all sampled traffic for all machines within 10 seconds can be used to calculate the proportion of recall rates below 80%.
[0078] By recording the recall rate through the embodiment of the present disclosure, it is convenient for users to check and analyze, and it is convenient for the statistics of the recall rate. By performing statistics on the recall rate, especially performing statistical analysis on the global recall rate, it is possible to calculate the overall recall rate and obtain the recall rate from a global perspective.
[0079] Furthermore, related services can be triggered based on the record information and statistical information about the recall rate, such as triggering the alarm logic to send prompt information to users so that users can find abnormal conditions of online services in the first place. For example, if the recall rate of a single machine service is lower than 70% for 1 minute, the alarm logic can be triggered to send an alarm SMS.
[0080] In addition, in the embodiments of the present disclosure, the recall rate record information and statistical information may be graphically displayed during the recall rate monitoring process for user reference.
[0081] Through the embodiments of the present disclosure, the recall rate of upward approximate neighbor calculation can be monitored in real time, with high performance and low cost, and this method has no impact on online services (online approximate neighbor calculation is performed), and the recall rate calculation result is highly reliable.
[0082] According to an embodiment of the present disclosure, the present disclosure also provides a processing device for nearest neighbor computing.
[0083] Figure 5A block diagram of a processing device for neighbor calculation according to an embodiment of the present disclosure is exemplarily shown.
[0084] like Figure 5 As shown, the processing device 500 for nearest neighbor calculation may include: a sampling module 510 , a true nearest neighbor calculation module 520 and a determination module 530 .
[0085] The sampling module 510 is used to sample the access traffic during the approximate neighbor calculation process to obtain at least one sampled traffic.
[0086] The real neighbor calculation module 520 is used to perform asynchronous real neighbor calculation on the at least one sampled flow in real time to obtain at least one corresponding real neighbor calculation result.
[0087] The determination module 530 is used to determine the recall rate of the approximate neighbor calculation based on the at least one true neighbor calculation result and the corresponding at least one approximate neighbor calculation result, wherein the at least one approximate neighbor calculation result is obtained by performing the approximate neighbor calculation on the at least one sampled flow.
[0088] As an optional embodiment, the approximate nearest neighbor calculation is performed by a first task; and the true nearest neighbor calculation is performed by a second task, wherein the first task and the second task are independent of each other.
[0089] As an optional embodiment, the determination module includes at least one of the following: a first determination unit, used to determine at least one instantaneous recall rate of the approximate neighbor calculation; a second determination unit, used to determine the average recall rate of the approximate neighbor calculation within a first preset time period; a third determination unit, used to determine the recall rate of the approximate neighbor calculation in a single scenario; a fourth determination unit, used to determine the recall rate of the approximate neighbor calculation in multiple scenarios.
[0090] As an optional embodiment, the device further includes: a recording module, configured to record the recall rate calculated by the approximate nearest neighbor to obtain record information about the recall rate.
[0091] As an optional embodiment, the device further includes: a statistical module, which is used to count the recall rates that meet preset conditions based on the record information to obtain statistical data on the recall rates.
[0092] As an optional embodiment, the device further includes: a sending module, configured to send corresponding prompt information in response to the statistical data of the recall rate exceeding a preset value within a second preset time period.
[0093] As an optional embodiment, the true neighbor calculation includes: neighbor calculation based on Euclidean distance calculation or cosine distance calculation; and / or the approximate neighbor calculation includes: neighbor calculation based on clustering method.
[0094] It should be understood that the embodiments of the device part of the present disclosure are the same or similar to the embodiments of the method part of the present disclosure, and the technical problems solved and the technical effects achieved are also the same or similar, and the present disclosure will not elaborate on them here.
[0095] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0096] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0097] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0098] Multiple components in the electronic device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0099] The computing unit 601 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 601 performs the various methods and processes described above, such as a processing method for neighbor calculation. For example, in some embodiments, the processing method for neighbor calculation may be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the processing method for neighbor calculation described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to execute the processing method for neighbor computing in any other appropriate manner (eg, by means of firmware).
[0100] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0101] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0102] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0103] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0104] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0105] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs that run on the corresponding computers and have a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server for a distributed system, or a server combined with a blockchain.
[0106] In the technical solution disclosed in the present invention, the recording, storage and application of the traffic data involved are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0107] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0108] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A method for determining system service anomalies in an industrial application scenario, wherein the industrial application scenario includes at least one of intelligent recommendation, intelligent search, knowledge graph, and user understanding, including: Performing approximate neighbor calculation on the access traffic in the system to obtain an approximate neighbor calculation result, wherein the access traffic includes access traffic in at least one of the application systems of intelligent recommendation, intelligent search, knowledge graph and user understanding; In the process of performing approximate nearest neighbor calculation on the access traffic, sampling the access traffic to obtain at least one sampled traffic; Performing asynchronous true neighbor calculation on the at least one sampled flow in real time to obtain at least one corresponding true neighbor calculation result; Determining a recall rate of the approximate neighbor calculation based on a similarity between the at least one true neighbor calculation result and the corresponding at least one approximate neighbor calculation result, wherein the recall rate reflects a deviation of the neighbor calculation result; According to the recall rate, determine whether the system service is abnormal, including: Counting at least one of the proportion of the recall rate that meets the preset condition and the duration of the recall rate being lower than the threshold as statistical data; In response to the statistical data triggering the alarm logic, an alarm message for prompting the system service abnormality is sent.
2. The method according to claim 1, wherein: Performing the approximate nearest neighbor calculation by a first task; and The true nearest neighbor calculation is performed by a second task, wherein the first task and the second task are independent of each other.
3. The method according to claim 1, wherein: Determining the recall of the approximate nearest neighbor calculation includes at least one of the following: determining at least one instantaneous recall of the approximate nearest neighbor calculation; Determining an average recall rate of the approximate neighbor calculation within a first preset time period; Determining a recall rate of the approximate nearest neighbor calculation in a single scenario; Determine the recall rate of the approximate neighbor calculation in multiple scenarios.
4. The method according to claim 3, further comprising: The recall rate calculated by the approximate nearest neighbor is recorded to obtain record information about the recall rate.
5. The method according to any one of claims 1 to 4, wherein: The real nearest neighbor calculation includes: nearest neighbor calculation directly based on Euclidean distance calculation or cosine distance calculation; The approximate nearest neighbor calculation includes: nearest neighbor calculation based on clustering method.
6. A system service anomaly determination device for an industrial application scenario, wherein the industrial application scenario includes at least one of intelligent recommendation, intelligent search, knowledge graph, and user understanding, including: An approximate neighbor calculation module, used to perform approximate neighbor calculation on the access traffic in the system to obtain an approximate neighbor calculation result, wherein the access traffic includes the access traffic in at least one of the application systems of intelligent recommendation, intelligent search, knowledge graph and user understanding; A sampling module, used for sampling the access traffic in the process of performing approximate nearest neighbor calculation on the access traffic to obtain at least one sampled traffic; A real neighbor calculation module, used for performing asynchronous real neighbor calculation on the at least one sampled flow in real time to obtain at least one corresponding real neighbor calculation result; as well as A recall rate determination module, configured to determine a recall rate of the approximate neighbor calculation based on a similarity between the at least one true neighbor calculation result and the corresponding at least one approximate neighbor calculation result, wherein the recall rate reflects a deviation of the neighbor calculation result; An abnormality determination module, used to determine whether the system service is abnormal according to the recall rate; The abnormality determination module is also used to count the proportion of recall rates that meet preset conditions and at least one of the duration of the recall rate below a threshold as statistical data; in response to the alarm logic triggered by the statistical data, an alarm message is sent to indicate that the system service is abnormal.
7. The device according to claim 6, wherein: Performing the approximate nearest neighbor calculation by a first task; and The true nearest neighbor calculation is performed by a second task, wherein the first task and the second task are independent of each other.
8. The device according to claim 6, wherein: The determination module includes at least one of the following: a first determining unit, configured to determine at least one instantaneous recall rate of the approximate nearest neighbor calculation; A second determining unit, configured to determine an average recall rate of the approximate nearest neighbor calculation within a first preset time period; A third determining unit, used to determine the recall rate of the approximate nearest neighbor calculation in a single scenario; The fourth determining unit is used to determine the recall rate of the approximate nearest neighbor calculation in multiple scenarios.
9. The apparatus according to claim 8, further comprising: The recording module is used to record the recall rate calculated by the approximate nearest neighbor to obtain record information about the recall rate.
10. The device according to any one of claims 6 to 9, wherein: The real nearest neighbor calculation includes: nearest neighbor calculation directly based on Euclidean distance calculation or cosine distance calculation; The approximate nearest neighbor calculation includes: nearest neighbor calculation based on clustering method.
11. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-5.
13. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
High-dimensional data accurate neighbor quick searching method based on euclidean distance
CN103279551A
Approximate nearest neighbor searching method and system of high dimensional data
CN105550368A