AI Application Abnormal Detection Method, Electronic Device, Storage Medium and Program Product
By hijacking the library to record the API call sequence of AI applications and detecting the execution time, the problem of difficulty in detecting exceptions by AI applications is solved, and non-invasive, low-cost and accurate abnormal detection is achieved, and the availability of AI applications is improved.
Patent Information
- Application Number
- CN202510142116.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-08
AI Technical Summary
It is difficult for AI applications to detect abnormal situations during operation during actual deployment, resulting in low availability.
By hijacking the library, intercept the AI application's calls to the API in the dynamic link library, record the API call sequence, and detect whether the AI application is in an abnormal state based on the execution time of the API call sequence.
It realizes non-invasive, low-cost and accurate abnormal detection of AI application, and improves the availability and reliability of AI applications.
Smart Images

Figure CN119576749B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to an AI application anomaly detection method, an electronic device, a storage medium, and a program product. Background Art
[0002] An AI (Artificial Intelligence) application refers to an application program that uses artificial intelligence technology to provide services. AI applications are generally divided into two major scenarios: training and inference. In the training scenario, a large amount of data is used to perform multiple rounds of iterative training on the AI application; in the inference scenario, the AI application is deployed to an electronic device and then provides services through an external interface.
[0003] When an AI application is actually deployed, it involves many components, and it is difficult to require users to add specified anomaly detection logic code to the AI application. Therefore, it is very difficult to detect anomalies during the operation of the AI application, resulting in low availability of the AI application. To improve the availability of the AI application, there is an urgent need to provide an AI application anomaly detection method so that effective handling measures can be taken in a timely manner when an AI application anomaly is detected. Summary of the Invention
[0004] Embodiments of this application provide an AI application anomaly detection method, an electronic device, a storage medium, and a program product. This method can detect AI application anomalies based on the execution duration of the API (Application Programming Interface) call sequence when the AI application provides a target service. The technical solution is as follows:
[0005] In a first aspect, an AI application anomaly detection method is provided. The method includes:
[0006] Obtain the target execution duration of the API call sequence. The API call sequence refers to a periodic sequence formed by the APIs in the hijacking library in the order of calls when the AI application provides a target service. The hijacking library is used to hijack the API call instructions of the AI application to the APIs in the dynamic link library. The hijacking library has the same APIs as the dynamic link library. The target execution duration is the execution duration of the API call sequence when the AI application provides any target service.
[0007] Obtain the reference execution duration of the API call sequence. The reference execution duration is the duration used to measure whether the AI application is in an abnormal state.
[0008] When the target execution duration exceeds the reference execution duration, determine that the AI application is in an abnormal state.
[0009] In a second aspect, an AI application anomaly detection device is provided. The device includes:
[0010] A first acquisition module, configured to acquire a target execution duration of an application programming interface (API) call sequence. The API call sequence refers to a periodic sequence formed by APIs in a hijacking library in the order of calls when the AI application provides a target service. The hijacking library is used to hijack the API call instructions of the AI application to dynamic link library APIs. The hijacking library has the same APIs as the dynamic link library. The target execution duration is the execution duration of the API call sequence when the AI application provides any target service.
[0011] A second acquisition module, configured to acquire a reference execution duration of the API call sequence. The reference execution duration is a duration used to measure whether the AI application is in an abnormal state.
[0012] A determination module, configured to determine that the AI application is in an abnormal state when the target execution duration exceeds the reference execution duration.
[0013] In a third aspect, an electronic device is provided, including a processor and a memory. The memory stores at least one program code. The at least one program code is configured to be called and executed by the processor to implement the artificial intelligence (AI) application anomaly detection method described in the first aspect.
[0014] In a fourth aspect, a computer-readable storage medium is provided. At least one computer program is stored in the computer-readable storage medium. When the at least one computer program is executed by a processor, the artificial intelligence (AI) application anomaly detection method described in the first aspect can be implemented.
[0015] In a fifth aspect, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the artificial intelligence (AI) application anomaly detection method described in the first aspect can be implemented.
[0016] The beneficial effects brought by the technical solutions provided in the embodiments of this application are:
[0017] In the embodiments of the present application, each API in the dynamic link library has definite semantic information. When an exception occurs during the process of an AI application calling an API, the called API can help users analyze and locate the exception of the AI application. Different APIs are called in the dynamic link library when different AI applications provide different services, while the same APIs are called in the dynamic link library when the same AI application provides the same service, and the API call sequence corresponding to the same AI application is fixed and has an obvious periodicity. To facilitate recording the calls of different AI applications to the APIs in the dynamic link library, in the embodiments of the present application, a hijacking library is deployed on the basis of the dynamic link library. The hijacking library has the same APIs as the dynamic link library. When an AI application needs to call an API in the dynamic link library to provide a target service, the hijacking library can hijack the API call instruction of the AI application to the API in the dynamic link library, so as to intercept the call to the API in the dynamic link library as a call to the same API in the hijacking library, and then determine the API call sequence based on the call order of the APIs in the hijacking library. The API call sequence is a periodic sequence formed by the APIs in the hijacking library in the call order when the AI application provides the target service. When it is necessary to perform anomaly detection on an AI application, obtain the target execution duration of the API call sequence when the AI application provides any target service once, and obtain the reference execution duration of the API call sequence. The reference execution duration is the duration used to measure whether the AI application is in an abnormal state. Then compare the target execution duration with the reference execution duration. If the target execution duration exceeds the reference execution duration, it can be determined that the AI application is in an abnormal state. The embodiments of the present application realize the anomaly detection of the AI application without intrusion based on the inherent API call sequence of the AI application providing the target service, and the detection result is more accurate. Description of the Drawings
[0018] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0019] Figure 1 It is a schematic diagram of the principle and overall structure of an AI application anomaly detection method provided by the embodiments of the present application;
[0020] Figure 2 It is a flowchart of an AI application anomaly detection method provided by the embodiments of the present application;
[0021] Figure 3 It is a flowchart of a method for determining an API call sequence provided by the embodiments of the present application;
[0022] Figure 4 It is a schematic diagram showing the transition between different state machines of an AI application provided by an embodiment of the present application;
[0023] Figure 5 It is a flowchart of a method for determining the target execution duration of an API call sequence provided by an embodiment of the present application;
[0024] Figure 6 It is a schematic structural diagram of an AI application anomaly detection device provided by an embodiment of the present application;
[0025] Figure 7 It shows a structural block diagram of an electronic device provided by an exemplary embodiment of the present application. Detailed implementation manners
[0026] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0027] It can be understood that the terms "each", "multiple", "any one", etc. used in the embodiments of the present application, multiple includes two or more, each refers to each one in the corresponding multiple, and any one refers to any one in the corresponding multiple. For example, multiple words include 10 words, and each word refers to each of these 10 words, and any one word refers to any one of the 10 words.
[0028] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0029] Before executing the embodiments of the present application, first, the nouns involved in the embodiments of the present application are explained.
[0030] API is a set of predefined functions that gives applications and developers the ability to directly access software and hardware resources without understanding the details of the internal working mechanism.
[0031] An online learning algorithm is an algorithm that can adaptively process data arriving in chronological order and adjust its own behavior decisions.
[0032] A watchdog is a program used to monitor the running status of other software or systems.
[0033] A dynamic link library is an implementation method for implementing shared function libraries in an operating system. A dynamic link library file is an unexecutable binary program file that allows programs to share the code and other resources necessary to perform special tasks. Dynamic linking provides a way for a process to call functions that are not part of its executable code. Using dynamic link libraries makes it easier to apply updates to individual modules without affecting other parts of the program.
[0034] An RPC (Remote Procedure Call Protocol) server refers to a protocol that requests services from a remote computer program over a network without the need to understand the underlying network technology. RPC adopts a client / server model. The requesting program is a client, and the service provider is a server. First, the calling process sends a call message with process parameters to the service process and then waits for a response message. On the server side, the process remains in a sleeping state until the call message arrives. When a call message arrives, the server obtains the process parameters, calculates the result, sends a reply message, and then waits for the next call message. Finally, the client calling process receives the reply message, obtains the process result, and then the calling execution continues.
[0035] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to obtain results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also involves researching the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0036] Artificial intelligence technology is an interdisciplinary subject that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0037] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields. For example, common applications include smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless cars, autonomous driving, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0038] AI applications, namely artificial intelligence applications, refer to application programs developed using artificial intelligence technology. The scope of AI applications is very wide, covering various fields. The following are some typical AI application scenarios:
[0039] Natural language processing
[0040] Chatbots: capable of natural language interaction with users, answering questions, providing services, etc.
[0041] Voice assistants: complete various tasks through voice commands.
[0042] Computer vision
[0043] Image recognition: recognize objects, faces, text, etc. in images
[0044] Video surveillance: use AI technology for intelligent surveillance to improve security.
[0045] Machine learning
[0046] Data analysis: process massive amounts of data, identify consumption patterns and trends, and provide support for decision-making.
[0047] AI applications are generally divided into two major scenarios: training and inference. In the training scenario, a large amount of data is used to perform multiple rounds of iterative training on the AI application to improve its usability; in the inference scenario, the AI application is deployed to a Web server / RPC server and then provides services through external interfaces. GPU (Graphics Processing Unit) is applied to the training and inference processes of AI applications due to its powerful parallel processing ability and high throughput.
[0048] However, the life cycle of AI applications is relatively long, and for heterogeneous computing units represented by GPUs, the probability of software and hardware errors is higher than that of CPU computing units. However, when an AI application is actually deployed, it involves many components, and it is very difficult to locate whether it is an error of the heterogeneous computing unit when the AI application is abnormal. Moreover, it is difficult for cloud service providers to require users to add abnormal detection logic code specified by the cloud service provider to the AI application. Therefore, it is very difficult to detect abnormal situations during the operation of AI applications, especially the abnormalities of heterogeneous computing units, resulting in low usability of AI applications.
[0049] Considering that when developing AI applications, developers need to control heterogeneous computing units through APIs provided by device manufacturers, and these APIs have clear semantic information, which can assist in analyzing and locating problems in case of anomalies. At the same time, in the training and inference scenarios of AI applications, the calls to APIs in dynamic link libraries have obvious periodic characteristics, and the API call sequences called when the same AI application provides the same service are fixed. In view of this, the embodiments of the present application provide an AI application anomaly detection method based on API call sequences. This method hijacks the library to intercept calls to APIs in dynamic link libraries, records the called APIs, and uses an adaptive online learning algorithm based on the called APIs to find the periodic rules of API calls, and then determines the API call sequence, so as to detect whether an AI application has an anomaly based on the API call sequence.
[0050] The embodiments of the present application perform anomaly detection on AI applications based on the API call sequences of AI applications and in combination with an online learning algorithm, which can help users identify anomaly events during the operation of AI applications without intrusion and at low cost, and have the advantages of non-intrusiveness, adaptability, low overhead, generality, dynamic switch, etc.
[0051] Non-intrusiveness means that it is possible to achieve non-intrusive fault detection without the user modifying the logic code of the AI application.
[0052] Adaptability means that the online learning algorithm can adapt to a variety of AI applications, can perform fault detection on a variety of AI applications, and realizes plug-and-play.
[0053] Low overhead means that it can reduce the detection and analysis overhead of faults and avoid affecting the performance of AI applications.
[0054] Generality means that the detection method has generality and does not need to be locked with hardware manufacturers.
[0055] Dynamic switch means that when anomaly detection needs to be performed on an AI application, the service corresponding to the AI application is provided and anomaly detection is performed on the AI application; when anomaly detection does not need to be performed on the AI application, the service corresponding to the AI application is provided without performing anomaly detection on the AI application.
[0056] The embodiments of the present application provide an AI application anomaly detection method. The AI application can be deployed in an electronic device, which can be a terminal, such as a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the electronic device can also be a server, such as an independent physical server, a server cluster or a distributed system composed of multiple physical servers, and a cloud server providing basic cloud computing services such as cloud services, cloud computing, cloud functions, cloud storage, network services, cloud communications, and artificial intelligence platforms. See Figure 1, to implement anomaly detection for AI applications, the electronic device may include a dynamic link library, a hijacking library, and a dynamic switch, etc.
[0057] Among them, the dynamic link library is a shared function library shared by each application in the electronic device. The dynamic link library includes multiple functions, and the multiple functions correspond to multiple APIs. By constructing the dynamic link library, the memory of the electronic device can be saved and the data reuse rate can be improved.
[0058] The hijacking library has the same API interface as the dynamic link library. During dynamic linking, it can hijack the API call instructions of the AI application to the APIs in the dynamic link library, that is, intercept the call to the APIs in the dynamic link library as a call to the same APIs in the hijacking library, and then replace the APIs in the call instructions. Specifically, when any AI application is started in the training scenario or the inference scenario, the electronic device loads the dynamic link library and the hijacking library into the AI application process. Then, during dynamic linking, it hijacks the API call instructions of the AI application to the APIs in the dynamic link library in the training scenario or the inference scenario, replaces the APIs in the call instructions with the storage addresses of these APIs in the AI application process in the dynamic link library, and then sends the replaced call instructions to the dynamic link library to implement the call of the AI application to the APIs in the dynamic link library in the training scenario or the inference scenario.
[0059] The dynamic switch is an interface for the anomaly detection function provided by the hijacking library to the outside. The dynamic switch can be a file system mount point, and it can turn on and off the anomaly detection function of the hijacking library in real time according to requirements, reduce the runtime overhead, and avoid affecting the normal operation of the AI application. When the dynamic switch is in the on state, the anomaly detection operation for the AI application can be performed; when the dynamic switch is in the off state, the anomaly detection operation for the AI application is not performed.
[0060] Taking any AI application as an example, the abnormal detection process of the above-mentioned AI application specifically includes: after any AI application is started in the training scenario or the inference scenario, the electronic device loads the dynamic link library and the hijacking library into the AI application process. If the dynamic switch is in the on state, the electronic device executes the AI application process, hijacks the API call instruction of the dynamic link library by the hijacking library in the training scenario or the inference scenario of the AI application, replaces the API in the call instruction with the storage address of the API in the dynamic link library in the AI application process, obtains the replaced call instruction, sends the replaced call instruction to the dynamic link library, and then records the call of the API in the hijacking library by the AI application in the training scenario or the inference scenario. Then, an online learning algorithm is used to analyze the recorded API to obtain a periodic API call sequence of the AI application in the training scenario or the inference scenario. Then, based on the periodic API call sequence, the abnormal situation in the AI application is detected. If the dynamic switch is in the off state, the electronic device executes the AI application process, hijacks the API call instruction of the dynamic link library by the hijacking library in the training scenario or the inference scenario of the AI application, replaces the API in the call instruction with the storage address of the API in the dynamic link library in the AI application process, obtains the replaced call instruction, and sends the replaced call instruction to the dynamic link library to implement the call of the API in the dynamic link library.
[0061] Further, when detecting the abnormal situation in the AI application based on the periodic API call sequence, a watchdog mechanism can be used to set an application abnormal timeout warning mechanism based on the API call sequence. For example, the execution duration of the API call sequence is set. If it is detected that the API call sequence does not complete within the set duration when the AI application provides services, it can be determined that the AI application is abnormal. If it is detected that the API call sequence completes within the set duration when the AI application provides services, the maximum execution duration of the API call sequence when the AI application is in the normal state can also be obtained. If it is detected that the execution duration of the API call sequence exceeds the maximum execution duration in the normal state when the AI application provides services, it can be determined that the AI application is abnormal.
[0062] Further, after detecting that the AI application is abnormal, a prompt message of the abnormal AI application can be reported through an external control interface, so that the user can take measures in time to handle the abnormality.
[0063] The embodiment of the present application provides a method for detecting abnormal AI applications. Taking the electronic device executing the embodiment of the present application as an example, see Figure 2 , the method flow provided by the embodiment of the present application includes:
[0064] 201. Obtain the target execution duration of the application programming interface API call sequence.
[0065] In the embodiments of the present application, a dynamic link library and a hijacking library are shared by various applications in the electronic device. For any AI application, the AI application is generally divided into a training scenario and an inference scenario. Regardless of the scenario in which the AI application is located, when the AI application is started, the electronic device will load the dynamic link library and the hijacking library into the AI application process corresponding to the AI application. The dynamic link library includes multiple APIs, and each API in the dynamic link library corresponds to a storage address in the AI application process. When the AI application runs in the training scenario or the inference scenario, in response to any acquisition request for a target service of the AI application, the electronic device can provide the target service by executing the AI application process. When providing the target service, the electronic device can also perform anomaly detection on the AI application. To implement anomaly detection of the AI application, the electronic device needs to obtain the target execution duration of the API call sequence.
[0066] Among them, the API call sequence refers to a periodic sequence formed by the APIs in the hijacking library in the order of calls when the AI application provides the target service. The hijacking library is used to hijack the API call instructions of the AI application to the APIs in the dynamic link library, and the hijacking library has the same APIs as the dynamic link library. Generally speaking, the dynamic link library can be shared by various applications in the electronic device and cannot be used to record the calls of a certain application to the APIs in the dynamic link library, and the APIs called by the AI application in different states are different. Therefore, in the embodiments of the present application, during the process of the AI application providing the target service, the APIs called by the AI application can be recorded based on the hijacking library, and the state machine of the AI application can be obtained. Furthermore, based on the APIs recorded by the hijacking library and the state machine of the AI application, the API call sequence can be determined. Among them, the state machine of the AI application includes an initialization state, a running state, an end state, etc. The initialization state refers to the state of the AI application before the API interfaces related to the target service in the hijacking library are first called after the AI application is started. The running state refers to the state of the AI application after a periodic sequence is recognized based on the called APIs. The end state refers to the state of the AI application when it is closed. The embodiments of the present application use the hijacking library to record the APIs called by the AI application and combine the states when the AI application calls these APIs, thereby providing a method for determining the API call sequence.
[0067] In the embodiments of the present application, based on the APIs recorded by the hijacking library and the state machine of the AI application, an online learning algorithm can be used to determine the API call sequence. Figure 3 The flowchart of the method for determining the API call sequence is shown, see Figure 3 and the method includes the following steps:
[0068] 301. When the AI application is started, determine that the AI application is in the initialization state.
[0069] In the embodiment of the present application, when the AI application is started, it can be defaulted that the AI application is in the initialization state.
[0070] 302. When the AI application is in the initialization state, record the APIs called in the hijacking library when the AI application provides the target service.
[0071] When the AI application is in the initialization state, each called API will be recorded. When the AI application is in the initialization state, the called APIs all belong to the APIs in the initialization stage and will not be used to determine the periodic API call sequence.
[0072] 303. When it is detected that the duration for which the API in the hijacking library has not been called by the AI application reaches the first preset duration, it is determined that the AI application is in the running state.
[0073] Among them, the first preset duration can be set by the technical personnel. When it is detected that the duration for which the API in the hijacking library has not been called by the AI application reaches the first preset duration, it can be considered that the initialization state is completed. In the training scenario, it can be considered that the input of the previous training batch is completed and the input of the next training batch has not been input, and the input gap between the two training batches is greater than the first preset duration. The APIs in the hijacking library will not be called during the input gap between the two training batches. At this time, it is considered that the initialization state is completed; in the inference scenario, it can be considered that the current inference request is received and the next inference request has not been received, and the reception gap between the two inference requests is greater than the first preset duration. The APIs in the hijacking library will not be called during the reception gap between the two inference requests. At this time, it is considered that the initialization state is completed.
[0074] In another embodiment of the present application, when the AI application is in the running state, if any API is called for the first time, it can be considered that the initialization state is not completed. At this time, the AI application will be reverted from the running state to the initialization state.
[0075] 304. When the AI application is in the running state, determine the API call sequence based on the APIs called in the hijacking library when the AI application provides the target service.
[0076] After the AI application is started in the embodiment of the present application, by determining the running state of the AI application, and then when the AI application is in the running state, determining the API call sequence based on the APIs called in the hijacking library when the AI application provides the target service, thereby providing a method for determining the API call sequence.
[0077] Specifically, when the AI application is in the running state, the large language model can be called to identify the APIs called in the hijacking library when the AI application provides the target service, so as to determine the API call sequence. It is also possible to determine the API call sequence based on the APIs called in the hijacking library when the AI application provides the target service and the duration during which the APIs in the hijacking library are not called. Of course, other methods can also be used to determine the API call sequence, which will not be elaborated one by one in the embodiments of this application. Taking the determination of the API call sequence based on the APIs called in the hijacking library when the AI application provides the target service and the duration during which the APIs in the hijacking library are not called as an example, if the duration during which the APIs in the hijacking library are not called reaches the second preset duration, and the APIs called in the hijacking library when the AI application provides the target service are periodic, then the periodic API call sequence is used as the API call sequence. Among them, the second preset duration can be set by the technical personnel, and this API call sequence can be marked as T. Generally speaking, the AI application in the embodiments of this application includes a training scenario and an inference scenario. In the training scenario, before any training batch input, a start mark can be added to the APIs to be hijacked by the hijacking library, or after any training batch input is completed, an end mark can be added to the APIs hijacked by the hijacking library. When the duration during which the APIs in the hijacking library are not called reaches the second preset duration, based on the start mark or end mark recorded in the hijacking library, the periodic API call sequence is used as the API call sequence to add a cycle start sequence; in the inference scenario, after the execution of the previous inference request is completed and before the next inference request is received, a start mark can be added to the APIs to be hijacked by the hijacking library, or after the execution of the next inference request is completed, an end mark can be added to the APIs to be hijacked by the hijacking library, and then the periodic API call sequence is used as the API call sequence to add a cycle start sequence. Assuming that the second preset duration is 10 seconds, when the AI application is in the running state, if it is detected that the duration during which the APIs in the hijacking library are not called reaches 10 seconds, and the obtained API sequence A1A2A3 called by the AI application is periodic, then A1A2A3 is used as the API call sequence. Further, when the AI application is in the running state, after the API call sequence is determined, the API call sequence each time the AI application provides the target service can be recorded based on the cycle start mark or cycle end mark of the learned API call sequence.
[0078] In another embodiment of the present application, taking the inference scenario as an example, if the electronic device receives a continuous plurality of inference requests, in response to the continuous plurality of inference requests, the hijacking library will record the API sequences called in multiple inferences and misidentify the API sequences called in multiple inferences as a cycle, resulting in inaccurate recorded API call sequences. To improve the accuracy of the API call sequence, when the AI application is in the running state, based on the APIs called in the hijacking library when the AI application provides the target service, after taking the periodic API call sequence as the API call sequence, the hijacking library continues to hijack the API call instructions for the APIs in the dynamic link library when the AI application provides the target service. If a new API call sequence is obtained and the first sequence length of the new API call sequence is less than the second sequence length of the API call sequence, the API call sequence is updated to the new API call sequence. Wherein, the new API call sequence can be marked as S, and the , wherein, , , …, are the APIs included in the new API call sequence. The sequence length of the API call sequence can be determined according to the number of APIs included in the call sequence. If an API call sequence is A1A2A3A1A2A3, it can be determined that the sequence length of the API call sequence is 6; if another API call sequence is A1A2A3, it can be determined that the sequence length of the API call sequence is 3. Assume that the new API call sequence S is A1A2A3, the already determined API call sequence T is A1A2A3A1A2A3, the first sequence length of the new API call sequence S is 3, the second sequence length of the already determined API call sequence T is 6, and the first sequence length is less than the second sequence length, then the API call sequence T can be updated to the new API call sequence S, that is, the new API call sequence A1A2A3 is determined as the API call sequence corresponding to the AI application.
[0079] In another embodiment of the present application, when the AI application is in the initialization state or the running state, when a specified API in the hijacking library is called, it can be considered that the AI application process terminates, and then the AI application is switched to the end state, thereby stopping the abnormal detection of the AI application. Wherein, the specified API refers to the API used to close the AI application. Further, after switching the AI application to the end state, a prompt message will also be sent to the upper-level management and control component to prompt that the AI application has stopped running.
[0080] Figure 4 shows the process of online learning / detection based on the state machine of the AI application. Through the online learning / detection process, the API call sequence of the AI application can be identified. See Figure 4, when the AI application starts, the AI application is in an initialization state. At this time, all the APIs called by the AI application to provide the target service belong to the APIs in the initialization stage. If the APIs in the hijacking library have not been called for a period of time, it can be considered that the initialization stage of the AI application is completed, and then the state machine of the AI application is converted to the running state. When the AI application is in the running state, if it is detected that a certain API in the hijacking library is called for the first time, it means that the initialization stage has not been completed. At this time, the running state is rolled back to the initialization state. In the running state, if the APIs in the hijacking library have not been called for a period of time, and the APIs called in the hijacking library when the AI application provides the target service are periodic, then the periodic API call sequence T is used as the API call sequence. If a new API call sequence is obtained in a certain cycle of the running state , if this new API call sequence has fewer APIs than the number of APIs in the API call sequence T, then the API call sequence T is updated to the new API call sequence . If a specified API in the hijacking library is called, it is considered that the AI application process ends. Then, the AI application is switched to the end state, and a prompt message is sent to the upper-level management and control component to prompt the upper-level management and control component that the AI application has stopped running.
[0081] Among them, the target execution duration is the execution duration of the API call sequence when the AI application provides any target service. Figure 5 Shows the method for determining the target execution duration of the API call sequence. See Figure 5 , this method includes the following steps:
[0082] 501. When the AI application provides any target service, generate at least one API call instruction, and each of the at least one API call instruction includes an API in the API call sequence.
[0083] Based on the determined API call sequence, when the AI application provides any target service, an API call instruction can be generated for each API in the API call sequence, obtaining at least one API call instruction, and each of the at least one API call instruction includes an API in the API call sequence.
[0084] 502. Based on the hijacking library, hijack at least one API call instruction, and replace the APIs in the at least one API call instruction hijacked by the hijacking library with the storage addresses of the corresponding APIs in the dynamic link library in the AI application process, obtaining at least one replaced API call instruction.
[0085] After generating at least one API call instruction, the hijacking library hijacks at least one API call instruction according to the execution order of the at least one API call instruction, and replaces the API in each hijacked API call instruction of the hijacking library with the storage address of the corresponding API in the dynamic link library in the AI application process, obtaining at least one replaced API call instruction.
[0086] 503. By executing the AI application process, call the resources stored at the storage addresses in at least one replaced API call instruction.
[0087] 504. Obtain the sum of the execution durations of at least one replaced API call instruction as the target execution duration.
[0088] The electronic device records the start execution time and the end execution time of at least one replaced API call instruction, then subtracts the start execution time from the end execution time of each replaced API call sequence to obtain the execution duration of each replaced API call sequence, and then adds up the execution durations of at least one replaced API call instruction to obtain the total execution duration, and further determines the total execution duration as the target execution duration.
[0089] In an embodiment of the present application, when providing any target service for the AI application, at least one API call instruction is generated based on each API included in the API call sequence, and the at least one API call instruction is hijacked and replaced by the hijacking library to record the API and its execution time called during the process of providing the target service for the AI application, obtaining the execution duration of each replaced API call instruction. By adding up the execution durations of each replaced API call instruction, the target execution duration when the AI application provides the any target service is obtained, thus providing a method for determining the target execution duration.
[0090] In another embodiment of the present application, considering that abnormal detection of the AI application will cause certain resource consumption, and the AI application does not need to perform abnormal detection in real time. It is only necessary to turn on the dynamic switch when the running speed of the AI application becomes slow and the user requests detection. When the AI application runs normally and the user does not request detection, detection can be avoided to reduce resource consumption. For this reason, before providing the target service in response to the acquisition request for any target service of the AI application, the electronic device can obtain the parameter value of the target parameter in the hijacking library. When the parameter value of the target parameter is a preset value, the target execution duration of the API call sequence is obtained to perform abnormal detection on the AI application. The target parameter is used to indicate whether to perform abnormal detection on the AI application. The target parameter can be the status parameter of the dynamic switch. The preset value is used to indicate performing abnormal detection on the AI application, and the preset value corresponds to the on state of the dynamic switch.
[0091] Further, when the parameter value of the target parameter corresponds to the off state of the dynamic switch, the electronic device only needs to execute the AI application process and call the resources stored at the storage addresses in at least one replaced API call instruction, without performing anomaly detection on the AI application.
[0092] 202. Obtain the reference execution duration of the API call sequence.
[0093] The reference execution duration is the duration used to measure whether the AI application is in an abnormal state. The reference execution duration can be the maximum execution duration of the API call sequence when the AI application provides a target service once in the normal state. The reference execution duration can be set by a technician based on experience, or multiple execution durations of the API call sequence when the AI application provides the target service multiple times in the normal state can be obtained, and the average value of the multiple execution durations can be calculated, and the obtained average execution duration can be used as the reference execution duration, or on the basis of the obtained average execution duration, a preset interference duration can be added to the average execution duration to obtain the reference execution duration.
[0094] 203. When the target execution duration exceeds the reference execution duration, determine that the AI application is in an abnormal state.
[0095] When the target execution duration exceeds the reference execution duration, it can be determined that the AI application is in an abnormal state. Further, when it is determined that the AI application is in an abnormal state, a prompt message can be sent to the upper-layer management and control component to prompt the abnormality of the AI application.
[0096] In another embodiment of the present application, a third preset duration for the execution of the API call sequence can also be set based on the watchdog mechanism. The third execution duration is the maximum execution duration for the completion of the API call sequence and can be set by a technician. Optionally, in the process of the AI application providing any target service in the embodiment of the present application, the execution duration of the API call sequence is recorded. If the API call sequence is executed within the third preset duration, the reference execution duration of the API call sequence can be obtained, and further judgment can be made based on the reference duration. If the API call sequence is not executed within the third preset duration when the AI application provides any target service, it can be determined that the AI application is in an abnormal state. Further, when it is determined that the AI application is in an abnormal state, a prompt message can be sent to the upper-layer management and control component to prompt the abnormality of the AI application.
[0097] In the embodiments of the present application, by hijacking the underlying dynamic link library and directly interacting with heterogeneous computing units, it does not depend on the upper-layer application framework and can be imperceptible to users. In addition, in the embodiments of the present application, by providing an interface in the file system and detecting the anomaly detection function according to requirements, the detection method is more flexible and the overhead of the electronic device is saved. In addition, since only the API call sequence is analyzed, the detection overhead is low. In addition, in the embodiments of the present application, there is no need for the AI application to be locked with the hardware manufacturer. Only the symbols in the hijacking library need to be adjusted, and no other parts need to be modified, so the compatibility is stronger.
[0098] Any combination of the above optional technical solutions can form an optional embodiment of the present application, which will not be elaborated here one by one.
[0099] Please refer to Figure 6 , which shows a schematic structural diagram of an AI application anomaly detection device provided by the embodiments of the present application. The device can be implemented by software, hardware, or a combination of both, and becomes all or part of an electronic device. The device includes:
[0100] A first acquisition module 601, configured to acquire a target execution duration of an application programming interface (API) call sequence, where the API call sequence refers to a periodic sequence formed by APIs in a hijacking library in an AI application according to the call order when providing a target service. The hijacking library is used to hijack the call instructions of the APIs in the dynamic link library by the AI application, and the hijacking library and the dynamic link library have the same APIs. The target execution duration is the execution duration of the API call sequence when the AI application provides any target service;
[0101] A second acquisition module 602, configured to acquire a reference execution duration of the API call sequence, where the reference execution duration is a duration used to measure whether the AI application is in an abnormal state;
[0102] A first determination module 603, configured to determine that the AI application is in an abnormal state when the target execution duration exceeds the reference execution duration.
[0103] In another embodiment of the present application, a first acquisition module 601 is configured to generate at least one API call instruction when the AI application provides any target service, and each of the at least one API call instructions includes one API in the API call sequence; hijack the at least one API call instruction based on the hijacking library, and replace the API in the at least one API call instruction hijacked by the hijacking library with the storage address of the corresponding API in the dynamic link library in the AI application process to obtain at least one replaced API call instruction; by executing the AI application process, call the API stored at the storage address in the at least one replaced API call instruction; and obtain the sum of the execution durations corresponding to the at least one replaced API call instruction as the target execution duration.
[0104] In another embodiment of the present application, the device further includes:
[0105] A loading module, configured to load the dynamic link library and the hijacking library into the AI application process when the AI application starts, and each API in the dynamic link library corresponds to a storage address in the AI application process.
[0106] In another embodiment of the present application, the device further includes:
[0107] A second determination module, configured to determine the API call sequence based on the APIs recorded by the hijacking library and the state machine of the AI application during the process of the AI application providing the target service.
[0108] In another embodiment of the present application, the second determination module is configured to determine that the AI application is in an initialization state when the AI application starts; record the APIs called in the hijacking library when the AI application provides the target service in the initialization state of the AI application; when it is detected that the duration for which the APIs in the hijacking library are not called by the AI application reaches a first preset duration, determine that the AI application is in a running state; and determine the API call sequence based on the APIs called in the hijacking library when the AI application provides the target service in the running state of the AI application.
[0109] In another embodiment of the present application, the second determination module is configured to, when the AI application is in a running state, if the duration for which the APIs in the hijacking library are not called reaches a second preset duration, and the APIs called in the hijacking library when the AI application provides the target service are periodic, determine the periodic API call sequence as the API call sequence.
[0110] In another embodiment of the present application, the device further includes:
[0111] A fallback module, configured to, when any API is called for the first time while the AI application is in a running state, fallback the AI application from the running state to the initialization state.
[0112] In another embodiment of the present application, the device further includes:
[0113] An update module, configured to, when a new API call sequence is obtained while the AI application is in a running state, and the first sequence length of the new API call sequence is less than the second sequence length of the API call sequence, update the API call sequence to the new API call sequence.
[0114] In another embodiment of the present application, the device further includes:
[0115] A switching module, configured to, when a specified API in the hijacking library is called while the AI application is in the initialization state or the running state, switch the AI application to the end state, where the specified API refers to an API for closing the AI application.
[0116] In another embodiment of the present application, the device further includes:
[0117] A third acquisition module, configured to acquire a parameter value of a target parameter in the hijacking library, where the target parameter is used to indicate whether to perform anomaly detection on the AI application;
[0118] The first acquisition module 601 is further configured to, when the parameter value of the target parameter is a preset value, perform an operation of acquiring a target execution duration of an application programming interface (API) call sequence.
[0119] In another embodiment of the present application, the device further includes:
[0120] The first acquisition module 601 is further configured to, during the process of the AI application providing any target service, if the API call sequence is completed within a third preset duration, perform an operation of acquiring a reference execution duration of the API call sequence, where the third execution duration is greater than the reference execution duration;
[0121] A third determination module, configured to, if the API call sequence is not completed within the third preset duration when the AI application provides any target service, determine that the AI application is in an abnormal state.
[0122] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, device, and unit can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0123] Figure 7 The block diagram of an electronic device 700 provided by an exemplary embodiment of the present application is shown. Generally, the electronic device 700 includes a processor 701 and a memory 702.
[0124] The processor 701 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 701 can also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state; the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 701 can be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 701 can also include an artificial intelligence processor, which is used to process computational operations related to machine learning.
[0125] The memory 702 can include one or more computer-readable storage media, and the computer-readable storage media can be non-temporary computer-readable storage media. For example, the non-temporary computer-readable storage media can be CD-ROM (Compact Disc Read-Only Memory), ROM, RAM (Random Access Memory), magnetic tape, floppy disk, and optical data storage devices, etc. At least one computer program is stored in the computer-readable storage media, and when the at least one computer program is executed, it can implement the above-mentioned artificial intelligence AI application anomaly detection method.
[0126] Of course, the above-mentioned electronic device may also include other components, such as an input / output interface, a communication component, etc. The input / output interface provides an interface between the processor and the peripheral interface module, and the above-mentioned peripheral interface module can be an output device, an input device, etc. The communication component is configured to facilitate wired or wireless communication between the electronic device and other devices, etc.
[0127] Those skilled in the art can understand that Figure 7 the structure shown in
[0128] An embodiment of the present application provides a computer-readable storage medium, in which at least one computer program is stored, and when the at least one computer program is executed by a processor, the above-mentioned artificial intelligence AI application anomaly detection method can be implemented.
[0129] An embodiment of the present application provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, the above-mentioned artificial intelligence AI application anomaly detection method can be implemented.
[0130] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0131] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for detecting anomalies in artificial intelligence (AI) applications, characterized in that: The method comprises: Obtaining a target execution duration of an application programming interface (API) call sequence, wherein the API call sequence refers to a periodic sequence of APIs in a hijacking library in accordance with a calling order when an AI application provides a target service, wherein the hijacking library is used to hijack the calling instructions of the AI application to the API in a dynamic link library, wherein the hijacking library and the dynamic link library have the same API, and wherein the target execution duration is the execution duration of the API call sequence when the AI application provides any target service, and wherein the process of determining the API call sequence comprises: when the AI application is started, determining that the AI application is in an initialization state; when the AI application is in the initialization state, recording the APIs called in the hijacking library when the AI application provides the target service; when it is detected that the duration for which the APIs in the hijacking library have not been called by the AI application reaches a first preset duration, determining that the AI application is in a running state; when the AI application is in the running state, determining the API call sequence based on the APIs called in the hijacking library when the AI application provides the target service; Obtaining a reference execution time of the API call sequence, where the reference execution time refers to a time used to measure whether the AI application is in an abnormal state; When the target execution time exceeds the reference execution time, it is determined that the AI application is in an abnormal state.
2. The method according to claim 1, characterized in that The step of obtaining the target execution time of the application programming interface (API) call sequence includes: When the AI application provides the target service at any one time, generating at least one API call instruction, wherein the at least one API call instruction respectively includes an API in the API call sequence; Based on the hijacking library hijacking the at least one API call instruction, the API in the at least one API call instruction hijacked by the hijacking library is replaced with the storage address of the corresponding API in the dynamic link library in the AI application process, thereby obtaining at least one replaced API call instruction; Calling the resource stored at the storage address in the at least one replaced API call instruction by executing the AI application process; The sum of the execution times corresponding to the at least one replaced API call instruction is obtained as the target execution time.
3. The method according to claim 2, characterized in that Before replacing the API in at least one API call instruction hijacked by the hijacking library with the storage address of the corresponding API in the dynamic link library in the AI application process to obtain at least one replaced API call instruction, the method further includes: When the AI application is started, the dynamic link library and the hijacking library are loaded into the AI application process, and each API in the dynamic link library corresponds to a storage address in the AI application process.
4. The method according to claim 1, characterized in that: The step of determining the API call sequence based on the API called in the hijacking library when the AI application provides a target service while the AI application is in the running state includes: When the AI application is in operation, if the time duration during which the API in the hijacking library has not been called reaches a second preset time duration, and the API in the hijacking library is called periodically when the AI application provides the target service, the periodic API call sequence is determined as the API call sequence.
5. The method according to claim 1, characterized in that The method further comprises: When the AI application is in the running state, if any API is called for the first time, the AI application is rolled back from the running state to the initialization state.
6. The method according to claim 1, characterized in that The step of hijacking the API in the library when the AI application is in the running state and taking the periodic API call sequence as the API call sequence based on the API called in the library when the AI application provides the target service also includes: When the AI application is in a running state, if a new API call sequence is acquired and a first sequence length of the new API call sequence is less than a second sequence length of the API call sequence, the API call sequence is updated to the new API call sequence.
7. The method according to claim 1, characterized in that The method further comprises: When the AI application is in an initialization state or a running state, when a specified API in the hijacking library is called, the AI application is switched to an end state, and the specified API refers to an API used to close the AI application.
8. The method according to any one of claims 1 to 7, characterized in that Before obtaining the target execution time of the API call sequence, the method further includes: Obtaining a parameter value of a target parameter in the hijacking library, where the target parameter is used to indicate whether to perform anomaly detection on the AI application; When the parameter value of the target parameter is a preset value, an operation of obtaining the target execution time of the application programming interface API call sequence is performed.
9. The method according to any one of claims 1 to 7, characterized in that The method further comprises: In the process of the AI application providing any target service, if the API call sequence is executed and completed within a third preset duration, an operation of obtaining a reference execution duration of the API call sequence is performed, and the third preset duration is greater than the reference execution duration; If the API call sequence is not executed to completion within the third preset time length, it is determined that the AI application is in an abnormal state.
10. An electronic device, characterized in that: It comprises a processor and a memory; the memory stores at least one program code; the at least one program code is used to be called and executed by the processor to implement the artificial intelligence AI application anomaly detection method as described in any one of claims 1 to 7.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one computer program, and when the at least one computer program is executed by the processor, it can implement the artificial intelligence AI application anomaly detection method according to any one of claims 1 to 7.
12. A computer program product, characterized in that The computer program product includes a computer program, and when the computer program is executed by a processor, it can implement the artificial intelligence AI application anomaly detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Monitoring alarm and traceability method and system based on call chain data
CN114531338A
Method and system for intercepting an application program interface
US6823460B1