Intelligent engine, intelligent engine cluster, distributed troubleshooting system and troubleshooting method
By introducing Sidecar containers into the intelligent engine, collecting and uploading on-site information of target abnormal events, the risks derived from abnormal events during the troubleshooting process in the existing technology are solved, and the stable operation and efficient troubleshooting of the intelligent engine are achieved.
Patent Information
- Application Number
- CN202311857843.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2043-12-29
AI Technical Summary
The prior art is prone to derive other abnormal events during the troubleshooting of artificial intelligence engines, increasing potential risks.
It adopts an intelligent engine architecture, including the engine main container and the Sidecar container. The engine main container runs the model engine process and monitors exceptions. After receiving the abnormal information, the Sidecar container collects target site information and uploads it to the diagnostic emergency subsystem to achieve non-invasive target site information collection.
Through the isolation between the main engine container and the sidecar container, the decoupling of the model engine process and troubleshooting can be achieved, and other abnormal events will be avoided and the operation stability of the intelligent engine will be improved.
Smart Images

Figure CN117806868B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent engine, an intelligent engine cluster, a distributed troubleshooting system and a troubleshooting method. Background Art
[0002] With the development of cloud computing and artificial intelligence technology, various artificial intelligence models have been trained and developed. At the same time, with the help of cloud native and microservice technologies, artificial intelligence service frameworks and development platforms are also developing rapidly. The engine hosting platform has become a key infrastructure for deploying and running various artificial intelligence models. However, due to the richness of artificial intelligence engines and the diversity of models, the variability of user input, and various problems that may exist in the engine itself, various anomalies may occur in artificial intelligence services during operation. These anomalies may affect the performance, accuracy and stability of the service, causing abnormal events such as service freezing, crashing, timeout, etc., which in turn affects the service quality of the entire application system, causing user usage anomalies and ultimately causing business losses.
[0003] In the prior art, on the one hand, troubleshooting is performed by directly embedding event capture logic in the model engine process, and on the other hand, troubleshooting is performed by starting a sub-process in the model engine process. However, both methods may lead to other abnormal events and increase potential risks. For example, engine processes interfere with each other, engine processes are coupled, and the isolation between engine processes is poor. In some cases, it is difficult for the model engine process to correctly handle the signal of the sub-process. If the sub-process does not handle the signal correctly, it may cause the signal loop to be thrown, and the model engine process cannot exit normally, resulting in zombie processes, etc. Summary of the invention
[0004] The present invention provides an intelligent engine, an intelligent engine cluster, a distributed troubleshooting system and a troubleshooting method, which are used to solve the defects of the prior art that other abnormal events are derived during troubleshooting, thereby increasing risks.
[0005] The present invention provides an intelligent engine, comprising an engine main container and at least one Sidecar container, wherein the engine main container is communicatively connected to the at least one Sidecar container, and the at least one Sidecar container is communicatively connected to a diagnostic emergency subsystem, wherein:
[0006] The engine main container is used to run the model engine process, provide artificial intelligence services to the outside, and send abnormal information to the at least one Sidecar container when a target abnormal event occurs in the model engine process;
[0007] The at least one Sidecar container is used to collect target site information corresponding to the target abnormal event corresponding to the engine main container when receiving the abnormal information, and upload the target site information to the diagnosis emergency subsystem.
[0008] According to the intelligent engine provided by the present invention, the engine main container includes an inference engine, a service framework and a monitoring module, wherein:
[0009] The inference engine is communicatively connected with the service framework, and an artificial intelligence model is deployed in the inference engine;
[0010] The service framework is used to run the artificial intelligence model in the inference engine based on the model engine process and provide artificial intelligence services to the outside world;
[0011] The monitoring module is communicatively connected to the at least one Sidecar container, and is used to monitor the target abnormal state corresponding to the model engine process; determine the target abnormal event corresponding to the target abnormal state, and the abnormal information corresponding to the target abnormal event; and send the abnormal information to the at least one Sidecar container based on the domain socket communication method.
[0012] According to the intelligent engine provided by the present invention, the monitoring module determines the target abnormal event corresponding to the target abnormal state, specifically for:
[0013] In the case where the target abnormal state indicates that a target abnormal signal is captured, matching the target abnormal signal with a preset registration signal in the monitoring module, and in the case where a preset registration signal corresponding to the target abnormal signal is matched, determining the crash event type corresponding to the preset registration signal as the target abnormal event corresponding to the target abnormal state;
[0014] When the target abnormal state indicates that the target abnormal signal has not been captured, and the number of request timeouts is greater than or equal to 1 and less than a preset threshold, determining that the target abnormal event corresponding to the target abnormal state is a timeout event;
[0015] When the target abnormal state indicates that the target abnormal signal is not captured and the number of request timeouts is equal to the preset threshold, it is determined that the target abnormal event corresponding to the target abnormal state is a stuck event, and the crash event type corresponding to the preset registration signal does not include the timeout event and the stuck event.
[0016] According to the intelligent engine provided by the present invention, the at least one Sidecar container is specifically used for:
[0017] Based on the exception information, determine the target process number and the target exception event corresponding to the model engine process in the engine main container; the engine main container shares a process namespace with the at least one Sidecar container;
[0018] Based on the target process number, collecting target scene information corresponding to the target abnormal event;
[0019] Based on the target site information, alarm information is generated, and the alarm information is sent to the diagnosis emergency subsystem.
[0020] According to the intelligent engine provided by the present invention, the at least one Sidecar container is further used for:
[0021] Classifying the target site information to determine at least two categories of sub-target site information;
[0022] Based on a preset mapping relationship, determining a target storage module corresponding to each type of sub-target on-site information; the preset mapping relationship includes a mapping relationship between on-site information type and storage module;
[0023] For each type of sub-target on-site information, the sub-target on-site information is sent to the corresponding target storage module, and each target storage module is deployed in the diagnosis emergency subsystem.
[0024] The present invention also provides an intelligent engine cluster, comprising: a cluster management executor and at least two intelligent engines as described in any one of the above items, each of the intelligent engines is communicatively connected to a diagnosis emergency subsystem and the cluster management executor, and the diagnosis emergency subsystem is communicatively connected to the cluster management executor, wherein:
[0025] The cluster management executor is used to receive the target emergency operation instructions sent by the diagnostic emergency subsystem and execute the target emergency operation instructions to achieve troubleshooting and recovery of the target intelligent engine, to which the target intelligent engine belongs.
[0026] The present invention also provides a distributed troubleshooting system, comprising: a diagnostic emergency subsystem and any one of the above-mentioned intelligent engine clusters, wherein:
[0027] The diagnostic emergency subsystem is used to determine the target emergency operation instructions corresponding to the target intelligent engine based on the alarm information uploaded by the target intelligent engine in the intelligent engine cluster, and send the target emergency operation instructions to the intelligent engine cluster.
[0028] According to the distributed troubleshooting system provided by the present invention, the diagnostic emergency subsystem is also used for:
[0029] Analyze the warning information to determine whether the target abnormal event corresponding to the warning information meets the emergency condition;
[0030] Based on the judgment result, a target emergency operation instruction corresponding to the warning information is generated.
[0031] According to the distributed troubleshooting system provided by the present invention, the diagnostic emergency subsystem is also used for:
[0032] In the case where the file type corresponding to the target field information is a minidump file, readable debugging information is generated based on the minidump file and the target symbol file corresponding to the target intelligent engine.
[0033] The present invention also provides a troubleshooting method, which is applied to any of the above-mentioned intelligent engines, comprising:
[0034] Run the model engine process, provide artificial intelligence services to the outside world, and send abnormal information to at least one Sidecar container when detecting a target abnormal event in the model engine process;
[0035] When the abnormal information is received, target scene information corresponding to the target abnormal event corresponding to the engine main container is collected, and the target scene information is uploaded to the diagnosis emergency subsystem.
[0036] The intelligent engine, intelligent engine cluster, distributed troubleshooting system and troubleshooting method provided by the present invention respectively deploy an independent engine main container and at least one Sidecar container in each intelligent engine, run the model engine process through the engine main container, provide artificial intelligence services to the outside, and monitor the running status of the model engine process in real time. When a target abnormal event is detected in the model engine process, abnormal information is sent to at least one Sidecar container. Without affecting the normal operation of the model engine process, at least one Sidecar container collects the target scene information corresponding to the target abnormal event, and uploads the target scene information to the diagnostic emergency subsystem. By making full use of cloud native technology and containerization technology, and utilizing the isolation between the engine main container and at least one Sidecar container, the decoupling between the model engine process and troubleshooting is achieved, thereby realizing non-intrusive collection of target scene information outside the model engine process, avoiding the derivation of other abnormal events, and improving the running stability of the intelligent engine. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0038] Figure 1 is a connection diagram of the intelligent engine provided by an embodiment of the present invention;
[0039] Figure 2 is a connection diagram of the engine main container provided by an embodiment of the present invention;
[0040] Figure 3 is a schematic diagram of the structure of abnormal information provided by an embodiment of the present invention;
[0041] Figure 4 is a schematic diagram of an interaction process of abnormal information provided by an embodiment of the present invention;
[0042] Figure 5 is a schematic diagram of the process of uploading target site information provided by an embodiment of the present invention;
[0043] Figure 6 is a connection diagram of an intelligent engine cluster provided by an embodiment of the present invention;
[0044] Figure 7 is a connection diagram of a distributed troubleshooting system provided by an embodiment of the present invention;
[0045] Figure 8 It is a schematic diagram of the process of formulating an emergency strategy provided by an embodiment of the present invention;
[0046] Fig. 9 is a flowchart of a troubleshooting method provided by an embodiment of the present invention;
[0047] Fig.10 is a schematic diagram of the structure of an obstacle removal device provided by an embodiment of the present invention;
[0048] Fig.11 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention.
[0049] Reference numerals:
[0050] 100: Intelligent engine; 110: Engine main container; 111: Reasoning engine; 112: Service framework; 113: Monitoring module; 120: Sidecar container; 200: Intelligent engine cluster; 210: Cluster management executor; 300: Distributed troubleshooting system; 310: Diagnostic emergency subsystem. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] In view of the problem that other abnormal events are generated during troubleshooting in the prior art, which increases the risk, an embodiment of the present invention provides an intelligent engine. Figure 1 is a connection diagram of the intelligent engine provided by an embodiment of the present invention, such as Figure 1 As shown, the smart engine 100 includes an engine main container 110 and at least one Sidecar container, the engine main container 110 is communicatively connected to the at least one Sidecar container, and the at least one Sidecar container is communicatively connected to the diagnostic emergency subsystem, wherein:
[0053] The engine main container 110 is used to run the model engine process, provide artificial intelligence services to the outside, and send abnormal information to the at least one Sidecar container when a target abnormal event occurs in the model engine process;
[0054] The at least one Sidecar container is used to collect the target site information corresponding to the target abnormal event corresponding to the engine main container 110 when receiving the abnormal information, and upload the target site information to the diagnosis emergency subsystem.
[0055] Specifically, in the intelligent engine 100, the engine main container 110 carries an artificial intelligence model. After receiving the service request corresponding to the user, the model engine process corresponding to the artificial intelligence model is run to execute the business logic in the model engine process to process the user's service request, so as to provide the artificial intelligence service corresponding to the artificial intelligence model to the outside. At the same time, while running the model engine process, the engine main container 110 monitors the running status of the model engine process in real time, determines whether the target abnormal event occurs in the model engine process, and sends abnormal information to at least one Sidecar container if the target abnormal event occurs. At the same time, at least one Sidecar container is deployed separately in the intelligent engine 100, and the at least one Sidecar container provides fault event management services, that is, the at least one Sidecar container monitors in real time whether the engine main container 110 sends abnormal information. If the abnormal information is received, the target scene information corresponding to the target abnormal event is collected according to the event type of the target abnormal event determined by the engine main container 110, and the non-intrusive target scene information is collected outside the model engine process without affecting the normal operation of the model engine process. Report to the diagnosis emergency subsystem to formulate corresponding emergency strategies to minimize the impact on the artificial intelligence service.
[0056] In the intelligent engine 100, both the model engine process and the fault event management service run in a cloud-native environment. Through the rapid deployment of each container, lightweight virtualization and resource isolation, etc., the advantages of containerization are fully utilized to improve the elasticity and scalability of the intelligent engine 100, while making the monitoring of target abnormal events and the collection of target on-site information more efficient and reliable. At the same time, the model engine process and the fault event management service are deployed in different containers respectively. The isolation between the containers decouples the model engine process from the fault event management service, so that the model engine process and the fault event management service can be independently deployed, maintained and upgraded without affecting each other, thereby improving the operational stability of the intelligent engine 100.
[0057] Optionally, the artificial intelligence services provided by the intelligent engine 100 may include intelligent interaction, speech recognition service, speech synthesis service, image recognition service, and video review service. When the artificial intelligence service is an intelligent interaction service, the artificial intelligence model in the engine main container 110 may process at least one of text data, audio data, image data, and video data; when the artificial intelligence service is a speech recognition service or a speech synthesis service, the artificial intelligence model in the engine main container 110 may process audio data; when the artificial intelligence service is an image recognition service or a video review service, the artificial intelligence model in the engine main container 110 may process image data and / or video data. The embodiment of the present invention does not limit this.
[0058] Optionally, at least one Sidecar container may be deployed in the intelligent engine 100, that is, when a Sidecar container is deployed in the intelligent engine 100, the Sidecar container is used to monitor the engine main container 110, collect target site information, and upload the target site information to the diagnostic emergency subsystem. When the number of Sidecar containers deployed in the intelligent engine 100 is greater than 1, each Sidecar container may handle different fault event management services respectively, and all Sidecar containers may determine the communication connection between the Sidecar container and the engine main container 110 according to the fault event management service handled by each Sidecar container, without the need to deploy all Sidecar containers to communicate with the engine main container 110. For example, three Sidecar containers can be deployed in the intelligent engine 100, wherein two Sidecar containers can be set to be connected to the engine main container 110 in communication, and the two Sidecar containers are connected to each other in communication, that is, Sidecar container A of the two Sidecar containers is used to monitor whether abnormal information is sent in the engine main container 110, and after Sidecar container A receives the abnormal information, it sends a notification to Sidecar container B, and Sidecar container B is used to collect the target site information corresponding to the target abnormal event in the model engine process. Another Sidecar container C is connected to the Sidecar container B and the diagnostic emergency subsystem in communication, and after Sidecar container B collects the target site information, it can receive the target site information sent by Sidecar container B, and send the target site information to the diagnostic emergency subsystem for analysis, and formulate emergency measures corresponding to the target abnormal event. By flexibly expanding the Sidecar container, the fault event management service can be adjusted by adding or reducing Sidecar containers without affecting the model engine process to deal with different types of abnormal events, thereby enhancing the customization and configurability of the intelligent engine 100.
[0059] Optionally, in the intelligent engine 100, the engine main container 110 and at least one Sidecar container are set to C / S (Client / Server) mode, wherein the engine main container 110 acts as a client and at least one Sidecar container acts as a server, and the client and the server respectively assume different functions. The communication connection between the engine main container 110 and at least one Sidecar container is not always maintained. A communication connection is established with at least one Sidecar container only when a target abnormal event occurs in the model engine process in the engine main container 110 and abnormal information needs to be sent to at least one Sidecar container. The C / S architecture simplifies the architecture of the intelligent engine 100, making the communication between the client and the server more intuitive and flexible, while also improving the scalability and maintainability of the intelligent engine 100.
[0060] Furthermore, Figure 2 1 is a connection diagram of the engine main container 110 provided in an embodiment of the present invention. Figure 2 As shown, the engine main container 110 includes an inference engine 111, a service framework 112 and a monitoring module 113, wherein:
[0061] The inference engine 111 is in communication with the service framework 112, and an artificial intelligence model is deployed in the inference engine 111;
[0062] The service framework 112 is used to run the artificial intelligence model in the inference engine 111 based on the model engine process and provide artificial intelligence services to the outside world;
[0063] The monitoring module 113 is communicatively connected to the at least one Sidecar container, and the monitoring module 113 is used to monitor the target abnormal state corresponding to the model engine process; determine the target abnormal event corresponding to the target abnormal state, and the abnormal information corresponding to the target abnormal event; and based on the domain socket communication method, send the abnormal information to the at least one Sidecar container.
[0064] Specifically, in the engine main container 110, by deploying the corresponding artificial intelligence model in the inference engine 111, and through the service framework 112, deploying, managing and running the artificial intelligence model, the corresponding artificial intelligence service is provided to the outside. The service framework 112 can be understood as a software framework, and the service framework 112 can be called by external applications or other services to receive external service requests and return results. In addition, a monitoring module 113 is also deployed in the engine main container 110. The monitoring module 113 is an SDK (Software Development Kit), and the monitoring module 113 monitors whether the running state of the model engine process is the target abnormal state. If the model engine process is in the target abnormal state, the target abnormal event corresponding to the target abnormal state can be further determined, and the abnormal information corresponding to the target abnormal event is triggered, and the abnormal information is sent to at least one Sidecar container in a domain socket communication manner.
[0065] Optionally, the service framework 112 may include TensorFlow Serving or PyTorch Serving, etc., which is not limited in this embodiment of the present invention.
[0066] Furthermore, the monitoring module 113 determines the target abnormal event corresponding to the target abnormal state, specifically for:
[0067] In the case where the target abnormal state indicates that a target abnormal signal is captured, the target abnormal signal is matched with a preset registration signal in the monitoring module 113, and in the case where the preset registration signal corresponding to the target abnormal signal is matched, the crash event type corresponding to the preset registration signal is determined as the target abnormal event corresponding to the target abnormal state;
[0068] When the target abnormal state indicates that the target abnormal signal has not been captured, and the number of request timeouts is greater than or equal to 1 and less than a preset threshold, determining that the target abnormal event corresponding to the target abnormal state is a timeout event;
[0069] When the target abnormal state indicates that the target abnormal signal is not captured and the number of request timeouts is equal to the preset threshold, it is determined that the target abnormal event corresponding to the target abnormal state is a stuck event, and the crash event type corresponding to the preset registration signal does not include the timeout event and the stuck event.
[0070] Specifically, the target abnormal event may occur at any time after the model engine process is started. The triggering reason may be that the developer has configured the wrong environment variables, startup parameters, or the environment is inconsistent, or even the user's input data error may cause the target abnormal event to occur. The target abnormal event may include different types of crash events, freeze events, or timeout events, among which:
[0071] 1) A crash event refers to an abnormal premature termination or termination of the model engine process due to an unrecoverable error or exception in the program or intelligent engine 100. In the crashed state, the program cannot be executed normally and exits, resulting in the unavailability of the artificial intelligence service. When the process is initialized, the monitoring module 113 in the embodiment of the present invention pre-calls the Init function of the SDK to register some preset registration signals related to the crash event, such as SIGSEGV (Segmentation fault), SIGABRT (Abort), SIGILL (Illegal instruction), SIGFPE (Floating point exception) and SIGBUS, etc. When a crash event occurs, the monitoring module 113 will capture the target exception signal, match the target exception signal with the above-mentioned preset registration signal in turn, and determine the crash event type corresponding to the matched preset registration signal as the target exception event.
[0072] 2) When the timeout event occurs, the artificial intelligence service corresponding to the model engine process is still available, but the service request response time of some users exceeds the preset threshold or does not receive a response, and some service requests may still be normal. In the embodiment of the present invention, the artificial intelligence model in the inference engine 111 can simultaneously process a preset threshold number of corresponding data to the maximum extent, that is, the engine main container 110 can simultaneously process a preset threshold number of service requests. A counting thread is set in the service framework 112. If the processing of the service request times out, the count can be added by 1 in the counting thread, that is, the number of request timeouts is added by 1. If the monitoring module 113 does not capture the target abnormal signal, it indicates that no crash event has occurred. At this time, the request timeout number can be obtained from the service framework 112. If the obtained request timeout number is greater than or equal to 1, but less than the preset threshold, it can be determined that the model engine process has a timeout event. At this time, the artificial intelligence service is still available, and the subsequent diagnosis emergency subsystem can determine whether to formulate an emergency strategy to restart and reply to the artificial intelligence service according to the severity of the timeout.
[0073] 3) When a deadlock event occurs, the model engine process is unavailable, the model engine process exists but will not exit, and all service requests are unresponsive when a deadlock event occurs, but the model engine process still exists, but cannot respond to any service request, and the model engine process falls into a blocked state. If the monitoring module 113 in the embodiment of the present invention does not capture the target abnormal signal, it indicates that no crash event has occurred. At this time, the request timeout number can be obtained from the service framework 112. If the obtained request timeout number is equal to the preset threshold, it can be determined that a deadlock event has occurred in the model engine process.
[0074] Furthermore, after determining the target abnormal event, abnormal information can be constructed according to the target abnormal event to notify the at least one Sidecar container to collect the target site information corresponding to the target abnormal event and perform event recording. Figure 3 is a schematic diagram of the structure of abnormal information provided by an embodiment of the present invention, such as Figure 3 As shown, the exception information includes a header and a data segment. The header includes a message type, a data length, and an event type. The event type is the type of the target exception event, i.e., a crash event, a stuck event, or a timeout event. The data segment may include a list of pipes (Pipe) and session IDs (IdentityDocument, identity identification). The session ID list may include at least one session ID. The session ID may be understood as a unique identifier corresponding to a service request. For example, when the target exception event is a timeout event, the session list may include at least one session ID corresponding to a service request that responds to a timeout.
[0075] also, Figure 4 is a schematic diagram of the interaction flow of abnormal information provided by an embodiment of the present invention, such as Figure 4 As shown, after the SDK is initialized and the SDK detects the target abnormal opinion, the engine main container 110 registers the pipeline and constructs the abnormal information. After sending the abnormal information to at least one Sidecar container, the pipeline is controlled to be in a blocked state. After at least one Sidecar container receives the abnormal information and parses the pipeline handle, the corresponding abnormal event processing logic is activated. After completing the target site information collection and upload, an instruction to close the pipeline is generated and sent to the engine main container 110. The engine main container 110 closes the pipeline after receiving the instruction or timing out. Through the pipeline, the engine main container 110 can monitor whether the at least one Sidecar container has completed the collection and upload of the target site information.
[0076] Further, after determining the abnormal information, since the engine main container 110 and at least one Sidecar container adopt a C / S architecture, when the engine main container 110 and at least one Sidecar container communicate, a domain socket (Unix Domain Socket, UDS) communication method can be used for communication. For example, after constructing the abnormal information, the engine main container 110 can send the abnormal information to at least one Sidecar container through a Socket. After at least one Sidecar container collects and uploads the target field information, the instruction to close the channel can be sent to the engine main container 110 through a Socket. In addition, since multiple intelligent engines 100 may be deployed on the same machine in a real environment, in order to prevent port conflicts, the abstract namespace method of UDS is adopted, and the Socket is named with a POD name. There is no need to create a Socket file. Just name a global name with the POD name, and the engine main container 110 or at least one Sidecar container can be linked according to the POD name.
[0077] Further, after at least one Sidecar container receives the abnormal information, an event record will be generated, which may include: event ID, availability zone, node name, host address, occurrence time, environment, stack, thread, coroutine address, startup command and some other context information, among which, event ID represents the unique identifier of the target abnormal event; availability zone represents the availability zone in which the target abnormal event occurs; node name and host address represent the service node where the target abnormal event occurs; occurrence time represents the time when the target abnormal event occurs; environment represents the environment in which the target abnormal event occurs, for example, the verification environment or the formal environment; stack, thread, coroutine address represents the file address uploaded to the storage module after the collection is completed; startup command and some other context information represent the startup command, environment variables and other context information of the crashed process. The embodiments of the present invention are not limited to this.
[0078] Furthermore, the at least one Sidecar container is specifically used for:
[0079] Based on the exception information, determine the target process number and the target exception event corresponding to the model engine process in the engine main container 110; the engine main container 110 shares a process namespace with the at least one Sidecar container;
[0080] Based on the target process number, collecting target scene information corresponding to the target abnormal event;
[0081] Based on the target site information, alarm information is generated, and the alarm information is sent to the diagnosis emergency subsystem.
[0082] Specifically, in order to ensure normal communication between the engine main container 110 and at least one Sidecar container, the at least one Sidecar container can capture identifiable target site information, and in an embodiment of the present invention, the engine main container 110 and at least one Sidecar container in the same intelligent engine 100 share a process namespace. In addition, in a UNIX system, using the mechanism of UNIX domain sockets, after the engine main container 110 and at least one Sidecar container both enable the SO_PASSCRED option, each time the engine main container 110 connects to at least one Sidecar container, the kernel in the UNIX system will attach information about the engine main container 110 to the message body, which may include PID (Process ID), UID (User ID), GID (Group ID), etc., that is, the kernel will attach the process ID in the engine main container 110 to the exception information. After the at least one Sidecar container receives the exception information, it can obtain the target process number corresponding to the model engine process from the shared namespace corresponding to the exception information, and based on the target process number, collect the target site information corresponding to the target exception event, and generate alarm information based on the target site information. Thereafter, the alarm information is sent to the diagnostic emergency subsystem to specify the corresponding emergency strategy.
[0083] Optionally, the target field information may include field stacks, threads, coroutines, request data, memory snapshots and logs, among which, capturing the field stack information of the model engine process when the target abnormal event occurs is convenient for locating the specific location where the target abnormal event occurs; for multi-threaded or coroutine systems, collecting information of each thread or coroutine can determine information such as the status and execution path of each thread or coroutine, which helps analyze concurrency problems; when memory-related problems occur, obtaining memory snapshots can determine memory usage, which helps analyze problems such as memory leaks.
[0084] In addition, the target scene information collected can be different depending on the type of target abnormal event. For example, when the target abnormal event is a crash event, the target scene information collected can include the scene stack, thread, coroutine, request data and log; when the target abnormal event is a stuck event, the target scene information collected can include the scene stack, thread, coroutine, request data and log; when the target abnormal event is a timeout event, the target scene information collected can include request data and log. Through differentiated information collection strategies, it is helpful to accurately grasp the operating status of the intelligent engine 100 and improve the efficiency and accuracy of troubleshooting.
[0085] Optionally, the alarm information may include engine status, abnormality type and possible cause, etc., which is not limited in the embodiment of the present invention.
[0086] Optionally, the CAP_ADMIN permission may be set in at least one Sidecar container to perform some operations that require special permissions. The operations may include calling setns(2), fanotify_init(2), and calling bpf, etc. This embodiment of the present invention does not limit this.
[0087] Furthermore, the at least one Sidecar container is further used to:
[0088] Classifying the target site information to determine at least two categories of sub-target site information;
[0089] Based on a preset mapping relationship, determining a target storage module corresponding to each type of sub-target on-site information; the preset mapping relationship includes a mapping relationship between on-site information type and storage module;
[0090] For each type of sub-target on-site information, the sub-target on-site information is sent to the corresponding target storage module, and each target storage module is deployed in the diagnosis emergency subsystem.
[0091] Specifically, after the at least one Sidecar container has collected the target field information corresponding to the target abnormal event, the target field information can be classified by the data upload unit in at least one Sidecar container or the Sidecar container responsible for data upload to obtain at least two types of sub-target field information, each type of sub-target field information can include: standard output log, minidump file or event record, etc. After classifying and obtaining at least two types of sub-target field information, for each type of sub-target field information, the type of the sub-target field information can be matched with the field information type in the preset mapping relationship according to the preset mapping relationship, and the storage module corresponding to the field information type matching the sub-target field information is determined as the target storage module of the sub-target field information, and then the sub-target field information is stored in the target storage module. By collecting streaming data and rich media data when the target abnormal event occurs, it is convenient for subsequent troubleshooting personnel or automated analysis tools in the diagnostic emergency subsystem to search on demand, which not only helps to reproduce and troubleshoot problems in real time, realizes comprehensive monitoring of the intelligent engine 100, but also provides data support for subsequent data analysis and mining.
[0092] For example, Figure 5 FIG. 1 is a schematic diagram of a process for uploading target site information provided by an embodiment of the present invention. Figure 5As shown, the storage module in the preset mapping relationship may include a search server (Elasticsearch), an object storage, and a database. Elasticsearch may store standard output logs, the object storage may store minidump files, and the database may store event records. If the target abnormal event is a crash event, the minidump file corresponding to the crash event may be uploaded to the object storage, and the event record corresponding to the crash event may be uploaded to the database.
[0093] The intelligent engine provided by the present invention deploys an independent engine main container and at least one Sidecar container 120 in each intelligent engine, respectively, runs the model engine process through the engine main container, provides artificial intelligence services to the outside, and monitors the running status of the model engine process in real time. When a target abnormal event is detected in the model engine process, abnormal information is sent to at least one Sidecar container 120. Without affecting the normal operation of the model engine process, at least one Sidecar container 120 collects the target scene information corresponding to the target abnormal event, and uploads the target scene information to the diagnostic emergency subsystem. By making full use of cloud native technology and containerization technology, and utilizing the isolation between the engine main container and at least one Sidecar container 120, the decoupling between the model engine process and troubleshooting is achieved, thereby realizing non-intrusive collection of target scene information outside the model engine process, avoiding the derivation of other abnormal events, and improving the running stability of the intelligent engine.
[0094] The present invention also provides an intelligent engine cluster. Figure 6 is a connection diagram of the intelligent engine cluster provided by an embodiment of the present invention, such as Figure 6 As shown, the intelligent engine cluster 200 includes: a cluster management executor 210 and at least two intelligent engines 100 as described in any one of the above items, each of the intelligent engines 100 is communicatively connected with a diagnosis and emergency subsystem 310 and the cluster management executor 210, and the diagnosis and emergency subsystem 310 is communicatively connected with the cluster management executor 210, wherein:
[0095] The cluster management executor 210 is used to receive the target emergency operation instructions sent by the diagnosis emergency subsystem 310 and execute the target emergency operation instructions to achieve troubleshooting and recovery of the target intelligent engine, which belongs to the intelligent engine 100.
[0096] Specifically, a cluster management executor 210 and multiple intelligent engines 100 can be deployed in the intelligent engine cluster 200. After a target abnormal event occurs in intelligent engine A, the target site information and alarm information corresponding to the target abnormal event can be uploaded to the diagnostic emergency subsystem. After the diagnostic emergency subsystem determines the corresponding emergency strategy and the target emergency operation instruction corresponding to the emergency strategy, the diagnostic emergency subsystem 310 can send the target emergency operation instruction to the cluster management executor 210 in the intelligent engine cluster 200. The cluster management executor 210 can execute the target emergency operation instruction and coordinate in the intelligent engine cluster 200 to ensure that the artificial intelligence service corresponding to the intelligent engine A can be quickly restored to a normal state.
[0097] Optionally, the types of artificial intelligence services corresponding to the multiple intelligent engines 100 included in the intelligent engine cluster 200 may be the same or different, and the embodiment of the present invention does not impose any limitation on this.
[0098] The embodiment of the present invention also provides a distributed troubleshooting system. Figure 7 FIG. 1 is a connection diagram of a distributed troubleshooting system provided by an embodiment of the present invention. Figure 7 As shown, the distributed troubleshooting system 300 includes: a diagnostic emergency subsystem 310 and the above-mentioned intelligent engine cluster 200, wherein:
[0099] The diagnostic emergency subsystem 310 is used to determine the target emergency operation instructions corresponding to the target intelligent engine based on the alarm information uploaded by the target intelligent engine in the intelligent engine cluster 200, and send the target emergency operation instructions to the intelligent engine cluster 200.
[0100] Specifically, the distributed troubleshooting system 300 includes a diagnostic emergency subsystem 310 and an intelligent engine cluster 200. When a target abnormal event occurs, the intelligent engine cluster 200 can collect target on-site information and generate alarm information to notify the diagnostic emergency subsystem 310. After receiving the alarm information, the diagnostic emergency subsystem 310 can automatically analyze the target abnormal event through the automated analysis tool in the system, formulate an emergency strategy and the target emergency operation instructions corresponding to the emergency strategy, and then send the target emergency operation instructions to the intelligent engine cluster 200, so that the intelligent engine in the intelligent engine cluster 200 where the target abnormal event occurs can quickly return to normal, minimizing the impact on the artificial intelligence service.
[0101] Furthermore, the diagnosis emergency subsystem 310 is also used for:
[0102] Analyze the warning information to determine whether the target abnormal event corresponding to the warning information meets the emergency condition;
[0103] Based on the judgment result, a target emergency operation instruction corresponding to the warning information is generated.
[0104] Specifically, Figure 8 FIG. 1 is a schematic diagram of a process for formulating an emergency strategy provided by an embodiment of the present invention. Figure 8 As shown, some emergency conditions and corresponding emergency strategies are preset in the diagnostic emergency subsystem 310, and the emergency strategy may include: restarting the engine, service degradation, diverting traffic, and rolling back the version, etc. After the diagnostic emergency subsystem 310 receives the alarm information, it can be determined whether the target abnormal event meets the emergency conditions. The emergency conditions may include: whether the cause of the target abnormal event is a problem with the intelligent engine itself or a problem related to the artificial intelligence service corresponding to the intelligent engine, the severity of the target abnormal event, whether the artificial intelligence service is unavailable, whether an emergency strategy needs to be made immediately, etc. According to the judgment result of whether the emergency conditions are met, the target emergency operation instruction corresponding to the alarm information can be generated. For example, if the target abnormal event is the stuck of the intelligent engine, the alarm information of the target abnormal event is compared with each emergency condition, and it is determined that the occurrence of the target abnormal event is a problem of the intelligent engine itself, and the corresponding artificial intelligence service is unavailable, which is more serious and needs to be made immediately. The emergency strategy can be generated as restarting the engine, and the target emergency operation instruction corresponding to the restart engine. For another example, if the target abnormal event is a crash event, that is, smart engine A crashes, after comparing with various emergency conditions, it can be determined that the occurrence of the crash event is a problem with the smart engine itself, but there is an smart engine B corresponding to the same type of artificial intelligence service in the smart engine cluster 200, that is, when the corresponding type of artificial intelligence service is still available, the artificial intelligence service corresponding to smart engine A can be downgraded, and the traffic corresponding to smart engine A can be diverted to smart engine B. Therefore, the determined emergency strategy is service downgrade and diversion to smart engine B, and the target emergency operation instructions corresponding to the emergency strategy are generated.
[0105] Furthermore, the diagnosis emergency subsystem 310 is also used for:
[0106] In the case where the file type corresponding to the target field information is a minidump file, readable debugging information is generated based on the minidump file and the target symbol file corresponding to the target intelligent engine.
[0107] Specifically, when the target abnormal event is a crash event, and the fault locating personnel manually triggers online simulation debugging, or the automated analysis tool in the diagnostic emergency subsystem 310 automatically triggers online simulation debugging, the diagnostic emergency subsystem 310 can extract the minidump file corresponding to the crash event from the corresponding object storage. The minidump file can be understood as an unreadable binary file. When performing online simulation debugging, the minidump file can be combined with the target symbol file exported from the monitoring module in the target intelligent engine to generate readable debugging information with program line numbers, providing real-time, visual debugging tools for fault locating personnel, facilitating fault locating personnel to simulate and observe system behavior in a production environment, facilitating fault locating personnel or automated analysis tools to perform positioning and diagnostic analysis, thereby shortening the troubleshooting time.
[0108] The distributed troubleshooting system provided by the embodiment of the present invention, when a target abnormal event occurs, collects the target site information corresponding to the target abnormal event through the intelligent engine cluster for troubleshooting, and combines the diagnostic emergency subsystem to perform diagnosis and formulate emergency strategies, thereby improving the distributed troubleshooting system's perception of the target abnormal event, the speed of response, the degree of automation and comprehensiveness of the troubleshooting process, reducing the cost of manual intervention, and at the same time, shortening the cycle from the occurrence of the target abnormal event to the final repair, thereby improving troubleshooting efficiency.
[0109] The embodiment of the present invention further provides a troubleshooting method, which is applied to any of the above-mentioned intelligent engines. Fig. 9 is a flowchart of a troubleshooting method provided by an embodiment of the present invention, such as Fig. 9 As shown, the method includes:
[0110] Step 910: Run the model engine process to provide artificial intelligence services to the outside world, and send exception information to at least one Sidecar container when a target abnormal event is detected in the model engine process.
[0111] Step 920: When the abnormal information is received, target scene information corresponding to the target abnormal event corresponding to the engine main container is collected, and the target scene information is uploaded to the diagnosis emergency subsystem.
[0112] The obstacle clearing device provided by the present invention is described below. The obstacle clearing device described below and the obstacle clearing method described above can be referenced to each other.
[0113] The embodiment of the present invention further provides a troubleshooting device, which is applied to any of the above-mentioned intelligent engines. Fig.10 is a schematic diagram of the structure of the obstacle removal device provided by an embodiment of the present invention, such as Fig.10As shown, the obstacle removal device 1000 includes: a monitoring unit 1010 and a collection unit 1020, wherein:
[0114] The monitoring unit 1010 is used to run the model engine process, provide artificial intelligence services to the outside, and send abnormal information to at least one Sidecar container when a target abnormal event occurs in the model engine process;
[0115] The collection unit 1020 is used to collect the target scene information corresponding to the target abnormal event corresponding to the engine main container when receiving the abnormal information, and upload the target scene information to the diagnosis emergency subsystem.
[0116] Fig.11 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention, such as Fig.11 As shown, the electronic device may include: a processor 1110, a communication interface 1120, a memory 1130 and a communication bus 1140, wherein the processor 1110, the communication interface 1120 and the memory 1130 communicate with each other through the communication bus 1140. The processor 1110 may call the logic instructions in the memory 1130 to execute the troubleshooting method, which includes:
[0117] Run the model engine process, provide artificial intelligence services to the outside world, and send abnormal information to at least one Sidecar container when detecting a target abnormal event in the model engine process;
[0118] When the abnormal information is received, target scene information corresponding to the target abnormal event corresponding to the engine main container is collected, and the target scene information is uploaded to the diagnosis emergency subsystem.
[0119] In addition, the logic instructions in the above-mentioned memory 1130 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0120] On the other hand, the present invention further provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, when the computer program is executed by a processor, the computer can execute the troubleshooting method provided by the above methods, the method includes:
[0121] Run the model engine process, provide artificial intelligence services to the outside world, and send abnormal information to at least one Sidecar container when detecting a target abnormal event in the model engine process;
[0122] When the abnormal information is received, target scene information corresponding to the target abnormal event corresponding to the engine main container is collected, and the target scene information is uploaded to the diagnosis emergency subsystem.
[0123] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the troubleshooting method provided by the above methods is implemented, and the method includes:
[0124] Run the model engine process, provide artificial intelligence services to the outside world, and send abnormal information to at least one Sidecar container when detecting a target abnormal event in the model engine process;
[0125] When the abnormal information is received, target scene information corresponding to the target abnormal event corresponding to the engine main container is collected, and the target scene information is uploaded to the diagnosis emergency subsystem.
[0126] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0127] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0128] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An intelligent engine, characterized in that: It includes an engine main container and at least one Sidecar container, wherein the engine main container is communicatively connected to the at least one Sidecar container, and the at least one Sidecar container is communicatively connected to the diagnostic emergency subsystem, wherein: The engine main container is used to run the model engine process, provide artificial intelligence services to the outside, and send abnormal information to the at least one Sidecar container when a target abnormal event occurs in the model engine process; The at least one Sidecar container is used to collect target site information corresponding to the target abnormal event corresponding to the engine main container when receiving the abnormal information, and upload the target site information to the diagnosis emergency subsystem; The target abnormal events include different types of crash events, freeze events or timeout events during the operation of the model engine process; The engine main container includes an inference engine, a service framework and a monitoring module, wherein: The inference engine is communicatively connected with the service framework, and an artificial intelligence model is deployed in the inference engine; The service framework is used to run the artificial intelligence model in the inference engine based on the model engine process and provide artificial intelligence services to the outside world; The monitoring module is communicatively connected to the at least one Sidecar container, and is used to monitor the target abnormal state corresponding to the model engine process; determine the target abnormal event corresponding to the target abnormal state, and the abnormal information corresponding to the target abnormal event; and send the abnormal information to the at least one Sidecar container based on the domain socket communication method.
2. The intelligent engine according to claim 1, characterized in that: The monitoring module determines the target abnormal event corresponding to the target abnormal state, specifically for: In the case where the target abnormal state indicates that a target abnormal signal is captured, matching the target abnormal signal with a preset registration signal in the monitoring module, and in the case where a preset registration signal corresponding to the target abnormal signal is matched, determining the crash event type corresponding to the preset registration signal as the target abnormal event corresponding to the target abnormal state; When the target abnormal state indicates that the target abnormal signal has not been captured, and the number of request timeouts is greater than or equal to 1 and less than a preset threshold, determining that the target abnormal event corresponding to the target abnormal state is a timeout event; When the target abnormal state indicates that the target abnormal signal is not captured and the number of request timeouts is equal to the preset threshold, it is determined that the target abnormal event corresponding to the target abnormal state is a stuck event, and the crash event type corresponding to the preset registration signal does not include the timeout event and the stuck event.
3. The intelligent engine according to claim 1 or 2, characterized in that: The at least one Sidecar container is specifically used for: Based on the exception information, determine the target process number and the target exception event corresponding to the model engine process in the engine main container; the engine main container shares a process namespace with the at least one Sidecar container; Based on the target process number, collecting target scene information corresponding to the target abnormal event; Based on the target site information, alarm information is generated, and the alarm information is sent to the diagnosis emergency subsystem.
4. The intelligent engine according to claim 3, characterized in that: The at least one Sidecar container is further used for: Classifying the target site information to determine at least two categories of sub-target site information; Based on the preset mapping relationship, determine the target storage module corresponding to each type of sub-target site information; The preset mapping relationship includes a mapping relationship between the field information type and the storage module; For each type of sub-target on-site information, the sub-target on-site information is sent to the corresponding target storage module, and each target storage module is deployed in the diagnosis emergency subsystem.
5. An intelligent engine cluster, characterized in that: include: A cluster management executor and at least two intelligent engines according to any one of claims 1 to 4, each of the intelligent engines being communicatively connected to a diagnostic emergency subsystem and the cluster management executor, and the diagnostic emergency subsystem being communicatively connected to the cluster management executor, wherein: The cluster management executor is used to receive the target emergency operation instructions sent by the diagnostic emergency subsystem and execute the target emergency operation instructions to achieve troubleshooting and recovery of the target intelligent engine, to which the target intelligent engine belongs.
6. A distributed troubleshooting system, characterized in that: include: The diagnostic emergency subsystem and the intelligent engine cluster as claimed in claim 5, wherein: The diagnostic emergency subsystem is used to determine the target emergency operation instructions corresponding to the target intelligent engine based on the alarm information uploaded by the target intelligent engine in the intelligent engine cluster, and send the target emergency operation instructions to the intelligent engine cluster.
7. The distributed troubleshooting system according to claim 6, characterized in that: The diagnostic emergency subsystem is also used for: Analyze the warning information to determine whether the target abnormal event corresponding to the warning information meets the emergency condition; Based on the judgment result, a target emergency operation instruction corresponding to the warning information is generated.
8. The distributed troubleshooting system according to claim 6, characterized in that: The diagnostic emergency subsystem is also used for: In the case where the file type corresponding to the target field information is a minidump file, readable debugging information is generated based on the minidump file and the target symbol file corresponding to the target intelligent engine.
9. A troubleshooting method, characterized in that: The intelligent engine as claimed in any one of claims 1 to 4 comprises: Run the model engine process, provide artificial intelligence services to the outside world, and send abnormal information to at least one Sidecar container when detecting a target abnormal event in the model engine process; When the abnormal information is received, target scene information corresponding to the target abnormal event corresponding to the engine main container is collected, and the target scene information is uploaded to the diagnosis emergency subsystem.
Citation Information
Patent Citations
Message processing system, method and device of application container engine and storage medium
CN112131023A
Container security protection method and system for biological information container cloud
CN117040797A