Information Acquisition Method and Device Based on Process Configuration
Through the information collection method based on process configuration, process examples are generated using predefined business process files, and user-defined content is dynamically called, which solves the problems of low efficiency and poor flexibility in the development of information collection tools in the existing technology, and realizes efficient information collection and exception handling.
Patent Information
- Application Number
- CN202111480223.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-12-06
AI Technical Summary
The existing information collection solutions have problems such as high program coupling, low development efficiency, poor flexibility, low scalability and difficult to detect abnormal situations.
The information collection method based on process configuration is adopted, and process instances are generated by pre-defined business process files, and user-defined implementation content is dynamically called when the process instance is running, so as to realize the flow control of information collection tasks, reduce hard coding work, and improve development efficiency.
It has achieved the improvement of the development efficiency of information collection tools, reduced the hard coding work caused by business modification, improved the flexibility and scalability of the system, and simplified exception handling.
Smart Images

Figure CN114185535B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of software technology, and particularly relates to an information collection method and device based on process configuration. Background Art
[0002] Building a security asset management platform to achieve visual security management of the entire life cycle of assets and establishing a comprehensive and dynamic asset inventory database for enterprises is the basis of information security. An excellent asset information collection tool is a necessary condition for establishing a sound asset database.
[0003] Currently, the defects of information collection solutions are as follows: (1) High program coupling and low development efficiency, resulting in a huge workload for developing collection tools in actual applications; (2) Highly customized collection tools with poor flexibility and low scalability; (3) Difficult to troubleshoot abnormal situations.
[0004] For example, a multi-source heterogeneous data collection and aggregation system and method based on a power system disclosed in the invention patent application with the application number 202011314586.0. However, this solution is oriented to the power system scenario. By providing a configuration interface for multi-source heterogeneous data sources, heterogeneous data sources can be flexibly and quickly configured through the interface, simplifying the development work of data sources. It mainly focuses on the collection and aggregation of non-real-time data, log data, and Internet data, as well as the storage between different data sources, and does not involve information collection in operating systems and software applications. Summary of the Invention
[0005] The present invention aims to provide an information collection method, device, tool, and storage medium based on process configuration to reduce program coupling and improve the development efficiency of information collection tools.
[0006] The present invention solves the above technical problems through the following technical means:
[0007] On the one hand, an embodiment of the present invention provides an information collection method based on process configuration. A corresponding business process file is predefined for each information collection task. The method includes:
[0008] Obtain the information collection task;
[0009] Generate a process instance corresponding to the information collection task according to the business process file corresponding to the information collection task;
[0010] When the process instance runs, execute the process instance at each process node using an execution strategy corresponding to the type of the process node to obtain an information collection result.
[0011] By adopting the idea of workflow, the business process file corresponding to each information collection task is predefined, and the business process file provides business capabilities. When an information collection task is obtained, a task instance corresponding to the information collection task is generated using the corresponding business process file. When the task instance runs, the work nodes are dynamically transferred according to the running data, so as to achieve the collection of most general information only by modifying the business process file. When the task instance runs, the user-defined implementation content is dynamically called, decoupling the user behavior from the public behavior, maximizing the control over the running transfer of the collection task, greatly reducing the hard-coding work caused by adapting to business modifications, and improving the development efficiency of the information collection tool.
[0012] Further, the information of the predefined business process file includes process basic definition, process nodes, node execution command list, and target nodes;
[0013] The attributes of the process basic definition include process ID, version number, timeout duration, and node parameter list;
[0014] The attributes of the process nodes include node type, node ID, result parsing api path, execution command list, and target node list;
[0015] The attributes of the node execution command list include command ID, command content, and command execution type;
[0016] The attributes of the target nodes include target node ID and connection expression.
[0017] Further, the method further includes:
[0018] Preload the business process file, generate a process definition object and add it to the cache list;
[0019] Start a hot loading monitoring thread to monitor the business process file, and when the change of the business process file is monitored, hot load the business process file.
[0020] Further, the node types include start node, login type node, execution type node, gateway type node, sub-process type node, and end node;
[0021] The execution type nodes include loop nodes and ordinary nodes, the gateway type nodes include ordinary gateway nodes and exclusive gateway nodes, and the sub-process type nodes include login sub-process nodes, ordinary sub-process nodes, and parallel sub-process nodes.
[0022] Further, the execution strategy corresponding to the start node is to directly transfer to the next process node;
[0023] The execution policy corresponding to the login type node is to construct a target login link;
[0024] The execution policy corresponding to the gateway type node is to obtain the node parameter list to be executed next according to the establishment conditions of the target node;
[0025] The execution policy corresponding to the execution type node is to traverse the node execution command list or the node parameter list, generate an execution command, and call an executor according to the execution command to obtain an execution result;
[0026] The execution policy corresponding to the sub - process type node is to suspend the main process, create a sub - task instance corresponding to the sub - process node according to the process ID bound by the sub - process node, and wake up the main process when all the sub - task instances are executed;
[0027] The execution policy corresponding to the end node is to mark the task instance as the end state.
[0028] Further, before executing the process instance with the execution policy corresponding to the process node type at each process node, it further includes:
[0029] Determine whether the current process node is configured with a node execution pre - processor;
[0030] If so, call the node execution pre - processor to re - define the attributes of the current process node;
[0031] If not, execute the process instance with the execution policy corresponding to the current process node type.
[0032] Further, when obtaining the execution result by executing the process instance with the execution policy corresponding to the execution type node, it further includes:
[0033] According to the result parsing api path configured by the current node, the tool proxy calls the method corresponding to the result parsing api path to process the execution result and obtain a structured result that meets the business requirements.
[0034] Further, after obtaining the information collection task, it further includes:
[0035] Judge the running status of the information collection task;
[0036] When the category of the information collection task is a stop task, set the information collection task to the stop state;
[0037] When the category of the information collection task is an execution task, determine that the information collection task is a valid task;
[0038] Generate a process instance corresponding to the valid task according to the business process document corresponding to the valid task;
[0039] Add the process instance to the cache list and insert the valid task into the database.
[0040] Further, the method further includes:
[0041] Monitor the running time of the valid tasks in the database, and determine that there are timeout tasks and non-timeout tasks among the valid tasks;
[0042] Resume the non-timeout tasks and clear the timeout tasks.
[0043] Further, the attributes of the process node further include an execution exception handling policy, and the method further includes:
[0044] Determine that the execution command is abnormal;
[0045] Call the execution exception handling policy for processing.
[0046] Further, the method further includes:
[0047] According to the command content and the command execution type, circularly match the command execution implementation class of the corresponding type through the template mode, and call the executor to obtain the execution result;
[0048] Start a result reading thread. After the command ends, trigger a command result callback event to implement result callback.
[0049] In a second aspect, an information collection device based on process configuration provided by an embodiment of the present invention pre-defines a corresponding business process document for each information collection task. The device includes:
[0050] An acquisition module, configured to acquire the information collection task;
[0051] A generation module, configured to generate a process instance corresponding to the information collection task according to the business process document corresponding to the information collection task;
[0052] An execution module, configured to execute the process instance by using an execution policy corresponding to the process node type at each process node during the running of the process instance to obtain an information collection result.
[0053] In a third aspect, an information collection tool provided by an embodiment of the present invention includes the above-mentioned device.
[0054] Fourthly, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned method is implemented.
[0055] Compared with the prior art, the advantages of the present invention are as follows:
[0056] (1) By adopting the idea of workflow, the business process file corresponding to each information collection task is predefined in advance, and the business capabilities are provided by the business process file. When an information collection task is obtained, a task instance corresponding to the information collection task is generated by using the corresponding business process file. When the task instance runs, the work nodes are dynamically transferred according to the running data, so as to realize the collection of most general information only by modifying the business process file. When the task instance runs, the user-defined implementation content is dynamically called, decoupling the user behavior from the public behavior, maximizing the control over the running transfer of the collection task, greatly reducing the hard coding work caused by adapting to business modifications, and improving the development efficiency of the information collection tool.
[0057] (2) By setting up a hot loading monitoring thread, when it is monitored that the business process file has been modified, the changed business process file is hot loaded to adapt to different information collection tasks, with flexible application and strong scalability.
[0058] (3) By configuring a pre-processor for node execution, it can be used to modify the process definition, with flexible use, further improving the software scalability and ensuring the security of information collection.
[0059] (4) After obtaining the execution result of the information collection task, the result parsing process is called to parse the execution result, and data conforming to the user-defined data structure can be obtained.
[0060] (5) By providing a configured node execution exception handling strategy, when it is determined that the execution command is abnormal, the abnormal command is processed, avoiding the problem that it is difficult to troubleshoot problems after an exception in the past, and also controlling the impact of the exception on the task execution.
[0061] The additional aspects and advantages of the present invention will be partly given in the following description, partly become obvious from the following description, or be understood through the practice of the present invention. Description of the Drawings
[0062] Figure 1 is a flowchart of an information collection method based on process configuration in an embodiment of the present invention;
[0063] Figure 2 is an overall flowchart of an information collection method based on process configuration in an embodiment of the present invention;
[0064] Figure 3 This is the structural diagram of an information collection device based on process configuration in an embodiment of the present invention. Detailed implementation manners
[0065] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0066] As Figure 1 shown, this embodiment discloses an information collection method based on process configuration. A corresponding business process file is predefined for each information collection task. The method includes the following steps:
[0067] S10. Obtain the information collection task.
[0068] S120. Generate a process instance corresponding to the information collection task according to the business process file corresponding to the information collection task.
[0069] It should be noted that the information collection task obtained in this embodiment may be from a third-party platform. After obtaining the information collection task, the parameters are first verified. After the verification passes, the business process file corresponding to the information collection task is obtained, a process instance corresponding to the information collection task is generated, the process instance is added to the cache, the information collection task information is inserted into the embedded database, and then the execution starts from the start node of the process.
[0070] It should be understood that the parameter verification in this embodiment refers to the non-empty and format verification of the target IP and the task process code, and the illegal tasks are filtered out.
[0071] S30. When the process instance is running, execute the process instance at each process node by using an execution strategy corresponding to the process node type to obtain an information collection result.
[0072] It should be noted that in this embodiment, by adopting the idea of workflow, the business process file is predefined to provide business capabilities. When the task instance is running, the work nodes are dynamically transferred according to the running data, so as to realize the collection of most common information only by modifying the business process file, reduce the development workload of the information collection tool, and improve the development efficiency. When the task instance is running, the user-defined implementation content is dynamically called, decoupling the user behavior from the public behavior, maximizing the control of the running transfer of the collection task, greatly reducing the hard coding work caused by adapting to business modifications, being flexible in application and having strong scalability.
[0073] It should be understood that in this embodiment, the business logic of the information collection process can be intuitively seen from the process definition file, without coupling all the logic into the code. The process logic can be modified at low cost or even zero cost, without the need to refactor the code everywhere as in the related technologies when modifying some logic. By pre-defining the business process file to provide business capabilities, the coupling between modules such as step execution, result parsing, and result return is reduced, and the user only needs to focus on the coding related to the business.
[0074] In one embodiment, the information of the pre-defined business process file includes process basic definition (process), process node (node), node execution command list (cmdInfo), and target node (targetRef);
[0075] The attributes of the process basic definition include process ID, version number, timeout duration, and node parameter list;
[0076] The attributes of the process node include node type, node ID, result parsing api path, execution command list, target node list, whether cache connection is required, and node execution pre-processor name;
[0077] The attributes of the node execution command list include command ID, whether encrypted transmission is required, command content, and command execution type. Among them, the command content supports dynamic parsing of commands using springEL expressions, substituting the task instance and node execution result as parameters, and the command execution type includes sh, wmi, nmap, etc., which are not specifically limited in this embodiment;
[0078] The attributes of the target node include target node ID and connection expression. Among them, the connection expression can adopt SpringEL expression, and whether the node is valid can be obtained after substituting the task execution result variable.
[0079] It should be noted that in this embodiment, the user can flexibly define the information collection and result parsing processes, achieving easy and flexible expansion of business functions.
[0080] In one embodiment, the method further includes:
[0081] Pre-load the business process file, generate a process definition object and add it to the cache list;
[0082] Start a hot load monitoring thread to monitor the business process file, and when the change of the business process file is detected, hot load the business process file.
[0083] It should be noted that when starting the information collection task in this embodiment, the loading component is used to obtain the process definition loading implementation class through Java reflection. According to the servicesJob.procDefFilePath attribute in the configuration file, the business process files in this configuration folder are scanned and loaded, a process definition object is generated and added to the cache, and a hot loading monitoring thread is started. When a change in the business process file is detected, the changed business process file is hot loaded to implement the modification of the business process file.
[0084] In one embodiment, when starting the information collection task, the scanning component is used to scan the classes and methods marked with @ExecResolverMapping through Java reflection, preload and construct the result parsing process object, and add it to the in-memory cache list.
[0085] In one embodiment, the types of the process nodes include start nodes, login nodes, execution nodes, gateway nodes, sub-process nodes, and end nodes;
[0086] The execution nodes include loop nodes and ordinary nodes, the gateway nodes include ordinary gateway nodes and exclusive gateway nodes, and the sub-process nodes include login sub-process nodes, ordinary sub-process nodes, and parallel sub-process nodes.
[0087] In one embodiment, as Figure 2 shown, the execution strategies corresponding to each type of process node are respectively:
[0088] (1) The execution strategy corresponding to the login node is to construct a target login link.
[0089] It should be noted that the login link of the target object is constructed through the login node, and the login node can be reused to directly log in to the target object, saving time and reducing memory consumption.
[0090] (2) The execution strategy corresponding to the gateway node is to obtain the node parameter list for the next step according to the establishment condition of the target node.
[0091] Among them, the ordinary gateway node will substitute the connection expression into the task variable to judge whether the target node is established, and obtain the node parameter list for the next step based on the established target node; the exclusive gateway will obtain the node parameter list for the next step based on the first established target node.
[0092] (3) The execution strategy corresponding to the execution node is to traverse the node execution command list or the node parameter list, generate an execution command, and call the executor according to the execution command to obtain an execution result.
[0093] Among them, ordinary nodes will traverse the command list, substitute task variables to generate the final execution command, call the corresponding type of executor asynchronously, and uniformly call the user-defined parsing through reflection and then proceed to the next step after obtaining the result.
[0094] It should be noted that if fragmented return is configured, the tool will report the result to the upstream platform through the event mode.
[0095] Loop nodes will traverse the node parameter list, generate multiple groups of command lists to loop and call the command executor, and uniformly call the user-defined result parsing process through reflection after obtaining the result asynchronously.
[0096] (4) The execution strategy corresponding to the sub-process type node is to suspend the main process, create a sub-task instance corresponding to the sub-process node according to the process ID bound by the sub-process node, and wake up the main process when all the sub-task instances are executed.
[0097] (5) The execution strategy corresponding to the end node is to mark the task instance as the end state.
[0098] In one embodiment, before each process node executes the process instance using the execution strategy corresponding to the process node type, it further includes:
[0099] Judge whether the current process node is configured with a node execution pre-processor;
[0100] If so, call the node execution pre-processor to re-define the attributes of the current process node;
[0101] If not, execute the process instance using the execution strategy corresponding to the current process node type.
[0102] It should be noted that in this embodiment, by configuring the node execution pre-processor, it can be used to modify the process definition, which is flexible to use, improves the software scalability, and ensures the security of information collection. It provides a hook for business developers to re-define the process during operation and ensures flexibility.
[0103] It should be noted that the node execution pre-processor is called before the node execution strategy, and then the node execution strategy needs to be called normally. It can perform post-processing on the definition of the process node and only affects the current task process instance.
[0104] In one embodiment, when using the execution strategy corresponding to the execution type node to execute the process instance and obtain the execution result, it further includes:
[0105] Parse the result parsing API path according to the current node configuration, and the tool proxy calls the method corresponding to the result parsing API path to process the execution result to obtain a structured result that meets business requirements.
[0106] It should be noted that the result parsing process is user-defined content. After obtaining the execution result of the information collection task, calling the result parsing process to parse the execution result can obtain data that conforms to the user-defined data structure.
[0107] In one embodiment, after obtaining the information collection task, it further includes:
[0108] Judge the running status of the information collection task;
[0109] When the category of the information collection task is a stop task, set the information collection task to a stopped state;
[0110] When the category of the information collection task is an execution task, determine that the information collection task is a valid task;
[0111] Generate a process instance corresponding to the valid task according to the business process file corresponding to the valid task;
[0112] Add the process instance to the cache list and insert the valid task into the database.
[0113] It should be noted that by reasonably using the cache, the efficiency of command execution is improved. For example, in the past, it took 10 minutes to collect the asset information of a linux4A host, but in this embodiment, it only takes 1 minute and 30 seconds, and the effect is significantly improved.
[0114] In one embodiment, the method further includes:
[0115] Monitor the running time of the valid tasks in the database, and determine that there are timeout tasks and non-timeout tasks among the valid tasks;
[0116] Resume the non-timeout tasks and clear the timeout tasks.
[0117] It should be noted that in this embodiment, the monitoring component can be set to query the non-timeout tasks in the embedded database job_info table and restart the tasks; start a scheduled task to periodically scan for timeout tasks and clear the timeout tasks.
[0118] In one embodiment, the attributes of the process node further include an execution exception handling strategy, and the method further includes:
[0119] Determine that the execution command is abnormal;
[0120] Invoke the execution exception handling policy for processing.
[0121] It should be noted that in this embodiment, by providing a configured node execution exception handling policy, when it is determined that the execution command is abnormal, the supported exception handling policies include skipping the current node, ending the current task, recursively ending the task, or supporting the configuration of a custom exception handling policy. When an exception occurs, it will trigger an event to report the node exception information, and more reasonably monitor the task execution.
[0122] It should be noted that the exception handling policy in this embodiment can be default. By default, it does not affect the continuous execution of the process, and it is necessary to customize the node execution pre-processor.
[0123] In one embodiment, the method further includes:
[0124] According to the command content and the command execution type, circularly match the command execution implementation class of the corresponding type through the template mode, and call the executor to obtain the execution result;
[0125] Start a result reading thread. After the command ends, trigger a command result callback event to implement result callback.
[0126] It should be noted that in this embodiment during task execution, the task execution tool can be built with a command execution tool to provide command execution and result callback functions for commands combined with different execution types and protocol types; the command execution types provided by the tool are: shell, curl, nmap, hydra, win, python, etc., and the protocol types are: SH, winrm, powerSession, telnet. The template mode ensures the flexible extension of the execution and protocol types.
[0127] In one embodiment, after the information collection task ends, the connection cached by the executor will be closed, the cached task instance will be removed, the task information stored in the database will be deleted, and an event to report the task result will be triggered. Third-party platform users use the method of event listening to obtain the collection result of the task execution. For example, when the platform service starts, it registers events of a specified type to this software tool. After the task execution ends, the tool will trigger an event, and the platform will then monitor the result.
[0128] The following illustrates the solution of this embodiment through several specific examples:
[0129] Example 1: Login and collection of Linux host device information within the network
[0130] (1) In the first step, the node constructs a login connection object by means of the implementation class (connPostProcessor) of the pre-processor hook interface for custom login connection nodes. After the executor logs in, it caches the connection and executes the sh command to obtain system information. After obtaining the original execution result, the tool calls the custom parsing method configured by the node to obtain a data structure that meets business requirements, and the tool reports data according to the node configuration. If an execution exception occurs, the task is set to the 'abnormal end' state according to the configured exception policy, and the tool will process it regularly.
[0131] (2) The second step is a parallel node. Multiple nodes execute simultaneously to obtain the startup items, patches, processes, ports, and service component lists of the target machine, report data according to the node configuration, and then continue to the next step.
[0132] (3) The third step is a parallel gateway node. According to the service component information obtained in the previous step and the establishment conditions of different target nodes, a list of nodes to be executed in the next step is obtained. For example, when the target assets have service components such as mysql, Nginx, and Oracle, the target node connection expression of the gateway node will be parsed by the spring EL expression to obtain true / false. For example, the mysql target node connection expression is [#{#result.componentMaps.containsKey('mysql')}]. If the parsed result is true, the mysql node will be executed.
[0133] (4) The fourth step is the execution of parallel loop nodes. Multiple hit loop nodes will concurrently call the executor to obtain detailed information such as the runpath and version of the components. After all nodes in this step are executed, the execution of the next node continues.
[0134] (5) In the fifth step, a custom parsing method is called to encapsulate the component details obtained in the previous step to obtain a data structure that meets business requirements.
[0135] (6) This step is the end node, and the task ends, reporting the overall task data.
[0136] The example process, especially the possibility of expanding the acquisition of component information, has many interconnections between steps, and the tool provides good support for all of them.
[0137] Example 2: Collection of network topology information discovery based on the login of a specified network device within a network
[0138] (1) In the first step, the node constructs a login connection object by implementing the pre-processor hook interface (connPostProcessor) through a custom login connection node execution, executes commands to obtain system information, and obtains manufacturer information after calling parsing; when an exception occurs during execution, the task is short-circuited according to the configured exception strategy and ends abnormally.
[0139] (2) The second step is the exclusive gateway node. According to the manufacturer information obtained in the first step, there is only one target node, and the next step is continued. If the manufacturer information obtained in the first step is 'HuaWei', it is routed to the corresponding node after processing in this step.
[0140] (3) The third step is to obtain the arp table of a specific manufacturer, execute the corresponding command to obtain the original result, parse the arp table of the device, and report the data to the platform user according to the node configuration to continue to the next node.
[0141] (4) The fourth step is to exclude the gateway node. According to the switch parameters configured in the process definition and the arp table information in the third step, the target node is obtained after expression parsing (nmap detection, end)
[0142] (5) The fifth step is to loop the node, loop the IP in the arp list to perform nmap detection, and return the overall data to the platform user after the concurrent execution obtains the results.
[0143] Because this example involves many device manufacturers and requires cyclic dynamic parameters to detect detailed information, traditional backend coding is not conducive to expansion. The gateway and sub-process components provided by this tool provide good support for this requirement.
[0144] Example 3: Scanning Internet-exposed assets by specifying an IP range and port range
[0145] (1) The first step is the gateway node. According to the task parameters (fast scan + ip is ipV4), if the conditions are met, the target node 'masscan scan' is executed. If the conditions are not met, it is routed to the 'nmap scan' node.
[0146] (2) Masscan scans and detects the survival ports of the target IP through commands. After parsing the results, a port list is obtained and set in the process instance context.
[0147] (3) nmap scan, scan the target IP's surviving port list details through commands (the command is substituted into the process instance variable and parsed to obtain the final execution command).
[0148] (4) The fourth step is the gateway node. According to the switch parameters of the process configuration and the port list information scanned by nmap, the parallel sub-process node is executed to detect the web fingerprint information when the conditions are met; otherwise, the process ends.
[0149] (5) Parallel sub - process node for detecting port web fingerprint information. Suspend the main process. The node traverses the port information scanned by nmap, and uses the thread pool method to generate 8 tasks (configurable) at a time. Parallelly execute the sub - process tasks for fingerprint detection. After a batch of sub - process tasks are completed, check if there are tasks to be executed. Until all tasks are completed, wake up the main process and continue to the next node.
[0150] (6) End of the process. Call the result parsing method configured in the process to reconstruct the data structure. After obtaining the data in the business - compliant format, return it to the platform user.
[0151] Since the upper limit of the port information volume detected in this example is relatively high, the execution method of the traditional method is not conducive to expansion and traffic control. The parallel sub - processes of this tool and the pooling control of sub - process concurrency well solve the above problems.
[0152] In addition, referring to Figure 3 , an information collection device based on process - based configuration is also proposed in an embodiment of the present invention. For each information collection task, a corresponding business process file is predefined. The device includes:
[0153] An acquisition module 10 for acquiring the information collection task.
[0154] A generation module 20 for generating a process instance corresponding to the information collection task according to the business process file corresponding to the information collection task.
[0155] An execution module 30 for, when the process instance is running, executing the process instance at each process node using an execution strategy corresponding to the type of the process node to obtain an information collection result.
[0156] It should be noted that in this embodiment, by adopting the workflow idea, a business process file is predefined to provide business capabilities. When a task instance is running, the work nodes are dynamically transferred according to the running data, so as to achieve the collection of most general information only by modifying the business process file. Dynamically call the user - defined implementation content when the task instance is running, decouple the user behavior from the public behavior, maximize the control of the running transfer of the collection task, greatly reduce the hard - coding work caused by adapting to business modifications, and has flexible application and strong scalability.
[0157] It should be understood that in this embodiment, the business logic of the information collection process can be directly seen from the process definition file, without coupling all the logic into the code. The process logic can be modified at low cost or even zero cost, without the need to reconstruct the code everywhere as in the related technologies. By pre-defining the business process file to provide business capabilities, the coupling between modules such as step execution, result parsing, and result return is reduced, and the user only needs to focus on the coding related to the business. In one embodiment, the information of the pre-defined business process file includes process basic definition (process), process node (node), node execution command list (cmdInfo), and target node (targetRef);
[0158] The attributes of the process basic definition include process ID, version number, timeout duration, and node parameter list;
[0159] The attributes of the process node include node type, node ID, result parsing api path, execution command list, target node list, whether to cache the connection, and node execution pre-processor name;
[0160] The attributes of the node execution command list include command ID, whether to encrypt the transmission, command content, and command execution type. Among them, the command content supports dynamic parsing of commands using springEL expressions, substituting the task instance and node execution result as parameters. The command execution types include sh, wmi, nmap, etc., which are not specifically limited in this embodiment;
[0161] The attributes of the target node include target node ID and connection expression. Among them, the connection expression can adopt the SpringEL expression, and it is determined whether the node is valid after substituting the task execution result variable.
[0162] It should be noted that in this embodiment, the user can flexibly define the information collection and result parsing processes, achieving easy and flexible expansion of business functions.
[0163] In one embodiment, the device further includes:
[0164] A loading component, which is used to obtain the process definition loading implementation class through java reflection, scan and load the business process files in this configuration folder according to the servicesJob.procDefFilePath attribute in the configuration file, generate a process definition object and add it to the cache, and start a hot loading monitoring thread. When it is detected that the business process file has changed, the changed business process file is hot loaded to implement the modification of the business process file.
[0165] The scanning component is used to scan the classes and methods marked with @ExecResolverMapping through Java reflection, preload and construct the result parsing process object, and add it to the in-memory cache list.
[0166] It should be noted that by reasonably using the cache, the efficiency of command execution is improved. For example, in the past, it took 10 minutes to collect the asset information of a Linux 4A host, but in this embodiment, it only takes 1 minute and 30 seconds, and the effect is significantly improved.
[0167] In one embodiment, the types of the process nodes include start nodes, login nodes, execution nodes, gateway nodes, sub-process nodes, and end nodes;
[0168] The execution nodes include loop nodes and ordinary nodes, the gateway nodes include ordinary gateway nodes and exclusive gateway nodes, and the sub-process nodes include login sub-process nodes, ordinary sub-process nodes, and parallel sub-process nodes.
[0169] In one embodiment, as Figure 2 shown, the execution strategies corresponding to each type of process node are respectively:
[0170] (1) The execution strategy corresponding to the login node is to construct a target login link.
[0171] It should be noted that by constructing the login link of the target object through the login node, and the login node can be reused to directly log in to the target object, saving time and reducing memory consumption.
[0172] (2) The execution strategy corresponding to the gateway node is to obtain the list of node parameters to be executed next according to the establishment condition of the target node.
[0173] Among them, the ordinary gateway node will substitute the connection expression into the task variable to judge whether the target node is established, and obtain the list of node parameters to be executed next based on the established target node; the exclusive gateway will obtain the list of node parameters to be executed next based on the first established target node.
[0174] (3) The execution strategy corresponding to the execution node is to traverse the list of node execution commands or the list of node parameters, generate an execution command, and call the executor according to the execution command to obtain an execution result.
[0175] Among them, the ordinary node will traverse the command list, substitute the task variable to generate the final execution command, call the corresponding type of executor, asynchronously obtain the result, and continue the next step after uniformly calling the user-defined parsing through reflection.
[0176] It should be noted that if fragmented return is configured, the tool will report the results to the upstream platform through the event mode.
[0177] The loop node will traverse the node parameter list, generate multiple sets of command lists to loop and call the command executor, and after obtaining the results asynchronously, uniformly call the user-defined result parsing process through reflection.
[0178] (4) The execution strategy corresponding to the sub-process type node is to suspend the main process, create a sub-task instance corresponding to the sub-process node according to the process ID bound by the sub-process node, and wake up the main process when all the sub-task instances are executed.
[0179] (5) The execution strategy corresponding to the end node is to mark the task instance as the end state.
[0180] In one embodiment, if a node execution pre-processor is configured in the process node attributes, the node execution pre-processor is called to re-define the attributes of the current process node. In this embodiment, by configuring the node execution pre-processor, it can be used to modify the process definition, is flexible to use, improves the software scalability, and ensures the security of information collection.
[0181] In one embodiment, the device further includes:
[0182] A result parsing module, which is used to process the execution result by the tool proxy calling the method corresponding to the result parsing api path configured by the current node according to the result parsing api path configured by the current node, and obtain a structured result that meets the business requirements.
[0183] It should be noted that the parsing method is user-defined content. After obtaining the execution result of the information collection task, the result parsing method is called to parse the execution result, and data that meets the user-defined data structure can be obtained.
[0184] In one embodiment, the device further includes a judgment module, which is used for:
[0185] Judge the running status of the information collection task;
[0186] When the category of the information collection task is a stop task, set the information collection task to the stop state;
[0187] When the category of the information collection task is an execution task, determine that the information collection task is a valid task;
[0188] Generate a process instance corresponding to the valid task according to the business process file corresponding to the valid task;
[0189] Add the process instance to the cache list and insert the valid tasks into the database.
[0190] In one embodiment, the device includes a monitoring component for:
[0191] Monitor the running time of the valid tasks in the database, and determine that there are timeout tasks and non-timeout tasks among the valid tasks;
[0192] Resume the non-timeout tasks and clear the timeout tasks.
[0193] It should be noted that in this embodiment, the monitoring component can be set to query the non-timeout tasks in the embedded database job_info table and restart the tasks; start a scheduled task to periodically scan for timeout tasks and clear the timeout tasks.
[0194] In one embodiment, the attributes of the process node further include an execution exception handling policy, and the device further includes an exception handling module, specifically used for:
[0195] Determine that the execution command is abnormal;
[0196] Call the execution exception handling policy for processing.
[0197] It should be noted that in this embodiment, by providing a configurable node execution exception handling policy, when it is determined that the execution command is abnormal, the supported exception handling policies include skipping the current node, ending the current task, recursively ending the task, or supporting the configuration of a custom exception handling policy. When an exception occurs, it will trigger an event to report the node exception information, and more reasonably monitor the task execution.
[0198] It should be noted that the exception handling policy in this embodiment can be default. When it is default, in order not to affect the continuous execution of the process, it is necessary to customize the node execution pre-processor.
[0199] In one embodiment, a built-in command execution tool provides command execution and result callback functions for commands combined with different execution types and protocol types; the command execution types provided by the tool are: shell, curl, nmap, hydra, win, python, etc., and the protocol types are: SH, winrm, powerSession, telnet. The flexible expansion of the execution and protocol types is ensured through the template mode.
[0200] In addition, an embodiment of the present invention also proposes an information collection tool, including the device as described above.
[0201] It should be noted that the information collection of this tool uses the java cross-platform programming language, which can be deployed separately as a service module or developed for horizontal expansion as a dependency, aiming to solve the existing pain points and provide a better tool.
[0202] In addition, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-described method is implemented.
[0203] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.
[0204] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0205] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples.
[0206] In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0207] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. An information collection method based on process configuration, characterized in that, Pre - define a corresponding business process file for each information collection task. The method includes: Obtain the information collection task; Generate a process instance corresponding to the information collection task according to the business process file corresponding to the information collection task; When the process instance is running, execute the process instance at each process node using an execution strategy corresponding to the type of the process node to obtain an information collection result; The information of the pre - defined business process file includes process basic definition, process nodes, a list of node execution commands, and target nodes. The process basic definition includes a list of node parameters. The node types include start nodes, login - type nodes, execution - type nodes, gateway - type nodes, sub - process - type nodes, and end nodes; The execution strategy corresponding to the gateway - type node is to obtain the list of node parameters for the next execution according to the establishment condition of the target node; The execution strategy corresponding to the execution - type node is to traverse the list of node execution commands or the list of node parameters to generate an execution command and call an executor according to the execution command to obtain an execution result; The execution strategy corresponding to the sub - process - type node is to suspend the main process, create a sub - task instance corresponding to the sub - process - type node according to the process ID bound to the sub - process - type node, and wake up the main process when all the sub - task instances are executed; Before executing the process instance at each process node using an execution strategy corresponding to the type of the process node, it further includes: Judge whether a node execution pre - processor is configured for the current process node; If so, call the node execution pre - processor to re - define the attributes of the current process node; If not, execute the process instance using an execution strategy corresponding to the type of the current process node.
2. The information collection method based on process configuration according to claim 1, wherein The method further includes: Pre - load the business process file, generate a process definition object and add it to the cache list; Start a hot - load monitoring thread to monitor the business process file, and when the change of the business process file is detected, hot - load the business process file.
3. The information collection method based on process configuration according to claim 1, characterized in that The attribute information of the process node includes the node type; The execution - type nodes include loop nodes and ordinary nodes. The gateway - type nodes include ordinary gateway nodes and exclusive gateway nodes. The sub - process - type nodes include login sub - process nodes, ordinary sub - process nodes, and parallel sub - process nodes; Among them, the execution strategy corresponding to the start node is to directly transfer to the next process node; The execution strategy corresponding to the login - type node is to construct a target login link; The execution strategy corresponding to the end node is to mark the task instance as the end state.
4. The information collection method based on process configuration according to claim 3, wherein The attribute of the process node further includes a result parsing API path. When using the execution strategy corresponding to the execution - type node to execute the process instance and obtain the execution result, it further includes: According to the result parsing API path configured for the current node, the tool proxy calls the method corresponding to the result parsing API path to process the execution result to obtain a structured result meeting business requirements.
5. The information collection method based on process configuration according to claim 1, wherein After obtaining the information collection task, it further includes: Judge the running status of the information collection task; When the category of the information collection task is a stop task, set the information collection task to the stop state; When the category of the information collection task is an execution task, determine that the information collection task is a valid task; Generate a process instance corresponding to the valid task according to the business process file corresponding to the valid task; Add the process instance to the cache list and insert the valid task into the database.
6. The information collection method based on process configuration according to claim 5, characterized in that, The method further includes: Monitor the running time of the valid tasks in the database and determine whether there are timeout tasks and non-timeout tasks among the valid tasks; If there are non-timeout tasks, resume the non-timeout tasks; If there are timeout tasks, clear the timeout tasks.
7. The information collection method based on process configuration according to claim 3, wherein The attributes of the process node further include an execution exception handling strategy, and the method further includes: Determine that the execution command is abnormal; Call the execution exception handling strategy for processing.
8. The information collection method based on process configuration according to claim 3, wherein, The method further includes: According to the content of the command and the execution type of the command, circularly match the command execution implementation class of the corresponding type through the template mode, and call the executor to obtain the execution result; Start a result reading thread. After the command ends, trigger a command result callback event to implement result callback.
9. An information collection device based on process configuration, characterized in that, For each information collection task, a corresponding business process file is predefined. The device includes: An acquisition module for acquiring the information collection task; A generation module for generating a process instance corresponding to the information collection task according to the business process file corresponding to the information collection task; An execution module for, when the process instance is running, executing the process instance at each process node by using an execution strategy corresponding to the type of the process node to obtain an information collection result; The information of the predefined business process file includes process basic definition, process nodes, a node execution command list, and target nodes. The process basic definition includes a node parameter list. The node types include start nodes, login type nodes, execution type nodes, gateway type nodes, sub-process type nodes, and end nodes; The execution strategy corresponding to the gateway type node is to obtain the next node parameter list to be executed according to the establishment condition of the target node; The execution strategy corresponding to the execution type node is to traverse the node execution command list or the node parameter list, generate an execution command, and call the executor according to the execution command to obtain an execution result; The execution strategy corresponding to the sub-process type node is to suspend the main process, create a sub-task instance corresponding to the sub-process type node according to the process ID bound by the sub-process type node, and wake up the main process when all the sub-task instances are executed; Before executing the process instance at each process node by using an execution strategy corresponding to the type of the process node, it further includes: Judge whether the current process node is configured with a node execution pre-processor; If so, call the node execution pre-processor to re-define the attributes of the current process node; If not, execute the process instance by using an execution strategy corresponding to the type of the current process node.
Citation Information
Patent Citations
Multi-source heterogeneous data acquisition and aggregation system and method based on power system
CN112433998A
Automatic service arrangement method and device based on parameter driving
CN110912724A