Method, apparatus, electronic device and computer readable medium for generating a treatment plan
Patent Information
- Application Number
- CN202410505046.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-25
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-04-25
AI Technical Summary
[0008]有鉴于此,本发明实施例提供一种生成处置方案的方法、装置、电子设备和计算机可读介质,以解决安全事件的分派效率和处置准确率下降的技术问题
[0057]上述发明中的一个实施例具有如下优点或有益效果:因为采用对多个数据源分别进行预处理,从而得到多个数据源对应的预处理文本,再对预处理文本进行切片,将切片和问题输入到对话式模型中,根据输出的回答进行聚类和排序,从而得到每个节点对应的目标答案的技术手段,所以克服了现有技术中安全事件的分派效率和处置准确率下降的技术问题。本发明实施例通过将多个数据源对应的切片文本作为对话式模型的输入,并对输出的结果进行聚类和排序,有助于提升安全事件的分派效率,从而提高安全事件的处置效率和处置准确率。
Smart Images

Figure CN118410166B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information security technology, and in particular to a method, apparatus, electronic device, and computer-readable medium for generating disposal solutions. Background Technology
[0002] With the continuous development of internet technology, malicious activities hidden within large amounts of normal network traffic are becoming increasingly difficult to detect. Hackers exploit high-risk system vulnerabilities and network attack methods to illegally obtain corporate data or cause system paralysis. Therefore, the rapid and effective analysis and handling of threats has become a crucial issue that urgently needs to be addressed in the field of information security.
[0003] In the process of realizing this invention, the inventors discovered at least the following problems in the prior art:
[0004] 1) As attack methods continue to evolve and the scale and complexity of data continue to increase, the efficiency of security incident dispatch and the accuracy of handling are declining. Traditional security support services can no longer meet the needs of security operations personnel in analyzing and handling security incidents.
[0005] 2) Automated incident dispatch and response to security incidents relies on built-in rules, which cannot cover diverse attack methods. Currently, standardized processes are formed through predefined scripts to automate the dispatch and response to different types of security incidents. If a security incident occurs that the rules cannot recognize, it will be impossible to respond quickly through automated dispatch, leading to missed optimal handling time and impacting handling efficiency.
[0006] 3) Operations personnel require extensive training in security operations tasks such as semantic understanding of alarm logs, incident analysis, and incident handling, resulting in high operational costs. Because alarm logs lack a unified format and semantic standard, operations personnel can only consult built-in knowledge bases and apply fixed log parsing rules. When encountering logs that cannot be matched, it is difficult to write parsing scripts independently. The same applies to incident analysis and handling; these tasks require significant experience and necessitate long-term training for security operations personnel.
[0007] 4) The security knowledge base lacks interactive response capabilities, resulting in insufficient security support services. The security knowledge base can only perform approximate answer searches. When handling complex incidents, operations and maintenance personnel may encounter situations requiring coordination of support services, leading to slow security incident response times. Summary of the Invention
[0008] In view of this, embodiments of the present invention provide a method, apparatus, electronic device, and computer-readable medium for generating disposal plans to address the technical problem of declining dispatch efficiency and disposal accuracy of security incidents.
[0009] To achieve the above objectives, according to one aspect of the present invention, a method for generating a disposal plan is provided, comprising:
[0010] Multiple data sources are collected, and each of the multiple data sources is preprocessed to obtain preprocessed text corresponding to the multiple data sources; wherein, the multiple data sources include alarm logs, voice data and image data;
[0011] For each data source, the preprocessed text is sliced to obtain multiple sliced texts corresponding to the data source.
[0012] For each node, the multiple slices of text corresponding to the multiple data sources and the question corresponding to the node are sequentially input into the conversational model to output multiple answers corresponding to the node; the multiple answers corresponding to the node are clustered to obtain each cluster; and the target cluster containing the most answers is selected from each cluster. The answers in the target cluster are sorted according to their frequency of occurrence, and the answer with the most frequency of occurrence is selected as the target answer corresponding to the node.
[0013] The target answers corresponding to each node are merged according to the order of the nodes to obtain the processing solution.
[0014] Optionally, the multiple data sources are preprocessed separately to obtain preprocessed text corresponding to the multiple data sources, including:
[0015] If the data source is an alarm log, then regular expressions are used to extract key information from each log in the data source, thereby obtaining the preprocessed text corresponding to the data source.
[0016] If the data source is voice data, an acoustic model is used to identify the data source to obtain the identification result, and then a text extraction model is used to extract the identification result to obtain the preprocessed text corresponding to the data source.
[0017] If the data source is image data, then an object detection model is used to extract the object detection region from the data source, the object detection region is cropped from the data source, and a text recognition model is used to recognize the text data in the object detection region, thereby obtaining the preprocessed text corresponding to the data source.
[0018] Optionally, the voice data may be voice data collected during a meeting triggered by the alarm log and / or call data triggered by the alarm log; the image data may be at least one of the following: image data of the operation and maintenance interface triggered by the alarm log, screenshot data of the system alarm corresponding to the alarm log, and screenshot data of chat history triggered by the alarm log.
[0019] Optionally, for each data source corresponding to preprocessed text, the preprocessed text is sliced to obtain multiple sliced texts corresponding to the data source, including:
[0020] If the data source is an alarm log, then each key piece of information extracted from the data source will be used as a slice of text corresponding to the data source, thereby obtaining multiple slices of text corresponding to the data source.
[0021] If the data source is voice data, the preprocessed text is sliced according to a preset file size to obtain multiple sliced texts corresponding to the data source.
[0022] If the data source is image data, then the preprocessed text corresponding to each image data is used as a slice text corresponding to the data source, thereby obtaining multiple slice texts corresponding to the data source.
[0023] Optionally, the multiple slices of text corresponding to the multiple data sources and the question corresponding to the node are sequentially input into the conversational model, thereby outputting multiple answers corresponding to the node, including:
[0024] For each data source, the multiple slices of text corresponding to the data source and the question corresponding to the node are sequentially input into the conversational model, thereby outputting multiple answers corresponding to the data source.
[0025] Optionally, multiple slices of text corresponding to the data source and the question corresponding to the node are sequentially input into the conversational model to output multiple answers corresponding to the data source, including:
[0026] For each slice of text corresponding to the data source, the slice of text and the question corresponding to the node are input into the conversational model together, thereby outputting the answer corresponding to the slice of text.
[0027] Optionally, the answers in the target cluster are sorted according to their frequency of occurrence, and the answer with the highest frequency is selected as the target answer corresponding to the node, including:
[0028] For each answer in the target cluster, the answers corresponding to multiple data sources are weighted and summed based on the weight corresponding to each data source to obtain the election value of the answer;
[0029] Sort the answers according to their election values, and select the answer with the highest election value as the target answer for the node.
[0030] Additionally, according to another aspect of the present invention, an apparatus for generating a treatment plan is provided, comprising:
[0031] The preprocessing module is used to collect data from multiple data sources and preprocess each of the multiple data sources to obtain preprocessed text corresponding to the multiple data sources; wherein, the multiple data sources include alarm logs, voice data and image data;
[0032] The slicing module is used to slice the preprocessed text corresponding to each data source, thereby obtaining multiple sliced texts corresponding to the data source.
[0033] The processing module is used to sequentially input multiple slices of text corresponding to the multiple data sources and the question corresponding to the node into the conversational model for each node, thereby outputting multiple answers corresponding to the node; clustering the multiple answers corresponding to the node to obtain various clusters; and selecting the target cluster containing the most answers from the various clusters, sorting the answers in the target cluster according to their occurrence frequency, and selecting the answer with the most occurrence frequency as the target answer corresponding to the node.
[0034] The merging module is used to merge the target answers corresponding to each node according to the order of each node, so as to obtain a solution.
[0035] Optionally, the preprocessing module is further configured to:
[0036] If the data source is an alarm log, then regular expressions are used to extract key information from each log in the data source, thereby obtaining the preprocessed text corresponding to the data source.
[0037] If the data source is voice data, an acoustic model is used to identify the data source to obtain the identification result, and then a text extraction model is used to extract the identification result to obtain the preprocessed text corresponding to the data source.
[0038] If the data source is image data, then an object detection model is used to extract the object detection region from the data source, the object detection region is cropped from the data source, and a text recognition model is used to recognize the text data in the object detection region, thereby obtaining the preprocessed text corresponding to the data source.
[0039] Optionally, the voice data may be voice data collected during a meeting triggered by the alarm log and / or call data triggered by the alarm log; the image data may be at least one of the following: image data of the operation and maintenance interface triggered by the alarm log, screenshot data of the system alarm corresponding to the alarm log, and screenshot data of chat history triggered by the alarm log.
[0040] Optionally, the slicing module is further configured to:
[0041] If the data source is an alarm log, then each key piece of information extracted from the data source will be used as a slice of text corresponding to the data source, thereby obtaining multiple slices of text corresponding to the data source.
[0042] If the data source is voice data, the preprocessed text is sliced according to a preset file size to obtain multiple sliced texts corresponding to the data source.
[0043] If the data source is image data, then the preprocessed text corresponding to each image data is used as a slice text corresponding to the data source, thereby obtaining multiple slice texts corresponding to the data source.
[0044] Optionally, the processing module is further configured to:
[0045] For each data source, the multiple slices of text corresponding to the data source and the question corresponding to the node are sequentially input into the conversational model, thereby outputting multiple answers corresponding to the data source.
[0046] Optionally, the processing module is further configured to:
[0047] For each slice of text corresponding to the data source, the slice of text and the question corresponding to the node are input into the conversational model together, thereby outputting the answer corresponding to the slice of text.
[0048] Optionally, the processing module is further configured to:
[0049] For each answer in the target cluster, the answers corresponding to multiple data sources are weighted and summed based on the weight corresponding to each data source to obtain the election value of the answer;
[0050] Sort the answers according to their election values, and select the answer with the highest election value as the target answer for the node.
[0051] According to another aspect of the present invention, an electronic device is also provided, comprising:
[0052] One or more processors;
[0053] Storage device for storing one or more programs.
[0054] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the above embodiments.
[0055] According to another aspect of the present invention, a computer-readable medium is also provided, on which a computer program is stored, which, when executed by a processor, implements the methods described in any of the above embodiments.
[0056] According to another aspect of the present invention, a computer program product is also provided, including a computer program that, when executed by a processor, implements the methods described in any of the above embodiments.
[0057] One embodiment of the above invention has the following advantages or beneficial effects: By preprocessing multiple data sources separately to obtain preprocessed text corresponding to multiple data sources, then slicing the preprocessed text, and inputting the slices and questions into a conversational model, clustering and sorting the output answers to obtain the target answer for each node, the technical problem of decreased efficiency in security event dispatch and handling accuracy in existing technologies is overcome. This invention, by using sliced text corresponding to multiple data sources as input to a conversational model and clustering and sorting the output results, helps to improve the efficiency of security event dispatch, thereby improving the efficiency and accuracy of security event handling.
[0058] The further effects of the aforementioned unconventional alternative methods will be explained below in conjunction with specific implementation methods. Attached Figure Description
[0059] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0060] Figure 1 This is a flowchart of a method for generating a disposal scheme according to an embodiment of the present invention;
[0061] Figure 2 This is a flowchart of the automated orchestration component and the conversational model according to an embodiment of the present invention;
[0062] Figure 3 This is a flowchart of a method for generating a disposal scheme according to a possible embodiment of the present invention;
[0063] Figure 4 This is a flowchart illustrating the preprocessing of various data sources according to an embodiment of the present invention;
[0064] Figure 5 This is a schematic diagram of an apparatus for generating a processing solution according to an embodiment of the present invention;
[0065] Figure 6 This is an exemplary system architecture diagram in which embodiments of the present invention can be applied;
[0066] Figure 7 This is a schematic diagram of the structure of a computer system suitable for implementing terminal devices or servers of the present invention. Detailed Implementation
[0067] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0068] It should be noted that the collection, analysis, use, transmission, and storage of user personal information involved in the technical solution of this invention all comply with relevant laws and regulations, are used for legitimate and reasonable purposes, and are not shared, disclosed, or sold outside of these legitimate uses, and are subject to supervision and management by regulatory authorities. Necessary measures should be taken to prevent unauthorized access to such personal information data, ensure that personnel authorized to access personal information data comply with relevant laws and regulations, and ensure the security of user personal information. Once this user personal information data is no longer needed, the risk should be minimized by restricting or even prohibiting data collection and / or deleting the data.
[0069] When applicable, including in certain relevant applications, data deidentification is used to protect user privacy, such as by removing specific identifiers (e.g., name, account, gender, date of birth, etc.), controlling the amount or specificity of stored data, controlling how data is stored, and / or other methods of deidentification.
[0070] The acquisition, transmission, storage, use, and processing of data in this application comply with relevant national laws and regulations. It should be noted that certain software, components, models, and other existing industry solutions may be mentioned in the embodiments of this application. These should be considered exemplary, intended only to illustrate the feasibility of implementing the technical solution of this application, and do not imply that the applicant has already used or necessarily used such solutions.
[0071] Figure 1 This is a flowchart of a method for generating a disposal scheme according to an embodiment of the present invention. As one embodiment of the present invention, such as... Figure 1 As shown, the method for generating a disposal plan may include:
[0072] Step 101: Collect multiple data sources and preprocess each of the multiple data sources to obtain preprocessed text corresponding to the multiple data sources; wherein, the multiple data sources include alarm logs, voice data and image data.
[0073] First, multiple different data sources are collected, such as alarm logs, voice data, and image data. Optionally, the voice data includes voice data collected during a meeting triggered by the alarm log and / or call data triggered by the alarm log. Optionally, the image data includes at least one of the following: image data of the operation and maintenance interface triggered by the alarm log, screenshot data of system alarms corresponding to the alarm log, and screenshot data of chat records triggered by the alarm log. Then, the multiple data sources are preprocessed separately, such as extracting keyword information from the alarm logs, converting voice data into text data, extracting text data from the image data, etc., to obtain the preprocessed text corresponding to each data source.
[0074] Step 102: For each data source, the preprocessed text is sliced to obtain multiple sliced texts corresponding to the data source.
[0075] In this step, for each preprocessed text corresponding to each data source obtained in step 101, each preprocessed text is sliced to obtain multiple sliced texts corresponding to each data source.
[0076] Optionally, if the data source is an alarm log, each key piece of information extracted from the data source is used as a slice of text corresponding to the data source, thereby obtaining multiple slices of text corresponding to the data source. If the data source is an alarm log, each key piece of information extracted from the alarm log is used as a slice of text, thereby obtaining multiple slices of text corresponding to the alarm log.
[0077] Optionally, if the data source is voice data, the preprocessed text is sliced according to a preset file size to obtain multiple sliced texts corresponding to the data source. If the data source is voice data, the text data corresponding to the voice data is sliced according to a preset file size (e.g., 1KB, 2KB, or 5KB) to obtain multiple sliced texts corresponding to the voice data.
[0078] Optionally, if the data source is image data, then the preprocessed text corresponding to each image data is used as a slice text corresponding to the data source, thereby obtaining multiple slice texts corresponding to the data source. If the data source is image data, then the text data corresponding to each image data is used as a slice text, thereby obtaining multiple slice texts corresponding to the image data.
[0079] Step 103: For each node, input the multiple slice texts corresponding to the multiple data sources and the question corresponding to the node into the conversational model in sequence to output the multiple answers corresponding to the node; cluster the multiple answers corresponding to the node to obtain each cluster; and select the target cluster containing the most answers from each cluster, sort the answers in the target cluster according to the frequency of occurrence, and select the answer with the most occurrences as the target answer corresponding to the node.
[0080] Slicing the preprocessed text can improve the accuracy of the conversational model's responses, thereby ensuring the accuracy of the final solution.
[0081] This invention extends from single alarm log data to multiple types of voice data and multiple types of image data. By using multiple types of data sources as input to the conversational model according to each slice, the accuracy of the conversational model's response can be improved, thereby ensuring the accuracy of the final handling solution.
[0082] Optionally, the multiple slices of text corresponding to the multiple data sources and the question corresponding to the node are sequentially input into the conversational model to output multiple answers corresponding to the node. This includes: for each data source, sequentially inputting the multiple slices of text corresponding to the data source and the question corresponding to the node into the conversational model to output multiple answers corresponding to the data source. Figure 2 As shown, the automated orchestration component interacts with the conversational model. Taking the attack detection node as an example, the automated orchestration component inputs multiple slice texts corresponding to multiple data sources obtained in step 102 and the question corresponding to the node (such as whether the alarm log contains attack behavior) into the conversational model. The conversational model outputs the answer corresponding to each slice text of each data source. For example, one answer output by the conversational model is: "The source address is a black market address, and the log result is successful, indicating an attack behavior exists." After receiving the answer output by the conversational model, the automated orchestration component first clusters these answers to obtain various clusters. Then, it selects the target cluster containing the most answers from each cluster, sorts the answers in the target cluster according to their frequency of occurrence, and finally selects the answer with the most occurrences as the target answer corresponding to the node.
[0083] Similarly, taking an emergency response node as an example, the automated orchestration component inputs multiple slice texts corresponding to multiple data sources obtained in step 102 and the question corresponding to the node (such as who should handle the alarm log) into the conversational model. The conversational model outputs the answer corresponding to each slice text of each data source. After receiving the answer output by the conversational model, the automated orchestration component first clusters these answers to obtain various clusters. Then, it selects the target cluster containing the most answers from each cluster. Next, it sorts the answers in the target cluster according to the frequency of occurrence. Finally, it selects the answer with the most occurrences as the target answer corresponding to the node.
[0084] Similarly, taking the node with the handling suggestion as an example, the automated orchestration component inputs multiple slice texts corresponding to multiple data sources obtained in step 102 and the question corresponding to the node (such as how to handle this alarm log) into the conversational model. The conversational model outputs the answer corresponding to each slice text of each data source. After receiving the answer output by the conversational model, the automated orchestration component first clusters these answers to obtain each cluster. Then, it selects the target cluster containing the most answers from each cluster. Next, it sorts the answers in the target cluster according to the frequency of occurrence. Finally, it selects the answer with the most occurrences as the target answer corresponding to the node.
[0085] Similarly, taking a suggested node in the process as an example, the automated orchestration component inputs multiple slice texts corresponding to multiple data sources obtained in step 102 and the question corresponding to the node (e.g., how should the system be optimized for this alarm log) into the conversational model. The conversational model outputs the answer corresponding to each slice text of each data source. For example, one answer output by the conversational model is: "Block the source address IP and log in to the target server to delete malicious files." After receiving the answer output by the conversational model, the automated orchestration component first clusters these answers to obtain various clusters. Then, it selects the target cluster containing the most answers from each cluster, sorts the answers in the target cluster according to their frequency of occurrence, and finally selects the answer with the most occurrences as the target answer corresponding to the node.
[0086] Optionally, multiple slice texts corresponding to the data source and the questions corresponding to the nodes are sequentially input into the conversational model to output multiple answers corresponding to the data source. This includes: for each slice text corresponding to the data source, inputting the slice text and the question corresponding to the node together into the conversational model to output the answer corresponding to the slice text. To ensure the accuracy of the answers output by the conversational model, each slice text and the question corresponding to that node are input into the conversational model together, and the conversational model outputs the answer corresponding to that slice text. Therefore, each slice text corresponds to one answer, and each node will have multiple answers. Then, for each node, the multiple answers corresponding to that node are clustered and sorted to obtain the target answer corresponding to that node, such as... Figure 2 As shown.
[0087] Therefore, this embodiment of the invention employs a cyclical questioning dialogic model to obtain answers for each node. Furthermore, each node can ask questions sequentially, which helps improve the accuracy of the dialogic model's responses, thereby ensuring the accuracy of the final solution. This embodiment of the invention leverages the contextual support of dialogic models, enabling more accurate question searching and solving the problem that security knowledge bases can only perform approximate answer searches.
[0088] Optionally, step 103 may include: for each answer in the target cluster, performing a weighted summation of the answers corresponding to multiple data sources based on the weights corresponding to each data source, thereby obtaining the election value of the answer; sorting the answers according to the election values, and selecting the answer with the largest election value as the target answer corresponding to the node. For each node, since each data source corresponds to multiple slice files with one answer, weights can be pre-configured for each data source. For each node, a weighted summation of the answers is performed. For example, if an answer appears four times, the election value of that answer = the weight of the first data source + the weight of the second data source + the weight of the second data source + the weight of the third data source; or, if an answer appears five times, the election value of that answer = the weight of the first data source + the weight of the first data source + the weight of the second data source + the weight of the second data source, and so on. Finally, the election values of the answers are sorted, and the answer with the largest election value is selected as the target answer for that node.
[0089] Step 104: Merge the target answers corresponding to each node according to the order of each node to obtain the processing solution.
[0090] like Figure 2As shown, the target answers corresponding to each node are sequentially pieced together according to the order of each node to obtain the final handling plan. Finally, the handling plan is assigned to help improve the efficiency and accuracy of security incident assignment.
[0091] Because it uses more data sources than existing technologies, the conversational model outputs responses that are closer to the actual situation, and the process of generating a response plan is seamless for the user; for example, the response plan is already issued after an emergency meeting. Compared to existing technologies, this invention not only generates accurate response plans but also significantly shortens alarm log response time, improves the quality of security support services, and reduces security operation costs. Therefore, this invention can automate alarm analysis, automate event dispatch, and provide real-time and professional response advice to security personnel, thereby improving the efficiency of threat analysis and security incident handling, as well as the problem resolution rate, in security operations.
[0092] Based on the various embodiments described above, it can be seen that the embodiments of the present invention solve the technical problem of decreased efficiency in security event dispatch and handling accuracy in the prior art by preprocessing multiple data sources separately to obtain preprocessed text corresponding to multiple data sources, then slicing the preprocessed text, inputting the slices and questions into a conversational model, and clustering and sorting the output answers to obtain the target answer corresponding to each node. The embodiments of the present invention, by using sliced text corresponding to multiple data sources as input to a conversational model and clustering and sorting the output results, help improve the efficiency in security event dispatch, thereby improving the efficiency and accuracy in handling security events.
[0093] Figure 3 This is a flowchart of a method for generating a processing solution according to a possible embodiment of the present invention. As another embodiment of the present invention, such as... Figure 3 As shown, the method for generating a disposal plan may include:
[0094] Step 301: Collect data from multiple data sources, including alarm logs, voice data, and image data. Specifically, the voice data includes voice data collected during a meeting triggered by the alarm log and / or call data triggered by the alarm log; the image data includes at least one of the following: image data of the operation and maintenance interface triggered by the alarm log, screenshot data of system alarms corresponding to the alarm log, and screenshot data of chat logs triggered by the alarm log.
[0095] If the data source is alarm logs, proceed to step 302; if the data source is voice data, proceed to step 303; if the data source is image data, proceed to step 304.
[0096] Step 302: Use regular expressions to extract key information from each log in the data source, thereby obtaining the preprocessed text corresponding to the data source.
[0097] Step 303: The data source is identified using an acoustic model to obtain the identification result. Then, the identification result is extracted using a text extraction model to obtain the preprocessed text corresponding to the data source.
[0098] Step 304: Extract the target detection region from the data source using a target detection model, crop the target detection region from the data source, and use a text recognition model to identify the text data in the target detection region, thereby obtaining the preprocessed text corresponding to the data source.
[0099] In the data preprocessing process, depending on the type of data source (log data, voice data, image data), there are three different processing methods. For example... Figure 4 As shown, if the data source is an alarm log, that is, log data, then regular expressions are used to extract the key information of each log in the alarm log. It can be equipped with extraction expressions such as JSON extraction, keyword extraction, error text extraction, and extraction expressions corresponding to the data structure of the security product to extract the key information of the log as the preprocessed text of the data source.
[0100] like Figure 4 As shown, if the data source is speech data, preprocessed text is obtained through an acoustic model and a text extraction model. Optionally, the acoustic model uses HMM+DNN, and the text extraction model uses RNN. After obtaining the preprocessed text, if the text volume is large, text slicing will be performed to ensure the accuracy of the subsequent output of the conversational model.
[0101] like Figure 4 As shown, if the data source is image data, the YOLO object detection model is used to detect the target detection region in the image, and then the target detection region is cropped from the image. The cropped target detection region is then input into a text recognition model (such as the TrOcr model), which recognizes the text data in the image, thereby obtaining the preprocessed text of the data source.
[0102] In embodiments of the present invention, such as Figure 4As shown, after steps 306 and 307, text correction can also be performed. Regardless of whether text data is extracted from speech data or images, the output text may contain errors. Therefore, an additional text correction model is introduced to correct these errors in the preprocessed text, ensuring the accuracy of the conversational model's output. This text correction model also employs the Transformer architecture (which includes multi-head self-attention mechanisms, residual connections, layer normalization, and other techniques). The advantage of the Transformer architecture is its ability to handle sequences of arbitrary length, making it suitable for the field of natural language processing.
[0103] Step 305: Each key log information extracted from the data source is used as a slice text corresponding to the data source, thereby obtaining multiple slice texts corresponding to the data source.
[0104] Step 306: Slice the preprocessed text according to the preset file size to obtain multiple sliced texts corresponding to the data source.
[0105] Step 307: Use the preprocessed text corresponding to each image data as a slice text corresponding to the data source, thereby obtaining multiple slice texts corresponding to the data source.
[0106] Step 308: For each node, input the multiple slice texts corresponding to the multiple data sources and the question corresponding to the node into the conversational model in sequence to output the multiple answers corresponding to the node; cluster the multiple answers corresponding to the node to obtain each cluster; and select the target cluster containing the most answers from each cluster, sort the answers in the target cluster according to the frequency of occurrence, and select the answer with the most occurrences as the target answer corresponding to the node.
[0107] Step 309: Merge the target answers corresponding to each node according to the order of each node to obtain the processing solution.
[0108] Furthermore, the specific implementation details of the method for generating a disposal plan in one of the reference embodiments of the present invention have been described in detail in the above-described method for generating a disposal plan, so the details will not be repeated here.
[0109] Figure 5 This is a schematic diagram of an apparatus for generating a processing solution according to an embodiment of the present invention. Figure 5As shown, the apparatus 500 for generating a disposal plan includes a preprocessing module 501, a slicing module 502, a processing module 503, and a merging model 504. The preprocessing module 501 collects data from multiple data sources and preprocesses each data source to obtain preprocessed text corresponding to those data sources. The multiple data sources include alarm logs, voice data, and image data. The slicing module 502 slices the preprocessed text corresponding to each data source to obtain multiple sliced texts. The processing module 503, for each node, sequentially inputs the multiple sliced texts corresponding to the multiple data sources and the question corresponding to the node into a conversational model to output multiple answers corresponding to the node. It then clusters the multiple answers corresponding to the node to obtain clusters, selects the target cluster containing the most answers from each cluster, sorts the answers in the target cluster according to their frequency of occurrence, and selects the answer with the highest frequency as the target answer corresponding to the node. The merging module 504 merges the target answers corresponding to each node according to the order of the nodes to obtain a disposal plan.
[0110] Optionally, the preprocessing module 501 is further configured to:
[0111] If the data source is an alarm log, then regular expressions are used to extract key information from each log in the data source, thereby obtaining the preprocessed text corresponding to the data source.
[0112] If the data source is voice data, an acoustic model is used to identify the data source to obtain the identification result, and then a text extraction model is used to extract the identification result to obtain the preprocessed text corresponding to the data source.
[0113] If the data source is image data, then an object detection model is used to extract the object detection region from the data source, the object detection region is cropped from the data source, and a text recognition model is used to recognize the text data in the object detection region, thereby obtaining the preprocessed text corresponding to the data source.
[0114] Optionally, the voice data may be voice data collected during a meeting triggered by the alarm log and / or call data triggered by the alarm log; the image data may be at least one of the following: image data of the operation and maintenance interface triggered by the alarm log, screenshot data of the system alarm corresponding to the alarm log, and screenshot data of chat history triggered by the alarm log.
[0115] Optionally, the slicing module 502 is further configured to:
[0116] If the data source is an alarm log, then each key piece of information extracted from the data source will be used as a slice of text corresponding to the data source, thereby obtaining multiple slices of text corresponding to the data source.
[0117] If the data source is voice data, the preprocessed text is sliced according to a preset file size to obtain multiple sliced texts corresponding to the data source.
[0118] If the data source is image data, then the preprocessed text corresponding to each image data is used as a slice text corresponding to the data source, thereby obtaining multiple slice texts corresponding to the data source.
[0119] Optionally, the processing module 503 is further configured to:
[0120] For each data source, the multiple slices of text corresponding to the data source and the question corresponding to the node are sequentially input into the conversational model, thereby outputting multiple answers corresponding to the data source.
[0121] Optionally, the processing module 503 is further configured to:
[0122] For each slice of text corresponding to the data source, the slice of text and the question corresponding to the node are input into the conversational model together, thereby outputting the answer corresponding to the slice of text.
[0123] Optionally, the processing module 503 is further configured to:
[0124] For each answer in the target cluster, the answers corresponding to multiple data sources are weighted and summed based on the weight corresponding to each data source to obtain the election value of the answer;
[0125] Sort the answers according to their election values, and select the answer with the highest election value as the target answer for the node.
[0126] It should be noted that the specific implementation details of the apparatus for generating the treatment scheme described in this invention have been described in detail in the method for generating the treatment scheme described above, so the details will not be repeated here.
[0127] Figure 6 An exemplary system architecture 600 is shown, in which a method or apparatus for generating a disposal scheme can be applied according to embodiments of the present invention.
[0128] like Figure 6As shown, system architecture 600 may include terminal devices 601, 602, and 603, a network 604, and a server 605. Network 604 serves as the medium for providing communication links between terminal devices 601, 602, and 603 and server 605. Network 604 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0129] Users can use terminal devices 601, 602, and 603 to interact with server 605 via network 604 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 601, 602, and 603, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social media platform software, etc. (for example only).
[0130] Terminal devices 601, 602, and 603 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0131] Server 605 can be a server that provides various services, such as a backend management server that supports shopping websites browsed by users using terminal devices 601, 602, and 603 (this is just an example). The backend management server can analyze and process data such as received item information query requests, and then feed the processing results back to the terminal devices.
[0132] It should be noted that the method for generating a processing solution provided in this embodiment of the invention is generally executed by server 605, and correspondingly, the device for generating the processing solution is generally located in server 605. The method for generating a processing solution provided in this embodiment of the invention can also be executed by terminal devices 601, 602, and 603, and correspondingly, the device for generating the processing solution can be located in terminal devices 601, 602, and 603.
[0133] It should be understood that Figure 6 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0134] The following is for reference. Figure 7 It shows a schematic diagram of the structure of a computer system 700 suitable for implementing a terminal device of the present invention. Figure 7 The terminal device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0135] like Figure 7As shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 702 or programs loaded from storage section 708 into random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the system 700. The CPU 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0136] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.
[0137] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this invention include a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 709, and / or installed from removable medium 711. When the computer program is executed by central processing unit (CPU) 701, it performs the functions defined above in the system of this invention.
[0138] It should be noted that the computer-readable medium shown in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0139] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer programs according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0140] The modules described in the embodiments of the present invention can be implemented in software or hardware. The described modules can also be located in a processor; for example, a processor can be described as including a preprocessing module, a slicing module, a processing module, and a merging module. The names of these modules do not necessarily limit the functionality of the module itself.
[0141] In another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs, and when the one or more programs are executed by the device, the device implements the following method: collecting multiple data sources, preprocessing the multiple data sources respectively to obtain preprocessed text corresponding to the multiple data sources; wherein the multiple data sources include alarm logs, voice data, and image data; for the preprocessed text corresponding to each data source, slicing the preprocessed text to obtain multiple sliced texts corresponding to the data source; for each node, sequentially inputting the multiple sliced texts corresponding to the multiple data sources and the question corresponding to the node into a conversational model to output multiple answers corresponding to the node; clustering the multiple answers corresponding to the node to obtain various clusters, and selecting the target cluster containing the most answers from the various clusters, sorting the answers in the target cluster according to their frequency of occurrence, and selecting the answer with the most frequency of occurrence as the target answer corresponding to the node; merging the target answers corresponding to each node according to the order of each node to obtain a disposal solution.
[0142] In another aspect, embodiments of the present invention also provide a computer program product, including a computer program that, when executed by a processor, implements the methods described in any of the above embodiments.
[0143] According to the technical solution of this invention, by preprocessing multiple data sources separately to obtain preprocessed text corresponding to multiple data sources, then slicing the preprocessed text, and inputting the slices and questions into a conversational model, and clustering and sorting the output answers to obtain the target answer corresponding to each node, this technical approach overcomes the technical problem of decreased dispatch efficiency and handling accuracy of security incidents in the prior art. This invention, by using sliced text corresponding to multiple data sources as input to a conversational model and clustering and sorting the output results, helps to improve the dispatch efficiency of security incidents, thereby improving the handling efficiency and accuracy of security incidents.
[0144] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for generating a disposal plan, characterized in that, include: Multiple data sources are collected, and each of the multiple data sources is preprocessed to obtain preprocessed text corresponding to the multiple data sources; wherein, the multiple data sources include alarm logs, voice data and image data; For each data source, the preprocessed text is sliced to obtain multiple sliced texts corresponding to the data source. For each node, the multiple slices of text corresponding to the multiple data sources and the question corresponding to the node are sequentially input into the conversational model to output multiple answers corresponding to the node; the multiple answers corresponding to the node are clustered to obtain each cluster; and the target cluster containing the most answers is selected from each cluster. The answers in the target cluster are sorted according to their frequency of occurrence, and the answer with the most frequency of occurrence is selected as the target answer corresponding to the node. According to the order of each node, the target answers corresponding to each node are merged to obtain a solution. The multiple slices of text corresponding to the multiple data sources and the questions corresponding to the nodes are sequentially input into the conversational model, thereby outputting multiple answers corresponding to the nodes, including: For each data source, the multiple slices of text corresponding to the data source and the question corresponding to the node are sequentially input into the conversational model, thereby outputting multiple answers corresponding to the data source; The multiple text slices corresponding to the data source and the questions corresponding to the nodes are sequentially input into the conversational model, thereby outputting multiple answers corresponding to the data source, including: For each slice of text corresponding to the data source, the slice of text and the question corresponding to the node are input into the conversational model together, thereby outputting the answer corresponding to the slice of text.
2. The method according to claim 1, characterized in that, Preprocessing is performed on the multiple data sources respectively to obtain preprocessed text corresponding to the multiple data sources, including: If the data source is an alarm log, then regular expressions are used to extract key information from each log in the data source, thereby obtaining the preprocessed text corresponding to the data source. If the data source is voice data, an acoustic model is used to identify the data source to obtain the identification result, and then a text extraction model is used to extract the identification result to obtain the preprocessed text corresponding to the data source. If the data source is image data, then an object detection model is used to extract the object detection region from the data source, the object detection region is cropped from the data source, and a text recognition model is used to recognize the text data in the object detection region, thereby obtaining the preprocessed text corresponding to the data source.
3. The method according to claim 2, characterized in that, The voice data is voice data collected during a meeting triggered by the alarm log and / or call data triggered by the alarm log; the image data is at least one of the following: image data of the operation and maintenance interface triggered by the alarm log, screenshot data of the system alarm corresponding to the alarm log, and screenshot data of chat history triggered by the alarm log.
4. The method according to claim 2, characterized in that, For each data source corresponding to preprocessed text, the preprocessed text is sliced to obtain multiple sliced texts corresponding to the data source, including: If the data source is an alarm log, then each key piece of information extracted from the data source will be used as a slice of text corresponding to the data source, thereby obtaining multiple slices of text corresponding to the data source. If the data source is voice data, the preprocessed text is sliced according to a preset file size to obtain multiple sliced texts corresponding to the data source. If the data source is image data, then the preprocessed text corresponding to each image data is used as a slice text corresponding to the data source, thereby obtaining multiple slice texts corresponding to the data source.
5. The method according to claim 1, characterized in that, Sort the answers in the target cluster according to their frequency of occurrence, and select the answer with the highest frequency as the target answer for the node, including: For each answer in the target cluster, the answers corresponding to multiple data sources are weighted and summed based on the weight corresponding to each data source to obtain the election value of the answer; Sort the answers according to their election values, and select the answer with the highest election value as the target answer for the node.
6. An apparatus for generating a treatment plan, characterized in that, include: The preprocessing module is used to collect data from multiple data sources and preprocess each of the multiple data sources to obtain preprocessed text corresponding to the multiple data sources; wherein, the multiple data sources include alarm logs, voice data and image data; The slicing module is used to slice the preprocessed text corresponding to each data source, thereby obtaining multiple sliced texts corresponding to the data source. The processing module is used to sequentially input multiple slices of text corresponding to the multiple data sources and the question corresponding to the node into the conversational model for each node, thereby outputting multiple answers corresponding to the node; clustering the multiple answers corresponding to the node to obtain various clusters; and selecting the target cluster containing the most answers from the various clusters, sorting the answers in the target cluster according to their occurrence frequency, and selecting the answer with the most occurrence frequency as the target answer corresponding to the node. The merging module is used to merge the target answers corresponding to each node according to the order of each node, so as to obtain a solution. The processing module is also used for: For each data source, the multiple slices of text corresponding to the data source and the question corresponding to the node are sequentially input into the conversational model, thereby outputting multiple answers corresponding to the data source; The processing module is also used for: For each slice of text corresponding to the data source, the slice of text and the question corresponding to the node are input into the conversational model together, thereby outputting the answer corresponding to the slice of text.
7. The apparatus according to claim 6, characterized in that, The preprocessing module is also used for: If the data source is an alarm log, then regular expressions are used to extract key information from each log in the data source, thereby obtaining the preprocessed text corresponding to the data source. If the data source is voice data, an acoustic model is used to identify the data source to obtain the identification result, and then a text extraction model is used to extract the identification result to obtain the preprocessed text corresponding to the data source. If the data source is image data, then an object detection model is used to extract the object detection region from the data source, the object detection region is cropped from the data source, and a text recognition model is used to recognize the text data in the object detection region, thereby obtaining the preprocessed text corresponding to the data source.
8. The apparatus according to claim 7, characterized in that, The voice data is voice data collected during a meeting triggered by the alarm log and / or call data triggered by the alarm log; the image data is at least one of the following: image data of the operation and maintenance interface triggered by the alarm log, screenshot data of the system alarm corresponding to the alarm log, and screenshot data of chat history triggered by the alarm log.
9. The apparatus according to claim 7, characterized in that, The slicing module is also used for: If the data source is an alarm log, then each key piece of information extracted from the data source will be used as a slice of text corresponding to the data source, thereby obtaining multiple slices of text corresponding to the data source. If the data source is voice data, the preprocessed text is sliced according to a preset file size to obtain multiple sliced texts corresponding to the data source. If the data source is image data, then the preprocessed text corresponding to each image data is used as a slice text corresponding to the data source, thereby obtaining multiple slice texts corresponding to the data source.
10. The apparatus according to claim 6, characterized in that, The processing module is also used for: For each answer in the target cluster, the answers corresponding to multiple data sources are weighted and summed based on the weight corresponding to each data source to obtain the election value of the answer; Sort the answers according to their election values, and select the answer with the highest election value as the target answer for the node.
11. An electronic device, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-5.
12. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-5.
Citation Information
Patent Citations
Fault handling method and device, medium and equipment
CN112446511A
Automatic root cause analysis positioning processing method for intelligent operation and maintenance
CN116225849A