Alarm study and judgment model training method and device and alarm study and judgment method and device
By fine-tuning the initial model and adjusting the preference data set, an alarm analysis model that meets the preferences of target users is trained, which solves the problem that the existing model is difficult to meet the preferences of individual users and achieves more efficient alarm analysis results.
Patent Information
- Application Number
- CN202510757308.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-06
- Publication Date
- 2025-09-23
AI Technical Summary
The existing alarm analysis model is difficult to fully meet the preferences of individual users, resulting in poor alarm analysis results.
By fine-tuning the initial model and adjusting the preferred data set, an alarm analysis model that meets the preferences of the target users is trained, and the trained model is confirmed after quality assessment.
The alarm analysis model improves the alarm analysis effect on target users, ensuring that the model outputs alarm analysis results that meet user preferences.
Smart Images

Figure CN120687767A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of alarm analysis and judgment, and in particular to an alarm analysis and judgment model training method and device, and an alarm analysis and judgment method and device. Background Art
[0002] Alarm analysis refers to the process of conducting in-depth analysis of alarm information after an alarm is triggered to determine its authenticity, severity, and potential causes. Alarm analysis is a critical component of network security operations, helping to quickly locate and resolve security incidents, minimizing potential risks and losses.
[0003] Currently, alarm analysis typically relies on artificial intelligence-based alarm analysis models. However, existing alarm analysis models focus on generalization across all users, making it difficult for them to fully meet the preferences of individual users. Consequently, when applied to individual users, these models struggle to produce results that meet their preferences, resulting in suboptimal alarm analysis.
[0004] Therefore, how to train an alarm analysis model that meets user preferences has become an urgent problem that needs to be solved. Summary of the Invention
[0005] This application proposes an alarm analysis model training method and device, an alarm analysis method and device, the main purpose of which is to train an alarm analysis model that meets user preferences.
[0006] In order to achieve the above objectives, this application mainly provides the following technical solutions:
[0007] In the first aspect, the present application provides an alarm analysis model training method, which may include: fine-tuning the initial model to be trained as the alarm analysis model to obtain a fine-tuned model with alarm analysis capabilities; performing preference adjustment on the fine-tuned model based on a preference data set to obtain a target model that meets the preferences of the target user, the target model being used to process alarm data to output alarm analysis data that meets the preferences of the target user, the preference data set including a plurality of preference data pairs, the preference data pairs including the alarm analysis data corresponding to the target user preferences and the alarm analysis data that does not meet the preferences of the target user; performing alarm analysis quality assessment on the target model from at least one quality assessment dimension; if it is assessed that the alarm analysis quality of the target model meets the requirements, the target model is determined as the trained alarm analysis model.
[0008] In the second aspect, the present application provides an alarm analysis method, which is applied to an alarm analysis system. The alarm analysis system is deployed with at least one alarm analysis model trained based on the alarm analysis model training method described in the first aspect. The alarm analysis method may include: if alarm data for any user is obtained, calling the alarm analysis model that meets the user's preferences to process the user's alarm data to obtain the alarm analysis data output by the model; based on the alarm analysis data output by the model, performing alarm handling on the user.
[0009] In a third aspect, the present application provides an alarm analysis model training device, which may include:
[0010] A fine-tuning module is used to fine-tune the initial model to be trained as an alarm analysis model to obtain a fine-tuned model with alarm analysis capabilities;
[0011] An adjustment module is configured to perform preference adjustment on the fine-tuning model based on a preference data set to obtain a target model that meets the preferences of the target user. The target model is configured to process the alarm data and output alarm analysis and judgment data that meets the preferences of the target user. The preference data set includes a plurality of preference data pairs, each of which includes the alarm analysis and judgment data that meets the preferences of the target user and the alarm analysis and judgment data that does not meet the preferences of the target user corresponding to the alarm data.
[0012] An evaluation module, configured to perform an alarm analysis and quality evaluation on the target model from at least one quality evaluation dimension;
[0013] The determination module is used to determine the target model as the trained alarm analysis model if it is assessed that the alarm analysis quality of the target model meets the requirements.
[0014] In a fourth aspect, the present application provides an alarm analysis device, which is applied to an alarm analysis system. The alarm analysis system is deployed with at least one alarm analysis model trained based on the alarm analysis model training method described in the first aspect. The alarm analysis device may include:
[0015] A calling module is used to, if alarm data for any user is obtained, call an alarm analysis model that meets the user's preferences to process the user's alarm data and obtain alarm analysis data output by the model;
[0016] The handling module is used to handle the alarm of the user based on the alarm analysis data output by the model.
[0017] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the alarm analysis model training method described in the first aspect, and / or the alarm analysis method described in the second aspect.
[0018] In a sixth aspect, an embodiment of the present application provides an electronic device, comprising: a memory for storing a program; a processor, coupled to the memory, for running the program to execute the alarm analysis model training method described in the first aspect, and / or the alarm analysis method described in the second aspect.
[0019] The alarm analysis model training method and device, alarm analysis method and device provided in this application can at least achieve the following effects: First, when training the alarm analysis model, the model is adjusted based on the preference data set that meets the target user's preferences, so that the model can process the alarm data and output alarm analysis data that meets the target user's preferences. The alarm analysis model trained in this way can perform alarm analysis on the target user in accordance with the target user's preferences, thereby improving the alarm analysis effect of the alarm analysis model on the target user. Second, after the model is adjusted based on the preference data set, the target model that meets the target user's preferences obtained by the preference adjustment is not directly used as the alarm analysis model. Instead, the target model is first evaluated for alarm quality, and only when it is evaluated that the alarm analysis quality of the target model meets the requirements is the target model determined as the trained alarm analysis model. The alarm analysis model obtained in this way has a higher alarm analysis quality, thereby further improving the alarm analysis effect of the alarm analysis model on the target user.
[0020] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0022] Figure 1 A flowchart of a method for training an alarm analysis model provided by one embodiment of the present application is shown;
[0023] Figure 2A flowchart of an alarm analysis method provided by an embodiment of the present application is shown;
[0024] Figure 3 A schematic structural diagram of an alarm analysis model training device provided by one embodiment of the present application is shown;
[0025] Figure 4 A schematic structural diagram of an alarm analysis and judgment model training device provided by another embodiment of the present application is shown;
[0026] Figure 5 A structural diagram of an alarm analysis and judgment device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION
[0027] The following describes exemplary embodiments of the present application in more detail with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present application and to fully convey the scope of the present application to those skilled in the art.
[0028] Alarm analysis is a critical component of network security operations. It distinguishes true security threats from false alarms, avoiding overreaction or underreaction in alarm handling and improving the accuracy and efficiency of alarm processing. Currently, the alarm analysis models relied upon focus solely on generalization across all users, without addressing the applicability to individual user preferences. This makes it difficult to conduct targeted alarm analysis for individual users when applied to these models, making it difficult to output alarm analysis results that align with individual user preferences. Consequently, this results in suboptimal alarm analysis for individual users.
[0029] After research, it was found that if the model is adjusted based on the preference data set that meets the target user's preferences during the training of the alarm analysis model, so that the model can process the alarm data and output alarm analysis data that meets the target user's preferences, then the trained alarm analysis model will be able to perform alarm analysis on the target user in accordance with the target user's preferences, thereby improving the alarm analysis effect of the alarm analysis model on the target user.
[0030] Based on the above findings, the embodiment of the present application specifically provides a technical solution for training an alarm analysis model, which specifically includes: fine-tuning the initial model to be trained as an alarm analysis model to obtain a fine-tuning model with alarm analysis capabilities; performing preference adjustment on the fine-tuning model based on a preference data set to obtain a target model that meets the preferences of the target user, the target model is used to process the alarm data to output alarm analysis data that meets the preferences of the target user, the preference data set includes a plurality of preference data pairs, the preference data pairs include the alarm analysis data corresponding to the alarm data that meets the preferences of the target user and the alarm analysis data that does not meet the preferences of the target user; the target model is evaluated for alarm analysis quality from at least one quality assessment dimension; if it is assessed that the alarm analysis quality of the target model meets the requirements, the target model is determined to be a trained alarm analysis model. The above-mentioned technical solution for alarm analysis model training can train an alarm analysis model that meets the preferences of the target user, so as to improve the alarm analysis effect of the alarm analysis model on the target user.
[0031] Correspondingly, based on the technical solution for training the above-mentioned alarm analysis model, the embodiment of the present application also provides a technical solution for alarm analysis, specifically including: deploying at least one alarm analysis model trained based on the technical solution for training the above-mentioned alarm analysis model in the alarm analysis system; if alarm data for any user is obtained, calling the alarm analysis model that meets the user's preferences to process the user's alarm data to obtain the alarm analysis data output by the model; and performing alarm handling on the user based on the alarm analysis data output by the model.
[0032] The model type of the alarm analysis model involved in the above-mentioned technical solutions for alarm analysis model training and alarm analysis can be flexibly selected based on business needs, and this embodiment does not limit this. Exemplary model types of alarm analysis models may include, but are not limited to: deep learning models (e.g., large language models).
[0033] Based on the above-mentioned technical solutions for alarm analysis model training and alarm analysis, this embodiment specifically provides an alarm analysis model training method and device, and an alarm analysis method and device. The following describes in detail the alarm analysis model training method and device, and the alarm analysis method and device provided in this embodiment.
[0034] The present application embodiment provides a method for training an alarm analysis model, such as Figure 1 As shown, the alarm analysis model training method provided in this embodiment may include at least the following steps 101 to 104:
[0035] 101. Fine-tune the initial model to be trained as the alarm analysis model to obtain a fine-tuned model with alarm analysis capabilities.
[0036] In some embodiments, when a model training instruction is received to instruct the training of an adapted alarm analysis model for any target user, the alarm analysis model training is started based on the alarm analysis model training method provided in this embodiment to train an alarm analysis model that meets the preferences of the target user.
[0037] When it is determined that an alarm analysis model that meets the preferences of the target user needs to be trained, the initial model to be trained as the alarm analysis model must first be selected so that the alarm analysis model can be trained based on the initial model. Training the alarm analysis model based on the initial model can achieve at least the following two effects: First, the initial model is a trained model. When training on the initial model, the features that have been learned by the initial model can be directly utilized to avoid training from scratch, thereby reducing the cost and time of training the alarm analysis model. Another is that compared to training the alarm analysis model from scratch, training the alarm analysis model based on the initial model requires fewer resources such as TPU (Tensor Processing Unit) for model training, thereby reducing the resource investment required for training the alarm analysis model.
[0038] The initial model to be trained as the alarm analysis model can be any of the following: 1. A pre-trained language model. 2. An alarm analysis model currently used by the target user that no longer meets the target user's current preferences. 3. An alarm analysis model that is universal for all users.
[0039] In some embodiments, after selecting an initial model to be trained as an alarm analysis model, a target alarm analysis dataset is obtained, and based on the target alarm analysis dataset, the initial model to be trained as an alarm analysis model is fine-tuned to obtain a fine-tuned model with alarm analysis capabilities. The implementation process of this step may include the following steps 101A to 101B:
[0040] 101A. Select at least some layers in the initial model as target layers, and freeze parameters of other layers in the initial model that are not selected as target layers.
[0041] The target layer selection strategy can include any of the following: First, the layers in the initial model that need to participate in alarm analysis and judgment are selected as target layers, so that these layers can have alarm analysis and judgment capabilities through fine-tuning. For example, the bottom layer in the initial model needs to retain common features, and its corresponding parameters cannot be modified, while the top layer in the initial model can learn alarm analysis and judgment, and its corresponding parameters can be modified. Therefore, the top layer in the initial model is selected as the target layer. Second, if the full learning potential of the initial model needs to be released, all layers in the initial model are selected as target layers, so that the parameters of all layers are updated during the fine-tuning process, and the model with the best performance is fine-tuned.
[0042] The target layer is the layer that needs to learn alarm analysis and judgment through parameter adjustment. Therefore, its corresponding parameters are not frozen to optimize its parameters during fine-tuning. Other layers in the initial model that were not selected as the target layer need to retain their original parameters. Therefore, the parameters of these other layers are frozen to prevent their parameters from being modified during the initial model fine-tuning process.
[0043] 101B. Fine-tune the initial model after freezing the parameters of other layers based on the target analysis dataset.
[0044] Fine-tuning the initial model is supervised fine-tuning. During fine-tuning, the target alarm analysis dataset is first obtained. This dataset includes multiple target alarm data and corresponding target alarm analysis data for each target alarm data. The target alarm data may include, but is not limited to, at least one of the following: device alarm messages and security event logs. Based on the target alarm analysis dataset, the initial model, after freezing the parameters of other layers, is fine-tuned. This ensures that the initial model has the alarm analysis capability to process the alarm data and output corresponding alarm analysis data.
[0045] Furthermore, in some embodiments, the method for determining the target alarm analysis data set may include the following steps: determining the target business category involved in the target user for whom the alarm analysis model is to be used; and determining the alarm analysis data set corresponding to the target business category as the target alarm analysis data set based on the correspondence between at least one preset business category and the alarm analysis data set. The target alarm analysis data set determined by this method is suitable for the target business category of the target user, so the alarm analysis model obtained through subsequent training can more effectively perform alarm analysis on the target user. The target business category can be expressed by a business name (for example, financial business).
[0046] Furthermore, in some embodiments, the alarm analysis data output by the alarm analysis model is used to locate and resolve security incidents. Based on this, the target alarm analysis data set in this embodiment includes multiple target alarm data and target alarm analysis data with a preset format corresponding to each target sample alarm data. The target alarm analysis data with a preset format is used to realize the ability of the fine-tuned initial model to output alarm analysis data with a preset format. The preset format defines the regularity of the alarm analysis data output by the alarm analysis model, which facilitates the rapid interpretation of the alarm analysis data based on the preset format, so as to more quickly perform subsequent positioning and resolution of security incidents. The preset format can be set based on the format requirements of the target user, which is not limited in this embodiment.
[0047] 101C. When the fine-tuning termination condition is met, fine-tuning is stopped, and the initial model after fine-tuning is stopped is obtained as the fine-tuning model.
[0048] The method for determining whether a fine-tuning termination condition is satisfied may include the steps of: after each round of fine-tuning, determining a target metric of the initial model after the current round of fine-tuning on a second alarm analysis dataset used to verify the fine-tuning; if it is determined that the target metric has not improved after N consecutive rounds of fine-tuning, determining that the fine-tuning termination condition is satisfied. The target metric may include, but is not limited to, at least one of the following: accuracy, loss value, etc.
[0049] When it is determined that the fine-tuning termination conditions are met, it means that the fine-tuned initial model already has the alarm analysis capability and can be used for subsequent alarm analysis model training. Therefore, fine-tuning is stopped, and the initial model after fine-tuning is obtained as the fine-tuning model for subsequent model training.
[0050] 102. Based on the preference data set, the fine-tuning model is adjusted for preferences to obtain a target model that meets the preferences of the target user. The target model is used to process the alarm data and output alarm analysis data that meets the preferences of the target user. The preference data set includes multiple preference data pairs. The preference data pairs include the alarm analysis data that meets the preferences of the target user and the alarm analysis data that does not meet the preferences of the target user.
[0051] The purpose of adjusting the preference of the fine-tuning model based on the preference data set is to train an alarm analysis model that meets the preferences of the target user, so that the alarm analysis model can perform targeted alarm analysis on the target user, thereby outputting alarm analysis results that meet the preferences of the target user and improving the alarm analysis effect of the alarm analysis model on the target user.
[0052] The preference data set is used to adjust the fine-tuning model so that the adjusted fine-tuning model outputs alarm analysis data that meets the target user's preferences after processing the alarm data. The preference data set includes multiple preference data pairs. For any preference data pair, the preference data pair includes two types of alarm analysis data, good and bad, corresponding to the alarm data. The high-quality alarm analysis data is the alarm analysis data that meets the target user's preferences, and the low-quality alarm analysis data is the alarm analysis data that does not meet the target user's preferences. The alarm analysis data that meets the target user's preferences and the alarm analysis data that does not meet the target user's preferences have different degrees of compliance with the target user's preferences, and the difference between the degrees of compliance between the two needs to reach a preset difference, so that the trained alarm analysis model can clearly output alarm analysis data that is more in line with the target user's preferences based on the difference.
[0053] In some embodiments, a method for adjusting the preferences of a fine-tuning model based on a preference data set to obtain a target model that meets the preferences of a target user may include at least the following method 1 and method 2.
[0054] Method 1, based on the preference data set, adjusts the preferences of the fine-tuning model to obtain a target model that meets the preferences of the target user. The specific process may include the following steps 102A to 102B:
[0055] 102A. Obtain a reward model trained based on a preference data set. The reward model is used to determine the degree to which the alarm analysis data output by the fine-tuning model conforms to the target user's preferences.
[0056] In some embodiments, the target user's preferences can be expressed through but not limited to at least one of the following preference contents: the clarity of the language logic of the alarm analysis data, the degree of understanding of the alarm data by the alarm analysis data, the comprehensiveness of the alarm analysis data, the practicality of the alarm analysis data in solving the problems involved in the alarm data, and the textual expression quality of the alarm analysis data.
[0057] The process of obtaining the preference data set may include: collecting different alarm data from a designated security product platform, and collecting security knowledge through crawlers; in the annotation platform, using different preset models for each alarm data to obtain alarm analysis data of different qualities corresponding to each alarm data; sorting the alarm analysis data of the same alarm data in order of quality from high to low, determining the alarm analysis data with the highest quality as the alarm analysis data that meets the target user's preferences for the corresponding alarm data, and determining the alarm analysis data with the lowest quality as the alarm analysis data that does not meet the target user's preferences.
[0058] After determining the preference data set, obtain the reward model trained based on the preference data set to evaluate the degree of conformity of the alarm analysis data output by the fine-tuning model with the target user's preferences through the reward model, so as to judge whether the preference adjustment of the fine-tuning model can be completed based on the degree of conformity.
[0059] 102B. Iteratively train the fine-tuning model, and after each round of iteration, perform the following steps B1 to B4:
[0060] B1. Call the fine-tuned model after the current iteration to process the first alarm data in the first alarm data set to obtain the first alarm analysis data output by the model.
[0061] The first alarm dataset is equivalent to a test set, which includes multiple first alarm data. After each iteration of training, the fine-tuned model after the current iteration is called to process the first alarm data in the first alarm dataset to obtain the first alarm judgment data output by the model. This first alarm judgment data is used to verify whether the fine-tuned model after the current iteration can output alarm judgment data that better meets the target user's preferences. Based on the verification results, the preference adjustment process is then concluded.
[0062] B2. Determine the degree to which the first alarm analysis data output by the model conforms to the target user's preferences based on the reward model.
[0063] After obtaining the first alarm analysis data output by the fine-tuning model after the current iteration, the following steps are performed for each alarm analysis data: the reward model is called to determine the degree of compliance of the alarm analysis data output by the fine-tuning model with the target user's preferences. The degree of compliance can be represented by a score. The higher the score, the greater the degree of compliance of the alarm analysis data with the target user's preferences, indicating that the alarm analysis data is more consistent with the target user's preferences.
[0064] B3. If the first alarm analysis data output by the judgment model meets the requirements for the degree of compliance with the target user's preferences, the fine-tuning model after the current iteration is determined as the target model.
[0065] After determining the corresponding degree of conformity of the first alarm analysis data output by the model based on the reward model, it is necessary to execute the step of judging whether the corresponding degree of conformity of the first alarm analysis data output by the model meets the requirements. The implementation method of this step may include the following process: judging whether the proportion of the first alarm analysis data with a degree of conformity not less than the second ideal degree of conformity in the first alarm analysis data output by the fine-tuning model after the current iteration is not less than the target proportion.
[0066] If the proportion of the first alarm analysis data whose degree of conformity is not less than the second ideal degree of conformity in the first alarm analysis data output by the fine-tuning model after the current iteration is not less than the target proportion, it means that the fine-tuning model after the current iteration can output alarm analysis data that meets the target user's preferences. At this time, the preference adjustment of the fine-tuning model can be terminated, so it is determined that the degree of conformity of the first alarm analysis data output by the model with the target user's preferences meets the requirements.
[0067] If the proportion of the first alarm analysis data whose degree of conformity is not less than the second ideal degree of conformity in the first alarm analysis data output by the fine-tuning model after the current iteration is less than the target proportion, it means that the fine-tuning model after the current iteration cannot output the alarm analysis data that meets the target user's preferences. At this time, it is necessary to continue to adjust the preferences of the fine-tuning model. Therefore, it is determined that the degree of conformity of the first alarm analysis data output by the model with the target user's preferences does not meet the requirements.
[0068] After executing the step of determining whether the corresponding degree of conformity of the first alarm analysis data output by the judgment model meets the requirements, if it is determined that the first alarm analysis data output by the judgment model meets the requirements for the target user's preferences, it means that the fine-tuning model after the current iteration already has the ability to output alarm analysis data that is more in line with the target user's preferences, and there is no need to continue to perform preference fine-tuning. Therefore, the fine-tuning model after the current iteration is determined as the target model, and the target model is the model to be determined as the alarm analysis model.
[0069] B4. If the first alarm analysis data output by the judgment model does not meet the requirements for the compliance with the target user's preferences, the parameters of the fine-tuning model after the current iteration are optimized, and the step of iteratively training the fine-tuning model is returned to execute.
[0070] If the first alarm analysis data output by the judgment model does not meet the requirements for the target user's preferences, it means that the fine-tuning model after the current iteration does not yet have the ability to output alarm analysis data that is more in line with the target user's preferences, and preference adjustment needs to be continued. Therefore, it is necessary to optimize the parameters of the fine-tuning model after the current iteration and return to the step of iterative training of the fine-tuning model.
[0071] There are two methods for optimizing the parameters of the fine-tuning model after the current iteration:
[0072] One is that the specific process of optimizing the parameters of the fine-tuning model after the current iteration may include the following steps: determining the first difference information between the degree of compliance of the first alarm analysis data output by the fine-tuning model after the current iteration with the target user's preference and the first ideal degree of compliance; optimizing the parameters of the fine-tuning model after the current iteration based on the first difference information.
[0073] The specific process of determining the first difference information between the degree of conformity of the first alarm analysis data output by the fine-tuning model after the current iteration to the target user's preference and the first ideal degree of conformity may include the following: (1) If the degree of conformity is expressed by a score, the average score of the degree of conformity of each output first alarm analysis data to the target user's preference is determined, and the difference between the average score and the score corresponding to the first ideal degree of conformity is determined as the first difference information. (2) Based on a preset loss function, the loss value is determined by the degree of conformity of the output first alarm analysis data to the target user's preference, and the difference between the loss value and the score corresponding to the first ideal degree of conformity is determined as the first difference information.
[0074] A first correspondence relationship applicable to each parameter is preset, the first correspondence relationship being used to indicate the correspondence between the first sample difference information and the parameter adjustment amplitude value. After determining the first difference information, the parameter to be optimized is determined, and the following steps are performed for each parameter to be optimized in the fine-tuning model after the current iteration: determining the first correspondence relationship applicable to the parameter, selecting the parameter adjustment amplitude value corresponding to the first difference information from the first correspondence relationship, and adjusting the parameter value of the parameter by a corresponding amplitude based on the selected parameter adjustment amplitude value.
[0075] Another is that the specific process of optimizing the parameters of the fine-tuning model after the current iteration may include the following steps: determining the conformity change trend data based on the conformity degree of the first alarm analysis data output by the fine-tuning model after the current iteration to the target user's preferences and the conformity degree of the first alarm analysis data output by the fine-tuning model of the first number of iterations adjacent to the current iteration to the target user's preferences; optimizing the parameters of the fine-tuning model after the current iteration based on the conformity change trend data.
[0076] A second corresponding relationship applicable to each parameter is preset, and the second corresponding relationship is used to indicate the corresponding relationship between the sample conformity change trend data and the parameter adjustment amplitude value. After determining the conformity change trend data, the parameter to be optimized is determined, and the following steps are performed for each parameter to be optimized in the fine-tuning model after the current iteration: determining the second corresponding relationship applicable to the parameter, selecting the parameter adjustment amplitude value corresponding to the conformity change trend data from the second corresponding relationship, and adjusting the parameter value of the parameter by a corresponding amplitude based on the selected parameter adjustment amplitude value.
[0077] At least one of the two aforementioned methods for optimizing the parameters of the fine-tuned model after the current iteration can be flexibly selected based on business needs, and this embodiment does not limit this. It should be noted that existing parameter optimization methods can also be used to optimize the parameters of the fine-tuned model after the current iteration. These existing parameter optimization methods will not be further described here.
[0078] After optimizing the parameters of the fine-tuning model after the current iteration, return to the step of iteratively training the fine-tuning model to continue to adjust the preferences for fine-tuning.
[0079] Furthermore, in some embodiments, considering that the current iterative model cannot output alarm analysis data that meets the target user's preferences, it may not be caused by parameters. Based on this situation, after determining that the first alarm analysis data output by the model does not meet the requirements for the target user's preferences, the alarm analysis model training method provided by this embodiment may also include the following steps: judging whether the number of first alarm analysis data with a degree of compliance in the target compliance range appears for a second number of consecutive iterations is not less than a third number; if so, replacing the alarm analysis data set based on which the iterative training fine-tuning model is based, and based on the replaced alarm analysis data set, returning to the step of iteratively training the fine-tuning model; if not, executing the step of optimizing the parameters of the fine-tuning model after the current iteration; the degree of compliance in the target compliance range cannot distinguish whether the alarm analysis data meets the target user's preferences.
[0080] If it is determined that the number of first alarm analysis data with a degree of compliance within the target compliance range appears for the second number of consecutive iterations is not less than the third number, it means that the data quality of the alarm analysis data set based on which the fine-tuning model is iteratively trained is not high, and it is difficult to achieve the fine-tuning model outputting alarm analysis data that is more in line with the target user's preferences based on the original alarm analysis data set. Therefore, the alarm analysis data set based on which the fine-tuning model is iteratively trained is replaced, and based on the replaced alarm analysis data set, the step of iteratively training the fine-tuning model is returned to, so as to use the new alarm analysis data set to re-adjust the preferences of the fine-tuning model.
[0081] If it is determined that the number of first alarm analysis data whose compliance level is within the target compliance level range does not appear in the second number of consecutive iterations and is not less than the third number, it means that the model cannot output alarm analysis data that meets the target user's preferences. This is most likely caused by insufficient parameter optimization. Therefore, the step of optimizing the parameters of the fine-tuning model after the current iteration is returned.
[0082] Method 2, based on the preference dataset, adjusts the preferences of the fine-tuning model to obtain a target model that meets the preferences of the target user. The specific process may include the following step 102C:
[0083] 102C. Iteratively train the fine-tuning model based on the preferred dataset, and after each round of iteration, perform the following steps C1 to C3:
[0084] C1. Determine whether the decrease in the probability that the fine-tuning model after the current iteration outputs alarm analysis data that does not conform to the target user's preferences relative to the corresponding reference model of the current iteration, and the increase in the probability that the fine-tuning model after the current iteration outputs alarm analysis data that conforms to the target user's preferences relative to the reference model, both meet the requirements.
[0085] The reference model corresponding to the current iteration can be selected from any of the following based on business needs: one is that the reference model corresponding to the current iteration is the fine-tuned model obtained after fine-tuning the initial model, and the fine-tuned model after the current iteration is compared with the original fine-tuned model obtained in step 101 to verify the alarm analysis and judgment capability of the fine-tuned model after the current iteration for user preferences. The other is that the reference model corresponding to the current iteration is the fine-tuned model after the previous iteration. The fine-tuned model after the current iteration is compared with the fine-tuned model after the previous iteration to observe the alarm analysis and judgment capability of the fine-tuned model after the current iteration for user preferences. Through such comparison, the alarm analysis and judgment capability of the fine-tuned model for user preferences can be continuously improved as the iteration proceeds.
[0086] The specific process of determining whether the probability that the fine-tuning model after the current iteration outputs alarm analysis data that does not conform to the target user's preferences is reduced relative to the reference model corresponding to the current iteration, and whether the probability that the fine-tuning model after the current iteration outputs alarm analysis data that conforms to the target user's preferences is increased relative to the reference model both meet the requirements may include the following steps (1) to (4).
[0087] (1) Calling the reference model and the fine-tuning model after the current iteration to process the second alarm data in the second alarm data set respectively, obtaining the first probability that the reference model outputs the alarm analysis data that meets the target user's preference for each second alarm data and the second probability that the reference model outputs the alarm analysis data that does not meet the target user's preference, and obtaining the third probability that the fine-tuning model after the current iteration outputs the alarm analysis data that meets the target user's preference for each second alarm data and the fourth probability that the fine-tuning model outputs the alarm analysis data that does not meet the target user's preference.
[0088] The second alarm data set is equivalent to a test set, which includes multiple second alarm data. After each iteration, the reference model and the fine-tuning model after the current iteration will be called to process the second alarm data in the second alarm data set respectively, so as to output the first probability of alarm analysis data that meets the target user's preferences for each second alarm data through the reference model and the second probability of outputting alarm analysis data that does not meet the target user's preferences, and obtain the third probability of the fine-tuning model after the current iteration outputting the alarm analysis data that meets the target user's preferences for each second alarm data and the fourth probability of outputting the alarm analysis data that does not meet the target user's preferences, to verify whether the fine-tuning model after the current iteration can output alarm analysis data that is more in line with the target user's preferences, so as to decide whether the preference adjustment process is ended based on the verification results.
[0089] (2) Based on the first probability, second probability, third probability, fourth probability and byte length of the alarm analysis data corresponding to each second alarm data, determine the loss data corresponding to the fine-tuning model after the current iteration.
[0090] In some embodiments, the loss data corresponding to the fine-tuning model after the current iteration can be determined by the following formula:
[0091]
[0092] Wherein, L represents the loss data corresponding to the fine-tuning model after the current iteration. D represents the second alarm data set. x represents a second alarm data in the second alarm data set, yw represents the alarm analysis data output by the model for the second alarm data that meets the target user's preferences, yl represents the alarm analysis data output by the model for the second alarm data that does not meet the target user's preferences, πθ(yw|x) represents the third probability that the fine-tuning model after the current iteration outputs the alarm analysis data that meets the target user's preferences for the second alarm data, πref(yw|x) represents the first probability that the reference model outputs the alarm analysis data that meets the target user's preferences for the second alarm data, πθ(yl|x ) represents the fourth probability that the fine-tuning model after the current iteration outputs the alarm analysis data that does not conform to the target user's preference for the second alarm data, πref(yl|x) represents the second probability that the reference model outputs the alarm analysis data that does not conform to the target user's preference for the second alarm data, Nw represents the byte length of the alarm analysis data that the fine-tuning model after the current iteration outputs the alarm analysis data that conforms to the target user's preference for the second alarm data, and Nl represents the byte length of the alarm analysis data that the fine-tuning model after the current iteration outputs the alarm analysis data that does not conform to the target user's preference for the second alarm data. σ is a sigmoid function, which is used to map the difference in "()" to a probability. β is a temperature hyperparameter, which is used to control the sensitivity to preference differences. E represents the expected value, which averages the subsequent expressions in "[]". Specifically, each second alarm data in the data set D performs the operations involved in the expressions in "[]", and then the expected value is the average of the results of the operations involved in the expressions in "[]" executed on each second alarm data in the data set D.
[0093] Furthermore, in some embodiments, if computational complexity needs to be reduced, the byte length of the alarm analysis data may be disregarded. The loss data corresponding to the fine-tuned model after the current iteration can be determined based on the first probability, second probability, third probability, and fourth probability corresponding to each second alarm data. Specifically, this process can be expressed by the following formula:
[0094]
[0095] The meaning of each part in the formula can be found in the above formula considering the byte length of the alarm analysis data, which will not be repeated here.
[0096] (3) If the loss data is determined to meet the loss requirements, then the requirements are met.
[0097] The loss data can be represented by a loss value. Based on the loss value of the fine-tuned model after the current iteration, it is determined whether the loss values of N consecutive rounds of preference adjustment have stabilized within the preset value range. If so, the loss data is determined to meet the loss requirement. If not, the loss data is determined to not meet the loss requirement.
[0098] If the loss data is determined to meet the loss requirements, then the probability that the fine-tuning model after the current iteration outputs alarm analysis data that does not meet the target user's preferences is reduced relative to the corresponding reference model of the current iteration, and the probability that the fine-tuning model after the current iteration outputs alarm analysis data that meets the target user's preferences is increased relative to the reference model. Both meet the requirements.
[0099] (4) If the loss data is determined to not meet the loss requirements, it is determined that the requirements are not met.
[0100] If the loss data is determined to not meet the loss requirements, it means that the fine-tuning model still has room for preference adjustment and preference adjustment still needs to be continued. Therefore, it is determined that the probability of the fine-tuning model after the current iteration outputting alarm analysis data that does not meet the target user's preference is reduced relative to the corresponding reference model of the current iteration, and the probability of the fine-tuning model after the current iteration outputting alarm analysis data that meets the target user's preference is increased relative to the reference model. Both do not meet the requirements.
[0101] C2. If all the requirements are met, the fine-tuned model after the current iteration is determined as the target model.
[0102] If it meets the requirements, it means that the fine-tuned model after the current iteration has the ability to output alarm analysis data that meets the preferences of the target user. Therefore, the fine-tuned model after the current iteration is determined as the target model.
[0103] C3. If none of the requirements are met, the parameters of the fine-tuning model after the current iteration are optimized, and the process returns to the step of iteratively training the fine-tuning model based on the preferred dataset.
[0104] It is determined that the probability that the fine-tuning model after the current iteration outputs alarm analysis data that does not meet the target user's preferences is reduced relative to the corresponding reference model of the current iteration, and the probability that the fine-tuning model after the current iteration outputs alarm analysis data that meets the target user's preferences is increased relative to the reference model. Both do not meet the requirements, indicating that the fine-tuning model still has room for preference adjustment and still needs to be adjusted. Therefore, the parameters of the fine-tuning model after the current iteration are optimized, and after optimization, the step of iteratively training the fine-tuning model based on the preference data set is returned to execute.
[0105] There are three methods for optimizing the parameters of the fine-tuning model after the current iteration:
[0106] First, the specific process of optimizing the parameters of the fine-tuning model after the current iteration may include the following steps: determining the second difference information between the loss data of the fine-tuning model after the current iteration and the ideal loss data; and optimizing the parameters of the fine-tuning model after the current iteration based on the second difference information.
[0107] Specifically, the loss data is the loss value after the current iteration of the fine-tuning model, and the ideal loss data is the ideal loss value indicating the convergence of the fine-tuning model.
[0108] Specifically, a third correspondence relationship applicable to each parameter is preset, and the third correspondence relationship is used to indicate the correspondence between the target sample difference information and the parameter adjustment amplitude value. Based on this, the specific process of optimizing the parameters of the fine-tuning model after the current iteration based on the second difference information may include: determining the second difference information, determining the parameters of the model that need to be optimized, and performing the following steps for each parameter that needs to be optimized in the fine-tuning model after the current iteration: determining the third correspondence relationship applicable to the parameter, selecting the parameter adjustment amplitude value corresponding to the second difference information from the third correspondence relationship, and adjusting the parameter value of the parameter by a corresponding amplitude based on the selected parameter adjustment amplitude value.
[0109] Second, the specific process of optimizing the parameters of the fine-tuning model after the current iteration may include the following steps: determining the changing trend data of the loss data based on the loss data of the fine-tuning model after the current iteration and the loss data of the fine-tuning model of the fourth number of iterations before the current iteration; optimizing the parameters of the fine-tuning model after the current iteration based on the changing trend data of the loss data.
[0110] A fourth corresponding relationship applicable to each parameter is preset, and the fourth corresponding relationship is used to indicate the corresponding relationship between the changing trend data of the sample loss data and the parameter adjustment amplitude value. After determining the changing trend data of the loss data, the parameter to be optimized is determined, and the following steps are performed for each parameter to be optimized in the fine-tuning model after the current iteration: determining the fourth corresponding relationship applicable to the parameter, selecting the parameter adjustment amplitude value corresponding to the compliance changing trend data from the fourth corresponding relationship, and adjusting the parameter value of the parameter by a corresponding amplitude based on the selected parameter adjustment amplitude value.
[0111] Third, a fifth corresponding relationship applicable to each parameter is preset, and the fifth corresponding relationship is used to indicate the corresponding relationship between the first difference interval, the second difference interval and the parameter adjustment amplitude value. Based on this, the specific process of optimizing the parameters of the fine-tuning model after the current iteration may include: determining the first difference between the probability that the fine-tuning model after the current iteration outputs alarm analysis data that does not conform to the target user's preferences and the probability that the reference model corresponding to the current iteration outputs alarm analysis data that does not conform to the target user's preferences, determining the second difference between the probability that the fine-tuning model after the current iteration outputs alarm analysis data that conforms to the target user's preferences and the probability that the reference model outputs alarm analysis data that conforms to the target user's preferences; for each parameter that needs to be optimized in the fine-tuning model after the current iteration, perform the following steps respectively: search for the first difference interval containing the first difference and the second difference interval containing the second difference, select the parameter adjustment amplitude corresponding to the first difference interval and the second difference interval found, and adjust the parameter value of the parameter by the corresponding amplitude based on the selected parameter call amplitude value.
[0112] At least one of the three methods for optimizing the parameters of the fine-tuned model after the current iteration can be flexibly selected based on business needs, and this embodiment does not limit this. It should be noted that existing parameter optimization methods can also be used to optimize the parameters of the fine-tuned model after the current iteration. These existing parameter optimization methods will not be described in detail here.
[0113] After optimizing the parameters of the fine-tuning model after the current iteration, return to the step of iteratively training the fine-tuning model based on the preference dataset to continue to adjust the preference of the fine-tuning model.
[0114] Furthermore, in some embodiments, the method for obtaining a preference dataset based on iterative training of the fine-tuning model in method 2 may include the following steps, if necessary: collecting multiple sample alarm data. For each sample alarm data, separately executing: calling a specified model to process the sample alarm data to obtain a plurality of inference chains corresponding to the alarm analysis data output by the model; screening a target inference chain from the inference chains, wherein the target inference chain includes an error step among the inference steps, where the error step is a step that causes the alarm analysis data to not conform to the target user's preferences; based on the correct steps preceding the error step in the target inference chain, determining subsequent correct steps for obtaining alarm analysis data that conforms to the target user's preferences; constructing all correct steps preceding the error step and the subsequent correct steps in order as alarm analysis data that conforms to the target user's preferences corresponding to the sample alarm data; and determining the error step and all correct steps preceding the error step in the target inference chain as alarm analysis data that does not conform to the target user's preferences corresponding to the sample alarm data. Based on the alarm analysis data that conforms to the target user's preferences and the alarm analysis data that does not conform to the target user's preferences corresponding to each sample alarm data, a preference dataset is formed. Such a preference data set enables the model to learn how to use the correct reasoning steps, and based on the learned correct reasoning steps, it can infer alarm analysis data that meets the preferences of the target user.
[0115] For example, the process of constructing the preference dataset described above is illustrated using sample alarm data A. The process specifically includes: invoking a specified model to process sample alarm data A, obtaining multiple inference chains corresponding to the model's output alarm analysis data. Each inference chain presents a series of inference steps: y = s1, s2, ..., sn. Then, a target inference chain 1 is selected from the inference chains. The target inference chain 1 includes an error step sk, which is a step that causes the alarm analysis data to not conform to the user's preferences. Then, based on all correct steps (s1 to sk-1) preceding the error step in the target inference chain, subsequent correct steps are determined for obtaining alarm analysis data that conforms to the user's preferences. All correct steps (s1 to sk-1) preceding the error step and subsequent correct steps are constructed, in order of steps, as the alarm analysis data corresponding to sample alarm data A that conforms to the target user's preferences. Then, the error step sk and all correct steps (s1 to sk-1) preceding the error step in the target inference chain are determined as the alarm analysis data corresponding to sample alarm data A that does not conform to the target user's preferences.
[0116] Furthermore, in order to further improve the effect of preference adjustment, after optimizing the parameters of the fine-tuning model after the current iteration in step C3, you may not return to the step of iteratively training the fine-tuning model based on the preference data set, but first perform the following steps: determine the fine-tuning model after the current iteration as the designated model, collect the alarm data in the preference data set used in the current iteration as sample alarm data, and execute the aforementioned method for obtaining the preference data set, and then form a new preference data set. After replacing the original preference data set with the new preference data set, return to the step of iteratively training the fine-tuning model based on the preference data set. In this way, the preference data set can be continuously updated during the iteration process, so that the fine-tuning model can continuously improve its alarm analysis and judgment capabilities that meet the preferences of the target user during the preference adjustment process.
[0117] Methods 1 and 2 for adjusting the preferences of the fine-tuned model based on the preference dataset to obtain the target model can be flexibly selected based on business needs and are not limited in this embodiment. For example, when the data volume in the preference dataset reaches a preset volume or training resources are limited, method 1 is used. When the data volume in the preference dataset does not reach the preset volume or training resources are sufficient, method 2 is used.
[0118] 103. Perform alarm analysis and quality assessment on the target model from at least one quality assessment dimension.
[0119] After determining the target model, it is necessary to perform an alarm analysis quality assessment on the target model from at least one quality assessment dimension to determine whether the target model is capable of high-quality alarm analysis.
[0120] The specific process of performing alarm analysis and quality assessment on the target model from at least one quality assessment dimension may include the following steps 103A to 103C:
[0121] 103A. Call the target model to process the alarm data in the test set and obtain the alarm analysis data output by the model.
[0122] The test set includes multiple alarm data and the preset alarm analysis data corresponding to each alarm data. The target model is called to process the alarm data in the test set to obtain the alarm analysis data output by the model. The alarm analysis data output by the model is used to verify the alarm analysis quality of the target model.
[0123] 103B. From the alarm analysis data output by the evaluation model in each quality evaluation dimension, the quality evaluation data corresponding to the target model in each evaluation dimension is obtained. The quality evaluation data is used to indicate the alarm analysis quality of the target model in the corresponding evaluation dimension.
[0124] The quality assessment dimensions may include but are not limited to at least one of the following: accuracy, false alarm rate, missed alarm rate, F1 value, completeness of the answer, naturalness of the language expression, and richness of the expression. For the analysis and judgment tasks, it is necessary to check the completeness of the thinking chain of the alarm analysis and judgment process, the clarity of the process description, the correctness of the conclusion, etc.
[0125] For each quality assessment dimension, the following steps are performed respectively: the evaluation model of the quality assessment dimension is called, and based on the preset alarm analysis and judgment data corresponding to each alarm data in the test set and the alarm analysis and judgment data corresponding to each alarm data in the test set output by the model, the quality assessment data corresponding to each evaluation dimension of the target model is obtained. The evaluation model is used to perform corresponding quality assessment using the evaluation logic of the corresponding quality assessment dimension.
[0126] 103C. Based on the quality assessment data corresponding to each assessment dimension, the target model is evaluated for alarm analysis quality to obtain an assessment result indicating whether the alarm analysis quality of the target model meets the requirements.
[0127] If the quality assessment data corresponding to each assessment dimension indicates that the alarm analysis quality of the target model in the corresponding assessment dimension reaches the corresponding alarm analysis quality threshold of each assessment dimension, then an assessment result is obtained indicating that the alarm analysis quality of the target model meets the requirements.
[0128] If the quality assessment data corresponding to any assessment dimension indicates that the alarm analysis quality of the target model in the corresponding assessment dimension does not reach the corresponding alarm analysis quality threshold of the assessment dimension, an assessment result is obtained indicating that the alarm analysis quality of the target model does not meet the requirements.
[0129] In some embodiments, the quality assessment dimensions may include objective category assessment dimensions and subjective category assessment dimensions, wherein the objective category assessment dimensions do not require human intervention, while the subjective category assessment dimensions require the intervention of evaluators. In this way, the objective and subjective dimensions can be combined to evaluate the alarm analysis quality of the target model, so as to improve the accuracy of the target model quality assessment.
[0130] For example, accuracy, false alarm rate, missed alarm rate, and F1 value are categorized as objective categories, while answer completeness, naturalness of language expression, and richness of expression are categorized as subjective categories. For each customer category's quality assessment dimension, the evaluation model for that quality assessment dimension is invoked. Based on the preset alarm assessment data corresponding to each alarm data in the test set and the alarm assessment data corresponding to each alarm data in the test set output by the model, quality assessment data corresponding to each assessment dimension of the target model is obtained. For each subjective category's quality assessment dimension, the following steps are performed separately: the preset alarm assessment data corresponding to each alarm data in the test set and the alarm assessment data corresponding to each alarm data in the test set output by the model are submitted to the evaluation terminal for that quality assessment dimension. The evaluation terminal establishes detailed scoring criteria and, based on the scoring criteria, guides the corresponding evaluators to obtain quality assessment data corresponding to the target model in that assessment dimension based on the data received by the evaluation terminal. By combining the quality assessment data corresponding to each objective category's quality assessment dimension and the quality assessment data corresponding to each subjective category's quality assessment dimension, the target model is evaluated for alarm assessment instructions, obtaining an evaluation result indicating whether the target model's alarm assessment quality meets the requirements.
[0131] 104. If it is assessed that the alarm analysis quality of the target model meets the requirements, the target model will be determined as the trained alarm analysis model.
[0132] If it is assessed that the alarm analysis quality of the target model meets the requirements, it means that the target model not only meets the target user's preferences, but also can perform high-quality alarm analysis. Therefore, the target model is determined as a trained alarm analysis model for use by users who are about to use the alarm analysis model.
[0133] The alarm analysis model training method provided in the embodiment of the present application can achieve at least the following effects: First, when training the alarm analysis model, the model is adjusted based on the preference data set that meets the target user's preferences, so that the model can process the alarm data and output alarm analysis data that meets the target user's preferences. The alarm analysis model trained in this way can perform alarm analysis on the target user that meets the target user's preferences, thereby improving the alarm analysis effect of the alarm analysis model on the target user. Second, after the model is adjusted based on the preference data set, the target model that meets the target user's preferences obtained by the preference adjustment is not directly used as the alarm analysis model. Instead, the target model is first evaluated for alarm quality, and only when it is evaluated that the alarm analysis quality of the target model meets the requirements is the target model determined as the trained alarm analysis model. The alarm analysis model obtained in this way has a high alarm analysis quality, thereby further improving the alarm analysis effect of the alarm analysis model on the target user.
[0134] In some embodiments of the present application, considering that the alarm analysis quality of the target model may not meet the requirements, based on this, after the target model is evaluated for alarm analysis quality from at least one quality assessment dimension in step 103, the alarm analysis model training method provided in this embodiment may also include the following steps: if it is assessed that the alarm analysis quality of the target model does not meet the requirements, then determine the target dimension that causes the alarm analysis quality to not meet the requirements. The alarm analysis quality of the target model in the target dimension does not meet the requirements. Each quality assessment dimension has a model associated with it (fine-tuning model or target model).
[0135] Specifically, if the target dimension is related to the fine-tuning model, a first operation is performed. After the first operation is completed, the process returns to step 101 to fine-tune the initial model trained as the alarm analysis model to obtain a fine-tuned model with alarm analysis capabilities, thereby obtaining a new fine-tuned model. The first operation may include at least one of the following: optimizing parameters of the initial model or replacing the alarm analysis dataset used to fine-tune the initial model.
[0136] Specifically, if the target dimension is related to the target model, a second operation is performed. After the second operation is completed, the process returns to step 102 to adjust the preferences of the fine-tuning model based on the preference dataset to obtain a target model that meets the preferences of the target user, thereby obtaining a new target model. The second operation may include at least one of the following: optimizing parameters of the fine-tuning model or replacing the preference dataset used as the basis for the preference adjustment.
[0137] Furthermore, the embodiment of the present application provides an alarm analysis method, which is applied to an alarm analysis system. The alarm analysis system is deployed with at least one alarm analysis model trained based on the above-mentioned alarm analysis model training method, such as Figure 2 As shown, the alarm analysis model training method provided in this embodiment may include at least the following steps 201 to 202:
[0138] 201. If alarm data for any user is obtained, the alarm analysis model that meets the user's preferences is called to process the user's alarm data to obtain the alarm analysis data output by the model.
[0139] The alarm analysis method provided in this embodiment is applied to the alarm analysis system. At least one alarm analysis model is deployed in the alarm analysis system. Each alarm analysis model has applicable user preferences. The alarm analysis model is used to process alarm data and output alarm analysis data that meets the corresponding user preferences.
[0140] The alarm analysis model adapted to the user's preferences is called to process the alarm data, resulting in the alarm analysis data output by the alarm analysis model. This alarm analysis data satisfies the user's preferences, making it easier for the user to quickly interpret the alarm analysis data and more quickly locate and resolve security incidents.
[0141] 202. Based on the alarm analysis data output by the model, alarm handling is carried out for users.
[0142] After obtaining the alarm analysis data, the alarm analysis data is analyzed to locate the security incidents corresponding to the alarm data, and corresponding handling strategies are formulated for the security incidents. Alarm handling of security incidents is carried out based on the handling strategies to eliminate security incidents and ensure user safety operations.
[0143] The alarm analysis method provided in this embodiment, upon obtaining alarm data for any user, invokes an alarm analysis model that conforms to the user's preferences to process the user's alarm data, obtaining alarm analysis data output by the model, and then performs alarm processing for the user based on the alarm analysis data output by the model. In this way, because the alarm data is processed using an alarm analysis model that conforms to the user's preferences, alarm analysis data that conforms to the user's preferences can be obtained. Based on this alarm analysis data, alarm analysis can be performed on the user in accordance with the user's preferences, thereby improving the effectiveness of alarm analysis for the user.
[0144] Furthermore, an embodiment of the present application also provides an alarm analysis model training device, such as Figure 3 As shown, the alarm analysis model training device provided in this embodiment may include:
[0145] A fine-tuning module 31 is used to fine-tune the initial model to be trained as an alarm analysis model to obtain a fine-tuned model with alarm analysis capabilities;
[0146] An adjustment module 32 is configured to perform preference adjustment on the fine-tuning model based on the preference data set to obtain a target model that meets the preferences of the target user. The target model is configured to process the alarm data and output alarm analysis and judgment data that meets the preferences of the target user. The preference data set includes a plurality of preference data pairs, each of which includes alarm analysis and judgment data that meets the preferences of the target user and alarm analysis and judgment data that does not meet the preferences of the target user, corresponding to the alarm data.
[0147] An evaluation module 33 is configured to perform an alarm analysis and quality evaluation on the target model from at least one quality evaluation dimension;
[0148] The determination module 34 is configured to determine the target model as a trained alarm analysis model if it is assessed that the alarm analysis quality of the target model meets the requirements.
[0149] The alarm analysis model training device provided by the embodiment of the present application can achieve at least the following effects: First, when training the alarm analysis model, the model is adjusted based on the preference data set that meets the preferences of the target user, so that the model can process the alarm data and output the alarm analysis data that meets the preferences of the target user. The alarm analysis model obtained by training in this way can perform alarm analysis on the target user in accordance with the preferences of the target user, thereby improving the alarm analysis effect of the alarm analysis model on the target user. Second, after the model is adjusted based on the preference data set, the target model that meets the preferences of the target user obtained by the preference adjustment is not directly used as the alarm analysis model. Instead, the target model is first evaluated for alarm quality, and only when it is evaluated that the alarm analysis quality of the target model meets the requirements is the target model determined as the trained alarm analysis model. The alarm analysis model obtained in this way has a high alarm analysis quality, thereby further improving the alarm analysis effect of the alarm analysis model on the target user.
[0150] In some embodiments of the present application, Figure 4 As shown, the adjustment module 32 may include:
[0151] An acquisition unit 321 is configured to acquire a reward model trained based on the preference data set, wherein the reward model is used to determine the degree to which the alarm analysis data output by the fine-tuning model conforms to the target user's preferences;
[0152] The first training unit 322 is used to iteratively train the fine-tuning model; after each round of iteration is completed, the fine-tuning model after the current iteration is called to process the first alarm data in the first alarm data set to obtain the first alarm analysis data output by the model, and the degree of compliance of the first alarm analysis data output by the model with the target user's preference is determined based on the reward model. If it is determined that the degree of compliance of the first alarm analysis data output by the model with the target user's preference meets the requirements, the fine-tuning model after the current iteration is determined as the target model. If it is determined that the degree of compliance of the first alarm analysis data output by the model with the target user's preference does not meet the requirements, the parameters of the fine-tuning model after the current iteration are optimized, and the step of iteratively training the fine-tuning model is returned.
[0153] In some embodiments of the present application, Figure 4 As shown, the first training unit 322 may include:
[0154] The first optimization subunit 3221 is used to determine the first difference information between the degree of compliance of the first alarm analysis data output by the fine-tuning model after the current iteration with the target user's preference and the first ideal degree of compliance; and optimize the parameters of the fine-tuning model after the current iteration based on the first difference information.
[0155] In some embodiments of the present application, Figure 4As shown, the first training unit 322 may include:
[0156] The second optimization sub-unit 3222 is used to determine the conformity degree change trend data based on the conformity degree of the first alarm analysis data output by the fine-tuning model after the current iteration to the target user's preferences and the conformity degree of the first alarm analysis data output by the fine-tuning model of the first number of iterations adjacent to the current iteration to the target user's preferences; and optimize the parameters of the fine-tuning model after the current iteration based on the conformity degree change trend data.
[0157] In some embodiments of the present application, Figure 4 As shown, the first training unit 322 may further include:
[0158] The first judgment sub-unit 3223 is used to judge whether the proportion of the first alarm analysis data whose compliance degree is not less than the second ideal compliance degree in the first alarm analysis data output by the fine-tuning model after the current iteration is not less than the target proportion. If not, it is judged that the compliance degree of the first alarm analysis data output by the model with the target user preference meets the requirements. If less, it is judged that the compliance degree of the first alarm analysis data output by the model with the target user preference does not meet the requirements.
[0159] In some embodiments of the present application, Figure 4 As shown, the adjustment module 32 may include:
[0160] The judgment unit 323 is used to judge whether, after the first training unit 322 determines that the first alarm analysis data output by the model does not meet the requirements for the target user's preference, the number of first alarm analysis data with a degree of compliance in the target compliance range appears to be not less than a third number in a second number of consecutive iterations; if so, the alarm analysis data set based on which the iterative training fine-tuning model is based is replaced, and based on the replaced alarm analysis data set, the first training unit 322 is triggered to return to the step of iteratively training the fine-tuning model; if not, the first training unit 322 is triggered to execute the step of optimizing the parameters of the fine-tuning model after the current iteration; the degree of compliance in the target compliance range cannot distinguish whether the alarm analysis data meets the target user's preference.
[0161] In some embodiments of the present application, Figure 4 As shown, the adjustment module 32 may include:
[0162] The second training unit 324 is used to iteratively train the fine-tuning model based on the preference data set; after each round of iteration is completed, it is determined whether the probability of the fine-tuning model after the current iteration outputting alarm analysis data that does not meet the target user's preference is reduced relative to the reference model corresponding to the current iteration, and the probability of the fine-tuning model after the current iteration outputting alarm analysis data that meets the target user's preference is increased relative to the reference model. If both meet the requirements, the fine-tuning model after the current iteration is determined as the target model; if both do not meet the requirements, the parameters of the fine-tuning model after the current iteration are optimized, and the step of iteratively training the fine-tuning model based on the preference data set is returned to execute; wherein, the reference model corresponding to the current iteration is the fine-tuning model obtained after fine-tuning the initial model, or the reference model corresponding to the current iteration is the fine-tuning model after the previous iteration.
[0163] In some embodiments of the present application, Figure 4 As shown, the second training unit 324 may include:
[0164] A calling subunit 3241 is configured to call the reference model and the fine-tuning model after the current iteration to respectively process the second alarm data in the second alarm data set, and obtain a first probability that the reference model outputs alarm analysis and judgment data that meets the target user's preferences for each second alarm data and a second probability that the reference model outputs alarm analysis and judgment data that does not meet the target user's preferences, and obtain a third probability that the fine-tuning model after the current iteration outputs alarm analysis and judgment data that meets the target user's preferences for each second alarm data and a fourth probability that the fine-tuning model outputs alarm analysis and judgment data that does not meet the target user's preferences;
[0165] Determine sub-unit 3242, which is used to determine the loss data corresponding to the fine-tuning model after the current iteration based on the first probability, second probability, third probability, fourth probability and byte length of the alarm analysis data corresponding to each second alarm data; if it is determined that the loss data meets the loss requirements, it is determined that all meet the requirements; if it is determined that the loss data does not meet the loss requirements, it is determined that not all meet the requirements.
[0166] In some embodiments of the present application, Figure 4 As shown, the second training unit 324 may include:
[0167] The third optimization subunit 3243 is used to determine second difference information between the loss data of the fine-tuning model after the current iteration and the ideal loss data; and optimize the parameters of the fine-tuning model after the current iteration based on the second difference information.
[0168] In some embodiments of the present application, Figure 4 As shown, the second training unit 324 may include:
[0169] The fourth optimization subunit 3244 is used to determine the change trend data of the loss data based on the loss data of the fine-tuning model after the current iteration and the loss data of the fine-tuning model of the fourth number of iterations adjacent to the current iteration; and optimize the parameters of the fine-tuning model after the current iteration based on the change trend data of the loss data.
[0170] In some embodiments of the present application, Figure 4 As shown, the alarm analysis model training device provided in this embodiment may also include:
[0171] A module 35 is set up to collect multiple sample alarm data; for each sample alarm data, the following are executed respectively: calling a specified model to process the sample alarm data to obtain a plurality of reasoning chains corresponding to the alarm analysis and judgment data output by the model; screening out a target reasoning chain from the reasoning chain, wherein the target reasoning chain includes an error step in the reasoning steps, and the error step is a step that causes the alarm analysis and judgment data to not conform to the target user's preference; based on all correct steps in the target reasoning chain that precede the error step, determining subsequent correct steps for obtaining the alarm analysis and judgment data that conforms to the target user's preference, and constructing all correct steps before the error step and the subsequent correct steps into the alarm analysis and judgment data that conforms to the target user's preference corresponding to the sample alarm data in accordance with the step sequence; determining the error step and all correct steps in the target reasoning chain that precede the error step as the alarm analysis and judgment data that does not conform to the target user's preference corresponding to the sample alarm data; forming a preference data set based on the alarm analysis and judgment data that conforms to the target user's preference and the alarm analysis and judgment data that does not conform to the target user's preference corresponding to each sample alarm data.
[0172] In some embodiments of the present application, Figure 4 As shown, the evaluation module 33 is specifically used to call the target model to process the alarm data in the test set to obtain the alarm analysis data output by the model; evaluate the alarm analysis data output by the model from each quality evaluation dimension to obtain the quality evaluation data corresponding to the target model in each evaluation dimension, and the quality evaluation data is used to indicate the alarm analysis quality of the target model in the corresponding evaluation dimension; based on the quality evaluation data corresponding to each evaluation dimension, the target model is evaluated for alarm analysis quality to obtain an evaluation result indicating whether the alarm analysis quality of the target model meets the requirements.
[0173] In some embodiments of the present application, Figure 4 As shown, the alarm analysis model training device provided in this embodiment may also include:
[0174] The iteration module 36 is used to determine the target dimension that causes the alarm analysis quality to not meet the requirements if it is assessed that the alarm analysis quality of the target model does not meet the requirements; if the target dimension is related to the fine-tuning model, a first operation is executed, and after the first operation is completed, the fine-tuning module 31 is triggered to return to execute the step of fine-tuning the initial model to be trained as the alarm analysis model to obtain a fine-tuning model with alarm analysis capabilities; the first operation includes at least one of the following: optimizing the parameters of the initial model, replacing the alarm analysis data set based on which the fine-tuning initial model is based; if the target dimension is related to the target model, a second operation is executed, and after the second operation is completed, the adjustment module 32 is triggered to return to execute the step of performing preference adjustment on the fine-tuning model based on the preference data set to obtain a target model that meets the preferences of the target user, and the second operation includes at least one of the following: optimizing the parameters of the fine-tuning model, replacing the preference data set based on which the preference adjustment is based.
[0175] In some embodiments of the present application, Figure 4 As shown, the fine-tuning module 31 can also be used to obtain a target alarm analysis data set, and based on the target alarm analysis data set, perform the steps of fine-tuning the initial model to be trained as an alarm analysis model to obtain a fine-tuned model with alarm analysis capabilities, wherein the target alarm analysis data set includes multiple target alarm data and target alarm analysis data with a preset format corresponding to each target sample alarm data, and the target alarm analysis data with a preset format is used to realize the ability of the fine-tuned initial model to output alarm analysis data with a preset format.
[0176] In the alarm analysis model training device provided in the embodiment of the present application, the detailed explanations used during the operation of each functional module can be found in the corresponding detailed explanations of the above-mentioned alarm analysis model training method embodiment, and will not be repeated here.
[0177] Furthermore, an embodiment of the present application also provides an alarm analysis device, which is applied to an alarm analysis system, wherein the alarm analysis system is deployed with at least one alarm analysis model trained based on the above-mentioned alarm analysis model training method, such as Figure 5 As shown, the alarm analysis and judgment device provided in this embodiment may include:
[0178] The calling module 41 is configured to, upon obtaining alarm data for any user, call an alarm analysis model that conforms to the user's preferences to process the user's alarm data and obtain alarm analysis data output by the model;
[0179] The handling module 42 is used to handle the alarm for the user based on the alarm analysis data output by the model.
[0180] The alarm analysis device provided in the embodiment of the present application, upon obtaining alarm data for any user, invokes an alarm analysis model that conforms to the user's preferences to process the user's alarm data, obtains alarm analysis data output by the model, and then performs alarm processing on the user based on the alarm analysis data output by the model. In this way, since the alarm data is processed using an alarm analysis model that is suitable for the user's preferences, alarm analysis data that conforms to the user's preferences can be obtained. Based on such alarm analysis data, an alarm analysis can be performed on the user that conforms to the user's preferences, thereby improving the effectiveness of the user's alarm analysis.
[0181] In the alarm analysis device provided in the embodiment of the present application, the detailed explanations used during the operation of each functional module can be found in the corresponding detailed explanations of the above-mentioned alarm analysis method embodiment, and will not be repeated here.
[0182] Furthermore, an embodiment of the present application also provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the device where the storage medium is located is controlled to execute the above-mentioned alarm analysis model training method, and / or the above-mentioned alarm analysis method.
[0183] Furthermore, an embodiment of the present application also provides an electronic device, which includes: a memory for storing a program; a processor, coupled to the memory, for running the program to execute the above-mentioned alarm analysis model training method, and / or the above-mentioned alarm analysis method.
[0184] Furthermore, an embodiment of the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the above-mentioned alarm analysis model training method, and / or the above-mentioned alarm analysis method.
[0185] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0186] It is understood that the relevant features of the above methods and devices can be referenced to each other. In addition, the terms "first" and "second" in the above embodiments are used to distinguish between the embodiments, and do not represent the advantages and disadvantages of the embodiments.
[0187] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0188] The algorithm and display provided herein are not inherently related to any particular computer, virtual system or other device. Various general-purpose systems can also be used together with the teachings based on this. According to the above description, it is obvious that the structure required for constructing such systems. In addition, the application is not directed to any specific programming language. It should be understood that various programming languages can be utilized to implement the content of the application described herein, and the above description of specific languages is for the purpose of disclosing the preferred embodiment of the application.
[0189] In addition, the memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0190] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0191] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data cutover device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data cutover device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0192] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data switching device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0193] These computer program instructions can also be loaded onto a computer or other programmable data switching device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0194] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0195] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0196] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0197] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0198] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0199] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for training an alarm analysis model, characterized in that: The method comprises: Fine-tune the initial model trained as an alarm analysis model to obtain a fine-tuned model with alarm analysis capabilities; Based on the preference data set, the fine-tuning model is adjusted for preferences to obtain a target model that meets the preferences of the target user. The target model is used to process the alarm data and output alarm analysis and judgment data that meets the preferences of the target user. The preference data set includes multiple preference data pairs, and the preference data pairs include the alarm analysis and judgment data that meets the preferences of the target user and the alarm analysis and judgment data that does not meet the preferences of the target user corresponding to the alarm data; Performing an alarm analysis and quality assessment on the target model from at least one quality assessment dimension; If it is assessed that the alarm analysis quality of the target model meets the requirements, the target model is determined as the trained alarm analysis model.
2. The method according to claim 1, characterized in that Based on the preference dataset, the fine-tuning model is adjusted to obtain a target model that meets the target user's preferences, including: Obtaining a reward model trained based on the preference data set, wherein the reward model is used to determine the degree to which the alarm analysis and judgment data output by the fine-tuning model conforms to the target user's preferences; Iteratively train the fine-tuned model; After each round of iteration is completed, the fine-tuning model after the current iteration is called to process the first alarm data in the first alarm data set to obtain the first alarm analysis data output by the model. The degree of compliance of the first alarm analysis data output by the model with the target user's preferences is determined based on the reward model. If it is determined that the degree of compliance of the first alarm analysis data output by the model with the target user's preferences meets the requirements, the fine-tuning model after the current iteration is determined as the target model. If it is determined that the degree of compliance of the first alarm analysis data output by the model with the target user's preferences does not meet the requirements, the parameters of the fine-tuning model after the current iteration are optimized, and the step of iteratively training the fine-tuning model is returned.
3. The method according to claim 2, characterized in that Optimizing the parameters of the fine-tuning model after the current iteration, including: determining first difference information between the degree of conformity of the first alarm analysis data output by the fine-tuning model after the current iteration with the target user's preference and a first ideal degree of conformity; and optimizing the parameters of the fine-tuning model after the current iteration based on the first difference information; and / or, Optimizing the parameters of the fine-tuning model after the current iteration, including: determining the conformity change trend data based on the conformity degree of the first alarm analysis data output by the fine-tuning model after the current iteration to the target user's preferences and the conformity degree of the first alarm analysis data output by the fine-tuning model of the first number of iterations adjacent to the current iteration to the target user's preferences; optimizing the parameters of the fine-tuning model after the current iteration based on the conformity change trend data.
4. The method according to claim 2, characterized in that The method further includes: determining whether the proportion of the first alarm analysis and judgment data having a degree of conformity not less than a second ideal degree of conformity in the first alarm analysis and judgment data output by the fine-tuning model after the current iteration is not less than a target proportion; if not, determining that the degree of conformity of the first alarm analysis and judgment data output by the model to the target user's preference meets the requirement; if less, determining that the degree of conformity of the first alarm analysis and judgment data output by the model to the target user's preference does not meet the requirement; and / or, After determining that the first alarm analysis data output by the model does not meet the requirements for the degree of compliance with the target user's preference, the method further includes: determining whether the number of first alarm analysis data with a degree of compliance in the target compliance range appears to be not less than a third number in a second consecutive iteration; if so, replacing the alarm analysis data set used as the basis for the iterative training of the fine-tuning model, and returning to the step of iteratively training the fine-tuning model based on the replaced alarm analysis data set; if not, executing the step of optimizing the parameters of the fine-tuning model after the current iteration; the degree of compliance in the target compliance range cannot distinguish whether the alarm analysis data meets the target user's preference.
5. The method according to claim 1, wherein Based on the preference dataset, the fine-tuning model is adjusted to obtain a target model that meets the target user's preferences, including: Iteratively training the fine-tuning model based on the preference dataset; After each round of iteration is completed, it is determined whether the probability of the fine-tuning model after the current iteration outputting alarm analysis data that does not meet the target user's preferences is reduced relative to the corresponding reference model of the current iteration, and the probability of the fine-tuning model after the current iteration outputting alarm analysis data that meets the target user's preferences is increased relative to the reference model. If both meet the requirements, the fine-tuning model after the current iteration is determined as the target model. If both do not meet the requirements, the parameters of the fine-tuning model after the current iteration are optimized, and the step of iteratively training the fine-tuning model based on the preference data set is returned to execute; The reference model corresponding to the current iteration is the fine-tuned model obtained after fine-tuning the initial model, or the reference model corresponding to the current iteration is the fine-tuned model after the previous iteration.
6. The method according to claim 5, characterized in that Determining whether the decrease in the probability of the fine-tuning model after the current iteration outputting alarm analysis data that does not conform to the target user's preferences relative to the reference model corresponding to the current iteration, and the increase in the probability of the fine-tuning model after the current iteration outputting alarm analysis data that conforms to the target user's preferences relative to the reference model both meet requirements, including: Calling the reference model and the fine-tuning model after the current iteration to respectively process the second alarm data in the second alarm data set, obtaining a first probability that the reference model outputs alarm analysis and judgment data that meets the target user's preferences for each second alarm data and a second probability that the reference model outputs alarm analysis and judgment data that does not meet the target user's preferences, and obtaining a third probability that the fine-tuning model after the current iteration outputs alarm analysis and judgment data that meets the target user's preferences for each second alarm data and a fourth probability that the fine-tuning model outputs alarm analysis and judgment data that does not meet the target user's preferences; Determine the loss data corresponding to the fine-tuning model after the current iteration based on the first probability, the second probability, the third probability, the fourth probability and the byte length of the alarm analysis data corresponding to each second alarm data; If the loss data is determined to meet the loss requirements, then the requirements are met; If the loss data is determined to not meet the loss requirements, it is determined that the requirements are not met.
7. The method according to claim 5, characterized in that Optimizing the parameters of the fine-tuning model after the current iteration, including: determining second difference information between the loss data of the fine-tuning model after the current iteration and the ideal loss data; and optimizing the parameters of the fine-tuning model after the current iteration based on the second difference information; and / or, Optimizing the parameters of the fine-tuning model after the current iteration, including: determining the change trend data of the loss data based on the loss data of the fine-tuning model after the current iteration and the loss data of the fine-tuning model of the fourth number of iterations before the current iteration; optimizing the parameters of the fine-tuning model after the current iteration based on the change trend data of the loss data.
8. The method according to claim 5, characterized in that The method further comprises: Collect multiple sample alarm data; For each sample alarm data, the following are executed respectively: calling a specified model to process the sample alarm data, and obtaining a plurality of reasoning chains corresponding to the alarm analysis and judgment data output by the model; screening out a target reasoning chain from the reasoning chain, wherein there is an error step in the reasoning steps included in the target reasoning chain, and the error step is a step that causes the alarm analysis and judgment data to not conform to the target user's preference; based on all correct steps preceding the error step in the target reasoning chain, determining subsequent correct steps for obtaining the alarm analysis and judgment data that conforms to the target user's preference, and constructing all correct steps preceding the error step and the subsequent correct steps into the alarm analysis and judgment data corresponding to the sample alarm data that conforms to the target user's preference in accordance with the step sequence; determining the error step and all correct steps preceding the error step in the target reasoning chain as the alarm analysis and judgment data corresponding to the sample alarm data that does not conform to the target user's preference; Based on the alarm analysis data that meets the target user's preferences and the alarm analysis data that does not meet the target user's preferences for each sample alarm data, a preference data set is formed.
9. The method according to any one of claims 1 to 8, characterized in that Performing an alarm analysis and judgment quality assessment on the target model from at least one quality assessment dimension, including: calling the target model to process the alarm data in the test set to obtain the alarm analysis and judgment data output by the model; evaluating the alarm analysis and judgment data output by the model from each quality assessment dimension to obtain the quality assessment data corresponding to each assessment dimension of the target model, the quality assessment data being used to indicate the alarm analysis and judgment quality of the target model in the corresponding assessment dimension; performing an alarm analysis and judgment quality assessment on the target model based on the quality assessment data corresponding to each assessment dimension to obtain an assessment result indicating whether the alarm analysis and judgment quality of the target model meets the requirements; and / or, The method also includes: if it is assessed that the alarm analysis quality of the target model does not meet the requirements, determining the target dimension that causes the alarm analysis quality to not meet the requirements; if the target dimension is related to the fine-tuning model, executing a first operation, and after the first operation is completed, returning to execute the step of fine-tuning the initial model to be trained as the alarm analysis model to obtain a fine-tuning model with alarm analysis capabilities; the first operation includes at least one of the following: optimizing the parameters of the initial model, replacing the alarm analysis data set based on which the initial model is fine-tuned; if the target dimension is related to the target model, executing a second operation, and after the second operation is completed, returning to execute the step of adjusting the preferences of the fine-tuning model based on the preference data set to obtain a target model that meets the preferences of the target user, and the second operation includes at least one of the following: optimizing the parameters of the fine-tuning model, and replacing the preference data set based on which the preference adjustment is made; and / or, The method also includes: obtaining a target alarm analysis and judgment data set, and based on the target alarm analysis and judgment data set, executing the step of fine-tuning the initial model to be trained as the alarm analysis and judgment model to obtain a fine-tuned model with alarm analysis and judgment capabilities, wherein the target alarm analysis and judgment data set includes multiple target alarm data and target alarm analysis and judgment data with a preset format corresponding to each target sample alarm data, and the target alarm analysis and judgment data with a preset format is used to realize the ability of the fine-tuned initial model to output alarm analysis and judgment data with a preset format.
10. An alarm analysis method, characterized in that: Applied to an alarm analysis and judgment system, the alarm analysis and judgment system deploys at least one alarm analysis and judgment model trained based on the alarm analysis and judgment model training method according to any one of claims 1 to 9, the method comprising: If alarm data for any user is obtained, the alarm analysis model that meets the user's preferences is called to process the user's alarm data to obtain the alarm analysis data output by the model; Based on the alarm analysis data output by the model, an alarm is taken for the user.
11. An alarm analysis model training device, characterized in that: The device comprises: A fine-tuning module is used to fine-tune the initial model to be trained as an alarm analysis model to obtain a fine-tuned model with alarm analysis capabilities; An adjustment module is configured to perform preference adjustment on the fine-tuning model based on a preference data set to obtain a target model that meets the preferences of the target user. The target model is configured to process the alarm data and output alarm analysis and judgment data that meets the preferences of the target user. The preference data set includes a plurality of preference data pairs, each of which includes the alarm analysis and judgment data that meets the preferences of the target user and the alarm analysis and judgment data that does not meet the preferences of the target user corresponding to the alarm data. An evaluation module, configured to perform an alarm analysis and quality evaluation on the target model from at least one quality evaluation dimension; The determination module is used to determine the target model as the trained alarm analysis model if it is assessed that the alarm analysis quality of the target model meets the requirements.
12. An alarm analysis and judgment device, characterized in that: Applied to an alarm analysis and judgment system, the alarm analysis and judgment system is deployed with at least one alarm analysis and judgment model trained based on the alarm analysis and judgment model training method according to any one of claims 1 to 9, the device comprising: A calling module is used to, if alarm data for any user is obtained, call an alarm analysis model that meets the user's preferences to process the user's alarm data and obtain alarm analysis data output by the model; The handling module is used to handle the alarm of the user based on the alarm analysis data output by the model.
13. A computer-readable storage medium, characterized in that The storage medium includes a stored program, wherein, when the program is running, the device where the storage medium is located is controlled to execute the alarm analysis model training method described in any one of claims 1 to 9, and / or the alarm analysis method described in claim 10.
14. An electronic device, characterized in that: The electronic device comprises: Memory, used to store programs; A processor, coupled to the memory, is used to run the program to execute the alarm analysis model training method described in any one of claims 1 to 9, and / or the alarm analysis method described in claim 10.
Citation Information
Patent Citations
Direct preference optimization method and device
CN118569348A
Network security question and answer method and system based on large model
CN118964581A
Generating digital event sequences utilizing a dynamic user preference interface to modify recommendation model reward functions
US20200033144A1