Task processing method, device, model and system, storage medium and product
By using a task processing model pre-trained with multi-stream data in the intelligent operation and maintenance scenario of wireless networks, the influence of intervention factors is learned, which solves the problem of insufficient accuracy of basic general-purpose large models and achieves more accurate task processing results.
Patent Information
- Application Number
- CN202410921832.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-10
- Publication Date
- 2026-01-13
AI Technical Summary
The existing basic general-purpose model does not consider the influence of intervention factors in the intelligent operation and maintenance scenario of wireless networks, resulting in large deviations in the processing results, poor accuracy, and affecting the effectiveness of subsequent decision-making.
A task processing model based on multi-stream data pre-training, including intervention stream, expectation stream, signal stream, and network element snapshot stream, is adopted. The influence of intervention parameter values is learned through a multi-stream attention mechanism to correct the learning direction of the task processing model.
It improves the accuracy and processing effect of the task processing model, and can better meet the processing results expected by users.
Smart Images

Figure CN121334718A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of communication technology, and in particular to a task processing method, apparatus, model and system, storage medium and product. Background Technology
[0002] One direction of network intelligence is using AI technology to support intelligent operation and maintenance (O&M) scenarios for various wireless networks. Therefore, in intelligent O&M scenarios for wireless networks, AI modeling techniques are used to construct models and then leverage these models to perform various O&M tasks. However, in actual O&M scenarios, tasks are diverse, and related technologies involve methods for training basic, general-purpose large models. This involves using data from various task scenarios to learn the relationships and temporal relationships between indicators to obtain high-dimensional representations at the network element level. Then, downstream tasks are fine-tuned based on these high-dimensional representations.
[0003] However, when using a basic general-purpose model to process tasks, significant model bias and user intervention are frequently encountered. These models primarily focus on capturing relationships between indicators and temporal sequences, neglecting the impact of these interventions on the task processing. Consequently, the results from these models tend to be biased and inaccurate. If subsequent decisions or tasks are made based on these inaccurate model results, the overall outcome will be unsatisfactory. Summary of the Invention
[0004] This disclosure provides a task processing method, apparatus, model and system, storage medium and product, which can improve the accuracy of task processing models and enhance task processing performance to a certain extent.
[0005] Firstly, this disclosure provides a task processing method, including:
[0006] Obtain a task processing model; the task processing model is pre-trained based on multi-stream data, the multi-stream data including at least: an intervention stream, the intervention stream being used to describe intervention parameter values;
[0007] The target task is processed using the task processing model described above.
[0008] Secondly, this disclosure provides a task processing apparatus, including:
[0009] An acquisition module is used to acquire a task processing model; the task processing model is obtained based on multi-stream data pre-training, and the multi-stream data includes at least: an intervention stream, which is used to describe intervention parameter values;
[0010] The processing module is used to process the target task using the task processing model.
[0011] Thirdly, this disclosure provides a task processing model, including multiple atomic modules, wherein the atomic modules include:
[0012] An atomic acquisition module is used to acquire the multi-stream data; the multi-stream data includes: a first intervention stream, a first network element snapshot stream, a desired stream, and a signal stream.
[0013] The mutual attention atom module is used to process the multi-stream data based on the mutual attention mechanism and the parameter tuning time to obtain the second intervention stream;
[0014] A multilayer perceptron atomic module is used to perform multilayer perceptron processing on the second intervention flow and the first network element snapshot flow to obtain the second network element snapshot flow;
[0015] The self-attention atom module is used to process the first network element snapshot stream based on the self-attention mechanism to obtain the third network element snapshot stream;
[0016] A fusion atom module is used to fuse the second network element snapshot stream and the third network element snapshot stream to obtain a fourth network element snapshot stream;
[0017] The restoration atom module is used to restore the snapshot stream of the fourth network element to obtain the state data of the second network element.
[0018] Fourthly, this disclosure provides a task processing system, including:
[0019] As described in the third aspect, the task processing model;
[0020] A task processing apparatus for performing the task processing method as described in any embodiment of the first aspect.
[0021] Fifthly, this disclosure provides an electronic device, including: a memory for storing computer-readable instructions; and a processor for executing the computer-readable instructions, causing the electronic device to perform the method as described in any embodiment of the first aspect.
[0022] In a sixth aspect, this disclosure provides a non-transitory computer-readable storage medium for storing computer-readable instructions that, when executed by a processor, cause the processor to perform the method as described in any embodiment of the first aspect.
[0023] In a seventh aspect, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the method as described in any embodiment of the first aspect.
[0024] This disclosure provides a task processing method, apparatus, model and system, storage medium, and product. The disclosure obtains a task processing model through pre-training on multi-stream data containing intervention streams, and utilizes this model to process a target task. The intervention streams describe intervention parameter values. Therefore, the pre-trained task processing model learns not only the relationships and temporal relationships between indicators, but also the impact of these intervention parameter values on the task processing process. In other words, intervention stream-based learning can correct the learning direction of the task processing model, preventing it from learning in the wrong direction. In summary, the technical solution provided by this disclosure can improve the accuracy of the task processing model and enhance task processing performance to a certain extent.
[0025] It should be understood that both the foregoing general description and the following detailed description are exemplary and intended to provide further illustration of the claimed technology. Attached Figure Description
[0026] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0027] Figure 1 A flowchart illustrating a task processing method provided in an embodiment of this disclosure;
[0028] Figure 2 This is a schematic diagram illustrating the relationship between indicator data and network element status, provided in an embodiment of the present disclosure.
[0029] Figure 3 This is a schematic diagram illustrating the construction of multi-stream data provided in an embodiment of this disclosure;
[0030] Figure 4 A schematic diagram of a signal flow matrix provided in an embodiment of this disclosure;
[0031] Figure 5 This is a schematic diagram of a multi-stream data relationship provided in an embodiment of the present disclosure;
[0032] Figure 6 A schematic diagram of a task processing model provided in an embodiment of this disclosure;
[0033] Figure 7 A schematic diagram of a target combination provided in an embodiment of this disclosure;
[0034] Figure 8 A schematic diagram illustrating another target combination provided in an embodiment of this disclosure;
[0035] Figure 9 A schematic diagram illustrating another target combination provided in an embodiment of this disclosure;
[0036] Figure 10 A schematic diagram illustrating another target combination provided in an embodiment of this disclosure;
[0037] Figure 11 A structural block diagram of a task processing device provided in an embodiment of this disclosure;
[0038] Figure 12 A structural block diagram of a task processing system provided in an embodiment of this disclosure;
[0039] Figure 13 A hardware block diagram of an electronic device provided in an embodiment of this disclosure;
[0040] Figure 14 This is a schematic diagram of a computer-readable storage medium provided in an embodiment of this disclosure. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this disclosure more apparent, exemplary embodiments according to this disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this disclosure, and not all embodiments of this disclosure. It should be understood that this disclosure is not limited to the exemplary embodiments described herein.
[0042] This disclosure applies to intelligent operation and maintenance scenarios for wireless networks. In this scenario, various tasks can be processed using AI models.
[0043] Specifically, this operation and maintenance scenario may involve various task types. For example, this disclosure provides several possible types, including but not limited to: time series prediction, root cause localization, causal inference, parameter tuning decision-making, etc., without exhaustive list.
[0044] Time-series prediction aims to predict key performance indicators (KPIs). For example, historical time-series KPI data can be used to predict target KPI values for the current moment or a future moment. In intelligent wireless network operation and maintenance scenarios, this predictive capability can be achieved through AI models. For instance, time-series KPI and MR data for network elements involved in work orders can be obtained, and AI models can be used to process historical KPI data up to a certain point in time to predict target KPI values for the current moment and beyond.
[0045] Root cause analysis aims to pinpoint the cause of network failures. For example, historical time-series metric data can be used to locate the cause of network failures. In intelligent wireless network operation and maintenance scenarios, AI models can be used to achieve this predictive capability. For instance, time-series metric data such as network element KPIs, MRs, and alarms related to work orders can be obtained. Based on this data, the cause of network failures can be determined. Network failures can involve multiple failure types, therefore root cause analysis can also be considered a multi-classification problem. Failure types can include, but are not limited to, at least one of the following: over-coverage, weak coverage, interference, and faults, without exhaustive list.
[0046] Causal inference, also known as parameter-adjusted causal inference, can be viewed as a time-series prediction task that incorporates human intervention. That is, it predicts the target indicator value at the current moment or at a future moment based on time-series indicator data before a certain moment and human intervention at that moment.
[0047] Parameter tuning decisions aim to provide guidance on when and how to tune parameters. Specifically, when a network experiences high load or other issues, load-related indicators can often be reduced through parameter tuning. How and when to tune parameters can be achieved through AI models. For AI models, the purpose is to predict how parameters need to be adjusted when a target indicator reaches a certain target range.
[0048] In actual wireless network maintenance scenarios, there may be many more types of maintenance tasks involved. This disclosure does not impose any special restrictions on these tasks, nor does it exhaustively list them.
[0049] Given the various task types involved in wireless network maintenance scenarios, the following two approaches are involved in utilizing AI technology to accomplish these tasks:
[0050] One implementation involves siloed, isolated modeling for each task type. This means that for each task type, a separate model is built and trained using data from its respective task scenario, and then deployed in its own scenario's environment. However, this isolated modeling approach suffers from low modeling efficiency, high requirements for data integrity, and the inability to transfer models.
[0051] Another approach, as described in the background section of this disclosure, is to integrate data from various task scenarios to construct and train a basic, general-purpose model for network operations and maintenance. Compared to isolated modeling solutions, this basic, general-purpose model can address issues such as low modeling efficiency and the inability to transfer model data. However, as the background section states, this basic, general-purpose model does not consider the impact of these intervention factors on the task processing, resulting in significant deviations and poor accuracy in its processing results. Consequently, subsequent decisions or task processing based on this model will also be ineffective due to the inaccurate model results.
[0052] To address the aforementioned problems in existing technologies, this disclosure provides a novel technical concept for a task processing solution: pre-training the task processing model based on a multi-stream attention mechanism. This allows the pre-trained model to learn not only the relationships and temporal relationships between indicators, but also the influence of these intervention parameter values on the task processing process. Under the influence of the intervention stream, the learning direction of the task processing model is corrected. This, to a certain extent, improves the accuracy of the task processing model and enhances the task processing effect. The details are explained below.
[0053] This disclosure provides a task processing method. Please refer to... Figure 1 , Figure 1 This is a flowchart illustrating a task processing method provided in an embodiment of this disclosure. Figure 1 As shown, the method includes:
[0054] S102, Obtain the task processing model; the task processing model is pre-trained based on multi-stream data, which includes at least: intervention stream, which describes the intervention parameter values.
[0055] When implementing this step, you can directly call up a pre-trained task processing model, or you can train and generate a task processing model.
[0056] It should be noted that the task processing model involved in this disclosure can be an independent model trained separately for each type of task (the model architecture can be customized based on the actual task), or it can be a basic general model trained comprehensively for multiple types of tasks (this will be explained in detail later using the wireless network operation and maintenance scenario as an example). Furthermore, this disclosure does not impose any particular restrictions on the types of tasks to which the task processing model is applicable. For example, it can be applied to operation and maintenance scenarios, and further, to wireless network operation and maintenance scenarios (please refer to the preceding text, which will not be repeated here). In addition, the task processing model can also be applied to various types of tasks involved in human-computer interaction scenarios, business processing scenarios, etc., without exhaustive listing. In the following text, for ease of explanation, the wireless network operation and maintenance scenario will be used as an example to illustrate this solution.
[0057] The task processing method provided in this disclosure utilizes a task processing model, which is pre-trained based on multi-stream data including intervention streams. Specifically, the multi-stream data involved in this disclosure refers to multiple data streams, each describing different information. Furthermore, the multi-stream data involved in this disclosure includes at least an intervention stream.
[0058] In this disclosure, intervention flow is used to describe intervention parameter values. Intervention parameter values refer to the parameter values used by the user when tuning the task processing model. In other words, intervention flow is the result of user intervention and adjustment, and can be used to reflect the impact of human intervention on the task processing model. In actual scenarios, the type of parameters manually intervened by the user is determined by the user. For example, in the wireless network operation and maintenance scenario, the human intervention parameter is the wireless network operation and maintenance parameter, which may include, but is not limited to, at least one of the following: reference signal power, coverage-based inter-frequency measurement A1 event threshold, coverage-based co-frequency measurement A3 event threshold, etc., without exhaustive list.
[0059] In addition to intervention streams, multi-stream data can also include other types of data. In one exemplary embodiment, besides intervention streams, multi-stream data can also include at least one of the following: expectation streams, signal streams, and network element snapshot streams. Of course, in a preferred embodiment, multi-stream data can include: intervention streams, expectation streams, signal streams, and network element snapshot streams.
[0060] In this disclosure, the expected flow is used to describe the value of a key performance indicator (KPI) at a target time. The target time is generally customizable; for example, it can be a specific moment or a preset duration after a certain condition (e.g., 24 hours after parameter tuning). Furthermore, the KPIs can also be customized based on the actual scenario and are generally defined as load-related indicators, such as the CCE occupancy rate of the PDCCH channel, the average utilization rate of the uplink pRB, the wireless call success rate, the wireless drop rate, and the handover success rate, etc., without exhaustive list. The KPI value refers to the value of the KPI at the target time; in other words, the expected flow indicates the value that the user expects the KPI to reach at the target time. In actual implementation scenarios, the expected flow can be customized by the user.
[0061] In this disclosure, the signal stream is used to record the parameter tuning time. It should be noted that the definition of the parameter tuning time may differ between the pre-training process and the application process of the task processing model, which will be explained in detail later.
[0062] In this disclosure, the network element snapshot stream is used to describe a high-dimensional representation of network element state data. The network element state data refers to the indicator data among various time-series data maintained by the network element. The indicators involved here can also be custom-designed, and the indicators of interest in the desired stream can be all or part of the indicators involved in the network element state data (for ease of distinction, denoted as task indicators). Compared to the network element state data, the network element snapshot stream is the high-dimensional representation result obtained after high-dimensional representation processing of the network element state data.
[0063] In practical implementation scenarios, the aforementioned multi-stream data can generally be obtained from the model's input data. However, the acquisition methods for some multi-stream data differ between model pre-training scenarios and model application scenarios, which will be detailed later.
[0064] S104 utilizes a task processing model to process the target task.
[0065] This step can be implemented in different ways. For example, the task processing model can be directly used to process the target task. Alternatively, the task processing model can be further adjusted and updated (e.g., by using task data from the same scenario as the target task for minor training adjustments) before using the adjusted model to process the target task. For example, if the task processing model is an independent model matching the type of the target task, it can be directly used for task processing; or, if the task processing model is a basic general model for a specific scenario, the basic general model can be split and combined into atomic modules, and then the combination of atomic modules related to the target task can be used to achieve task processing. This embodiment will be explained in more detail later.
[0066] In summary, regardless of the method used to process the target task, since the task processing model is pre-trained on multi-stream data, including intervention streams, it can not only learn static network relationships but also capture dynamic relationships caused by human intervention. This allows the task processing to achieve results that better meet user expectations. Therefore, the technical solution provided in this disclosure can, to a certain extent, improve the accuracy of the task processing model and enhance task processing effectiveness.
[0067] The following section first explains how to acquire multi-stream data. Taking multi-stream data including intervention stream, expectation stream, signal stream, and network element snapshot stream as an example.
[0068] Multistream data can be obtained from the model's input data, which can be the state data of one or more network elements. A network element can be any network function in a wireless communication scenario, such as the AMF (Advanced Management Function) or SMF (Search Engine Function) network element in the core network. Each network element can maintain its own corresponding cell's indicator data, typically in the form of time-series structured data. For example, the indicator data involved in this disclosure may include, but is not limited to, at least one of the following: cell KPI data, cell MR (Mean Average Rank), intervention parameters, resource data, etc., without exhaustive list. It should be understood that the types of indicator data maintained by different network elements can be the same or different, and can be customized based on the actual scenario.
[0069] For example, please refer to Figure 2 , Figure 2 This is a schematic diagram illustrating the relationship between indicator data and network element status, provided in an embodiment of this disclosure. Firstly, as... Figure 2The table shown illustrates the parameter values (i.e., indicator data) of various indicators maintained by the network element at different time slices. It should be understood that the indicators involved in this disclosure may include, but are not limited to, those mentioned above. Figure 2 The part shown, Figure 2 This is for illustrative purposes only. Furthermore, to protect data security, some parameter values have been obfuscated. Figure 2 The specific values are not relevant. In this disclosure, the length of a time slice can be customized. For example, Figure 2 The system maintains metric data in 15-minute time slices. Additionally, a time slice can be set to 30 minutes, 1 hour, etc., without exhaustive listing.
[0070] In this disclosure, the indicator data of k time slices is defined as a network element state (or network element state data, denoted as token); where k is an integer greater than 0. In other words, k represents the number of time slices contained in a network element state, and any network element state includes indicator data of k time slices. The number of indicator data (i.e., the number of indicators) is denoted as s. Thus, a network element state includes parameter values of k*s indicator data. Please refer to [reference needed]. Figure 2 , Figure 2 The value of k is 4, meaning that the status of a network element includes indicator data for 4 time slices. This is for ease of identification. Figure 2 Three network element status data are shown, distinguished by different colors. Furthermore, the k time slices contained in a network element status should be consecutive time slices, meaning one network element status corresponds to a time interval of four time slices. However, this disclosure does not impose any particular restrictions on the starting position of any time interval.
[0071] Based on this, for a given cell, if it contains indicator data for K time slices, then the cell can have m network element states, where m = K / k. Thus, the cell (and its corresponding network element state data) can be represented as: {token1, token2, ..., tokenm}. The model's input data can include network element state data from multiple cells. To facilitate differentiation between different cells, a cls identifier can be added to the network element state data. Any two cells will have different cls identifiers, thus distinguishing the cells corresponding to each network element state data. This cls identifier can be added before each network element state data as a prompt. In this case, the network element state data corresponding to a cell can be represented as: {cls, token1, token2, ..., tokenm}.
[0072] In this disclosure, the input data of the task processing model can be organized in batches. Each batch can include network element status data of multiple cells. After normalizing the network element status data of each cell, input data in a predetermined format can be obtained.
[0073] In a specific embodiment of this disclosure, the input data can satisfy the following format: (batchsize, m, k, s); where batchsize represents the number of samples, m represents the number of network element states, k represents the number of time slices contained in a network element state, and s represents the number of indicators; any network element state includes indicator data for k time slices. Figure 2 The following example illustrates the specific steps. Figure 2 The indicator data is maintained in 15-minute time slices. If one hour is taken as one network element status, then k is 4. In this case, the input data meets the following format: (batchsize, m, 4, s). After normalizing the network element status data of each cell according to this format, it can be used as the input of the task processing model for subsequent processing (model pre-training or model application).
[0074] The multi-stream data involved in this disclosure can be obtained from the input data. However, based on the different model structures of the task processing model, the multi-stream data can be obtained through the data processing process within the model. Alternatively, the multi-stream data can be preprocessed before inputting the input data into the model and then input into the task processing model.
[0075] In one exemplary embodiment, the input data is processed during the actual data processing of the task processing model to extract multi-stream data. In this case, the task processing model includes at least an acquisition atom module. The acquisition atom module is used to acquire the multi-stream data; the multi-stream data includes: a first intervention stream, a first network element snapshot stream, a desired stream, and a signal stream. The terms "first" and "second" are only used to distinguish different objects with the same name and have no actual physical meaning.
[0076] Alternatively, in another possible embodiment, the input data can be further processed in advance to obtain multi-stream data, which can then be used as input to the task processing model to execute subsequent processes. In this case, the method further includes: extracting multi-stream data from the input data, wherein the multi-stream data includes: a first intervention stream, a first network element snapshot stream, a desired stream, and a signal stream.
[0077] Specifically, extracting multi-stream data from the input data may include (or the acquisition atom module may be specifically used for): performing high-dimensional characterization processing on the first network element state data in the input data to obtain the first network element snapshot stream; and extracting the first intervention stream, the desired stream, and the signal stream from the input data.
[0078] The input data conforms to the aforementioned (batchsize, m, k, s) format. Based on this, this disclosure further defines the format of multi-stream data to facilitate subsequent processing. Specifically, it includes:
[0079] A network element snapshot stream is used to describe a high-dimensional representation of network element status data. Here, it is necessary to perform high-dimensional representation processing on the first network element status data in the input data to obtain the first network element snapshot stream. This disclosure does not impose any particular restrictions on the format of the network element snapshot stream.
[0080] An intervention flow is used to describe intervention parameter values. For example, an intervention flow can be a combination of intervention parameters, and it can satisfy the following format: (batchsize, m', k, parameter size); where m' represents the number of network element states in a preset time interval, and parameter size represents the number of intervention parameters. For example, in... Figure 2 In the scenario shown, if m = 512 and m' = 24 (i.e., the preset time interval is 24 hours), then the intervention stream within 24 hours can be taken from the input data (batchsize, 512, 4, s). The intervention stream can be represented as (batchsize, 24, 4, parameter size). This at least indicates that within 24 hours, human intervention was performed on parameter size of the s indicators out of the 512 network element status data.
[0081] The expected flow describes the value of the key metrics at a target time. For example, the expected flow can satisfy the following format: (batchsize, m', k, key metrics size); where m' represents the number of network element states in the preset time interval, and key metrics size represents the number of key metrics.
[0082] A signal stream is used to record parameter tuning times. For example, the signal stream can be a matrix satisfying the format (batchsize, x, x), where x represents the number of parameter tuning operations within a preset time interval. For instance, if parameter tuning is performed at a 15-minute granularity, then in a sequence like this... Figure 2 In the scenario shown, the number of parameter adjustments within the preset time interval of 24 hours is: x = 24 * 60 / 15 = 96. At this time, the signal flow can be represented as a matrix of (batchsize, 96, 96).
[0083] Furthermore, the input data differs in different processing scenarios, and the methods for extracting multi-stream data from the input data can also vary. The scenarios involved here include model pre-training scenarios and model application scenarios. In model pre-training scenarios, the input data is training data, specifically, historical task data. In model application scenarios, the input data is actual task data, such as the task data for the target task.
[0084] In other words, multi-stream data can be extracted from either training data (model pre-training scenario) or task data (model application scenario). Furthermore, to ensure the pre-trained task processing model better reflects real-world task scenarios, the training data used in the model pre-training process is generally historical task data. For example, in a wireless network operation and maintenance scenario, time-series data maintained by each network element within the past month (the historical interval can be customized; this is just an example, such as the past week, the past day, or the 24-48 hours prior to the current moment) can be obtained as training data to pre-train the task processing model. In the model application scenario, the task processing model can then process the relevant task data at the current moment.
[0085] The following explanation will be provided using two scenarios as examples.
[0086] In one embodiment of this disclosure, when the input data is training data, the step of extracting the first intervention stream, the expected stream, and the signal stream from the input data includes:
[0087] A1, from the input data, obtain the intervention parameters in the preset time interval after the first moment to obtain the first intervention stream.
[0088] The term "first moment" describes the time when a network anomaly occurs. This disclosure does not limit the causes or manifestations of network anomalies; when a network anomaly occurs, it is possible to adjust the parameters in the wireless network, i.e., perform parameter tuning. In one exemplary embodiment, a network anomaly can specifically be: the network experiencing a high load state.
[0089] The preset time interval can be customized without any special restrictions. Taking the example of a preset time interval of 24 hours mentioned earlier, this means obtaining the intervention parameters within 24 hours after the time of the anomaly occurrence from the input data and constructing them into the format described in the aforementioned intervention flow, thus obtaining the first intervention flow.
[0090] In the model training scenario, after a period of time following the occurrence of high load (i.e., within a preset time interval), human intervention to adjust parameters will inevitably occur. Thus, the intervention parameters within the preset time interval following the occurrence of high load can be extracted from the input data, and the first intervention flow can be constructed according to the aforementioned format.
[0091] A2, from the input data, obtain the target index value after a preset time interval following the first moment, and obtain the desired flow.
[0092] The expected flow indicates the desired value of a target metric after human intervention. During model training, when a wireless network experiences high load, network parameters are typically tuned after the high load occurs to avoid impacting network operation. This tuning inherently has a certain lag. Therefore, constructing the expected flow requires a certain time interval (i.e., a preset time interval) after the high load event. Thus, the values of the target metric (considered as the target metric values) are obtained after the preset time interval following the first high load event, and the expected flow is constructed. For example, the expected flow can be constructed by extracting data on the target metric 24 hours after the high load event.
[0093] A3, based on the second time point, determine the signal flow.
[0094] In this context, the second time point is the actual parameter tuning time; the first time point is before the second time point. The training data is historical task data, meaning parameter tuning has already occurred. Therefore, the signal stream can be directly constructed based on the actual parameter tuning time. The first time point is before the second time point because, in real-world scenarios, parameter tuning generally has a certain lag (or relative lag) compared to the time when a network anomaly occurs. Thus, the actual parameter tuning time (i.e., the second time point) is generally after the network anomaly occurs (i.e., the first time point). The actual parameter tuning time can be customized in real-world scenarios; for example, it could be the early morning of each day, or other custom fixed or variable times, which will not be elaborated further.
[0095] For easier understanding, please refer to Figure 3 , Figure 3 This is a schematic diagram illustrating the construction of multi-stream data provided in an embodiment of this disclosure. Figure 3 The diagram specifically compares the multi-stream data construction methods in model training and application scenarios. For ease of understanding, Figure 3 Taking a specific scenario of high network load as an example, the high load moment is considered the first moment, while the actual parameter tuning moment is considered the second moment. Furthermore, Figure 3 Taking a preset time interval of 24 hours as an example, and the signal flow is represented as 96*96, it means that the signal flow is a matrix that satisfies the format (batchsize, 96, 96).
[0096] During the model training phase, please refer to the following: Figure 3 The left side. (For example) Figure 3As shown, the multi-stream data is extracted based on two key time points: the first time point and the second time point. During the model training phase, when the wireless network experiences high load, the metric values of the indicators of interest 24 hours later can be extracted from the training data to construct the desired stream. The actual parameter tuning time (i.e., the second time point) generally occurs within 24 hours after the high load time point (i.e., the preset time interval). Thus, the signal stream is constructed based on the actual parameter tuning time point. The intervention stream, on the other hand, can be constructed by taking the high load time point (i.e., the first time point) as the starting point and obtaining the intervention parameters within 24 hours from the first time point (including the actual parameter tuning time point).
[0097] In this way, multi-stream data in the scenario can be constructed based on the training data in a way that is more in line with the real parameter tuning scenario. Therefore, training the task processing model based on this data is more conducive to the task processing model learning the relationship between the intervention stream and the parameter tuning time, and is more conducive to improving the accuracy of the task processing model.
[0098] In another embodiment of this disclosure, when the input data is task data, the step of extracting the first intervention stream, the expected stream, and the signal stream from the input data includes:
[0099] B1, from the input data, obtain the intervention parameters in the preset time interval before the first moment to obtain the first intervention stream.
[0100] In the model application scenario, it is assumed that the first time point and the second time point are the same. Therefore, in this scenario, the first time point (i.e. the high-load time point) is used as the main time point to obtain multi-stream data.
[0101] The intervention stream (including the first intervention stream) consists of intervention parameter values. In the model application scenario, these parameters have already undergone at least one parameter tuning within a preset time interval before the high-load moment (i.e., the first moment). Therefore, in the model application scenario, the intervention parameters for a period of time before the high-load moment in the task data are obtained and the first intervention stream is constructed according to the aforementioned format. For example, intervention parameters within 24 hours before the high-load moment can be extracted from the task data to construct the first intervention stream.
[0102] B2, from the input data, obtain the target index value after a preset time interval following the first moment, to obtain the desired flow.
[0103] The implementation details here can be found in section B1, and will not be repeated here.
[0104] B3, Based on the first time point, determine the signal flow. The first time point describes the moment when a network anomaly occurred.
[0105] As mentioned earlier, in the application scenario of the model, it is assumed that the first time point is the same as the second time point. Therefore, the signal flow can be directly constructed based on the first time point (e.g., the high load time point).
[0106] At this point, you can refer to Figure 3 The right side. (For example) Figure 3 As shown, the multi-stream data is extracted based on the key time point of the first moment, which is... Figure 2 The context involves two points: the high-load moment and the actual parameter tuning moment (both referring to the same time point). During model application, when the wireless network experiences a high load, intervention parameters within the 24 hours prior to the high-load moment can be extracted from the task data to construct the first intervention stream. The desired stream, consistent with the training scenario, requires a certain time interval from the actual parameter tuning moment. Therefore, the indicator values of each key metric can be obtained from the task data 24 hours after the high-load moment to construct the desired stream. In this scenario, the actual parameter tuning moment is considered the high-load moment (i.e., considered as the moment when parameter tuning occurs immediately upon an anomaly). At this point, a signal stream can be constructed based on the high-load moment.
[0107] Furthermore, as mentioned above, the signal flow in this disclosure can be specifically in matrix form. In one exemplary embodiment, the signal flow can be a parameter tuning matrix, where the position of the target identifier in any row of the parameter tuning matrix indicates the index position of the intervention value. The target identifier can be a custom identifier; for example, if the signal flow matrix consists of only 0 and 1 elements, the target identifier can be 1. Thus, the position of 1 in each row indicates the index position of the intervention value. Of course, the signal flow matrix can also have various other custom element configurations. This disclosure does not limit or exhaustively list these, as long as each row of the signal flow matrix contains only one target identifier, and the position of this target identifier indicates the index position of the intervention value.
[0108] In this case, when determining the signal flow based on a specified time, it can be achieved as follows: First, obtain the parameter tuning information at the specified time from the input data, and then generate the parameter tuning matrix based on the parameter tuning information.
[0109] The parameter tuning matrix follows the format: (batchsize, x, x); where batchsize represents the number of samples and x represents the number of parameter tuning attempts within the preset time interval. Refer to the previous text for further details.
[0110] The specified time includes: the first time (i.e. the first time of the model application scenario, corresponding to B3 above) or the second time (i.e. the second time of the model training scenario, corresponding to A3 above).
[0111] The parameter tuning information indicates whether parameter tuning occurred at the specified time and the parameter tuning intervention value. Parameter tuning means artificially intervening in the network parameters, i.e., generating intervention parameters. For any parameter tuning time and any parameter, if parameter tuning occurred at that time, the parameter tuning intervention value needs to be further determined; if parameter tuning did not occur at that time, the parameter value at that time remains unchanged, and the parameter value from the previous time is taken. Thus, for ease of processing, an intervention value vector (or sequential set, the form is not limited, taking a vector as an example) can be preset. In this way, the position of the parameter value in the signal flow matrix indicates which value in the preset intervention value vector is taken. Thus, the parameter value after intervention tuning can be obtained by simply obtaining the dot product of the signal flow matrix and the preset intervention value vector.
[0112] For easier understanding, please refer to Figure 4 , Figure 4 A schematic diagram of a signal flow matrix provided in an embodiment of this disclosure, such as... Figure 4 As shown, the signal flow matrix is still constructed based on elements 0 and 1. The parameter tuning information obtained from the input data can be represented as... Figure 4 The leftmost vector in the vector (denoted as vector 1) indicates whether parameter tuning has occurred at the current position (each position indicates a parameter tuning time, and thus can also be understood as a parameter tuning time). Therefore, a value of 0 indicates that no parameter tuning has occurred at that position, and a value of 1 indicates that parameter tuning has occurred.
[0113] The second vector from the left (denoted as vector 2) indicates which position's intervention value the current position should take as the intervention value. For example... Figure 4 As shown, no parameter tuning occurred at the first position of vector 1 (denoted as tuning time 1, which can also be understood as the initial time corresponding to the initial situation). Therefore, no intervention parameter value needs to be taken at tuning time 1. Consequently, the first position of vector 2 takes a value of 0, which means that there is no need to take the assumed intervention value (i.e., ... Figure 4 The value is taken from the third vector from the left (denoted as vector 3). Parameter tuning occurs at the second position of vector 1 (denoted as tuning time 2), indicating that tuning time 2 needs to take a value from vector 3. Therefore, the value of vector 2 at the second position is not empty; specifically, the value of vector 2 at this position is 1, meaning that tuning time 2 needs to select the value of the first position in vector 3 as the intervention value for tuning time 2. No parameter tuning occurs at the third position of vector 1 (denoted as tuning time 3). At this time, the parameter at tuning time 3 remains unchanged, still being the parameter value of the previous position (i.e., tuning time 2). Thus, the value of the third position in vector 2 is 1 (consistent with the second position). And so on. Based on vector 1 and vector 2, a system can be constructed as follows... Figure 4 The signal flow matrix mentioned in the middle.
[0114] In this signal flow matrix, each row corresponds to a parameter tuning moment. Each row contains only one element 1 (other identifiers can be used in actual scenarios). The position of element 1 in the row indicates which position of the intervention value is taken from the preset intervention value (i.e., the right vector 3) at that parameter tuning moment.
[0115] In practical processing, since the value of one position will continue the value of the previous position when the parameters are not adjusted, in order to minimize the superposition error, such as Figure 4 As shown, this disclosure fixes the value of the first row in the signal flow matrix (corresponding to the parameter tuning time 1 mentioned above) to (1, 0, 0...), meaning that the subsequent processing is based on the assumed first intervention value in the initial case. Based on this, the position of element 1 in each of the remaining rows needs to exclude the influence of the fixed value of the first row. Therefore, when determining the position of element 1 in each row, the sorting needs to start from the second position. Correspondingly, the value of the first position of vector 3 is also a pre-set value and does not participate in the sorting of subsequent value positions.
[0116] For example, in the second row of the signal flow matrix (corresponding to parameter tuning time 2), element 1 is located in the second position of that row. After excluding the influence of the fixed initial value (i.e., sorting from the second position onwards), element 1 in the second row of the signal flow matrix is located in the first position of that row. This means that the intervention value at parameter tuning time 2 is the value of the first position in vector 3 excluding the first fixed value. Therefore, the intervention value after parameter tuning at time 2 is 5. Similarly, the third and fourth rows of the signal flow also indicate taking the intervention value at the first position excluding the first fixed value, so the intervention values after parameter tuning at times 3 and 4 are also 5. The fifth row of the signal flow indicates taking the intervention value at the fourth position excluding the first fixed value, so the intervention value after parameter tuning at time 5 is the value of the fourth position in vector 3 excluding the first fixed value, which is 9.
[0117] Based on the above processing, this disclosure enables the construction and extraction of multi-stream data (multi-stream training data, multi-stream application data). For ease of understanding, the following describes... Figure 5 Explain the relationships between the multi-stream data in detail. Figure 5 This is a schematic diagram of a multi-stream data relationship provided in an embodiment of this disclosure. Figure 5 The following example uses a 24-hour preset time interval as an illustration.
[0118] like Figure 5As shown, the intervention flow at time t is related to the network element snapshot flow before time t, the intervention flow before time t, and the expected flow at time t+24h and before. For example, when t is 3, the intervention flow is C. Then, the intervention flow C is related to the network element snapshot flow before time 3 (i.e., 1 and 2), the intervention flow before time C (A and B), the expected flow at time c+24h, and the expected flow before time c+24h (a+24h and b+24h).
[0119] like Figure 5 As shown, the network element snapshot flow at time t is related to the network element snapshot flows before time t and the intervention flow at time t. Taking t as 3 as an example, the network element snapshot flow 3 at this time is related to the network element snapshot flows before 3 (i.e., 1 and 2) and the intervention flow C corresponding to this time.
[0120] Based on such Figure 5 The multi-stream data relationships shown can be used to construct a general foundational model based on these relationships, and pre-trained to obtain a task processing model to handle various tasks. Furthermore, as mentioned earlier, independent models for each task can also be constructed based on the aforementioned multi-stream data.
[0121] Based on this, this disclosure also provides a task processing model, which may include multiple atomic modules. Different atomic modules can be split and combined. By splitting and combining atomic modules, the processing requirements of different task types can be met.
[0122] Figure 6 This is a schematic diagram of a task processing model provided in an embodiment of this disclosure. Figure 6 As shown, the task processing model includes:
[0123] Atom acquisition module 610 is used to acquire the multi-stream data; the multi-stream data includes: a first intervention stream, a first network element snapshot stream, a desired stream, and a signal stream.
[0124] The mutual attention atom module 620 is used to process the multi-stream data based on the mutual attention mechanism and the parameter tuning time to obtain the second intervention stream;
[0125] The multilayer perceptron atomic module 630 is used to perform multilayer perceptron processing on the second intervention flow and the first network element snapshot flow to obtain the second network element snapshot flow.
[0126] The self-attention atom module 640 is used to process the first network element snapshot stream based on the self-attention mechanism to obtain the third network element snapshot stream;
[0127] The fusion atom module 650 is used to fuse the second network element snapshot stream and the third network element snapshot stream to obtain the fourth network element snapshot stream;
[0128] The restoration atom module 660 is used to restore the snapshot stream of the fourth network element to obtain the state data of the second network element.
[0129] In such Figure 6 The task processing model shown can include three main parts: The first part includes the acquisition atom module 610, which is mainly used to realize the mapping processing from network element state to network element snapshot (i.e., high-dimensional representation of network element state); The second part includes: mutual attention atom module 620, multilayer perceptron atom module 630, self-attention atom module 640, and fusion atom module 650. These modules can be combined based on the actual task scenario. For the entire task processing model, this part is mainly used to realize the generation of intervention flow and the generation of network element snapshot from front to back (i.e., generation of network element snapshot flow); The third part includes: restoration atom module 660, which is mainly used to realize the restoration of network element snapshot.
[0130] In this process, the first part (acquiring atom module 610) and the third part (restoring atom module 660) are used to realize the forward and reverse conversion between the network element state and the network element snapshot. Thus, in a preferred embodiment, the high-dimensional representation processing involved in the acquisition atom module 610 and the restoration processing involved in the restoration atom module 660 correspond to each other.
[0131] The following details the function and capabilities of each atomic module.
[0132] In this disclosure, the atomic acquisition module 610 is used to acquire multi-stream data from the input data. For example, the atomic acquisition module is specifically used to: perform high-dimensional representation processing on the first network element state data in the input data to obtain the first network element snapshot stream; and extract the intervention stream, the desired stream, and the signal stream from the input data. The acquisition of the network element snapshot stream, i.e., the process of performing high-dimensional representation processing on the first network element state data in the input data, can be performed by, for example... Figure 6 The atomic acquisition module 610 shown processes the input data through a transformer encoder (i.e., dimensionality transformation processing), a flatten layer (for dimensionality transformation processing), and a linear layer (i.e., a linear transformation processing), mapping the input data to a size of (batchsize, m, embedding_len) to obtain a high-dimensional representation of the input data, i.e., a network element snapshot stream. Here, embedding_len refers to the embedding (feature) length, which can be customized in actual scenarios. Figure 6 The example used is embedding_len of 1024. For details on the specific implementation methods for extracting the intervention stream, the expected stream, and the signal stream in different scenarios, please refer to the previous text; they will not be repeated here.
[0133] The input to the atomic module 610 is the input data that satisfies the (batchsize, m, k, s) format mentioned above; the output is multi-stream data. The output method of the multi-stream data is not limited. In one embodiment of this disclosure, the output data of the atomic module 610 can be: individual multi-stream data, and / or, a combination of at least two types of multi-stream data. For example, Figure 6 One possible implementation is illustrated: the desired flow, network element snapshot flow, and intervention flow are concatenated together to form a combined data set of m'*(1024 + 4 * attention index size + 4 * parameter size). The concatenation method is not limited; for example, the concat function can be used to combine multi-flow data. The concatenated combined data can be referenced from the combined data section shown in the mutual attention atom module 620. Furthermore, in another possible embodiment, the output of the acquisition atom module 610 can be multi-flow data, while the aforementioned combined data can be obtained by the mutual attention atom module 620.
[0134] Based on the existence of multi-stream data, such as Figure 5 As shown in the data relationship, this disclosure connects the mutual attention atomic module 620 after the acquisition atomic module 610 to acquire the second intervention stream from the multi-stream data by utilizing the mutual attention mechanism and the parameter tuning time.
[0135] Specifically, multi-stream data possesses, for example, Figure 5 The relationship shown indicates that different data streams have certain temporal correlations. Mutual attention mechanisms perform well in capturing the correlations between different sequences. Therefore, this disclosure uses mutual attention mechanisms to capture and learn the correlations between multi-stream data. Specifically, it learns the mutual influence relationship between the expected stream, the network element snapshot stream, and the intervention stream on the intervention stream itself in real-world scenarios, so that the intervention stream can be expressed more accurately in subsequent model applications.
[0136] On the other hand, if the signal stream is used to record the parameter tuning time, then the parameter tuning time can be characterized by the signal stream. As mentioned above... Figure 4 As explained in the description of the signal flow, the elements in the parameter tuning matrix corresponding to the intervention flow indicate the index positions for taking values in the intervention flow. Therefore, by obtaining the dot product of the intervention flow and the signal flow, the intervention flow data corresponding to the parameter tuning time can be obtained. In this process, the intervention flow obtained through the mutual attention mechanism can be used as... Figure 4 Vector 3 in the example.
[0137] In summary, the mutual attention atom module 620 is based on the principle of mutual attention, learning the correlation between multi-stream data, and learning the impact of this correlation on the intervention stream.
[0138] Specifically, the mutual attention atom module 620 is used for:
[0139] S1, acquire the combined data of the expected flow, the first network element snapshot flow, and the first intervention flow.
[0140] Combination Figure 6 Note that, at this point, the style of the combined data can be referenced from the part shown in the combined data section, which can be represented as a combined data of m'*(1024+4*size of the indicator to be monitored+4*size of the parameter).
[0141] As mentioned earlier, the combined data can be directly generated by the acquisition atom module 610 or the mutual attention atom module 620. In one possible implementation, the acquisition atom module 610 can directly output the combined data, which the mutual attention atom module 620 can then directly receive. Alternatively, if the acquisition atom module 610 outputs multi-stream data, the mutual attention atom module 620 performs a combination processing (e.g., using the `concat` function) on the desired stream, the first network element snapshot stream, and the first intervention stream in the multi-stream data to obtain the combined data. The processing method for the combined data is described above and will not be repeated here.
[0142] S2, the combined data and the first intervention stream are processed using the mutual attention mechanism to obtain the third intervention stream.
[0143] The input to the mutual attention mechanism includes two data sets: combined data and a first intervention stream. The first intervention stream comes from the output of the acquisition atom module 610. The output of the mutual attention mechanism can specifically be a third intervention stream. The third intervention stream is used to characterize the prediction result after the multi-stream data influences the intervention stream.
[0144] like Figure 6 As shown, for the combined data (expected flow-token-intervention flow (A, B...m')) and the first intervention flow (A, B...m') (it should be understood that the intervention flow in the combined data is the first intervention flow), under the action of the mutual attention mechanism, it can be regarded as an upper triangular matrix, and finally output the third intervention flow (B', C'...m'). In this case, the first position of the third intervention flow is empty. In order to complete the third intervention flow, data needs to be added to the first position of the third intervention flow.
[0145] In real-world scenarios, any custom data can be used to supplement this position, such as random data, data with other preset rules, or custom fixed values, etc., without exhaustive list. However, considering that random data may cause superposition errors, this disclosure provides a preferred solution to further reduce superposition errors: the first intervention parameter of the first intervention stream is filled into the first position of the third intervention stream, thus forming the fourth intervention stream, i.e., S3 in this embodiment.
[0146] S3, combine the first intervention parameter of the first intervention stream with the third intervention stream to obtain the fourth intervention stream.
[0147] In practice, the first intervention parameter A of the first intervention stream is combined with the third intervention stream (B', C', ..., m') to obtain the fourth intervention stream, which can be represented as (A', B', C', ..., m'). This fourth intervention stream is actually the final processing result of the mutual attention mechanism, and can also be used to characterize the prediction result after multi-stream data influences the intervention stream.
[0148] S4, obtain the dot product between the fourth intervention flow and the signal flow to obtain the second intervention flow.
[0149] When this step is implemented, the fourth intervention flow is used as... Figure 4 In the illustrated embodiment, vector 3 is used to obtain the dot product between vector 3 and the signal flow, thus yielding the second intervention flow. Combined with... Figure 4 As illustrated, this step means that for the parameter tuning moments in the signal stream (i.e., moments when parameters change), the currently predicted intervention value is used; while for the parameterless moments in the signal stream (i.e., moments when parameters remain unchanged), the previously predicted intervention parameter value is used. Thus, the resulting second intervention stream, in addition to considering the correlation between multiple data streams, also comprehensively considers whether parameter tuning (parameter changes) occurred at the tuning moments, in order to comprehensively learn the generation of intervention values.
[0150] In multi-stream data relationships, the intervention stream can directly affect the network element snapshot stream. Therefore, after predicting the intervention stream, it is necessary to further determine the impact of the intervention stream on the network element snapshot stream, that is, to obtain the network element snapshot stream affected by the intervention stream (denoted as the second network element snapshot stream). In this disclosure, this purpose is achieved through a multilayer perceptron (MLP).
[0151] In other words, in this disclosure, the multilayer perceptron atomic module 630 is used to process the second intervention stream and the first network element snapshot stream using the multilayer perceptron to obtain the second network element snapshot stream. Each neuron layer of the multilayer perceptron consists of many neurons, where the input layer receives input features, the output layer provides the final prediction result, and the hidden layers in between are used to extract features and perform nonlinear transformations. Each neuron receives the output of the previous layer, performs a weighted sum and an activation function operation to obtain the output of the current layer. Through continuous iterative training, the multilayer perceptron can automatically learn the complex relationships between input features and make predictions on new data.
[0152] like Figure 6 As shown, the multilayer perceptron atomic module 630 can be connected after the acquisition atomic module 610 and the mutual attention atomic module 620. Thus, the input data of the multilayer perceptron atomic module 630 includes: the second intervention stream output by the mutual attention atomic module 620 and the first network element snapshot stream output by the acquisition atomic module 610. To maintain consistency with the data volume (m' items) of the intervention stream, the first network element snapshot stream processed here can be a portion of the data in the first network element snapshot stream output by the acquisition atomic module 610, specifically represented as: {cls, token1, token2, ..., tokenm'-1}. The output of the multilayer perceptron atomic module 630 remains m' second network element snapshots, and at this time, the second network element snapshot stream can be represented as {cls, token1, token2, ..., tokenm'}.
[0153] Furthermore, when applying this task processing model to handle some possible tasks, the mutual attention atom module 620 may not be necessary, or it can be combined and connected after the acquisition atom module 610. In this case, the input data of the multilayer perceptron atom module 630 is the multi-stream data output by the acquisition atom module 610 (including: the first intervention stream and the first network element snapshot stream).
[0154] Furthermore, it should be noted that this disclosure does not impose any particular restrictions on the type or structure of the MLP, and custom designs are permissible in practical scenarios. For example, the activation function used in the MLP may include, but is not limited to, the rectified linear unit (ReLU). Additionally, the sigmoid function, tanh function, etc., can also be used as activation functions, without exhaustive list.
[0155] like Figure 4In the multi-flow relationship shown, for the element snapshot flow, any element snapshot between different elements affects the current element snapshot. Therefore, in order to learn the correlation between different element snapshots in the element snapshot flow, this disclosure also employs a self-attention mechanism to learn and predict the element snapshot flow. This is implemented through the self-attention atom module 640.
[0156] That is, in this disclosure, the self-attention atom module 640 is used to process the first network element snapshot stream based on the self-attention mechanism to obtain the third network element snapshot stream.
[0157] Self-attention, or self-attention mechanism, focuses on the relationships between elements in the input sequence, connecting different positions of a single sequence to compute the representation of the same sequence. Thus, self-attention not only considers the features of each network element snapshot in the network element snapshot stream but also comprehensively considers the relevant semantic information of the context.
[0158] The input data of the self-attention atom module 640 is a first network element snapshot stream, which can be represented as {cls, token1, token2, ..., tokenm}. This stream can originate from the output of the acquisition atom module 610; that is, the self-attention atom module 640 can also be connected after the acquisition atom module 610. The output of the self-attention atom module 640 is also m third network element snapshots, which can then be represented as {cls, token1, token2, ..., tokenm}.
[0159] like Figure 6 As shown, after the above processing, the multilayer perceptron atom module 630 and the self-attention atom module 640 can obtain different network element snapshot streams. In order to facilitate subsequent processing, this disclosure further provides a fusion atom module 650 in the second part to fuse the outputs of the two and obtain the fused fourth network element snapshot stream.
[0160] In practical implementation, the second and third network element snapshots at the corresponding times in the network element snapshot streams can be added together and processed through a linear layer to obtain the fused fourth network element snapshot stream. For ease of understanding, Figure 6 This fusion process is represented by "+".
[0161] It should be understood that during task processing, if the multilayer perceptron atomic module 630 and the self-attention atomic module 640 exist simultaneously, a fusion atomic module 650 is required to fuse the results of the two. However, if some tasks only involve one of the multilayer perceptron atomic module 630 and the self-attention atomic module 640, then there is no need to add a fusion atomic module 650.
[0162] Based on the aforementioned processing in Part 2, it is possible to predict and generate network element snapshot streams under the influence of multi-stream data.
[0163] In this disclosure, the third part of the task processing model is used to restore the network element snapshot stream generated in the second part into network element status data in (batchsize, m, k, s) format. In other words, the restoration atom module 660 in this disclosure is used to perform restoration processing on the fourth network element snapshot stream to obtain the second network element status data.
[0164] It should be understood that the fourth network element snapshot stream referred to here refers to the output of the second part. In actual application scenarios, it can be specifically the fourth network element snapshot stream output by the fusion atomic module 650, the third network element snapshot stream output by the self-attention atomic module 640, or the second network element snapshot stream output by the multilayer perceptron atomic module 630 (the combination of atomic modules is different in different scenarios).
[0165] As mentioned earlier, in its specific implementation, the reduction processing method corresponds to the high-dimensional characterization processing method of the atomic module 610. Based on this, as follows: Figure 6 As shown, when the acquisition atom module 610 uses a transformer encoder + flatten layer + linear layer to implement high-dimensional representation processing, the restoration atom module 660 can use a linear layer + unflatten layer to implement restoration processing. That is, the restoration atom module 660 can first perform dimensional transformation on the network element snapshot stream through a linear layer, and then split it into network element states that meet the aforementioned format, finally outputting the second network element state data in the format of (batchsize, m, k, s).
[0166] In summary, it can be achieved through methods such as Figure 6 The task processing model shown learns the interrelationships between multi-stream data and uses this to predict different task types such as network element status, network element snapshot streams, and intervention streams, meeting the actual needs of operation and maintenance scenarios. This will be explained in detail later.
[0167] Based on this, the task processing method provided in this disclosure may further include the following steps:
[0168] Obtain training samples;
[0169] The task processing model is trained using the training samples and the target loss function until the preset training conditions are met.
[0170] As mentioned earlier, training data can be historical task data. How to obtain multi-stream data based on historical task data, and how each atomic module processes the data, are all explained in the previous text and will not be repeated here.
[0171] Based on this, the task processing model described in S102 can be obtained through multiple rounds of iterative learning until the preset training conditions are met. This disclosure does not impose any particular restrictions on the preset training conditions. For example, the preset training conditions can be customized by at least one of the following: training duration, number of iterative training rounds, and range of the target loss function. No exhaustive list is provided.
[0172] The target loss function can also be customized in real-world scenarios. Based on the task processing model provided above, this disclosure further presents a preferred implementation: the target loss function is a weighted sum of a first loss function and a second loss function; wherein, the first loss function is used to characterize the degree of difference between the first network element state data and the second network element state data; and the second loss function is used to characterize the degree of difference between the first intervention flow and the second intervention flow.
[0173] At this point, please refer to Figure 6 The target loss function in this disclosure consists of two parts: a first loss function and a second loss function. Specifically, the first loss function can be used to indicate the degree of difference between the model's input data (i.e., the first network element state, also known as the original input index) and the output data (i.e., the second network element state, also known as the output index). The second loss function can be used to indicate the degree of difference between the predicted intervention stream (i.e., the second intervention stream) generated in the second part and the actual intervention stream (i.e., the first intervention stream).
[0174] In this disclosure, both the first and second loss functions are used to characterize the degree of difference between data. In practical scenarios, this can be characterized by at least one method, such as mean-square error (MSE), similarity, root mean squared error (RMSE), or mean absolute error (MAE), without exhaustive enumeration. Taking MSE as an example to characterize the degree of difference between data, the first loss function can be expressed as: MSE(predict_indicators, input_indicators), where predict_indicators represents the state of the second network element and input_indicators represents the state of the first network element; the second loss function can be expressed as: MSE(predict_treatment_flow, true_treatment_flow), where predict_treatment_flow represents the second intervention flow and true_treatment_flow represents the first intervention flow.
[0175] Furthermore, it should be noted that considering the large volume of network data and the potential for some data loss in real-world scenarios, null value removal is necessary before determining the aforementioned loss function. Specifically, the locations where data in the first network data set, the second network data set, the first intervention stream, and the second intervention stream is empty are identified and removed before calculating the corresponding loss function.
[0176] The target loss function is the weighted sum of the first loss function and the second loss function. This disclosure does not impose any particular restrictions on the weighting method between the two. For example, it may include, but is not limited to, weighted sum, weighted average, etc., without exhaustive list.
[0177] In one exemplary embodiment, the target loss function can satisfy the following formula:
[0178] total_loss=α*MSE(predict_treatment_flow,true_treatment_flow)+
[0179] β*MSE(predict_indicators,input_indicators)
[0180] Where total_loss represents the target loss function, and α and β are weighting coefficients (which can be customized).
[0181] Thus, based on the aforementioned objective loss function, during the training process of the task processing model, the model parameters can be adjusted through backpropagation of the objective loss function, ultimately training a task processing model that meets the requirements.
[0182] Based on this, the trained task processing model can be deployed, and then the pre-trained task processing model can be called to process specific tasks.
[0183] As mentioned above, the task processing model provided in this disclosure (such as...) Figure 6 As shown, this is a common basic model. In real-world scenarios, in addition to directly using this task processing model for task processing, the atomic modules in the task processing model can also be combined based on the actual task type to obtain a suitable atomic module combination (denoted as target combination or target combination model), and the target combination can be used to process the corresponding task.
[0184] In other words, S104 can utilize the task processing model to process the target task in the following ways:
[0185] In one embodiment of this disclosure, the target task is processed directly using a task processing model.
[0186] In another embodiment of this disclosure, a target combination of atomic modules can be determined in the task processing model based on the type of the target task; and the target task can be processed using the target combination.
[0187] It should be noted that in the implementation method of using a combination of atomic modules to achieve task processing, before processing the target task using the target combination, the following steps may also be included: Supervised Fine-Tuning (SFT) of the target combination using a small amount of labeled data of the same type as the target task. In other words, based on a common basic model, the target combination is fine-tuned using task data of the same type, making the target combination more suitable for the target task, which is beneficial to improving the model's prediction performance.
[0188] The following sections will explain how to implement the task processing process for each of the several task types in the wireless operation and maintenance scenarios described above.
[0189] In this disclosure, when the target task is a time series prediction task, the target combination (model) is determined to include: acquisition atomic module, self-attention atomic module, and restoration atomic module.
[0190] At this point, you can refer to Figure 7 , Figure 7 A schematic diagram of a target combination provided in an embodiment of this disclosure, such as... Figure 7 As shown, for time series prediction tasks, methods such as... can be reused. Figure 6 The diagram illustrates the basic structure of the task processing model. Based on this, since the temporal prediction task does not require consideration of the expectation stream, the parts generating the expectation stream and the second part related to the intervention stream (e.g., mutual attention atomic modules, multilayer perceptron atomic modules) in the multi-stream data can be removed. The intervention streams involved in the process are set to the same parameters as the previous moment. Thus, in this task scenario, only three atomic modules are needed: the acquisition atomic module (where the prompt is set to temporal prediction at the cls position), the self-attention atomic module, and the restoration atomic module. This is sufficient to process the corresponding task, thereby determining the target combination. It should be understood that the input to the target combination model is the input to the acquisition atomic module, the output of the acquisition atomic module is the input to the self-attention atomic module, the output of the self-attention atomic module is the input to the restoration atomic module, and the output of the restoration atomic module is the output of the target combination model. This output can then be used to indicate the attention index value for the future moment indicated by the temporal prediction task.
[0191] In this disclosure, when the target task is a root cause localization task, determining the target combination includes: the acquisition atom module and the self-attention atom module.
[0192] At this point, you can refer to Figure 8 , Figure 8 A schematic diagram illustrating another target combination provided in an embodiment of this disclosure, such as... Figure 8 As shown, for root cause localization tasks, methods such as... can be reused. Figure 6 The diagram illustrates the basic structure of the task processing model. Based on this, since the root cause localization task does not require consideration of the expectation flow and intervention flow, the part related to the intervention flow can be removed. In this task scenario, the intervention flow and the expectation flow can be set to None. Thus, in this task scenario, only the atomic module (where the prompt is set to root cause localization at the cls position) and the self-attention atomic module need to be acquired to process the corresponding task, thereby determining the target combination. It should be understood that the input to the target combination model is the input to the acquisition atomic module, the output of the acquisition atomic module is the input to the self-attention atomic module, and the output of the self-attention atomic module is the output of the target combination model.
[0193] It is important to note that in root cause localization scenarios, since the output of the target combination is a third-level network element snapshot stream, it is insufficient to achieve the requirement of pinpointing a specific cause to a particular category. Therefore, in practical implementations, the following steps can be further included: connecting a classification layer after the target combination to obtain a target processing object; and using the target processing object to process the root cause localization task. Please refer to [reference needed]. Figure 8 This disclosure also connects an MLP module 810 after the target combination. That is, after obtaining a high-dimensional representation such as a network element snapshot using the first part (acquisition atom module) and the second part (self-attention atom module) of the general model, a lightweight fully connected classification network is connected after the high-dimensional representation, which enables fast processing of the root cause localization classification model.
[0194] In this disclosure, when the target task is a causal inference task, the target combination is determined to include: the acquisition atom module, the multilayer perceptron atom module, the self-attention atom module, the fusion atom module, and the reduction atom module.
[0195] At this point, you can refer to Figure 9 , Figure 9 A schematic diagram illustrating another target combination provided in an embodiment of this disclosure, such as... Figure 9 As shown, for causal inference tasks, methods such as... can be reused. Figure 6The basic structure of the task processing model is shown. Since the causal inference task can be viewed as a time-series prediction task with human intervention, the impact of the intervention flow on the network element snapshot flow needs to be considered. Therefore, the objective combination includes a multilayer perceptron atomic module, and the intervention flow in this multilayer perceptron atomic module can be set to a tuned value. Furthermore, this task still does not need to consider the expected flow, which can be set to None. Based on this, through the processing of the objective combination (where the prompt is set to causal inference at the cls position), the focus indicator value indicating the future time indicated by the time-series prediction task can be output.
[0196] In this disclosure, when the target task is a parameter tuning decision task, the target combination is determined to include: the acquisition atomic module and the mutual attention atomic module.
[0197] At this point, you can refer to Figure 10 , Figure 10 A schematic diagram illustrating another target combination provided in an embodiment of this disclosure, such as... Figure 10 As shown, for parameter tuning decision-making tasks, methods such as... can be reused. Figure 6 The diagram illustrates the basic structure of the task processing model. This task scenario primarily focuses on how to intervene to quickly restore abnormal indicators to normal values. Therefore, the target combination for this task scenario mainly involves the first part (where the `cls` position sets `prompt` to the parameter tuning decision) and the second part related to the intervention flow. The output of the target combination is the output of the mutual attention atomic module, which indicates the parameter tuning value for the parameter tuning decision. In a practical implementation, the expected flow can be set to the target value one day later, the signal flow to the parameter tuning time, and the intervention flow to the current parameter value. Based on this, the model ultimately outputs the parameter tuning value for the parameter tuning decision.
[0198] In addition to the four tasks mentioned above, the task processing methods provided in this disclosure can also be applied to the processing of other types of tasks, including, but not limited to, at least one of the following: anomaly detection, disturbance discovery, etc. Furthermore, the task processing model provided in this disclosure can also support general modeling for various network domains, covering wireless networks, core networks, transmission networks, etc.
[0199] It should be noted that although only some atomic modules of the task processing model were used in the above embodiments to implement the target task processing, the entire task processing model was trained based on multi-stream data. Thus, even if some atomic modules do not participate in the actual target task processing during the actual task processing process, the task processing model (or target combination) itself has already learned the correlation between multi-stream data. Therefore, the corresponding data processing is still based on these learned correlations during the actual task processing process. These correlations between multi-stream data can still be applied to any task process to improve the task processing effect.
[0200] In summary, this disclosure addresses the shortcomings of current intelligent operation and maintenance algorithms for wireless networks. Isolated modeling suffers from low efficiency, strong data dependency, and poor model versatility and performance. General modeling also has limitations in supporting downstream tasks, particularly intervention-type downstream tasks where prompts cannot directly provide support. This disclosure innovatively introduces multiple flows, including intervention flow, expectation flow, and signal flow, and proposes a generative network pre-trained large model based on multi-flow attention (i.e., as mentioned above). Figure 6 The task processing model shown.
[0201] Furthermore, in the process of introducing multi-stream data such as intervention stream, expectation stream, and signal stream, the parameter part is selected as the intervention stream, the focus indicator after the next 24 hours (i.e. the preset time interval mentioned above) is selected as the expectation stream, and the parameter adjustment time is selected as the signal stream and represented in matrix form. While learning the static relationship of the network, the dynamic relationship caused by the intervention is captured, which is more accurate.
[0202] Furthermore, this disclosure proposes a generative pre-trained large-scale model based on multi-stream attention in the field of network operations and maintenance. This task processing model consists of three parts: obtaining network element snapshots (i.e., the first part mentioned above), the base model (i.e., the second part mentioned above), and network element snapshot recovery (i.e., the third part mentioned above). In practical applications, this task processing model can be adopted by splitting and combining atomic modules. During this process, a prompt can be set at the CLS position based on the target task type. Additionally, different model modules can be selected for SFT fine-tuning training to make the target combination more suitable for the specific task and achieve better processing results.
[0203] In addition, this disclosure also uses different data construction methods to construct training data and application data (i.e., the task data mentioned above). This data construction method is helpful to alleviate the lag problem of intervention-type parameter tuning to a certain extent.
[0204] In summary, compared to isolated modeling solutions in existing technologies, this disclosure not only improves modeling efficiency but also reduces dependence on data, enhancing the model's versatility and performance. Compared to general modeling in existing technologies, this disclosure introduces multi-stream data such as intervention streams, expectation streams, and signal streams, and utilizes the model to capture the dynamic relationships of the network under the influence of intervention. Furthermore, it proposes a generative network pre-trained large-scale model architecture based on multi-stream attention, which can more universally support intervention-related business scenarios. For example, when applied to wireless network operation and maintenance scenarios, this solution can handle various types of tasks such as time series prediction, root cause localization, causal inference, and parameter tuning decisions, promoting intelligent diagnosis and intelligent operation and maintenance of wireless networks.
[0205] This disclosure also provides a task processing apparatus. Figure 11 This is a structural block diagram of a task processing device provided in an embodiment of the present disclosure, such as... Figure 11 As shown, the task processing device 1100 includes:
[0206] The acquisition module 1110 is used to acquire a task processing model; the task processing model is obtained based on multi-stream data pre-training, and the multi-stream data includes at least: an intervention stream, which is used to describe intervention parameter values;
[0207] The processing module 1120 is used to process the target task using the task processing model.
[0208] In one exemplary embodiment, the multi-stream data includes:
[0209] The intervention flow;
[0210] Expected flow, used to describe the value of the key indicator at the target time;
[0211] Signal stream, used to record parameter tuning times;
[0212] Network element snapshot streams are used to describe high-dimensional representations of network element status data.
[0213] In one exemplary embodiment, the task processing model includes multiple atomic modules, each atomic module comprising:
[0214] An atomic acquisition module is used to acquire the multi-stream data; the multi-stream data includes: a first intervention stream, a first network element snapshot stream, a desired stream, and a signal stream.
[0215] The mutual attention atom module is used to process the multi-stream data based on the mutual attention mechanism and the parameter tuning time to obtain the second intervention stream;
[0216] A multilayer perceptron atomic module is used to perform multilayer perceptron processing on the second intervention flow and the first network element snapshot flow to obtain the second network element snapshot flow;
[0217] The self-attention atom module is used to process the first network element snapshot stream based on the self-attention mechanism to obtain the third network element snapshot stream;
[0218] A fusion atom module is used to fuse the second network element snapshot stream and the third network element snapshot stream to obtain a fourth network element snapshot stream;
[0219] The restoration atom module is used to restore the snapshot stream of the fourth network element to obtain the state data of the second network element.
[0220] In one exemplary embodiment, the mutual attention atom module is specifically used for:
[0221] Obtain combined data of the desired flow, the first network element snapshot flow, and the first intervention flow;
[0222] The combined data and the first intervention stream are processed using the mutual attention mechanism to obtain a third intervention stream;
[0223] The first intervention parameter of the first intervention stream is combined with the third intervention stream to obtain the fourth intervention stream;
[0224] The second intervention flow is obtained by obtaining the dot product between the fourth intervention flow and the signal flow.
[0225] In one exemplary embodiment, the atom acquisition module is specifically used for:
[0226] The first network element status data in the input data is subjected to high-dimensional representation processing to obtain the first network element snapshot stream; wherein, the high-dimensional representation processing corresponds to the processing method of the restoration processing;
[0227] Extract the intervention stream, the desired stream, and the signal stream from the input data;
[0228] The input data satisfies the following format: (batchsize, m, k, s); where batchsize represents the number of samples, m represents the number of network element states, k represents the number of time slices contained in a network element state, and s represents the number of indicators; any network element state includes indicator data for k time slices.
[0229] In one exemplary embodiment, when the input data is training data, the atom acquisition module is specifically used for:
[0230] From the input data, intervention parameters within a preset time interval after the first moment are obtained to obtain the first intervention stream;
[0231] From the input data, the target index value after a preset time interval following the first moment is obtained to obtain the desired flow;
[0232] The signal flow is determined based on the second time point;
[0233] Wherein, the first moment is used to describe the moment when the network anomaly occurs; the second moment is the actual parameter tuning moment; the first moment is before the second moment.
[0234] In one exemplary embodiment, when the input data is task data, the atomic acquisition module is specifically used for:
[0235] From the input data, intervention parameters within a preset time interval prior to the first moment are obtained to obtain the first intervention stream;
[0236] From the input data, the target index value after a preset time interval following the first moment is obtained to obtain the desired flow;
[0237] Based on the first moment, the signal flow is determined;
[0238] The first moment is used to describe the moment when the network anomaly occurs.
[0239] In one exemplary embodiment, the signal stream is a parameter tuning matrix, and the position of the target identifier in any row of the parameter tuning matrix is used to indicate the index position of the intervention value; the atom acquisition module is specifically used for:
[0240] The parameter tuning information at a specified time is obtained from the input data. The parameter tuning information is used to indicate whether parameter tuning occurs at the specified time and the parameter tuning intervention value. The specified time includes: a first time or a second time.
[0241] Based on the parameter tuning information, the parameter tuning matrix is generated;
[0242] The parameter tuning matrix satisfies the format: (batchsize, x, x); where batchsize represents the number of samples and x represents the number of parameter tuning times in the preset time interval.
[0243] In one exemplary embodiment, the task processing device 1100 includes a training module ( Figure 11 (Not shown), the training module, specifically used for:
[0244] Obtain training samples;
[0245] The task processing model is trained using the training samples and the target loss function until the preset training conditions are met.
[0246] Wherein, the target loss function is the weighted sum of the first loss function and the second loss function;
[0247] The first loss function is used to characterize the degree of difference between the first network element state data and the second network element state data;
[0248] The second loss function is used to characterize the degree of difference between the first intervention flow and the second intervention flow.
[0249] In one exemplary embodiment, the processing module 1120 is specifically used for:
[0250] Based on the type of the target task, the target combination of atomic modules is determined in the task processing model;
[0251] The target task is processed using the target combination.
[0252] In one exemplary embodiment, the processing module 1120 is specifically used for:
[0253] When the target task is a time-series prediction task, the target combination is determined to include: acquiring atomic modules, self-attention atomic modules, and restoring atomic modules;
[0254] When the target task is a root cause localization task, the target combination is determined to include: the acquisition atom module and the self-attention atom module;
[0255] When the target task is a causal inference task, the target combination is determined to include: the acquisition atom module, the multilayer perceptron atom module, the self-attention atom module, the fusion atom module, and the restoration atom module;
[0256] When the target task is a parameter tuning decision task, the target combination is determined to include: the acquisition atomic module and the mutual attention atomic module.
[0257] In one exemplary embodiment, the processing module 1120 is specifically used for:
[0258] After the target combination is connected to the classification layer, the target processing object is obtained;
[0259] The root cause localization task is processed using the target processing object.
[0260] This disclosure also provides a task processing model, please refer to... Figure 6 And related explanations. This task processing module includes multiple atomic modules, which include:
[0261] An atomic acquisition module is used to acquire the multi-stream data; the multi-stream data includes: a first intervention stream, a first network element snapshot stream, a desired stream, and a signal stream.
[0262] The mutual attention atom module is used to process the multi-stream data based on the mutual attention mechanism and the parameter tuning time to obtain the second intervention stream;
[0263] A multilayer perceptron atomic module is used to perform multilayer perceptron processing on the second intervention flow and the first network element snapshot flow to obtain the second network element snapshot flow;
[0264] The self-attention atom module is used to process the first network element snapshot stream based on the self-attention mechanism to obtain the third network element snapshot stream;
[0265] A fusion atom module is used to fuse the second network element snapshot stream and the third network element snapshot stream to obtain a fourth network element snapshot stream;
[0266] The restoration atom module is used to restore the snapshot stream of the fourth network element to obtain the state data of the second network element.
[0267] This disclosure also provides a task processing system. Figure 12 A structural block diagram of a task processing system provided in this disclosure embodiment is shown below. Figure 12 As shown, the task processing system 1200 includes:
[0268] Task processing model 1210 as described in any of the preceding embodiments;
[0269] The task processing device 1220 is used to execute the task processing method as described in any of the preceding embodiments.
[0270] This disclosure also provides an electronic device. Figure 13 This is a hardware block diagram of an electronic device provided according to an embodiment of the present disclosure. The electronic device 1300 according to an embodiment of the present disclosure includes at least a processor; and a memory for storing computer-readable instructions. When the computer-readable instructions are loaded and executed by the processor, the processor performs the task processing method described in any of the preceding embodiments of the present disclosure.
[0271] Figure 13 The illustrated electronic device 1300 specifically includes a central processing unit (CPU) 1301, a graphics processing unit (GPU) 1302, and a memory 1303. These units are interconnected via a bus 1304. The CPU 1301 and / or GPU 1302 can function as the aforementioned processor, and the memory 1303 can function as the aforementioned memory storing computer-readable instructions. Furthermore, the electronic device 1300 may also include a communication unit 1305, a storage unit 1306, an output unit 1307, an input unit 1308, and an external device 1309, all of which are also connected to the bus 1304.
[0272] Figure 14 This is a schematic diagram of a computer-readable storage medium provided in an embodiment of this disclosure. (As shown...) Figure 14As shown, a computer-readable storage medium 1400 according to an embodiment of the present disclosure stores computer-readable instructions 1401 thereon. When the computer-readable instructions 1401 are executed by a processor, the task processing method described with reference to the above figures according to any embodiment of the present disclosure is performed. The computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, optical disk, magnetic disk, etc.
[0273] This disclosure further provides a computer program product, including a computer program that, when executed by a processor, implements the task processing method described in any of the preceding embodiments of this disclosure.
[0274] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0275] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0276] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0277] Additionally, as used herein, the “or” used in a list of items beginning with “at least one” indicates a separate list, such that a list of, for example, “at least one of A, B, or C” means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word “exemplary” does not imply that the described example is preferred or better than other examples.
[0278] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.
[0279] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.
[0280] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0281] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A task processing method, characterized in that, The method includes: Obtain a task processing model; the task processing model is pre-trained based on multi-stream data, the multi-stream data including at least: an intervention stream, the intervention stream being used to describe intervention parameter values; The target task is processed using the task processing model described above.
2. The method according to claim 1, characterized in that, The multi-stream data includes: The intervention flow; Expected flow, used to describe the value of the key indicator at the target time; Signal stream, used to record parameter tuning times; Network element snapshot streams are used to describe high-dimensional representations of network element status data.
3. The method according to claim 1, characterized in that, The task processing model includes multiple atomic modules, and the atomic modules include: An atomic acquisition module is used to acquire the multi-stream data; the multi-stream data includes: a first intervention stream, a first network element snapshot stream, a desired stream, and a signal stream. The mutual attention atom module is used to process the multi-stream data based on the mutual attention mechanism and the parameter tuning time to obtain the second intervention stream; A multilayer perceptron atomic module is used to perform multilayer perceptron processing on the second intervention flow and the first network element snapshot flow to obtain the second network element snapshot flow; The self-attention atom module is used to process the first network element snapshot stream based on the self-attention mechanism to obtain the third network element snapshot stream; A fusion atom module is used to fuse the second network element snapshot stream and the third network element snapshot stream to obtain a fourth network element snapshot stream; The restoration atom module is used to restore the snapshot stream of the fourth network element to obtain the state data of the second network element.
4. The method according to claim 3, characterized in that, The mutual attention atom module is specifically used for: Obtain combined data of the desired flow, the first network element snapshot flow, and the first intervention flow; The combined data and the first intervention stream are processed using the mutual attention mechanism to obtain a third intervention stream; The first intervention parameter of the first intervention stream is combined with the third intervention stream to obtain the fourth intervention stream; The second intervention flow is obtained by obtaining the dot product between the fourth intervention flow and the signal flow.
5. The method according to claim 3, characterized in that, The atom acquisition module is specifically used for: The first network element status data in the input data is subjected to high-dimensional representation processing to obtain the first network element snapshot stream; wherein, the high-dimensional representation processing corresponds to the processing method of the restoration processing; Extract the first intervention stream, the desired stream, and the signal stream from the input data; The input data satisfies the following format: (batchsize, m, k, s); where batchsize represents the number of samples, m represents the number of network element states, k represents the number of time slices contained in a network element state, and s represents the number of indicators; any network element state includes indicator data for k time slices.
6. The method according to claim 5, characterized in that, When the input data is training data, the step of extracting the first intervention stream, the expected stream, and the signal stream from the input data includes: From the input data, intervention parameters within a preset time interval after the first moment are obtained to obtain the first intervention stream; From the input data, the target index value after a preset time interval following the first moment is obtained to obtain the desired flow; The signal flow is determined based on the second time point; Wherein, the first moment is used to describe the moment when the network anomaly occurs; the second moment is the actual parameter tuning moment; the first moment is before the second moment.
7. The method according to claim 5, characterized in that, When the input data is task data, the step of extracting the first intervention stream, the expected stream, and the signal stream from the input data includes: From the input data, intervention parameters within a preset time interval prior to the first moment are obtained to obtain the first intervention stream; From the input data, the target index value after a preset time interval following the first moment is obtained to obtain the desired flow; The signal flow is determined based on the first moment; The first moment is used to describe the moment when the network anomaly occurs.
8. The method according to claim 5, characterized in that, The signal stream is a parameter tuning matrix, and the position of the target identifier in any row of the parameter tuning matrix is used to indicate the index position of the intervention value; Extracting the first intervention stream, the desired stream, and the signal stream from the input data includes: The parameter tuning information at a specified time is obtained from the input data. The parameter tuning information is used to indicate whether parameter tuning occurs at the specified time and the parameter tuning intervention value. The specified time includes: a first time or a second time. Based on the parameter tuning information, the parameter tuning matrix is generated; The parameter tuning matrix satisfies the format: (batchsize, x, x); where batchsize represents the number of samples and x represents the number of parameter tuning times in the preset time interval.
9. The method according to any one of claims 1-8, characterized in that, The method further includes: Obtain training samples; The task processing model is trained using the training samples and the target loss function until the preset training conditions are met. Wherein, the target loss function is the weighted sum of the first loss function and the second loss function; The first loss function is used to characterize the degree of difference between the first network element state data and the second network element state data; The second loss function is used to characterize the degree of difference between the first intervention flow and the second intervention flow.
10. The method according to any one of claims 1-8, characterized in that, The process of processing the target task using the task processing model includes: Based on the type of the target task, the target combination of atomic modules is determined in the task processing model; The target task is processed using the target combination.
11. The method according to claim 10, characterized in that, The determination of the target combination of atomic modules in the task processing model based on the type of the target task includes: When the target task is a time-series prediction task, the target combination is determined to include: acquiring atomic modules, self-attention atomic modules, and restoring atomic modules; When the target task is a root cause localization task, the target combination is determined to include: the acquisition atom module and the self-attention atom module; When the target task is a causal inference task, the target combination is determined to include: the acquisition atom module, the multilayer perceptron atom module, the self-attention atom module, the fusion atom module, and the restoration atom module; When the target task is a parameter tuning decision task, the target combination is determined to include: the acquisition atomic module and the mutual attention atomic module.
12. The method according to claim 10, characterized in that, When the target task is a root cause localization task, the step of processing the target task using the target combination includes: After the target combination is connected to the classification layer, the target processing object is obtained; The root cause localization task is processed using the target processing object.
13. A task processing device, characterized in that, The device includes: An acquisition module is used to acquire a task processing model; the task processing model is obtained based on multi-stream data pre-training, and the multi-stream data includes at least: an intervention stream, which is used to describe intervention parameter values; The processing module is used to process the target task using the task processing model.
14. A task processing model, characterized in that, It includes multiple atomic modules, the atomic modules including: An atom acquisition module is used to acquire multi-stream data; the multi-stream data includes: a first intervention stream, a first network element snapshot stream, a desired stream, and a signal stream; The mutual attention atom module is used to process the multi-stream data based on the mutual attention mechanism and the parameter tuning time to obtain the second intervention stream; A multilayer perceptron atomic module is used to perform multilayer perceptron processing on the second intervention flow and the first network element snapshot flow to obtain the second network element snapshot flow; The self-attention atom module is used to process the first network element snapshot stream based on the self-attention mechanism to obtain the third network element snapshot stream; A fusion atom module is used to fuse the second network element snapshot stream and the third network element snapshot stream to obtain a fourth network element snapshot stream; The restoration atom module is used to restore the snapshot stream of the fourth network element to obtain the state data of the second network element.
15. A task processing system, characterized in that, include: The task processing model as described in claim 13; A task processing apparatus for performing the task processing method as described in any one of claims 1-12.
16. An electronic device, characterized in that, include: Memory, used to store computer-readable instructions; as well as A processor for executing the computer-readable instructions, causing the electronic device to perform the method as described in any one of claims 1-12.
17. A non-transitory computer-readable storage medium, characterized in that, Used to store computer-readable instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-12.
18. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-12.