Task processing method, device, equipment, readable storage medium and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-11
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本申请的至少一个实施例提供了一种任务处理方法、装置、设备、可读存储介质及程序产品,用于解决现有技术中在面对用户投诉驱动的工单、能效约束与覆盖优化等多目标矛盾并存的情况下,无法在通过智能体进行参数调整的同时保证网络多方面的均衡性的问题
[0040]与现有技术相比,本申请实施例提供的任务处理方法、装置、设备、可读存储介质及程序产品,通过所述第二智能体对所述待处理任务进行识别,确定所述候选动作和所述效果评估值;从而使得所述第一智能体根据所述候选动作和所述效果评估值构建所述影响度矩阵,然后确定所述用于表征不同所述候选动作之间的关联关系的跨目标耦合度;从而根据所述跨目标耦合度和所述影响度矩阵确定所述待处理任务的响应主体,并最终根据所述候选动作以及所述效果评估值确定所述执行策略,并向所述响应主体发送所述执行策略。解决了现有技术中面对用户投诉驱动的工单、能效约束与覆盖优化等多目标矛盾并存的情况下,无法在通过智能体进行参数调整的同时保证网络多方面的均衡性的问题。
Smart Images

Figure CN122554875A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of mobile communication operation and maintenance technology, specifically to a task processing method, apparatus, device, readable storage medium, and program product. Background Technology
[0002] In recent years, with the rapid development of 5G and future 6G networks, the scale and complexity of cellular networks have continued to increase. The traditional network operation and maintenance model that relies on human experience can no longer meet the needs of real-time and refined operation and maintenance. Instead, it is necessary to comprehensively coordinate multiple tasks such as complaint handling, energy efficiency optimization, network optimization, outsourced maintenance execution and energy-saving control.
[0003] In existing technologies, knowledge distillation and hierarchical replay pool design have alleviated the problems of low training efficiency and inter-task interference in multi-task, multi-agent systems. However, their application scenarios are mainly concentrated in the model training stage, and their ability to handle dynamic conflicts across domains in actual network operation and maintenance is limited. Especially when faced with multiple conflicting objectives such as user complaint-driven work orders, energy efficiency constraints, and coverage optimization, it is impossible to ensure the balance of various aspects of the network while adjusting parameters through agents. Summary of the Invention
[0004] At least one embodiment of this application provides a task processing method, apparatus, device, readable storage medium, and program product to solve the problem in the prior art that, when faced with multiple conflicting objectives such as user complaint-driven work orders, energy efficiency constraints, and coverage optimization, it is impossible to ensure the balance of various aspects of the network while adjusting parameters through an intelligent agent.
[0005] To solve the above-mentioned technical problems, this application is implemented as follows:
[0006] In a first aspect, embodiments of this application provide a task processing method applied to a first intelligent agent, comprising:
[0007] The system receives multiple candidate actions and effect evaluation values generated by second agents for a task to be processed; the candidate actions are actions that adjust network parameters; and the effect evaluation values are predictions of the effects achieved by executing the candidate actions.
[0008] Construct an influence matrix; the influence matrix is the rate of change of the effect evaluation value as a function of the candidate action;
[0009] Based on the influence matrix, the cross-target coupling degree is determined; the cross-target coupling degree is used to characterize the correlation between different candidate actions.
[0010] The response agent for the task to be processed is determined based on the cross-target coupling degree and the influence matrix; the response agent includes at least one of the second agents.
[0011] Based on the candidate actions and the effect evaluation values, an execution strategy is sent to the response entity.
[0012] Further, based on the candidate action and the effect evaluation value, an execution strategy is sent to the response subject, including:
[0013] When the response subject includes multiple second agents, the effect estimate is subjected to Bayesian weighted fusion to obtain multiple first target actions;
[0014] The degree of violation of the plurality of first target actions is determined; the degree of violation is determined according to the safety constraint corresponding to each action; the safety constraint is a threshold corresponding to the effect evaluation value;
[0015] Based on the degree of violation, a judgment result is determined on whether there is a conflict among the plurality of first target actions;
[0016] Based on the judgment result, the execution strategy is determined and sent to the response subject.
[0017] Further, based on the judgment result, the execution strategy is determined, including:
[0018] In the event of a conflict among the multiple first target actions, the approval rating of the second agent for each candidate action is obtained; the approval rating is the result of the second agent's judgment on the probability of successfully executing the candidate action.
[0019] The execution strategy is determined based on the acceptance level, the domain preference vector, and the violation level; the domain preference vector is determined according to the type of the second agent.
[0020] Further, the execution strategy is determined based on the acceptance level, the domain preference vector, and the violation degree, including:
[0021] The weighted consensus score is obtained by weighting and summing the recognition level and the domain preference vector.
[0022] In the case of a second target action with a violation degree of zero and a weighted consensus score greater than a first preset threshold, the second target action and the second agent corresponding to the second target action are determined as the execution strategy.
[0023] Furthermore, in the absence of the second target action, the method further includes:
[0024] Construct a conflict matrix, which represents the opposition strength between the candidate actions; the opposition strength is determined based on the violation degree.
[0025] Based on the intensity of the opposition, determine the global conflict;
[0026] If the global conflict exceeds a second preset threshold, the execution strategy is determined through an arbitration model.
[0027] The arbitration model is related to target utility, retention point, and priority; the target utility is the weighted average of the effect evaluation values when the target action is performed; the retention point is the threshold of the weighted average; and the priority is determined according to the application scenario of the task to be processed.
[0028] Furthermore, the method also includes:
[0029] The degree of irreconcilability is calculated based on the degree of acceptance, the security constraints, and the weighted consensus score.
[0030] If the incompatibility is greater than a second preset threshold, a capacity expansion suggestion is generated; the capacity expansion suggestion includes at least one of the following: target cell, type of base station to be added, coverage gain, and energy consumption boundary.
[0031] Secondly, embodiments of this application provide a task processing apparatus, including:
[0032] A receiving module is used to receive candidate actions and effect evaluation values generated by multiple second agents for the task to be processed; the candidate actions are actions that adjust network parameters; the effect evaluation values are prediction results of the effect achieved by executing the candidate actions;
[0033] A construction module is used to construct an influence matrix; the influence matrix is the rate of change of the effect evaluation value as a function of the candidate action;
[0034] The first determining module is used to determine the cross-target coupling degree based on the influence degree matrix; the cross-target coupling degree is used to characterize the correlation between different candidate actions;
[0035] The second determining module is used to determine the response subject of the task to be processed based on the cross-target coupling degree and the influence matrix; the response subject includes at least one of the second intelligent agents;
[0036] The sending module is used to send the execution strategy to the response subject based on the candidate action and the effect evaluation value.
[0037] Thirdly, embodiments of this application provide a task processing device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the task processing method described above.
[0038] Fourthly, embodiments of this application provide a computer-readable storage medium storing a program, which, when executed by a processor, implements the steps of the task processing method described above.
[0039] Fifthly, embodiments of this application provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the task processing method described above.
[0040] Compared with existing technologies, the task processing method, apparatus, device, readable storage medium, and program product provided in this application identify the task to be processed by a second intelligent agent, determine the candidate actions and the effect evaluation value; thereby enabling the first intelligent agent to construct the influence matrix based on the candidate actions and the effect evaluation value, and then determine the cross-target coupling degree used to characterize the correlation between different candidate actions; thereby determining the response subject of the task to be processed based on the cross-target coupling degree and the influence matrix, and finally determining the execution strategy based on the candidate actions and the effect evaluation value, and sending the execution strategy to the response subject. This solves the problem in existing technologies where, in the face of multiple conflicting objectives such as user complaint-driven work orders, energy efficiency constraints, and coverage optimization, it is impossible to ensure the balance of various aspects of the network while adjusting parameters through an intelligent agent. Attached Figure Description
[0041] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the scope of this application. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:
[0042] Figure 1 This is a schematic diagram illustrating an application scenario according to an embodiment of this application;
[0043] Figure 2 This is a schematic diagram illustrating the steps of the task processing method according to an embodiment of the present application when applied to a first intelligent agent;
[0044] Figure 3 This is a schematic diagram of the module of the task processing device according to an embodiment of this application;
[0045] Figure 4This is a schematic diagram of the structure of a task processing device according to an embodiment of this application. Detailed Implementation
[0046] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, without limiting the number of objects; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, "A or B" covers three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0047] The term "instruction" in this application can be either a direct instruction (or explicit instruction) or an indirect instruction (or implicit instruction). A direct instruction can be understood as one in which the sender explicitly informs the receiver of specific information, the operation to be performed, or the requested result, etc.; an indirect instruction can be understood as one in which the receiver determines the corresponding information based on the instruction sent by the sender, or makes a judgment and determines the operation to be performed or the requested result, etc., based on the judgment result.
[0048] It is worth noting that the technologies described in this application are not limited to Long Term Evolution (LTE) / LTE-Advanced (LTE-A) systems, but can also be used in other wireless communication systems, such as Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Frequency Division Multiple Access (FDMA), Orthogonal Frequency Division Multiple Access (OFDMA), Single-carrier Frequency-Division Multiple Access (SC-FDMA), or other systems. The terms "system" and "network" in this application are often used interchangeably, and the described technologies can be used in the systems and radio technologies mentioned above, as well as in other systems and radio technologies. The following description describes New Radio (NR) systems for illustrative purposes, and the term NR is used in most of the following description; however, these technologies can also be applied to systems other than NR systems, such as 6th Generation (6G) communication systems.
[0049] Figure 1This diagram illustrates a block diagram of a wireless communication system applicable to embodiments of this application. The wireless communication system includes a terminal 11 and a network device 12. The terminal 11 can be a mobile phone, tablet computer, laptop computer, notebook computer, personal digital assistant (PDA), handheld computer, netbook, ultra-mobile personal computer (UMPC), mobile internet device (MID), augmented reality (AR), virtual reality (VR) device, robot, wearable device, flight vehicle, vehicle user equipment (VUE), shipboard equipment, pedestrian user equipment (PUE), smart home (home devices with wireless communication capabilities, such as refrigerators, televisions, washing machines, or furniture), game console, personal computer (PC), ATM, or self-service machine, etc. Wearable devices include: smartwatches, smart bracelets, smart headphones, smart glasses, smart jewelry (smart bracelets, smart chains, smart rings, smart necklaces, smart anklets, smart anklets, etc.), smart wristbands, smart clothing, etc. Among these, in-vehicle devices can also be referred to as in-vehicle terminals, in-vehicle controllers, in-vehicle modules, in-vehicle components, in-vehicle chips, or in-vehicle units, etc. It should be noted that the specific type of terminal 11 is not limited in this application embodiment. Network device 12 may include access network devices or core network devices, wherein access network devices may also be referred to as Radio Access Network (RAN) devices, radio access network functions, or radio access network units. Access network devices may include base stations, Wireless Local Area Network (WLAN) access points (APs), or Wireless Fidelity (WiFi) nodes, etc.In this context, a base station may be referred to as a Node B (NB), an Evolved Node B (eNB), a Next Generation Node B (gNB), a New Radio Node B (NR Node B), an Access Point, a Relay Base Station (RBS), a Serving Base Station (SBS), a Base Transceiver Station (BTS), a Radio Base Station, a Radio Transceiver, a Basic Service Set (BSS), an Extended Service Set (ESS), a Home Node B (HNB), a Home Evolved Node B, a Transmission Reception Point (TRP), or any other suitable term in the relevant field, as long as the same technical effect is achieved. The base station is not limited to any specific technical terminology. It should be noted that in this application embodiment, only a base station in an NR system is used as an example for introduction, and the specific type of base station is not limited.
[0050] Core network equipment may include, but is not limited to, at least one of the following: core network node, core network function, Mobility Management Entity (MME), Access and Mobility Management Function (AMF), Session Management Function (SMF), User Plane Function (UPF), Policy Control Function (PCF), Policy and Charging Rules Function (PCRF), Edge Application Server Discovery Function (EASDF), Unified Data Management (UDM), Unified Data Repository (UDR), Home Subscriber Server (HSS), Centralized network configuration (CNC), Network Repository Function (NRF), Network Exposure Function (NEF), Local NEF (or L-NEF), Binding Support Function (BSF), and Application Function. Function (AF), etc. It should be noted that the embodiments of this application only use the core network equipment in the NR system as an example for introduction, and do not limit the specific type of core network equipment.
[0051] To enable those skilled in the art to better understand the embodiments of this application, the following description is provided first:
[0052] With the rapid development of 5G and future 6G networks, the scale and complexity of cellular networks are constantly increasing. Traditional network operation and maintenance models that rely on human experience are no longer able to meet the demands for real-time and refined operations. Therefore, academia and industry generally adopt multi-agent collaborative methods based on intelligent agents to assist in network operation and maintenance and optimization.
[0053] As described in the background section, in the prior art, when faced with multiple conflicting objectives such as user complaint-driven work orders, energy efficiency constraints, and coverage optimization, it is impossible to ensure the balance of various aspects of the network while adjusting parameters through an intelligent agent. This seriously affects the user experience. To solve at least one of the above problems, this application provides a task processing method that can reduce or avoid the occurrence of the above situations and improve the user experience.
[0054] This application provides a task processing method and apparatus. The method and apparatus are based on the same concept, and since the principles by which the method and apparatus solve problems are similar, their implementations can be mutually referenced; repeated details will not be repeated.
[0055] like Figure 2 As shown in the embodiment of this application, a task processing method, when applied to a first intelligent agent, includes the following steps:
[0056] Step 201: Receive multiple candidate actions and effect evaluation values generated by second agents for the task to be processed; the candidate actions are actions that adjust network parameters; the effect evaluation values are prediction results of the effect achieved by executing the candidate actions;
[0057] Step 202, construct an influence matrix; the influence matrix is the rate of change of the effect evaluation value as a function of the candidate action;
[0058] Step 203: Determine the cross-target coupling degree based on the influence degree matrix; the cross-target coupling degree is used to characterize the correlation between different candidate actions;
[0059] Step 204: Determine the response agent of the task to be processed based on the cross-target coupling degree and the influence matrix; the response agent includes at least one of the second agents;
[0060] Step 205: Send the execution strategy to the response subject based on the candidate action and the effect evaluation value.
[0061] The task processing method of this application embodiment identifies the task to be processed by a second intelligent agent, determines the candidate actions and the effect evaluation value, and then enables the first intelligent agent to construct the influence matrix based on the candidate actions and the effect evaluation value. Next, it determines the cross-target coupling degree used to characterize the correlation between different candidate actions. Based on the cross-target coupling degree and the influence matrix, it determines the response subject of the task to be processed, and finally determines the execution strategy based on the candidate actions and the effect evaluation value, and sends the execution strategy to the response subject. This solves the problem in the prior art where, in the face of multiple conflicting objectives such as user complaint-driven work orders, energy efficiency constraints, and coverage optimization, it is impossible to ensure the balance of various aspects of the network while adjusting parameters through an intelligent agent.
[0062] It should be noted that the first intelligent agent is the master intelligent agent;
[0063] The second intelligent agent is a slave intelligent agent, including: a complaint intelligent agent, a network optimization intelligent agent, an energy efficiency optimization intelligent agent, a maintenance agent, and an energy-saving intelligent agent.
[0064] Optionally, the candidate actions include: adjusting the base station power, adjusting the base station antenna downtilt angle, switching the carrier on and off, and whether or not to access the station.
[0065] The performance evaluation value can be a weighted value of parameters such as network coverage area, cell overlap coverage rate, base station energy consumption, and energy efficiency.
[0066] Optionally, after receiving the task to be processed, the second agent generates the candidate action and the effect evaluation value through the second policy library.
[0067] Specifically, the sub-objectives and constraint weights are output through the second strategy library:
[0068] ;
[0069] in, The candidate action;
[0070] For multi-index vectors (the effect estimation refers to...) The predicted change;
[0071] For uncertainty estimation (covariance);
[0072] This is the second strategy library;
[0073] This represents the local observation of the second agent at time t;
[0074] For semantic sub-targets (the effect evaluation value), for example, a gain of ≥6dB for weak coverage points in a certain area;
[0075] for For multi-objective weights;
[0076] These are hard constraint thresholds (such as energy consumption growth ≤8%, overlap coverage ≤10%, etc.).
[0077] It should be noted that the first intelligent agent maintains the queue of tasks to be processed Q (e.g., work orders, alarms, or complaints) and the first policy library:
[0078]
[0079] in, The first strategy library;
[0080] The global network state at time t is .
[0081] The first intelligent agent and the second intelligent agent communicate using structured messages; Domain confidence; system maintenance level replay buffer They store task semantics and policy-effect pairs respectively for continuous learning and backtracking evaluation.
[0082] Optionally, an influence matrix is constructed, including:
[0083] ;
[0084] in, The influence matrix represents the rate of change of the effect evaluation value as a function of the candidate action.
[0085] For adjusting the vector (i.e., the candidate actions, such as power, downtilt angle, carrier switching, whether to go to the station, etc.);
[0086] This is a KPI vector.
[0087] Optionally, determining the cross-target coupling degree based on the influence matrix includes:
[0088] ;
[0089] in, The sensitive block for the main target (the value of the influence matrix);
[0090] Sensitive blocks for other targets;
[0091] Prevent division by zero.
[0092] It should be noted that the main objective is related to the task to be processed; the other objectives are all the effect evaluation values other than the main objective.
[0093] For example, if the task to be processed is a user complaint about weak signal, then the main objective is coverage of the target area.
[0094] Optionally, determining the response subject of the task to be processed based on the cross-target coupling degree and the influence matrix includes:
[0095] ;
[0096] in, The third preset threshold;
[0097] The fourth preset threshold;
[0098] and, It equals 0 or 1.
[0099] In this embodiment of the application, when At that time, the responding entity is a single second intelligent agent;
[0100] when At that time, the responding entity is a plurality of the second intelligent agents.
[0101] Furthermore, when the responding subject is a single second agent, the responding subject generates a first target action through the first policy library and executes the first target action independently;
[0102] When the responding entity consists of multiple second intelligent agents, a multi-agent collaborative process is carried out through the first intelligent agent.
[0103] For example, in a 5G macro base station equipment room, a power module failure occurs. The system first triggers task acceptance via the alarm channel. The first intelligent agent determines the task attributes: the alarm points to a clear hardware failure, and its impact on cross-domain indicators such as coverage, energy efficiency, interference, and user perception is negligible, belonging to a loosely coupled single-domain problem. Based on this, the first intelligent agent activates the single-agent mode, directly assigning the second intelligent agent (the maintenance agent) to generate a work order and organize on-site handling. After the maintenance agent completes standardized operations such as on-site inspection, power module replacement, and reset, it sends back the execution results. Since this type of failure does not involve network reconstruction on the parameter side, the first intelligent agent does not schedule other agents such as complaint, optimization, energy efficiency, or energy saving to intervene, avoiding unnecessary coordination overhead. After the fault is cleared, the site's services and KPIs return to the baseline level, and the task is archived in a closed loop. This embodiment shows that for hardware alarms with strong directional characteristics and small cross-domain impact, a single maintenance agent can independently complete the diagnosis and repair. The first intelligent agent only undertakes the responsibility of judgment and work order dispatch, thereby achieving a minimized decision-making chain and the fastest recovery time.
[0104] Optionally, based on the candidate action and the effect evaluation value, an execution strategy is sent to the response subject, including:
[0105] When the response subject includes multiple second agents, the effect estimate is subjected to Bayesian weighted fusion to obtain multiple first target actions;
[0106] The degree of violation of the plurality of first target actions is determined; the degree of violation is determined according to the safety constraint corresponding to each action; the safety constraint is a threshold corresponding to the effect evaluation value;
[0107] Based on the degree of violation, a judgment result is determined on whether there is a conflict among the plurality of first target actions;
[0108] Based on the judgment result, the execution strategy is determined and sent to the response subject.
[0109] The effect estimate is subjected to Bayesian weighted fusion, including:
[0110] ;
[0111] Furthermore, a feasible solution set (the first target action) is generated based on the Bayesian weighted fusion. ;
[0112] And further calculate the degree of violation of the first target action.
[0113] Specifically, the degree of violation is calculated using the following formula:
[0114] .
[0115] Optionally, based on the degree of violation, the determination of whether there is a conflict among the plurality of first target actions includes:
[0116] When the violation rate is 0, there is no conflict among the plurality of first target actions;
[0117] The response entity executes its corresponding first target action.
[0118] Optionally, based on the degree of violation, the determination of whether there is a conflict among the plurality of first target actions includes:
[0119] When the violation degree is greater than 0, there is a conflict among the plurality of first target actions.
[0120] Optionally, determining the execution strategy based on the judgment result includes:
[0121] In the event of a conflict among the multiple first target actions, the approval rating of the second agent for each candidate action is obtained; the approval rating is the result of the second agent's judgment on the probability of successfully executing the candidate action.
[0122] The execution strategy is determined based on the acceptance level, the domain preference vector, and the violation level; the domain preference vector is determined according to the type of the second agent.
[0123] It should be noted that when the response subject includes multiple second agents, the first agent constructs a multi-objective optimization function based on the semantic information of the task to be processed;
[0124] The loss of the effect estimate is determined based on the first target action, the effect estimate, and the multi-objective optimization function.
[0125] Optionally, the objective optimization function is:
[0126] ;
[0127] in, The loss term is the estimated effect value.
[0128] Optionally, the execution strategy is determined based on the acceptance level, the domain preference vector, and the violation degree, including:
[0129] The weighted consensus score is obtained by weighting and summing the recognition level and the domain preference vector.
[0130] In the case of a second target action with a violation degree of zero and a weighted consensus score greater than a first preset threshold, the second target action and the second agent corresponding to the second target action are determined as the execution strategy.
[0131] Optionally, the first agent calculates the approval rating of the candidate action:
[0132] ;
[0133] in, For the Sigmoid function;
[0134] The domain preference vector is determined based on the attributes of the second agent;
[0135] The level of recognition mentioned above.
[0136] Furthermore, the dual of the loss term is calculated and determined as the collaborative gain of the multiple second agents:
[0137] ;
[0138] in, For aggregation utility (and loss) (pair).
[0139] Optionally, the weighted consensus score is:
[0140] ;
[0141] Among them, there exists Make and At that time, determine This is the second target action;
[0142] Determine the second agent corresponding to the second target action, and execute the second agent.
[0143] Optionally, in the absence of the second target action, the method further includes:
[0144] Construct a conflict matrix, which represents the opposition strength between the candidate actions; the opposition strength is determined based on the violation degree.
[0145] Based on the intensity of the opposition, determine the global conflict;
[0146] If the global conflict exceeds a second preset threshold, the execution strategy is determined through an arbitration model.
[0147] The arbitration model is related to target utility, retention point, and priority; the target utility is the weighted average of the effect evaluation values when the target action is performed; the retention point is the threshold of the weighted average; and the priority is determined according to the application scenario of the task to be processed.
[0148] Optionally, the elements in the conflict matrix represent the candidate actions and the degree of opposition between them.
[0149] Optionally, the method further includes:
[0150] Build a global conflict:
[0151] ;
[0152] exist Or there are no feasible candidates At that time, the execution strategy is determined through an arbitration model.
[0153] Optionally, the arbitration model is:
[0154] ;
[0155] in, , is the target utility corresponding to the target action;
[0156] The reserved point;
[0157] This indicates the lexicographical priority set according to the business scenario (e.g., "priority of complaints during critical periods > coverage > interference > energy consumption").
[0158] It should be noted that the arbitration model can be transformed into an additive form through logarithmic transformation and solved using methods such as augmented Lagrange / projected gradient.
[0159] ;
[0160] If there exists Satisfying constraints and achieving ,but The joint action determined by the first intelligent agent;
[0161] Furthermore, the execution strategy is to execute the joint action through the second agent corresponding to the joint action.
[0162] In this embodiment of the application, the joint action output by the first intelligent agent... After it takes effect, in the window Continuously collect KPI changes If the "Alarm Cleared / Goal Achieved" event is triggered. Then apply the backtracking operator. The temporary compensation parameters are restored to the baseline to ensure that resources are not occupied for a long time and to avoid structural side effects.
[0163] Optionally, if the arbitration model fails to be solved (i.e. the solution cannot converge), the multi-objective weights are adjusted.
[0164] Specifically, ;
[0165] Where β is a hyperparameter;
[0166] Furthermore, the other objectives are changed from hard constraints to penalized soft constraints in order to expand the feasible domain and improve solvability.
[0167] Optionally, the method further includes:
[0168] The degree of irreconcilability is calculated based on the degree of acceptance, the security constraints, and the weighted consensus score.
[0169] If the incompatibility is greater than a second preset threshold, a capacity expansion suggestion is generated; the capacity expansion suggestion includes at least one of the following: target cell, type of base station to be added, coverage gain, and energy consumption boundary.
[0170] Optionally, the degree of irreconcilability is:
[0171] ;
[0172] Where γ is a hyperparameter;
[0173] when If the existing network resources are deemed insufficient, the first intelligent agent automatically generates a "planned site construction / expansion" suggestion (including target cell, suggested site type, estimated coverage gain and energy consumption boundary), and simultaneously issues a temporary emergency plan (such as a slight increase in neighboring cell power + a slight adjustment in downtilt angle) to ensure that short-term risks are controllable.
[0174] In this embodiment of the application, in a university campus scenario, the complaint agent receives a user's weak signal complaint for "a certain university in a certain city in a certain province and district." It then combines the location coordinates with the information of the relevant cell to complete a preliminary assessment and form a coverage improvement suggestion for the target area: increase the power of the "HF-Hefei University of Science and Technology South Campus-HHF-31" sector from 14.8 dBm to 17.8 dBm and adjust the electronic downtilt angle from 6° to 2°. It is expected that the coverage of the target area can be improved from about -112 dBm to about -100 dBm. The complaining agent relayed the preliminary plan and expected results to the first agent. Based on this, the first agent initiated a multi-agent collaborative process and solicited opinions from relevant second agents: the network optimization agent assessed the coverage and interference of the "power +3 dB, downtilt angle -4°" combination, pointing out that this plan would increase the overlap coverage from 8% to 11%, exceeding the established threshold; the energy efficiency optimization agent calculated the overall station energy consumption change, determining that the energy consumption increase was approximately 10.8%, exceeding the 8% threshold; the energy conservation agent explained from the perspective of network-wide energy conservation scheduling and performance constraints that continuous high-power operation would conflict with existing energy conservation strategies. Based on the feedback from all parties, the first agent confirmed that the preliminary plan had a significant conflict between coverage improvement and energy efficiency / energy conservation constraints, and thus entered the arbitration stage: while maintaining the achievability of the complaint target, the adjustment range was narrowed, proposing a compromise solution, increasing the sector power only to 16.8 dBm and adjusting the electronic downtilt angle to 4°, and again seeking approval from each of the second agents. After network optimization and energy efficiency optimization review, the compromise solution can improve the coverage of the target area to approximately -105 dBm, with the overlap coverage rate controlled within the threshold, and the overall station energy consumption increase to approximately 6.8%, which does not exceed the threshold. The energy-saving intelligent agent confirms that the solution is compatible with the current green energy schedule. Based on this, the first intelligent agent determines the final implementation plan, issues it to the maintenance side for implementation, and requires continuous monitoring of the effect. After implementation, the on-site feedback of parameter changes and effectiveness records is returned, the complaint indicators reach the improvement target, the KPI stabilizes within the controllable range, and the task is archived in a closed loop. This embodiment fully demonstrates the closed-loop process of "complaint intelligent agent proposing a preliminary solution - the first intelligent agent organizing multi-agent review - cross-domain conflict exposure - the first intelligent agent arbitrating and generating a compromise solution - review and approval by all parties - implementation and verification," demonstrating that under the condition of multiple objectives being mutually constrained, a final solution acceptable to all parties and with controllable risks can be obtained through master-slave collaboration.
[0175] Furthermore, to continuously improve the quality of the strategy, this application introduces hierarchical reinforcement learning for strategy iteration without changing the master-slave decision-making logic.
[0176] Specifically, the first policy library of the first agent adopts a high-level Actor-Critic, with the goal of maximizing constrained long-term returns.
[0177] ;
[0178] in, For the reason Derivative rewards.
[0179] From agent policy We employ an uncertainty-based advantage actor-critic approach to minimize prediction residuals and constrain covariance. .
[0180] Cross-task migration is achieved through a hierarchical replay buffer, and semantic sub-target distillation is used to compress decision priors for frequent scenarios into lightweight quantum policies, thereby improving online response efficiency.
[0181] The various methods described above are based on embodiments of this application. Apparatus for implementing the above methods will now be provided.
[0182] like Figure 3 As shown in the figure, this application embodiment also provides a task processing device 300, including:
[0183] The receiving module 301 is used to receive multiple candidate actions and effect evaluation values generated by the second agents for the task to be processed; the candidate actions are actions that adjust network parameters; the effect evaluation values are prediction results of the effect achieved by executing the candidate actions;
[0184] Construction module 302 is used to construct an influence matrix; the influence matrix is the rate of change of the effect evaluation value as a function of the candidate action;
[0185] The first determining module 303 is used to determine the cross-target coupling degree based on the influence degree matrix; the cross-target coupling degree is used to characterize the correlation between different candidate actions;
[0186] The second determining module 304 is used to determine the response subject of the task to be processed based on the cross-target coupling degree and the influence matrix; the response subject includes at least one of the second intelligent agents;
[0187] The sending module 305 is used to send the execution strategy to the response subject based on the candidate action and the effect evaluation value.
[0188] The task processing device of this application embodiment identifies the task to be processed by the second intelligent agent, determines the candidate actions and the effect evaluation value; thereby enabling the first intelligent agent to construct the influence matrix based on the candidate actions and the effect evaluation value, and then determine the cross-target coupling degree used to characterize the correlation between different candidate actions; thereby determining the response subject of the task to be processed based on the cross-target coupling degree and the influence matrix, and finally determining the execution strategy based on the candidate actions and the effect evaluation value, and sending the execution strategy to the response subject. This solves the problem in the prior art where, in the face of multiple conflicting objectives such as user complaint-driven work orders, energy efficiency constraints, and coverage optimization, it is impossible to ensure the balance of various aspects of the network while adjusting parameters through an intelligent agent.
[0189] The task processing device in the embodiments of this application, such as Figure 4 As shown, it includes a transceiver 410, a processor 400, a memory 420, and a program or instructions stored in the memory 420 and executable on the processor 400; when the processor 400 executes the program or instructions, it implements the various processes of the above-described task processing method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0190] The transceiver 410 is used to receive and send data under the control of the processor 400.
[0191] Among them, Figure 4 In this context, the bus architecture may include any number of interconnected buses and bridges, specifically linking various circuits together, represented by one or more processors (processor 400) and memory (memory 420). The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 410 may be multiple elements, including transmitters and receivers, providing a unit for communicating with various other devices over a transmission medium. The processor 400 is responsible for managing the bus architecture and general processing, and the memory 420 may store data used by the processor 400 during operation.
[0192] This application also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the various processes of the above-described task processing method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0193] This application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, they implement the various processes of the above-described task processing method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.
[0194] If the solution involves implicit data collection such as user location and internet access behavior, add a statement similar to the following to the specification:
[0195] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in this disclosed technical solution all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to safeguard user personal information security and network security.
[0196] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0197] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0198] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A task processing method, applied to a first intelligent agent, characterized in that, include: The system receives multiple candidate actions and performance evaluation values generated by second agents for the task to be processed; the candidate actions are actions that adjust network parameters. The effect evaluation value is the predicted result of the effect achieved by performing the candidate action; Construct an influence matrix; The influence matrix represents the rate of change of the effect evaluation value as a function of the candidate action. Based on the influence matrix, determine the cross-target coupling degree; The cross-target coupling degree is used to characterize the association relationship between different candidate actions; The response agent for the task to be processed is determined based on the cross-target coupling degree and the influence matrix; the response agent includes at least one of the second agents. Based on the candidate actions and the effect evaluation values, an execution strategy is sent to the response entity.
2. The method of claim 1, wherein, Based on the candidate actions and the effect evaluation values, an execution strategy is sent to the response entity, including: When the response subject includes multiple second agents, the effect estimate is subjected to Bayesian weighted fusion to obtain multiple first target actions; The degree of violation of the plurality of first target actions is determined; the degree of violation is determined according to the safety constraint corresponding to each action; the safety constraint is a threshold corresponding to the effect evaluation value; Based on the degree of violation, a judgment result is determined on whether there is a conflict among the plurality of first target actions; Based on the judgment result, the execution strategy is determined and sent to the response subject.
3. The method of claim 2, wherein, Based on the judgment result, the execution strategy is determined, including: In the event of a conflict among the multiple first target actions, the approval rating of the second agent for each candidate action is obtained; the approval rating is the result of the second agent's judgment on the probability of successfully executing the candidate action. The execution strategy is determined based on the acceptance level, the domain preference vector, and the violation level; the domain preference vector is determined according to the type of the second agent.
4. The method of claim 3, wherein, The execution strategy is determined based on the acceptance level, the domain preference vector, and the violation level, including: The weighted consensus score is obtained by weighting and summing the recognition level and the domain preference vector. In the case of a second target action with a violation degree of zero and a weighted consensus score greater than a first preset threshold, the second target action and the second agent corresponding to the second target action are determined as the execution strategy.
5. The method of claim 4, wherein, In the absence of the second target action, the method further includes: Construct a conflict matrix, which represents the opposition strength between the candidate actions; the opposition strength is determined based on the violation degree. Based on the intensity of the opposition, determine the global conflict; If the global conflict exceeds a second preset threshold, the execution strategy is determined through an arbitration model. The arbitration model is related to target utility, retention point, and priority; the target utility is the weighted average of the effect evaluation values when the target action is performed; the retention point is the threshold of the weighted average; and the priority is determined according to the application scenario of the task to be processed.
6. The method of claim 4, wherein, The method further includes: The degree of irreconcilability is calculated based on the degree of acceptance, the security constraints, and the weighted consensus score. If the incompatibility is greater than a second preset threshold, a capacity expansion suggestion is generated; the capacity expansion suggestion includes at least one of the following: target cell, type of base station to be added, coverage gain, and energy consumption boundary.
7. A task processing apparatus characterized by comprising: include: A receiving module is used to receive candidate actions and effect evaluation values generated by multiple second agents for the task to be processed; the candidate actions are actions that adjust network parameters. The effect evaluation value is the predicted result of the effect achieved by performing the candidate action; The building block is used to construct the influence matrix; The influence matrix represents the rate of change of the effect evaluation value as a function of the candidate action. The first determining module is used to determine the cross-target coupling degree based on the influence degree matrix; The cross-target coupling degree is used to characterize the association relationship between different candidate actions; The second determining module is used to determine the response subject of the task to be processed based on the cross-target coupling degree and the influence matrix; the response subject includes at least one of the second intelligent agents; The sending module is used to send the execution strategy to the response subject based on the candidate action and the effect evaluation value.
8. A task processing device characterized by comprising: include: Transceiver, processor, memory, and programs or instructions stored in the memory and executable on the processor; When the processor executes the program or instructions, it implements the steps of the method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 6.
10. A computer program product, characterised in that, Includes computer instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 6.