Method and device for determining broadband maintenance strategy and electronic equipment

By introducing a mechanism in broadband maintenance strategy where a first agent generates the strategy and a second agent evaluates it, the problem of the lack of real-time response and verification of LLM agents is solved, high-quality dynamic optimization decision-making is achieved, and the real-time performance and accuracy of maintenance strategy are improved.

CN121814583APending Publication Date: 2026-04-07CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing home broadband maintenance solutions based on Large Language Model (LLM) agents lack the ability to respond in real time to external feedback from users, engineers, and customer service, resulting in an inability to dynamically optimize decisions based on actual scenarios. Furthermore, the lack of an effective verification mechanism makes it difficult to guarantee the quality of decisions.

Method used

The first intelligent agent receives user fault reports, generates broadband maintenance strategies, and the second intelligent agent conducts the first round of evaluation, including task decomposition evaluation and effectiveness evaluation. Once the evaluation score exceeds a preset threshold, the strategy is determined as the target strategy and pushed to the user terminal.

Benefits of technology

It enables effective verification of external feedback from users and other sources, ensuring the quality of generated broadband maintenance strategies. It can dynamically optimize decisions based on actual scenarios, improving the real-time nature and accuracy of decisions and reducing user operation failures and repeated on-site service issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121814583A_ABST
    Figure CN121814583A_ABST
Patent Text Reader

Abstract

The invention discloses a broadband maintenance strategy determination method and apparatus, and an electronic device. The method comprises the following steps: receiving user fault reporting information through a first agent; generating a first broadband maintenance strategy at least according to the user fault reporting information through the first intelligent agent; a first round of evaluation is carried out on the first broadband maintenance strategy through the second intelligent agent, an evaluation score corresponding to the first round of evaluation is obtained, and multiple evaluation modes of the first round of evaluation include task decomposition evaluation and effectiveness evaluation; the validity evaluation is at least used for checking relevance between the first broadband maintenance strategy and an effective broadband maintenance decision of a historical target application scene, and the historical target application scene is a scene in which a similarity index of an application scene corresponding to the user fault reporting information is smaller than a preset similarity threshold; and under the condition that the evaluation score corresponding to the first round of evaluation is greater than a preset confidence threshold, determining the first broadband maintenance strategy as a target broadband maintenance strategy and pushing the target broadband maintenance strategy to the user side.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and more specifically, to a method, apparatus, and electronic device for determining a broadband maintenance strategy. Background Technology

[0002] Among related technologies, home broadband maintenance solutions based on Large Language Model (LLM) agents suffer from two main technical shortcomings. First, the LLM decision-making mechanism lacks real-time responsiveness to external feedback from users, engineers, and customer service, hindering dynamic optimization of decisions (broadband maintenance strategies) based on actual scenarios. For example, when user self-service operations fail or engineers find fault diagnoses incorrect, the system fails to adjust its judgment promptly. Second, the decisions generated by LLM lack effective verification mechanisms, making it difficult to guarantee decision quality. Inaccurate decisions can lead to two consequences: for simple faults, they may increase the probability of user operation failures; for complex faults, they may result in repeated on-site visits by engineers. These problems not only exacerbate user dissatisfaction but also waste human resources, failing to fundamentally solve the inefficiency of the traditional maintenance model (user report fault → manual intervention → on-site troubleshooting). Therefore, there is an urgent need to develop an improved architecture to enhance the intelligence level of broadband (e.g., home broadband) maintenance.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides a method, apparatus, and electronic device for determining broadband maintenance strategies, which at least solves the technical problem that the lack of an effective verification mechanism for external feedback from users and other sources makes it impossible to dynamically optimize broadband maintenance strategies according to actual scenarios, and the quality of broadband maintenance strategies is difficult to guarantee.

[0005] According to one aspect of the embodiments of this application, a method for determining a broadband maintenance strategy is provided, comprising: receiving user fault reporting information through a first intelligent agent, wherein the first intelligent agent is at least used to generate a broadband maintenance strategy corresponding to the user fault reporting information; generating a first broadband maintenance strategy through the first intelligent agent at least based on the user fault reporting information; performing a first round of evaluation on the first broadband maintenance strategy through a second intelligent agent to obtain an evaluation score corresponding to the first round of evaluation, wherein the multiple evaluation methods of the first round of evaluation include task decomposition evaluation and effectiveness evaluation, wherein the task decomposition evaluation is at least used to examine the quality of the sub-tasks decomposed in the first broadband maintenance strategy, and the effectiveness evaluation is at least used to examine the correlation between the first broadband maintenance strategy and effective broadband maintenance decisions of historical target application scenarios, wherein the historical target application scenario is a scenario whose similarity index with the application scenario corresponding to the user fault reporting information is less than a preset similarity threshold; and, if the evaluation score corresponding to the first round of evaluation is greater than a preset similarity threshold, determining the first broadband maintenance strategy as the target broadband maintenance strategy and pushing it to the user terminal.

[0006] According to another aspect of the embodiments of this application, a broadband maintenance strategy determination apparatus is also provided, comprising: a receiving module, configured to receive user fault reporting information through a first intelligent agent, wherein the first intelligent agent is at least configured to generate a broadband maintenance strategy corresponding to the user fault reporting information; a generating module, configured to generate a first broadband maintenance strategy through the first intelligent agent at least based on the user fault reporting information; an evaluation module, configured to perform a first round of evaluation on the first broadband maintenance strategy through a second intelligent agent to obtain an evaluation score corresponding to the first round of evaluation, wherein the multiple evaluation methods of the first round of evaluation include task decomposition evaluation and effectiveness evaluation, wherein the task decomposition evaluation is at least used to examine the quality of the sub-tasks decomposed in the first broadband maintenance strategy, and the effectiveness evaluation is at least used to examine the correlation between the first broadband maintenance strategy and the effective broadband maintenance decisions of historical target application scenarios, wherein the historical target application scenarios are scenarios whose similarity index with the application scenario corresponding to the user fault reporting information is less than a preset similarity threshold; and a determination module, configured to determine the first broadband maintenance strategy as a target broadband maintenance strategy and push it to the user terminal if the evaluation score corresponding to the first round of evaluation is greater than a preset similarity threshold.

[0007] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, wherein a program is stored in the non-volatile storage medium, wherein the method for determining the broadband maintenance strategy is controlled by the device where the non-volatile storage medium is located when the program is running.

[0008] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the program executes the above-described method for determining broadband maintenance strategies during runtime.

[0009] According to another aspect of the embodiments of this application, a computer program product is also provided, including computer instructions, which, when executed by a processor, implement the above-described method for determining broadband maintenance strategies.

[0010] In this embodiment, a first intelligent agent receives user fault reporting information, wherein the first intelligent agent is at least used to generate a broadband maintenance strategy corresponding to the user fault reporting information; the first intelligent agent generates a first broadband maintenance strategy based at least on the user fault reporting information; a second intelligent agent performs a first round of evaluation on the first broadband maintenance strategy to obtain an evaluation score corresponding to the first round of evaluation, wherein the first round of evaluation includes multiple evaluation methods such as task decomposition evaluation and effectiveness evaluation. The task decomposition evaluation is at least used to verify the quality of the sub-tasks decomposed in the first broadband maintenance strategy, and the effectiveness evaluation is at least used to verify the relationship between the first broadband maintenance strategy and the effective broadband maintenance decisions of historical target application scenarios. The system determines the broadband maintenance strategy based on its relevance to user reports. The target application scenario is one where the similarity index between the application scenario and the user's reported fault information is less than a preset similarity threshold. If the evaluation score in the first round of evaluation exceeds a preset threshold, the first broadband maintenance strategy is designated as the target strategy and pushed to the user. This process involves a first intelligent agent receiving user reports and generating the first broadband maintenance strategy, followed by a second intelligent agent performing a first-round evaluation. If the evaluation score exceeds the preset threshold, the first broadband maintenance strategy is designated as the target strategy and pushed to the user. This achieves effective evaluation and verification of external feedback from users, ensuring the quality of the generated broadband maintenance strategy. It also solves the technical problem of failing to dynamically optimize broadband maintenance strategies based on actual scenarios due to the lack of an effective verification mechanism for external feedback, thus hindering the quality assurance of broadband maintenance strategies. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0012] Figure 1 This is a hardware structure block diagram of a computer terminal for implementing a method for determining broadband maintenance strategies, according to an embodiment of this application.

[0013] Figure 2 This is a flowchart of a method for determining a broadband maintenance strategy according to an embodiment of this application;

[0014] Figure 3This is a schematic diagram of the architecture of a broadband maintenance strategy determination system provided according to an embodiment of this application;

[0015] Figure 4 This is a schematic diagram illustrating the operation process of a broadband maintenance strategy determination system according to an embodiment of this application;

[0016] Figure 5 This is a flowchart of another method for determining a broadband maintenance strategy according to an embodiment of this application;

[0017] Figure 6 This is a schematic diagram of a device for determining a broadband maintenance strategy according to an embodiment of this application. Detailed Implementation

[0018] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0019] The information collected in this application embodiment is information and data authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with the relevant laws, regulations and standards of the relevant regions, and necessary confidentiality measures have been taken. It does not violate public order and good morals, and provides corresponding operation entry points for users to choose to authorize or reject the automated decision results. If the user chooses to reject, the process will proceed to the expert decision-making process.

[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0021] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained below:

[0022] Artificial Intelligence (AI) is a comprehensive discipline that studies, develops, and applies theories, methods, technologies, and application systems to simulate, extend, and expand human intelligence. Its core objective is to enable machines to possess intelligent behaviors similar to humans, such as perception, reasoning, learning, and decision-making, thereby allowing them to autonomously or assistedly complete complex tasks.

[0023] Large Language Model (LLM): A deep learning model trained on massive amounts of text data. It learns the syntax, semantics, knowledge and logic of human language through a large number of parameters (usually billions to trillions). It can understand and generate natural language-like text and complete complex language tasks such as translation, question answering, creation and reasoning.

[0024] Domain-specific LLM: Domain-specific large-scale models refer to specialized large-scale models formed by optimizing, training, injecting knowledge, and adapting functions for specific industries or professional fields (such as healthcare, finance, law, education, manufacturing, etc.) based on the general LLM technical framework. Its core characteristics are that it has a deeper knowledge reserve, more accurate professional understanding, and more domain-specific task processing capabilities.

[0025] Prompts are textual instructions, questions, contextual descriptions, or examples that users input into the model to guide it in generating expected output. They are the core medium for human-large language model interaction, enabling the model to understand requirements and respond accordingly by clearly defining task objectives, providing background information, or setting output formats.

[0026] Fine-tuning is a technique that uses a pre-trained model (a model already trained on large-scale general data) as a foundation for secondary training on a specific task or small dataset. Its core purpose is to "adapt" the pre-trained model to a specific scenario, retaining general knowledge while learning the detailed rules of the specific task, thereby rapidly improving performance on the target task. Essentially, fine-tuning freezes most of the parameters of the pre-trained model (or only fine-tunes some layers), updates a small number of parameters with labeled data from the target task, allowing the model to remember the rules of the specific task while retaining its general capabilities.

[0027] Reinforcement Learning from Human Feedback (RLHF) is a machine learning paradigm that combines traditional reinforcement learning with human subjective judgment. Its core is to utilize direct human evaluations of agent behavior, such as preference ranking, rating, or correction, to dynamically adjust the model's optimization objective, making the agent's behavior patterns more aligned with human expectations.

[0028] Batch-Constrained Q-learning (BCQ) is a classic offline reinforcement learning algorithm designed to address scenarios where agents are trained only on a fixed historical dataset (offline data) and do not interact with the environment in real time. Its core idea is to strictly constrain the policy output range, ensuring that the model learns only behavioral patterns validated in offline data, thus avoiding performance degradation caused by reliance on unobserved data.

[0029] Low-Rank Adaptation (LoRA) is a parameter-efficient fine-tuning technique. Its core idea is to freeze most of the parameters of a pre-trained model and train only a small number of newly added low-rank matrix parameters to achieve efficient fine-tuning of large models, thereby reducing computational and storage costs while maintaining model performance.

[0030] In related technologies, home broadband maintenance solutions based on Large Language Model (LLM) agents suffer from two main technical drawbacks. First, the LLM decision-making mechanism lacks real-time responsiveness to external feedback from users, engineers, and customer service, making it impossible to dynamically optimize decisions (broadband maintenance strategies) based on actual scenarios. For example, when user self-service operations fail or engineers discover incorrect fault diagnoses, the system fails to adjust its judgment in a timely manner. Second, the decisions generated by LLM lack an effective verification mechanism, making it difficult to guarantee decision quality. Inaccurate decisions can have two consequences: for simple faults, they may increase the probability of user operation failures; for complex faults, they may lead to problems such as repeated on-site visits by engineers. Therefore, the lack of an effective verification mechanism for external feedback from users and other sources results in the inability to dynamically optimize broadband maintenance strategies based on actual scenarios, making it difficult to guarantee the quality of broadband maintenance strategies. To address this problem, this application provides relevant solutions, which are detailed below.

[0031] According to an embodiment of this application, an embodiment of a method for determining a broadband maintenance strategy is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0032] The methods and embodiments provided in this application can be executed on a computer terminal or similar computing device. Figure 1 A hardware block diagram of a computer terminal for implementing a method for determining broadband maintenance strategies is shown. Figure 1As shown, the computer terminal 10 may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0033] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10. As involved in the embodiments of this application, the data processing circuits serve as processor control (e.g., selection of a variable resistor termination path connected to an interface).

[0034] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the broadband maintenance strategy determination method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned broadband maintenance strategy determination method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0035] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0036] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10.

[0037] In the above operating environment, this application provides an embodiment of a method for determining a broadband maintenance strategy. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0038] like Figure 2 The diagram shown is a flowchart of a method for determining a broadband maintenance strategy according to an embodiment of this application, including:

[0039] Step S202: Receive user fault report information through the first intelligent agent.

[0040] In the technical solution provided in step S202, the first intelligent agent is at least used to generate a broadband maintenance strategy corresponding to the user's fault report information.

[0041] In some embodiments of this application, the first intelligent agent may be named an engineer intelligent agent. The system executing the broadband maintenance strategy determination method of the embodiments of this application includes a first intelligent agent (engineer intelligent agent) and a second intelligent agent (named inspector intelligent agent). Users report user fault information (broadband fault) to the system's first intelligent agent through various channels (such as telephone, online customer service, self-service fault reporting platform, etc.), and this operation serves as the starting point of the entire intelligent maintenance process. At this time, the first intelligent agent will determine the number of failures. Set it to 0.

[0042] Step S204: The first intelligent agent generates a first broadband maintenance strategy based at least on the user's fault report information.

[0043] In the technical solution provided in step S204, there are multiple ways to implement the first intelligent agent generating a first broadband maintenance strategy based at least on user fault reporting information. For example: the first intelligent agent acquires multi-dimensional data, including user equipment information, historical broadband maintenance data, operator backend data, and network status data; the first intelligent agent standardizes the user fault reporting information to obtain standardized user fault reporting information; the multi-dimensional data and standardized user fault reporting information are determined as input data; the broadband maintenance strategy prediction model of the first intelligent agent analyzes the input data to obtain the first broadband maintenance strategy. The first broadband maintenance strategy includes a thought process and a preliminary decision. The preliminary decision includes multiple initial sub-tasks for resolving the fault problem indicated in the user fault reporting information and the task content of each initial sub-task. The thought process is used to demonstrate the process of obtaining the preliminary decision. The implementation process of step S204 is explained in detail below.

[0044] Since user fault reports are verbal descriptions, after the first intelligent agent receives these reports, its intent perception module standardizes them. This standardization process involves accurately extracting key information from the user's fault report to obtain standardized user fault reports. Standardized user fault reports include, but are not limited to: fault symptoms (e.g., network outages, slow network speeds, frequent disconnections, abnormal device connections); specific user needs (e.g., requests for quick repairs, understanding the cause of the fault, scheduling on-site service); and preliminary contextual information (any additional details provided by the user during the fault report, such as attempted troubleshooting steps and device status descriptions). Next, the data perception module within the first intelligent agent acquires multi-dimensional data, including user device information, historical broadband maintenance data, operator backend data, and network status data (specifically, real-time network status diagnostic data).

[0045] After collecting sufficient data, multi-dimensional data and standardized user fault reports are identified as input data, which is then processed by the decision processing module within the first intelligent agent. The core of this module is to leverage the broadband maintenance strategy prediction model of the first intelligent agent (e.g., the domain LLM brain) to conduct deep analysis and reasoning on the input data: First, the standardized user fault reports are deeply analyzed to identify key fault descriptions, such as "cannot connect to the internet," "weak signal," and "frequent disconnections." Natural Language Processing (NLP) technology is used to analyze the user's actual needs and the possible causes behind the fault. For example, if a user mentions "cannot connect to the internet, optical modem is showing a red light," the domain LLM of the first intelligent agent will identify that the user is primarily experiencing broadband connection interruptions, and the source of the fault is likely the optical modem. The next step is to analyze user equipment information, such as model, brand, and version number, to determine whether the problem lies with the optical modem, router, or terminal device. Network status diagnostic tools are used to analyze network status data, including signal strength, packet loss rate, latency, and other network indicators, to identify potential network-level obstacles. Historical fault records with similar descriptions are searched from historical broadband maintenance data, and their processing procedures and results are extracted to provide a reference for the current fault. Based on effective solutions from historical fault records, combined with current network status and equipment information, possible fault points and optimal response strategies are inferred. For example, the combination of "optical modem red light + optical attenuation exceeding -28dB" is often diagnosed as an optical line problem in past cases and can be used as part of the current scenario analysis. The domain LLM brain constructs a thought chain based on the above analysis. The thought chain explains the reasoning process behind generating the initial decision. According to the thought chain, a series of initial sub-tasks are determined as the initial decision. The initial decision includes multiple initial sub-tasks to resolve the fault indicated by the user's fault report, as well as the specific task content corresponding to each initial sub-task. Each initial sub-task has specific goals and operational guidelines to ensure the feasibility and relevance of the solution. The thought process (thought chain) including the generation of the initial decision and the initial decision constitute the first broadband maintenance strategy for the fault indicated by the user's fault report.

[0046] After generating the first broadband maintenance strategy, to verify its accuracy, step S206 is performed, where a second intelligent agent conducts a first-round evaluation of the first broadband maintenance strategy, obtaining the corresponding evaluation score. In this step, the second intelligent agent (the verifyer agent) conducts a comprehensive and multi-dimensional evaluation of the preliminary decision generated by the engineer agent to ensure the real-time performance, reliability, and accuracy of the decision.

[0047] In the technical solution provided in step S206, the multiple evaluation methods of the first round of evaluation include task decomposition evaluation and effectiveness evaluation. Task decomposition evaluation is used at least to examine the quality of the sub-tasks decomposed in the first broadband maintenance strategy. Effectiveness evaluation is used at least to examine the correlation between the first broadband maintenance strategy and the effective broadband maintenance decisions of historical target application scenarios. Historical target application scenarios are scenarios whose similarity index with the application scenario corresponding to the user fault reporting information is less than a preset similarity threshold.

[0048] In the technical solution provided in step S206, the second intelligent agent performs a first round of evaluation on the first broadband maintenance strategy. There are several ways to obtain the evaluation score corresponding to the first round of evaluation. For example, the second intelligent agent can evaluate the first broadband maintenance strategy using task decomposition evaluation to determine a first initial score, a second initial score, a third initial score, a fourth initial score, and a fifth initial score. The first initial score is used at least to quantify whether the multiple initial sub-tasks of the preliminary decision cover all aspects of the fault problem; the second initial score is used at least to quantify the rationality of the dependencies and execution order between the multiple initial sub-tasks; the third initial score is used at least to quantify the rationality of the task granularity of the multiple initial sub-tasks; and the fourth initial score... The scores are used to quantify the operational feasibility of multiple initial subtasks. The fifth initial score is used to quantify the suitability of the fault-handling tools indicated in the multiple initial subtasks to the fault problems. The set of the first, second, third, fourth, and fifth initial scores is determined as the evaluation score corresponding to the task decomposition evaluation method. The first broadband maintenance strategy is evaluated using an effectiveness evaluation by a second intelligent agent to determine the sixth initial score, which is used to determine the correlation between the first broadband maintenance strategy and effective broadband maintenance decisions in historical broadband maintenance data. The evaluation score corresponding to the first round of evaluation is determined based on the evaluation score corresponding to the task decomposition evaluation method and the sixth initial score. The following details the process of obtaining the evaluation score corresponding to the first round of evaluation by a second intelligent agent on the first broadband maintenance strategy.

[0049] First, a task decomposition evaluation is performed: This module aims to examine the quality of the task decomposition after the first agent generates the thought chain. It utilizes the domain LLM brain of the second agent, the examiner agent, to evaluate the task breakdown from multiple dimensions and return a task decomposition quality score. Evaluation dimensions include: completeness score. (Corresponding to the first initial score or first score mentioned above), logical rationality score (Corresponding to the second initial score or second score mentioned above), task decomposition granularity scoring (Corresponding to the third initial score or third score mentioned above), feasibility score (Corresponding to the fourth initial score or fourth score above), tool selection rationality score (Corresponding to the fifth initial score or the fifth score mentioned above).

[0050] The first initial score is used to quantify whether multiple subtasks cover all aspects of the fault problem. This first initial score can also be named a completeness score. Completeness aims to determine whether each initial subtask formed by the decomposition fully covers all objectives, requirements, and constraints of the original task (the fault problem proposed by the user), ensuring no key aspects are omitted. The completeness assessment uses a large-model-based decision evaluation method. The evaluation method for the generated task decomposition and completeness is as follows: the input data is analyzed through the target evaluation model (the domain LLM of the examiner agent) in the second agent to generate an initial reference decision (referring to the process of generating the first broadband maintenance strategy, which will not be elaborated here). This initial reference decision includes multiple initial reference subtasks. The number of times the preset text appears in the multiple initial subtasks and the number of times the preset text appears in the multiple initial reference subtasks are determined, and the first target score is determined based on the first and second target counts. For example, this can be expressed by the following formula:

[0051] After the inspector agent is trained on the integrity assessment (the assessment process for obtaining the first initial score) in the task decomposition evaluation, the task decomposition integrity assessment score generated during system (the system executing the broadband maintenance strategy determination method of the embodiments of this application) is... (That is, the first initial score mentioned above) can be expressed as:

[0052]

[0053] in, This indicates the number of times the preset text appears as the first target in multiple initial subtasks. This indicates the number of times the preset text appears as the second target in multiple initial reference subtasks.

[0054] When the aforementioned inspector agent is trained for integrity assessment, it is achieved through fine-tuning using Reinforcement Learning from Human Feedback (RLHF). During training, based on a large amount of historical input data, a set of tasks generated by the inspector (corresponding to initial reference subtasks) is produced. This set is then compared with a set of subtasks provided by experts corresponding to the historical input data (also known as the verification subtask set, which is the accurate set of tasks corresponding to the historical input data; the verification subtask set refers to a set of predefined, verified subtasks or operation steps). A loss function is calculated based on this loss function, thereby fine-tuning the inspector agent. The corresponding loss function is:

[0055]

[0056] in, This represents the loss function used for fine-tuning the large tester model (the domain LLM of the tester agent). The number of times a specific text (i.e., the preset text) appears in the subtask set (also known as the validation subtask set) provided by experts. This involves calculating the frequency of occurrences of a specific text (i.e., the preset text, located within the set of verification subtasks) in the set of subtasks generated by the verifier. Specifically, this can be achieved through fine-tuning the Reinforcement Learning from Human Feedback (RLHF) process.

[0057] The second initial score is used to quantify the rationality of the dependencies and execution order among multiple initial subtasks. In essence, it involves a logical evaluation of the multiple initial subtasks, examining whether the order and dependencies between them conform to logical rules and actual business processes. The implementation is as follows: A first initial directed acyclic graph (DAG) is constructed based on the multiple initial subtasks. Each node in the first DAG corresponds to an initial subtask, and the edges between nodes represent the execution order. The set of edges in the first DAG is determined as the initial task dependency set, and the second initial score is determined at least based on this set. After the verifier agent has completed training on the logical evaluation in the task decomposition assessment, the task decomposition logical evaluation score is generated during system runtime. (That is, the second initial fraction mentioned above) can be expressed as:

[0058]

[0059] in, , This represents the set of task dependencies generated by the engineer agent (i.e., the initial set of task dependencies mentioned above). The task dependency set generated by the tester agent (i.e., the initial reference task dependency set, which is constructed based on multiple initial reference subtasks to form an initial reference directed acyclic graph, where each node in the initial reference directed acyclic graph corresponds to an initial reference subtask, and the edges between nodes in the initial reference directed acyclic graph represent the execution order between nodes; the set of edges in the initial reference directed acyclic graph is determined as the initial reference task dependency set). represent and The intersection of these factors. The task dependency set generated by the verifier is used as the ground truth to evaluate the logicality of the tasks generated by the engineer agent. This indicates the precision corresponding to the logical evaluation. This represents the recall rate corresponding to logical evaluation.

[0060] When training the aforementioned tester agent for logical reasoning evaluation, a large-model-based decision evaluation method is used to assess logical reasoning. To quantify the logical consistency of the tester agent's subtask dependency predictions during RLHF fine-tuning, a logical reasoning loss function oriented towards a subtask graph structure is designed. This function models tasks as directed acyclic graphs (DAGs), treats subtasks as nodes, and the order of subtasks as edges between nodes. The DAG of the engineer agent's decision transformation is denoted as... Furthermore, a logistic loss function is used for calculation. This method uses the dependency edge set as the basic unit and quantifies the logistic bias of the subtask sequence by comparing the overlap between predicted and actual dependencies using the F1 score. The F1 score used for RLHF fine-tuning ( And the loss function corresponding to the training of the tester agent for logical evaluation. (i.e., the reward function) can be expressed as follows:

[0061]

[0062] in, The set of inter-task dependency outcomes predicted by the model (the goal evaluation model in the second agent) (i.e., the set of edges of the directed acyclic graph constructed based on the reference decisions predicted by the goal evaluation model in the second agent). and Different tasks in the reference decision, as predicted by the objective evaluation model in any two second agents; This is a set of accurate and true dependency results between tasks, used for verification. To and and The corresponding task; This represents the intersection of dependencies correctly predicted by the objective evaluation model in the second agent. This indicates the accuracy corresponding to the logical evaluation of the training process. This represents the recall rate corresponding to the logical evaluation of the training process.

[0063] The third initial score is used to quantify the reasonableness of the task granularity of multiple initial subtasks. The task decomposition granularity reasonableness assessment aims to evaluate whether the level of coarseness in the subtask breakdown is appropriate, avoiding both overly coarse granularity leading to difficulty in execution and overly fragmented granularity increasing management costs. The total number of initial subtasks is determined; the second agent predicts the total number of target subtasks for solving the fault problem based on domain knowledge (the total number of target subtasks is the reasonable number of subtasks corresponding to solving the fault problem (i.e., the optimal number of subtasks)); the third score is determined based on the total number of tasks and the total number of target subtasks. After the verifier agent has been trained for granularity reasonableness assessment, the task decomposition granularity reasonableness assessment score is generated during system runtime. (That is, the third initial fraction) can be expressed as:

[0064]

[0065] in, This represents the number of subtasks generated by the engineer agent (the first agent) (i.e., the total number of tasks corresponding to the multiple initial subtasks mentioned above). This represents the number of reasonable subtasks corresponding to the task determined by the inspector agent based on domain knowledge (i.e., the total number of the aforementioned target subtasks). The number of reasonable subtasks determined by the inspector agent (the second agent) is used as a benchmark to evaluate the reasonableness of the task decomposition granularity generated by the engineer agent.

[0066] In the above steps, when the second agent predicts the total number of target sub-tasks for solving the fault problem based on domain knowledge (the total number of target sub-tasks is the reasonable number of sub-tasks corresponding to solving the fault problem (i.e., the optimal number of sub-tasks)), the domain knowledge specifically refers to professional knowledge in the field of home broadband repair, covering the cause analysis of common broadband faults, standardized steps of repair processes, handling techniques for specific faults, and lessons learned from past repair cases. This knowledge is encoded and integrated into the domain LLM model (e.g., a Transformer-based model) of the tester agent. The domain LLM model in the second agent takes user fault reporting information as input and, guided by its internal domain knowledge, predicts the total number of target sub-tasks for solving the fault problem.

[0067] When training the aforementioned tester agent to evaluate the reasonableness of task decomposition granularity, a large-model-based decision evaluation method is used. To quantify the appropriateness of the subtask decomposition granularity during the RLHF fine-tuning process, a granularity reasonableness loss function oriented towards the number of subtasks is used. This function measures the reasonableness deviation of the task decomposition granularity by comparing the deviation between the actual number of generated subtasks and the optimal number of subtasks. The granularity reasonableness loss function corresponding to the tester agent's training for task decomposition granularity reasonableness evaluation is described below. It can be represented as follows:

[0068]

[0069] in, The actual number of subtasks generated for the model (i.e., the total number of subtasks generated by the first agent during training for the training set (multiple historical user fault reports)). The optimal number of subtasks (also known as the reasonable number of subtasks) corresponding to a task determined by historical data or domain knowledge. This evaluation method can also use a scoring system where a large tester model (a domain LLM for the tester agent) trained by human reviewers is used to score the clarity of description, the specificity of objectives, and the concreteness of operational steps for each subtask generated by the first agent during training. For example, a scoring criterion, such as 1-5 points, can be defined, and the average score calculated. The loss function in this case... as follows:

[0070]

[0071] in This represents the total number of tasks required to solve the problem (the total number of sub-tasks generated by the first agent during training for the training set (multiple historical user fault reports)). The sum of scores for all tasks graded by humans, here. Indicates the first i The scores from human evaluations of each task are used. Training the tester agent using human scoring would be more accurate, but data acquisition would be more complex. Therefore, the first choice is to use... The loss function fine-tuned in RHLF, which acts as the inspector agent.

[0072] The fourth initial score is used to quantify the operational feasibility of at least several initial subtasks. It is determined through a feasibility assessment, which aims to ensure that each initial subtask has a clear, actionable objective, a defined implementation method, and quantifiable metrics. The feasibility assessment employs a programmed decision-making method, for example, by: determining the total number of initial subtasks; determining the complexity score for each initial subtask; and determining the fourth initial score based on the total number of initial tasks and the complexity scores of all initial subtasks. The complexity score for each initial subtask is determined based on the reasonableness score of its operational steps, a preset weight corresponding to the reasonableness score, and a fifth initial score (using the decisions generated by the engineer agent as input, counting the number of steps and tools involved in each initial subtask, and calculating the complexity score). The reasonableness score of the operational steps is determined based on the actual number of operational steps of the initial subtask, the baseline number of operational steps corresponding to the task level of the initial subtask, and a preset range of operational step fluctuations. Therefore, the feasibility assessment score... (That is, the fourth initial fraction) can be expressed by the following formula:

[0073]

[0074] Among them, here It is the number of subtasks; It is a subtask (the initial subtask). The complexity score; A score is awarded for the reasonableness of the number of steps (i.e., the reasonableness score for the number of operation steps). Specifically, the fewer steps a task requires, the lower its feasibility. In specific scenarios, the reasonable range for generating a task's number of steps can be determined by the task's difficulty. The higher the difficulty, the lower the score for tasks with fewer steps can be, in order to increase the impact of high-step tasks on reflection. The weights (preset weights corresponding to the reasonableness score of the number of operation steps) are determined by... The settings are based on the importance of the operational steps and tools in the field. The selection of a reasonable score for the tool used in subsequent testing (i.e., the fifth initial score) will not be discussed in detail here. The method for determining the reasonableness score of the number of steps is described below:

[0075]

[0076] in, This represents the actual number of steps in the generated task (corresponding to the actual number of operation steps in the initial subtask). The task difficulty is categorized into 1-5 levels, with higher levels indicating greater difficulty. The specific task difficulty classification can be implemented using the user intent recognition module within the engineer's intelligent agent. The representative difficulty is The baseline reasonable number of steps (i.e., the baseline number of operation steps corresponding to the task level of the above subtasks) is calculated using the following formula: (in Based on the number of steps, The base number of steps added for each increase in difficulty level, for example (and so on) This indicates a reasonable range of step fluctuations (i.e., the aforementioned preset range of operation step fluctuations, for example...). This indicates that the baseline number of steps is allowed. (fluctuation within a range).

[0077] The fifth initial score is used to quantify the fit between the fault-handling tools indicated in multiple initial sub-tasks and the fault problems. It is determined through a tool selection rationality assessment, which checks whether the tools selected for each sub-task are suitable for its objectives, application scenarios, resource constraints, and execution requirements. For example, the fifth initial score can be determined as follows: Obtain the first target tool selection set in the first broadband maintenance strategy, where the first target tool selection set is the set of tools indicated in the task content of multiple initial sub-tasks in the first broadband maintenance strategy; obtain the second target tool selection set in the reference decision, where the second target tool selection set is the set of tools indicated in the task content of multiple initial reference sub-tasks in the initial reference decision, where the reference decision is obtained by analyzing the input data and real-time feedback data through a target evaluation model (domain LLM) in the second agent; the fifth initial score is determined at least based on the first target tool selection set and the second target tool selection set. The tool selection rationality assessment score is generated during system runtime after the tester agent has been trained. (That is, the fifth initial fraction) can be expressed as:

[0078]

[0079] in, Selecting the appropriate rationality assessment for the tool Fraction, , Indicates the precision of tool selection. Recall rate, representing tool selection This represents the set of tool selections generated by the engineer agent (corresponding to the first target tool selection set). The tool selection set generated by the inspector agent (corresponding to the second target tool selection set). This represents the intersection of two tool selection sets. The reasonableness of the tool selections generated by the engineer agent is evaluated using the task dependency set generated by the verifier as the true value.

[0080] When training the second agent's ability to evaluate the rationality of tool selection, a large model-based evaluation decision-making method is adopted. In the RLHF fine-tuning stage, label pairs defining the scope of tools are extracted from the returned decisions of the examiner agent's domain LLM brain. <tool>< / tool> The tool information included is used in supervised learning through a list of tools used by industry experts, presented as labeled data. Similar to the evaluation of logicality, the F1 score can be used to measure the bias between the tool selection of the examiner's LLM brain and that of the expert. Therefore, the loss function used to construct the F1 score and for fine-tuning the RLHF is... The relevant set in can be represented as follows: The toolset of results needed for predictions by the model (domain LLM of the test agent agent). Representative tools; This is the actual set of tools and results needed for practical use. The intersection of dependencies that represent correct predictions.

[0081] The set of the first initial score, the second initial score, the third initial score, the fourth initial score, and the fifth initial score is determined as the evaluation score corresponding to the task decomposition evaluation method.

[0082] The effectiveness of the first broadband maintenance strategy is evaluated by a second intelligent agent to determine the sixth initial score. This involves: acquiring historical log data, which contains multiple log records corresponding to all historical target application scenarios. Each log record includes at least a problem description, an effective broadband maintenance decision, and the corresponding processing flow record. A second directed acyclic graph (DAG) set is constructed based on the historical log data. This DAG set contains multiple second DAGs, with each log record in the historical log data corresponding to one second DAG. Nodes in the second DAGs represent events or operation steps recorded in the log records, and edges represent temporal and dependency relationships between nodes. Correlation analysis is performed between the first and second initial DAGs to obtain the sixth initial score. The process of calculating the sixth initial score is detailed below.

[0083] The significance of effectiveness evaluation lies in identifying potential low-quality, ineffective, or even erroneous tendencies in current decision-making through backtracking and correlation analysis of historical data. This provides a basis for optimizing decisions and ultimately ensures that the output decisions are more effective and better suited to real-world needs. This module will use a programmatic evaluation method. The evaluation of decisions generated by the engineer agent in this module consists of three main steps.

[0084] Step 1: Relevant Record Search: By using the keyword decomposition method in the real-time feedback assessment module to search the execution records for user-reported issues, the search can retrieve execution record information corresponding to all historical target application scenarios. The information structure of the execution record is {Location, Problem Description, Agent Processing Flow (including thought process and decision-making process), User Feedback, On-site Engineer Feedback, and Customer Service Intervention Records}. Further, irrelevant information such as location and problem description is removed, and the collected record information set is denoted as... (i.e., historical log data).

[0085] Step 2: Record Graphification: Construct a second set of directed acyclic graphs based on historical log data. The second set of directed acyclic graphs contains multiple second directed acyclic graphs, and each data record in the historical log data corresponds to a second directed acyclic graph. Each data record This generates a directed acyclic graph. ,in For a set of nodes, Let be the edge set. Finally, the set of all records in the directed acyclic graph is denoted as . Among them, for the node set Any data record Extract key sub-tasks (operation steps) or events (such as "fault detection", "parameter adjustment", "feedback verification", etc.) from the intelligent agent processing flow (processing flow record), user feedback, and field engineer feedback fields, and use them as nodes in the corresponding second directed acyclic graph. ,in (i takes values ​​from 1 to k, where k represents the corresponding data record) The total number of events or operation steps represents the specific processing steps (operation steps) or events involved in the record. For edge sets... Based on the temporal and dependency relationships of the processing flow in record r, define directed edges between nodes: if the operation steps It occurred in the record Previously and for The prerequisite is to establish a directed edge. ,express Depends on If the record r explicitly mentions the logical dependencies between events (e.g., "feedback from the field engineer" needs to be based on "preliminary diagnostic results of the agent"), then the corresponding directed edge is directly established.

[0086] Steps three and four together constitute the process of correlation analysis: Step three, correlation calculation: In the logical evaluation of the task decomposition and evaluation module, the decision generated by the engineer agent has already been graphed (i.e., the first initial directed acyclic graph). Therefore, the decision directed acyclic graph (first initial directed acyclic graph) can be... and all records of the directed acyclic subgraph Perform correlation analysis. and The correlation can be expressed using the F1 score ( )express:

[0087]

[0088] in, Generate a set of dependency results between tasks in the decision-making process (i.e., the set of edges in the first directed acyclic graph) for the engineer agent (the first agent). Represents a subtask; This represents the set of task node dependencies after the record graph is visualized (i.e., the set of edges of the second directed acyclic graph in the second directed acyclic graph set). The tasks representing decision-making and recording depend on each other. Therefore, multiple F1 scores can be obtained through calculation. A set consisting of ) .

[0089] Step 4, Validity Prediction: From The highest F1 score was selected. The corresponding historical log data records Furthermore, the directed acyclic subgraph corresponding to this record is defined as follows: And then segmented, finally More subgraphs can be decomposed. The segmentation is based on the subgraph. The task nodes (graph nodes) involve four state nodes: "Decision Regeneration," "Decision Invalid," "Customer Service Intervention," and "Decision Valid." Each subgraph represents a specific stage or result of the processing flow. The next step will be... and Repeat the correlation calculation in step three for multiple subgraphs, and select the subgraph corresponding to the highest F1 score. Different scores can be pre-set for the four types of state nodes. The F1 score is then multiplied by the score corresponding to the tail node to obtain the final score output by this module to the decision module. (i.e., the sixth initial score mentioned above). In some embodiments of this application, validity prediction can also serve as a bypass. In this case, different scores are no longer pre-set for the four state nodes. If the end node of the subgraph is "Decision Valid", the decision generated by the engineer agent (first broadband maintenance strategy) can directly ignore the evaluation results of other modules and be immediately provided to the user. If the end node is "Decision Regeneration" or "Decision Invalid", the decision generated by the engineer cannot directly ignore the evaluation results of other modules and must be policy generated. If the end node is "Customer Service Intervention", the customer service intervention situation can be used as reflection content and provided to the engineer agent for decision regeneration. It should be noted that if If the highest F1 score is too low, the process cannot proceed to step four. This is because the validity prediction acts as a bypass, largely determining whether a decision can be passed immediately. To ensure the confidence level of the bypass effect, the F1 score of this module must be at least higher than the threshold (0.6) to proceed to step four.

[0090] Finally, the evaluation score for the first round of evaluation is determined based on the evaluation score corresponding to the task decomposition evaluation method and the sixth initial score: the average of multiple scores in the evaluation score corresponding to the task decomposition evaluation method is determined as the first average, and the weighted average of the first average and the sixth initial score (or the average can be taken directly without weighting) is determined as the evaluation score for the first round of evaluation.

[0091] If the evaluation score in the first round of evaluation is less than or equal to a preset confidence threshold (e.g., 80), the first intelligent agent adjusts the first broadband maintenance strategy a preset number of times (the number of times the decision needs to be regenerated after each adjustment). Add one, then we have If the evaluation score of the first broadband maintenance strategy after each adjustment is less than or equal to a preset information threshold (i.e., the number of decision generation times exceeds the corresponding threshold), a customer service intervention request is initiated. Each time the first intelligent agent adjusts the first broadband maintenance strategy, it uses Domain LLM based on the evaluation results and detailed reflection information to adjust the broadband maintenance strategy, generating a new first broadband maintenance strategy. The evaluation results include the results of the entire evaluation process during the first round of evaluation (task decomposition completeness, logic, granularity rationality, feasibility, and tool selection rationality, etc.). Detailed reflection information consists of further refined and targeted improvement suggestions from the examiner intelligent agent based on the above evaluation results. The reflection information not only points out problems in the first broadband maintenance strategy but also includes specific guidance on how to adjust the strategy, such as supplementing missing steps, clarifying ambiguous operation descriptions, and selecting more suitable tools, thereby supporting the engineer intelligent agent in decision optimization. If, in any of the preset number of adjustments, the evaluation score of the first broadband maintenance strategy is greater than the preset information threshold, then proceed to step S208.

[0092] Step S208: If the evaluation score corresponding to the first round of evaluation is greater than the preset threshold, the first broadband maintenance strategy is determined as the target broadband maintenance strategy and pushed to the user terminal.

[0093] In the technical solution provided in step S208, the first broadband maintenance strategy is determined as the target broadband maintenance strategy and pushed to the user terminal. The user will operate according to the instructions of the target broadband maintenance strategy and provide feedback after the operation is completed. If the user feedback indicates that the broadband fault has been resolved, the decision in the target broadband maintenance strategy can correctly guide the fault location and repair, and the process terminates. If the user feedback indicates that the broadband fault has not been resolved, that is, if the target broadband maintenance strategy has failed to resolve the fault, real-time feedback data after the user operates according to the target broadband maintenance strategy is obtained (the user operates according to the target broadband maintenance strategy by calling the engineer's intelligent agent to use the operation execution tool. The real-time feedback data includes the execution status of the operation execution tool on the first broadband maintenance strategy, changes in the fault phenomenon, newly emerging anomalies or other relevant information, and feedback from the on-site engineer (if involved), such as equipment inspection results, network status updates, etc.). The broadband maintenance strategy prediction model of the first intelligent agent regenerates the second broadband maintenance strategy based on at least the real-time feedback data. The second maintenance strategy includes a thought chain and The target decision-making process comprises multiple sub-tasks for problem-solving and the specific content of each sub-task: parsing real-time feedback data to extract key information and fault-related parameters. This information includes the specific results of user operations, changes in environmental factors (such as signal source location adjustments or device restart status), further descriptions of the fault, or other unstructured feedback (such as verbal descriptions from users). The next step is to integrate this information with the content of the first broadband maintenance strategy to form a comprehensive input data packet. Based on this input data packet, the fault description is reconstructed. Based on this reconstructed fault description, the thought chain reasoning process is updated to more accurately locate potential fault points, while also considering new information in the real-time feedback and environmental changes. Based on the updated thought chain reasoning process, a second broadband maintenance strategy is generated (i.e., re-triggering the data perception and decision generation process and performing a new round of verification to form an optimization closed loop). At this point, the number of failed fault location decisions is recorded. At the same time, the number of times the decision is regenerated (Reset). A second round of evaluation is conducted on the second broadband maintenance strategy using a second intelligent agent, yielding an evaluation score. This second round evaluation employs various methods, including task decomposition evaluation, real-time feedback evaluation, and effectiveness evaluation. Real-time feedback evaluation is used to verify the real-time response capability of the second broadband maintenance strategy. If the evaluation score in the second round exceeds a preset threshold, the second broadband maintenance strategy is designated as the new target broadband maintenance strategy and pushed to the user. If the user reports that the second broadband maintenance strategy fails to resolve the issue, the data perception and decision generation process is retried, and a new round of verification is performed, forming an optimization loop and determining the number of failed fault location attempts. Increment by 1 again until the number of failures in fault location is reached. If the broadband fault persists beyond the corresponding threshold, the inspector agent needs to request intervention from professional customer service to help locate the problem. The professional customer service agent will be provided with the following information: the engineer agent's thought process: a complete thought chain and reasoning process; and iterative decision records: decisions generated in each iteration, their evaluation results, and reflection information. This information is simultaneously provided to the professional customer service agent to assist in problem localization and resolution, avoiding starting from scratch and providing valuable expert experience data for subsequent LLM incremental training and fine-tuning.

[0094] If the evaluation score in the second round of evaluation is less than or equal to a preset threshold, the second broadband maintenance strategy is adjusted a preset number of times by the first intelligent agent (the number of times the decision needs to be regenerated after each adjustment). Add one, then we have If the evaluation score of the second broadband maintenance strategy after each adjustment is less than or equal to a preset threshold (i.e., the number of decision generation times exceeds the corresponding threshold), a customer service intervention request is initiated. Each time the first agent adjusts the second broadband maintenance strategy, domain LLM is used to adjust the strategy based on the evaluation results and detailed reflection information, generating a new second broadband maintenance strategy. The evaluation results include the results of the entire evaluation process during the second round of evaluation (task decomposition completeness, logic, granularity rationality, feasibility, and tool selection rationality, etc.). Detailed reflection information consists of further refined and targeted improvement suggestions from the examiner agent based on the above evaluation results. The reflection information not only points out problems in the second broadband maintenance strategy but also includes specific guidance on how to adjust the strategy, such as supplementing missing steps, clarifying ambiguous operation descriptions, and selecting more suitable tools, thereby supporting the engineer agent in decision optimization. If, in any of the preset number of adjustments, the evaluation score of the second broadband maintenance strategy is greater than the preset threshold, then the second broadband maintenance strategy is determined as the new target broadband maintenance strategy and pushed to the user end.

[0095] There are several ways to obtain the evaluation score for the second round of evaluation of the second broadband maintenance strategy by a second intelligent agent. For example, the second intelligent agent can evaluate the second broadband maintenance strategy by task decomposition to determine a first score, a second score, a third score, a fourth score, and a fifth score. The first score is used to quantify whether multiple sub-tasks cover all aspects of the fault problem; the second score is used to quantify the rationality of the dependencies and execution order between multiple sub-tasks (multiple sub-tasks in the second broadband maintenance strategy); the third score is used to quantify the rationality of the task granularity of multiple sub-tasks; the fourth score is used to quantify the operational feasibility of multiple sub-tasks; and the fifth score is used to quantify the appropriate handling of the indicated actions in multiple sub-tasks. The suitability of fault management tools to fault problems is assessed. The set of first, second, third, fourth, and fifth scores is determined as the evaluation score corresponding to the task decomposition evaluation method. A sixth score is determined by evaluating the second broadband maintenance strategy using real-time feedback evaluation through a second intelligent agent. This sixth score quantifies whether the second broadband maintenance strategy responds to real-time user feedback data. A seventh score is determined by evaluating the second broadband maintenance strategy using effectiveness evaluation through a second intelligent agent. This seventh score determines the correlation between the second broadband maintenance strategy and effective broadband maintenance decisions in historical broadband maintenance data. The evaluation score for the second round of evaluation is determined based on the evaluation score corresponding to the task decomposition evaluation method, as well as the sixth and seventh scores. Both the second and first broadband maintenance strategies are generated by the first intelligent agent. The difference is that the first broadband maintenance strategy is the first maintenance strategy generated specifically for broadband faults, while the second broadband maintenance strategy is a strategy generated again after a broadband maintenance strategy has already been provided to the user once, and the user's feedback on the problem was not resolved. Therefore, the method for evaluating the first broadband maintenance strategy in step S206 is also applicable to evaluating the second broadband maintenance strategy. Besides the task decomposition evaluation and effectiveness evaluation performed in step S206 (the specific implementation methods for task decomposition evaluation and effectiveness evaluation are described in step S206 and will not be repeated here), real-time feedback evaluation is required for both the second broadband maintenance strategy and any subsequently generated broadband maintenance strategies (excluding the first broadband maintenance strategy generated initially). That is, three evaluations are required: task decomposition evaluation, real-time feedback evaluation, and effectiveness evaluation. In this way, if the user feedback issue remains unresolved, the method in this embodiment re-acquires the latest feedback data and generates a new broadband maintenance strategy. Furthermore, real-time feedback evaluation is introduced during the evaluation process to dynamically optimize the broadband maintenance strategy based on the actual scenario, thereby improving the accuracy and effectiveness of the generated broadband maintenance strategy.

[0096] For the second broadband maintenance strategy, the first score is determined as follows: The input data and real-time feedback data are analyzed by the target evaluation model in the second intelligent agent to generate a reference decision, which includes multiple reference sub-tasks. The first occurrence count of the preset text in the multiple sub-tasks and the second occurrence count of the preset text in the multiple reference sub-tasks are determined, and the first score is determined based on the first occurrence count and the second occurrence count. The method for determining the first score is the same as the method for determining the first initial score in step S206 above, and the formula used is also the same (adaptive substitution is sufficient; the initial reference sub-task is replaced with a reference sub-task, and the initial sub-task is replaced with a sub-task), which will not be elaborated further here.

[0097] The second score is determined as follows: A first directed acyclic graph (DAG) is constructed based on multiple subtasks, where each node in the first DAG corresponds to a subtask, and the edges between nodes in the first DAG represent the execution order of the nodes. The set of edges in the first DAG is determined as the task dependency set, and the second score is determined at least based on the task dependency set. A reference DAG is constructed based on multiple reference subtasks, where each node in the reference DAG corresponds to a reference subtask, and the edges between nodes in the reference DAG represent the execution order of the nodes. The set of edges in the reference DAG is determined as the reference task dependency set, and the second score is determined based on the reference task dependency set and the task dependency set. The method for determining the second score is the same as the method for determining the second initial score in step S206 above, and the formula used is also the same (adaptive substitution is sufficient; the initial task dependency set is replaced with the task dependency set, and the initial reference task dependency set is replaced with the reference task dependency set), which will not be elaborated further here.

[0098] The third score is determined as follows: the total number of tasks in the multiple sub-tasks is determined; the total number of target sub-tasks for solving the fault problem is predicted by the second agent based on domain knowledge; and the third score is determined based on the total number of tasks and the total number of target sub-tasks. The method for determining the third score is the same as the method for determining the third initial score in step S206 above, and the formula used is also the same (adaptive substitution is sufficient, replacing the total number of tasks in the multiple initial sub-tasks with the total number of tasks in the multiple sub-tasks), which will not be repeated here.

[0099] The fourth score is determined as follows: The total number of subtasks is determined; the complexity score for each subtask is determined, and the fourth score is determined based on the total number of tasks and the complexity scores of all subtasks. The complexity score for each subtask is determined based on the reasonableness score of its operation steps, the preset weight corresponding to the reasonableness score, and the fifth score. The reasonableness score of the operation steps is determined based on the actual number of operation steps of the subtask, the baseline number of operation steps corresponding to the task level to which the subtask belongs, and the preset range of operation step fluctuations. The method for determining the fourth score is the same as that for determining the fourth initial score in step S206 above, and the formula used is also the same (adaptive substitution is sufficient), so it will not be repeated here.

[0100] The fifth score is determined as follows: First, a first tool selection set is obtained from the second broadband maintenance strategy, where the first tool selection set is the set of tools indicated in the task content of multiple sub-tasks within the second broadband maintenance strategy; second, a second tool selection set is obtained from the reference decision, where the second tool selection set is the set of tools indicated in the task content of multiple reference sub-tasks within the reference decision, wherein the reference decision is obtained by analyzing input data and real-time feedback data through a target evaluation model in the second agent; the fifth score is determined based at least on the first and second tool selection sets. The method for determining the fifth score is the same as the method for determining the fifth initial score in step S206 above, and the formula used is also the same (adaptive substitution is sufficient, replacing the first target tool selection set with the first tool selection set, and the second target tool selection set with the second tool selection set), which will not be elaborated further here.

[0101] There are several ways to evaluate the effectiveness of the second broadband maintenance strategy using a second intelligent agent to obtain the seventh score. For example: acquiring historical log data, which contains multiple log records corresponding to all historical target application scenarios, with each log record including at least a problem description, an effective broadband maintenance decision, and the corresponding processing flow record; constructing a second directed acyclic graph (DAG) set based on the historical log data, where the DAG set contains multiple DAGs, with each log record in the historical log data corresponding to one DAG, nodes in the DAG representing events or operation steps recorded in the log records, and edges representing temporal and dependency relationships between nodes; and performing correlation analysis based on the first and second DAGs to obtain the seventh score. The process of determining the seventh score by evaluating the second broadband maintenance strategy through the second intelligent agent is the same as the method of determining the sixth initial score by evaluating the effectiveness in step S206 above, and the formula used is also the same (adaptive substitution is sufficient, for example, the first initial directed acyclic graph is replaced with the first directed acyclic graph), which will not be elaborated here.

[0102] The second broadband maintenance strategy is evaluated using real-time feedback by a second intelligent agent to obtain a sixth score. This process includes: extracting keywords from the real-time feedback data to obtain a first keyword set; and determining the sixth score based on the first keyword set, the first broadband maintenance strategy, and the second broadband maintenance strategy. The following details the process of obtaining the sixth score by evaluating the second broadband maintenance strategy using real-time feedback by a second intelligent agent.

[0103] Real-time feedback evaluation is used to verify the real-time responsiveness of the second broadband maintenance strategy. The real-time feedback evaluation module aggregates immediate feedback from users and on-site engineers to validate the decisions (second broadband maintenance strategy) generated by the engineer's agent. Feedback and decisions are generated synchronously and can be directly used as reference text to calculate the correlation between decisions and feedback items, thus enabling the use of a programmatic evaluation method. Therefore, the real-time feedback evaluation module is designed using a programmatic decision-making method. This module divides the feedback evaluation into three steps: feedback item keyword extraction, keyword matching degree calculation, and real-time feedback coverage accuracy score calculation.

[0104] Step 1: Keyword Extraction for Feedback Items: The goal of this step is to extract keywords for each user feedback item in the real-time feedback data. Example feedback: "I have checked my home optical modem; it is on, and the green light is constantly on." After intent recognition, the domain-specific large model of the second agent directly outputs {user, optical modem, power, green}. Since the general large model performs well in keyword extraction, and domain training does not weaken this ability, we utilize the LLM brain's capabilities and leverage prompt word engineering to have it return keyword information in JSON format ({"Keyword list (KeyWords):[...]}, where {"KeyWords":[...]} is a JSON data structure representing a set containing keywords) to extract keywords.

[0105] Step 2: Keyword Match Ratio Calculation: This step uses the keywords extracted in Step 1 to calculate the matching rate between different keyword entries and the decisions generated by the engineer agent (first agent). The Keyword Match Ratio (KMR) can be expressed as follows:

[0106]

[0107] in, Representing the The keyword set of any one feedback item; Representing the The first of the feedback items Any one keyword; The decision that the engineer agent regenerates differs from the old decision. The first strategy item (a strategy item is an operational step in decision-making, i.e., a readjustment step, the first...) (A single item can be any distinguishable strategy item). The set of policy entries that differ from the old decisions in the regenerated decisions of the engineer agent (i.e., the set of operational steps that differ between two consecutive broadband maintenance policies). The set of decisions regenerated by the engineer agent (corresponding to the target decision in the second broadband maintenance strategy). The set of decisions generated by the engineer agent in the previous round for the unsuccessful location of the fault (corresponding to the preliminary decisions of the second broadband maintenance strategy).

[0108] Step 3: Real-time feedback coverage and accuracy score calculation: This step is obtained from Step 2. Coverage and accuracy are calculated to further calculate the F1 score for real-time feedback. The expression for calculating Feedback Coverage (FC) is as follows:

[0109]

[0110] The formula for calculating Feedback Revision Accuracy (FRA) is as follows:

[0111]

[0112] in Represents coverage. The decision items representing the modifications covered a set of keyword information from the feedback; It is the set of decision modification items made by the engineer intelligent agent, and The meaning is the same, and it can also be used. To represent. Real-time rating. (That is, the sixth score mentioned above) can be represented by the Real-time Feedback Score (RFS):

[0113]

[0114] When determining the evaluation score for the second round of evaluation based on the evaluation score corresponding to the task decomposition evaluation method, as well as the sixth and seventh scores, first calculate the average of the first to fifth scores in the evaluation score corresponding to the task decomposition evaluation method, and determine this average as the task decomposition evaluation score. Then, calculate the weighted average of the task decomposition evaluation score, the sixth score, and the seventh score (or simply take the average) as the evaluation score for the second round of evaluation.

[0115] Example 1: Taking a user's fault report "Broadband cannot connect to the internet, optical modem indicator light is red" as an example, the user reports the fault and intent perception: The user reports the fault through the APP, "Broadband cannot connect to the internet, optical modem indicator light is red." The engineer's intelligent agent's intent perception module extracts key information: the fault phenomenon is "internet outage," the device status is "optical modem red light," and the initial judgment is an optical path fault. Data perception: The data perception module calls the toolset: the network status diagnostic tool detects an optical attenuation value of -32dB (normal value <-28dB); the operator's backend interface obtains that the splitter port status is "interrupted"; the memory module retrieves the user's historical records: there have been no similar faults in the past 3 months, and the device model is model A. Decision generation: The engineer's intelligent agent's LLM brain, based on thought chain reasoning, generates a preliminary decision including the following operation steps: "1. On-site engineer checks the fiber optic connector; 2. If the connector is normal, check the splitter port; 3. Replace the faulty port and restart the optical modem." Decision Evaluation: Task Decomposition Evaluation: Completeness score 90, Logicality 95, Granularity Reasonableness 85, Feasibility 90, Tool Selection Reasonableness 95, Overall score 88. Effectiveness Evaluation: The effective decision graph for the "Optical Modem Red Light + Excessive Optical Attenuation" scenario in historical records has an F1 score of 0.92 compared to the current decision graph, matching the "Port Failure" effective mode, with an effectiveness score of 92. The decision verification module sets a threshold of 80 points. The overall score (88+92) / 2=90 points > 80 points, so the decision is approved. The system pushes the decision to the on-site engineer, who replaces the splitter port, and the fault is repaired.

[0116] Example 2: User reports fault: "The WiFi signal in my study is extremely weak, and my laptop frequently disconnects after connecting to the internet, while the signal in the living room is normal." The engineer's intelligent agent's intent perception module extracts key information: the fault symptom is "weak WiFi signal in the study, frequent disconnections from the laptop," and the environmental characteristic is "normal signal in the living room (signal difference compared to the study)." The initial judgment is uneven WiFi signal coverage or channel interference. Initial decision generation and evaluation (failed): The engineer's intelligent agent generates a decision, intent perception: extracting key information—the fault symptom is "weak WiFi signal in the study, disconnections," and the environmental difference is "normal signal in the living room." Data perception: calling tools to obtain data: router model; historical records: no similar faults in the past month, router placement is on the TV cabinet in the living room; remote detection: 2.4GHz signal strength in the study -78dBm (weak), 5GHz signal -85dBm (extremely weak); 2.4GHz signal in the living room -55dBm (normal). Decision output: The engineer's intelligent agent's LLM brain generates the initial decision: "1. Move the router from the TV cabinet in the living room to the center of the living room; 2. Restart the router and observe the signal in the study." The inspector's intelligent agent evaluation (failed) task decomposition evaluation module: completeness score 60, logic score 70, granularity reasonableness score 80, feasibility score 75, tool selection reasonableness score 70. Effectiveness evaluation module: score 80. Total module score: (60+70+80+75+70+80) / 6=72.5 points (threshold set at 70 points), evaluation passed. 3) User feedback and initial regeneration decision (still failed) User feedback: After operating according to the initial decision, the reply was: "The router has been moved to the center of the living room, but after restarting, the signal in the study is still very weak, and the mobile phone cannot even find 5GHz." Engineer's intelligent agent regeneration decision (first regeneration)

[0117] Based on user feedback and reflection, the LLM brain adjusted its decision: "1. Open the router management app, check the 2.4GHz channel occupancy rate, and switch to a less interfered channel (such as 1, 6, or 11); 2. Adjust the router antenna angle to 45°; 3. Recommend that devices in the study prioritize connecting to the 2.4GHz band (strong penetration)." The second evaluation of the intelligent agent (still failed) task decomposition evaluation module: completeness 80 points, logic 85 points, granularity reasonableness 75 points, feasibility 70 points, tool selection reasonableness 70 points. Total module score: (80+85+75+70+70) / 5=76 points. Real-time feedback evaluation module: Extracted user feedback keywords "cannot find 5GHz" and "weak signal," with a KMR matching rate of 1.0 with the decision "switch to 2.4GHz," coverage FC=1.0, correction accuracy FRA=0.8, and a real-time score of 88 points. Effectiveness Assessment Module: The success rate of "channel + frequency band optimization" in resolving signal differences in the historical records is 65% (F1 score 0.65), with a score of 65. Decision Verification Module: The comprehensive score (76+88+65) / 3≈76 points (the threshold is still 70 points, but the verification module detected "vague description of feasible steps", triggering additional verification), and the assessment failed. Reflection Information Output: "The specific name of the 'router management APP' and the channel switching operation steps are not clearly defined. It is recommended to supplement the tool guide for model adaptation." 4) Secondary Regeneration Decision and Assessment (Passed) The engineer's intelligent agent generates accurate decisions by combining the latest reflection information: "1. Download the router management APP, log in and go to 'WiFi settings' → 'channel selection', check the 2.4GHz channel occupancy rate, and prioritize channel 1 or 11; 2. Adjust the two antennas of the router to vertical (90°) and horizontal (0°) respectively; 3. Connect the study room device to the WiFi with the name A, and avoid using the 5GHz band." The inspector agent underwent three evaluations (passed): Task decomposition evaluation module: all dimensions scored ≥85 points, fully covering the fault points, with specific and operable steps, and tool selection matching the device model. Real-time feedback evaluation module: the decision fully responded to the user's feedback of "5GHz cannot be found", KMR=1.0, and the real-time score was 95 points. Effectiveness evaluation module: matched historical successful cases (success rate 92%), with a score of 92 points. Decision verification module: the comprehensive score (85+95+92) / 3≈91 points>70 points, and the evaluation passed. 5) Fault resolution: After the user performed the second regeneration decision operation, the feedback was: "No more disconnections, problem solved", and the process terminated.

[0118] The method in this application embodiment involves a reflective interactive design process between the tester agent and the engineer agent. It provides a closed-loop optimization of the "fault response - decision generation - decision evaluation - decision regeneration" process for home broadband solutions, addressing practical problems in home broadband maintenance such as inaccurate fault location, high user operation failure rates, repeated on-site engineer visits, and resource waste caused by a lack of real-time feedback and effective verification mechanisms. It also provides a multi-dimensional evaluation capability for the quality of the engineer agent's decision-making task decomposition (from the perspectives of completeness, logic, granularity, feasibility, and tool selection rationality), resolving issues such as incomplete decision-making task decomposition, logical confusion, inappropriate step detail, infeasible operations, or poor tool adaptability, leading to broken fault handling processes, low execution efficiency, and resource mismatch.

[0119] Figure 3 This is a schematic diagram of the architecture of a broadband maintenance strategy determination system according to an embodiment of this application. Figure 3 The broadband maintenance strategy determination system shown is used to execute the broadband maintenance strategy determination method shown in steps S202-S208 of the embodiments of this application. Figure 3 The system architecture shown involves four entities: the engineer agent (first agent), the inspector agent (second agent), the memory module, and the toolset. Both the engineer agent (first agent) and the inspector agent (second agent) are based on Domain LLM. It's important to note that the relationship between the Domain LLM brain and the agents can be twofold: multiple agents can share a single Domain LLM brain, or multiple agents can independently own and maintain their own Domain LLM brain. The engineer agent's agent application contains four modules: the awareness perception module, the data perception module, the decision processing module, and the LLM brain interaction module. The functions of each module are as follows: Intent Perception Module: Extracts key information based on user feedback, including fault symptoms and specific user needs. Data Perception Module: Collects user device information, historical data, and operator backend data to support subsequent decision-making. Decision Processing Module: Based on the collected data and information, performs in-depth analysis and reasoning to generate a thought process and provides a preliminary decision. LLM Brain Interaction Module: Used to interact with LLM (Shared LLM), defining interaction behaviors with domain LLM, prompt templates for different interaction behaviors, return formats, and return data parser.

[0120] The agent agent application for the inspector agent comprises six modules: Task Decomposition Evaluation Module, Real-time Feedback Evaluation Module, Effectiveness Evaluation Module, Data Awareness Module (shared with the engineer agent; the two agents share data through the Data Awareness Module), LLM Brain Interaction Module, and Verification Decision Module. The functions of each module are as follows: Task Decomposition Evaluation Module: Evaluates the completeness and depth of the engineer agent's thinking and decision-making framework from the perspectives of completeness, logic, and granularity. Real-time Feedback Evaluation Module: Composed of a user feedback response evaluation module and a field engineer feedback evaluation module. Used to evaluate whether the engineer agent's thinking and decision-making results meet real-time requirements. User Feedback Response Evaluation Module: Evaluates whether the engineer agent has made sufficient considerations and decisions based on real-time feedback provided by users. Field Engineer Feedback Evaluation Module: Evaluates whether the engineer agent has made sufficient considerations and decisions based on real-time feedback provided by field engineers. Effectiveness Evaluation Module: Composed of a historical execution record evaluation module and a professional customer service record evaluation module. Historical Execution Record Evaluation Module: By retrieving historical execution records highly relevant to the target scenario from short-term and long-term memory, this module determines whether decisions are related to effective or ineffective decisions in similar historical scenarios, thus identifying low-quality and ineffective decisions. Professional Customer Service Record Evaluation Module: By retrieving professional customer service intervention records highly relevant to the current scenario from long-term memory, this module determines whether decisions are related to effective or ineffective decisions in similar historical customer service intervention scenarios, thus identifying low-quality and ineffective decisions. Decision Validation Module: This module summarizes and validates the scores from each evaluation module, ultimately determining whether the engineer agent's decisions pass the validation of each module. LLM Brain Interaction Module: This module interacts with the LLM, defining the domain LLM interaction behaviors with the agent, prompt templates for different interaction behaviors, return formats, and a return data parser.

[0121] The decision verification module integrates scores from three evaluation modules (task decomposition evaluation, real-time feedback evaluation, and effectiveness evaluation) and dynamically adjusts the verification threshold based on the problem difficulty and the number of times the decision needs to be reconsidered. The specific process is as follows: Score aggregation and verification; Verification failure evaluation: If the score is below the threshold, the system will feed back the evaluation results and detailed reflection information to the engineer agent. The latter will use its LLM (Limited Learning Model) for self-correction and generate a new decision. At this point, the decision generation count needs to be incremented by one. Decision validation: Decisions that pass rigorous evaluation are analyzed, and the fault location results and solutions are ultimately returned to the user, ensuring decision quality. Validated decisions are implemented by the engineer agent using the operation execution tool. After the decision is executed, the user judges whether the decision successfully resolved the fault. Based on user feedback data, the system takes different measures: Fault resolved: If the decision correctly guides fault location and repair, the process terminates. Fault not resolved: If the decision fails to resolve the problem, the system collects the latest feedback, re-triggers the data perception and decision generation process, and performs a new round of validation, forming an optimization loop. At this point, At the same time Reset to zero. The decision verification module can dynamically set the verification score threshold based on factors such as the difficulty of the user's problem and the number of times the decision has been reconsidered, thereby dynamically adjusting the confidence level of the decisions generated by the engineer agent in different problem scenarios.

[0122] In some embodiments of this application, a crucial aspect of the verification decision module is the total verification score threshold (i.e., the aforementioned preset confidence threshold). ,in It is a function for calculating threshold scores. It is the total number of times the engineer agent regenerates decisions. This is a user-submitted question. It's important to note that you can address the question by... Input should be performed after a difficulty assessment. Furthermore, the decision (broadband maintenance strategy) generated by the engineer's intelligent agent will only be provided to the user when the assessment scores of each module and the total verification score threshold meet the following conditions:

[0123] ,

[0124] in The sum of the scores for the module evaluation. This function calculates the scores (total score, average score, etc.) of three modules. The formula uses a weighted average, but other methods can be used in practice. For example, the task decomposition module includes five scores such as completeness and logicality, which can be represented as follows: .in This function represents the calculation of scores for each item in the task breakdown. These are weight matrices, corresponding to the task decomposition module, real-time feedback evaluation module, and effectiveness evaluation module, respectively. The total score represents the sum of the scores in the task decomposition module. The score represents the real-time feedback evaluation module. This represents the score for the validity assessment module. It's important to note that, in addition to setting thresholds for the total score and average score, each module can also set thresholds for individual section scores.

[0125] In addition to the score threshold, this module needs to set a threshold for the number of times the strategy needs to be regenerated. The number of times the engineer agent's decision is regenerated satisfies the constraint. In such cases, the inspector agent needs to request intervention from professional customer service to help locate the problem. Simultaneously, we need to limit the number of times engineer-generated decisions fail to resolve the issues. This module also needs to set a threshold to limit the number of failed decision-making troubleshooting attempts. .when In such cases, the inspector agent needs to request intervention from professional customer service to help the agent locate the problem.

[0126] The memory module is composed of short-term memory and long-term memory, serving as... Figure 3 The system described above is the core support for storing and retrieving information. Short-term memory focuses on storing real-time data from the current session, covering detailed descriptions of user fault reports, the engineer's real-time decision-making and reasoning process, real-time feedback from users and on-site engineers, and temporary equipment operating parameters. This ensures the continuity of multi-round interactions, supports the agent in dynamically adjusting decisions based on the latest scenario, and provides real-time information for the inspector agent's evaluation. Long-term memory stores historical data and domain knowledge, including past user fault records, equipment lifecycle information, historical decision validity tags, a professional maintenance knowledge base, and customer service intervention cases. By providing historical experience references for the agent, it supports personalized decision-making and provides a data foundation for the inspector agent to verify the effectiveness of decisions, helping the system learn from history to optimize subsequent processing.

[0127] The toolset serves as a bridge between the intelligent agent and its external environment and system, integrating various functional modules to expand the agent's practical capabilities. Network status diagnostic tools can collect technical parameters such as network speed and optical attenuation in real time; device information query tools can retrieve hardware models and configuration details; and operator backend data interfaces can access internal data such as account status and bandwidth. These tools provide the intelligent agent with comprehensive perception data, supporting accurate fault location. Operation execution tools can translate decisions into actual actions, such as remotely restarting devices and generating work orders; feedback collection and analysis tools can transform verbal feedback from users and engineers into structured information and input it into a memory module for use in the evaluation process; and historical data retrieval tools can quickly match similar cases to assist in decision verification. These tools overcome the limitations of the native capabilities of large language models, constructing a closed loop of data acquisition, operation execution, and feedback processing, providing key technical support for collaborative work and reflective optimization by the intelligent agent.

[0128] Figure 4 This is a schematic diagram illustrating the operation process of a broadband maintenance strategy determination system according to an embodiment of this application. Figure 3In the three evaluation modules of the inspector agent in the system shown, each evaluation item can be evaluated as follows: Figure 4 The diagram illustrates two approaches to decision evaluation: model-based and program-based. Model-based decision evaluation is suitable for tasks where quantification is difficult. This method leverages the deep reasoning capabilities of domain LLMs to generate high-quality reflective content, aiding the entire system in reflection and decision repair. It is suitable for evaluation tasks where the evaluation logic is difficult to extract and requires significant reflective cues, such as analyzing completeness and logical coherence in task decomposition. However, a challenge lies in extracting the inherent evaluation logic from large amounts of data, processes, and knowledge, requiring substantial data collection, cleaning, and labeling. Program-based decision evaluation is suitable for evaluation tasks where the logic is easily extracted. Compared to model-based methods, this approach is easier to implement and requires less complex data engineering, but it exhibits poorer scalability and reflective capabilities. This method is suitable for evaluation tasks where reflection is not required and the evaluation logic is relatively clear, such as task decomposition feasibility assessment. Both model-based and program-based decision evaluation methods start with structured records (e.g., user's previous fault reports, fault handling results, decision execution feedback, detailed equipment specifications and maintenance history), existing processes (established and widely accepted operating procedures or protocols), expert knowledge, and technical documentation. In the fine-tuning phase of model-based decision evaluation, the inspector agent domain LLM is fine-tuned using a loss function (reward function) and optimization methods (e.g., Batch-Constrained Q-learning, BCQ) through RLHF. During system operation, the following steps are executed: 1. User fault reporting; 2. Engineer agent domain LLM generating strategy; 3. Inspector agent domain LLM evaluating strategy; 4. Returning score and reflection. In the program design phase, the logic is extracted from data to obtain the evaluation logic, then the program is designed, resulting in an evaluation program for evaluation. During system operation, the following steps are executed: 1. User fault reporting; 2. Engineer agent generating strategy; 3. Evaluation program evaluating strategy; 4. Returning score.

[0129] The task decomposition and evaluation module, which requires support from the domain LLM brain, needs to be fine-tuned using RLHF. For example... Figure 4The large-model-based evaluation approach outlines three steps to construct a domain LLM with evaluation capabilities. Step 1: Determine the reinforcement learning reward model. The reward function can directly utilize the loss functions mentioned in the previous steps. Step 2: Determine the optimization method: Since reinforcement learning training is performed on a static dataset, the behavior cloning method of the Batch-Constrained Q-learning (BCQ) algorithm in offline reinforcement learning is chosen for reinforcement training of the domain LLM. Behavior cloning allows the LLM to first mimic the distribution of "historically effective behaviors" in offline data through supervised learning. Step 3: Determine the fine-tuning mechanism: Low-Rank Adaptation (LoRA) is commonly used for fine-tuning in RLHF, therefore, LoRA is used as the fine-tuning mechanism.

[0130] Figure 5 This is a flowchart of another method for determining a broadband maintenance strategy according to an embodiment of this application. First, 1. The user reports a fault (e.g., Figure 5 The user indicated that broadband internet access was unavailable. This indicates the number of times the decision-making fault location process failed. 0. 2. User intent perception. 3. Data perception. Number of times the decision needs to be regenerated. =0. 4. Use the engineer agent to generate decisions (corresponding to the process of the first agent's broadband maintenance strategy). 5. The inspector agent performs decision evaluation: 5.1 Task decomposition evaluation, 5.2 Effectiveness evaluation, and 5.3 Real-time feedback evaluation when the number of failures in decision fault location is greater than 0. After evaluation, if the number of failures in decision fault location and the number of decision regenerations are both less than the corresponding thresholds, 6.1.a Decision verification is performed. If 6.1.c verification passes, the decision is returned. If the problem is solved, the problem-solving process ends in 6.2.a. If the problem is not solved, then in 6.2.b the problem is not solved, the number of failures in decision fault location is incremented by 1, and the entire process is restarted from data perception (collecting the latest feedback, re-triggering the data perception and decision generation process, and performing a new round of verification to form an optimization closed loop). At this time, At the same time (Reset). If the decision verification fails in step 6.1.a, proceed to step 6.1.b. If the verification fails, repeat step 4. Decision Generation. If either the number of failed decision fault location attempts or the number of decision regeneration attempts is greater than or equal to the corresponding threshold, professional customer service will intervene (the decision verification module will proactively request intervention from professional customer service. Through expert customer service intervention, the system greatly improves the efficiency and accuracy of manual intervention while shortening the user's waiting time for problem resolution).

[0131] Figure 6This is a schematic diagram of a device for determining a broadband maintenance strategy according to an embodiment of this application.

[0132] The receiving module 602 is used to receive user fault reporting information through a first intelligent agent, wherein the first intelligent agent is at least used to generate a broadband maintenance strategy corresponding to the user fault reporting information.

[0133] The generation module 604 is used to generate a first broadband maintenance strategy by the first intelligent agent based at least on user fault reporting information.

[0134] The evaluation module 606 is used to conduct a first round of evaluation of the first broadband maintenance strategy through the second intelligent agent and obtain the evaluation score corresponding to the first round of evaluation. The first round of evaluation includes multiple evaluation methods such as task decomposition evaluation and effectiveness evaluation. The task decomposition evaluation is used to at least verify the quality of the sub-tasks decomposed in the first broadband maintenance strategy. The effectiveness evaluation is used to at least verify the correlation between the first broadband maintenance strategy and the effective broadband maintenance decisions of historical target application scenarios. The historical target application scenarios are scenarios whose similarity index with the application scenarios corresponding to user fault reporting information is less than a preset similarity threshold.

[0135] The determination module 608 is used to determine the first broadband maintenance strategy as the target broadband maintenance strategy and push it to the user terminal when the evaluation score corresponding to the first round of evaluation is greater than the preset information threshold.

[0136] It should be noted that, Figure 6 The broadband maintenance strategy determination device shown is used to perform... Figure 2 The method for determining the broadband maintenance strategy shown is therefore... Figure 2 The relevant explanations in the method for determining the broadband maintenance strategy also apply to the device for determining the broadband maintenance strategy, and will not be repeated here.

[0137] It should be noted that each module in the above-mentioned broadband maintenance strategy determination device can be a program module (for example, a set of program instructions to implement a certain function) or a hardware module. For the latter, it can be manifested in the following forms, but is not limited to them: each of the above modules is manifested as a processor, or the functions of each of the above modules are implemented by a processor.

[0138] This application also provides a non-volatile storage medium, which includes a stored program, wherein, when the program is running, it controls the device where the non-volatile storage medium is located to execute the method for determining the broadband maintenance strategy of any of the above embodiments.

[0139] This application also provides an electronic device, which includes a processor for running a program, wherein the method for determining a broadband maintenance strategy according to any of the above embodiments is executed when the program is running.

[0140] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the method for determining the broadband maintenance strategy of any of the above embodiments.

[0141] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0142] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0143] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0144] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0145] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0146] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for determining a broadband maintenance strategy, characterized in that, include: The system receives user fault reports through a first intelligent agent, wherein the first intelligent agent is at least used to generate a broadband maintenance strategy corresponding to the user fault reports. The first intelligent agent generates a first broadband maintenance strategy based at least on the user's fault report information. The first broadband maintenance strategy is evaluated in the first round by the second intelligent agent to obtain the evaluation score corresponding to the first round of evaluation. The evaluation methods in the first round of evaluation include task decomposition evaluation and effectiveness evaluation. The task decomposition evaluation is used at least to check the quality of the sub-tasks decomposed in the first broadband maintenance strategy. The effectiveness evaluation is used at least to check the correlation between the first broadband maintenance strategy and the effective broadband maintenance decisions of historical target application scenarios. The historical target application scenarios are scenarios whose similarity index with the application scenario corresponding to the user fault reporting information is less than a preset similarity threshold. If the evaluation score corresponding to the first round of evaluation is greater than the preset threshold, the first broadband maintenance strategy will be determined as the target broadband maintenance strategy and pushed to the user terminal.

2. The method according to claim 1, characterized in that, The step of generating a first broadband maintenance strategy through a first intelligent agent based at least on the user's fault reporting information includes: The first intelligent agent acquires multi-dimensional data, including user equipment information, historical broadband maintenance data, operator backend data, and network status data. The user fault reporting information is standardized by the first intelligent agent to obtain standardized user fault reporting information. The multi-dimensional data and the standardized user fault reporting information are determined as input data; The input data is analyzed by the broadband maintenance strategy prediction model of the first intelligent agent to obtain the first broadband maintenance strategy. The first broadband maintenance strategy includes a thought chain and a preliminary decision. The preliminary decision includes multiple initial sub-tasks for solving the fault problem indicated in the user's fault report information and the task content of each initial sub-task. The thought chain is used to demonstrate the process of obtaining the preliminary decision.

3. The method according to claim 2, characterized in that, The method further includes: If the target broadband maintenance strategy fails to resolve the fault, obtain real-time feedback data after the user performs operations according to the target broadband maintenance strategy. The broadband maintenance strategy prediction model of the first intelligent agent regenerates the second broadband maintenance strategy based at least on the real-time feedback data. The second intelligent agent performs a second round of evaluation on the second broadband maintenance strategy to obtain the evaluation score corresponding to the second round of evaluation. The various evaluation methods of the second round of evaluation include the task decomposition evaluation, real-time feedback evaluation, and effectiveness evaluation. The real-time feedback evaluation is at least used to verify the real-time response capability of the second broadband maintenance strategy. If the evaluation score corresponding to the second round of evaluation is greater than the preset information threshold, the second broadband maintenance strategy will be determined as the new target broadband maintenance strategy.

4. The method according to claim 3, characterized in that, The second round of evaluation of the second broadband maintenance strategy by the second intelligent agent yields an evaluation score, including: The second intelligent agent evaluates the second broadband maintenance strategy using task decomposition assessment to determine a first score, a second score, a third score, a fourth score, and a fifth score. The first score is used to quantify whether multiple sub-tasks cover all aspects of the fault problem. The second score is used to quantify the rationality of the dependencies and execution order among the multiple sub-tasks. The third score is used to quantify the rationality of the task granularity of the multiple sub-tasks. The fourth score is used to quantify the operational feasibility of the multiple sub-tasks. The fifth score is used to quantify the adaptability of the fault handling tools indicated in the multiple sub-tasks to the fault problem. The set of the first score, the second score, the third score, the fourth score, and the fifth score is determined as the evaluation score corresponding to the task decomposition evaluation method; The second intelligent agent evaluates the second broadband maintenance strategy using real-time feedback to determine a sixth score, wherein the sixth score is used to quantify whether the second broadband maintenance strategy responds to the user's real-time feedback data. The second intelligent agent evaluates the effectiveness of the second broadband maintenance strategy and determines a seventh score, wherein the seventh score is used to determine the correlation between the second broadband maintenance strategy and effective broadband maintenance decisions in historical broadband maintenance data. The evaluation score for the second round of evaluation is determined based on the evaluation score corresponding to the task decomposition evaluation method, as well as the sixth score and the seventh score.

5. The method according to claim 4, characterized in that, The first score is determined in the following manner: The input data and the real-time feedback data are analyzed by the target evaluation model in the second intelligent agent to generate a reference decision, wherein the reference decision includes multiple reference sub-tasks; The first occurrence number of the preset text in the plurality of subtasks and the second occurrence number of the preset text in the plurality of reference subtasks are determined, and the first score is determined based on the first occurrence number and the second occurrence number.

6. The method according to claim 4, characterized in that, The second score is determined in the following manner: A first directed acyclic graph is constructed based on the multiple subtasks, wherein each node in the first directed acyclic graph corresponds to a subtask, and the edges between the nodes in the first directed acyclic graph represent the execution order between the nodes. The set of edges of the first directed acyclic graph is determined as the task dependency set, and the second score is determined at least based on the task dependency set.

7. The method according to claim 4, characterized in that, The third fraction is determined in the following manner: Determine the total number of tasks for the plurality of subtasks; The second intelligent agent predicts the total number of target sub-tasks to solve the fault problem based on domain knowledge. The third score is determined based on the total number of tasks and the total number of target sub-tasks.

8. The method according to claim 4, characterized in that, The fourth fraction is determined in the following manner: Determine the total number of tasks for the plurality of subtasks; The complexity score corresponding to each of the multiple subtasks is determined, and the fourth score is determined based on the total number of tasks and the complexity scores of all subtasks. The complexity score of each subtask is determined based on the reasonableness score of the number of operation steps of the subtask, the preset weight corresponding to the reasonableness score of the number of operation steps, and the fifth score. The reasonableness score of the number of operation steps is determined based on the actual number of operation steps of the subtask, the baseline number of operation steps corresponding to the task level to which the subtask belongs, and the preset fluctuation range of operation steps.

9. The method according to claim 4, characterized in that, The fifth fraction is determined in the following manner: Obtain the first tool selection set in the second broadband maintenance strategy, wherein the first tool selection set is a set of tools indicated in the task content of the multiple sub-tasks in the second broadband maintenance strategy; Obtain a second tool selection set in the reference decision, wherein the second tool selection set is a set of tools indicated in the task content of multiple reference subtasks in the reference decision, wherein the reference decision is obtained by analyzing the input data and the real-time feedback data through the target evaluation model in the second agent; The fifth score is determined based at least on the first tool selection set and the second tool selection set.

10. The method according to claim 4, characterized in that, The evaluation of the second broadband maintenance strategy by the second intelligent agent using real-time feedback assessment to obtain a sixth score includes: Keyword extraction is performed on the real-time feedback data to obtain a first keyword set; The sixth score is determined based on the first set of keywords, the first broadband maintenance strategy, and the second broadband maintenance strategy.

11. The method according to claim 6, characterized in that, The second intelligent agent evaluates the effectiveness of the second broadband maintenance strategy, resulting in a seventh score, which includes: Obtain historical log data, wherein the historical log data contains multiple log records corresponding to all historical target application scenarios, and each log record includes at least a problem description, effective broadband maintenance decisions, and corresponding processing flow records; A second directed acyclic graph set is constructed based on the historical log data. The second directed acyclic graph set contains multiple second directed acyclic graphs. Each log record in the historical log data corresponds to a second directed acyclic graph. The nodes in the second directed acyclic graph are the events or operation steps recorded in the log record. The edges in the second directed acyclic graph are the temporal relationships and dependencies between nodes. The seventh score is obtained by performing correlation analysis based on the first and second directed acyclic graphs.

12. The method according to claim 1, characterized in that, The method further includes: when the evaluation score corresponding to the first round of evaluation is less than or equal to the preset information threshold, adjusting the first broadband maintenance strategy a preset number of times through the first intelligent agent, and initiating a customer service intervention request when the evaluation score of the first broadband maintenance strategy after each adjustment is less than or equal to the preset information threshold.

13. A device for determining a broadband maintenance strategy, characterized in that, include: A receiving module is used to receive user fault reporting information through a first intelligent agent, wherein the first intelligent agent is at least used to generate a broadband maintenance strategy corresponding to the user fault reporting information. The generation module is used to generate a first broadband maintenance strategy by the first intelligent agent based at least on the user fault reporting information; The evaluation module is used to perform a first round of evaluation on the first broadband maintenance strategy through a second intelligent agent to obtain an evaluation score corresponding to the first round of evaluation. The first round of evaluation includes multiple evaluation methods such as task decomposition evaluation and effectiveness evaluation. The task decomposition evaluation is used at least to check the quality of the sub-tasks decomposed in the first broadband maintenance strategy. The effectiveness evaluation is used at least to check the correlation between the first broadband maintenance strategy and the effective broadband maintenance decisions of historical target application scenarios. The historical target application scenarios are scenarios whose similarity index with the application scenario corresponding to the user fault reporting information is less than a preset similarity threshold. The determination module is used to determine the first broadband maintenance strategy as the target broadband maintenance strategy and push it to the user terminal when the evaluation score corresponding to the first round of evaluation is greater than a preset information threshold.

14. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores a program, wherein when the program is executed, it controls the device containing the non-volatile storage medium to execute the method for determining the broadband maintenance strategy as described in any one of claims 1 to 12.

15. An electronic device, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, executes the method for determining a broadband maintenance strategy as described in any one of claims 1 to 12.

16. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the method for determining the broadband maintenance strategy as described in any one of claims 1 to 12.