Large-scale model inference optimization method and system for operation and maintenance technical services
By dynamically prioritizing and allocating resources for operation and maintenance service requests, the problem of improper task processing in the operation and maintenance system is solved, and efficient and high-quality response of operation and maintenance services is achieved.
Patent Information
- Application Number
- CN202511178626.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-08-22
AI Technical Summary
Existing operation and maintenance systems lack flexibility in handling operation and maintenance service requests to adapt to different operation and maintenance scenarios and needs. This results in critical and urgent tasks not being processed in a timely manner, low resource utilization, and an inability to fully leverage the advantages of large-scale models in operation and maintenance technical services.
By acquiring a set of operation and maintenance service requests, prioritizing them dynamically based on the requirement description information and equipment operation scenario identifiers, generating a priority list of operation and maintenance inference tasks, allocating computing resources and operation and maintenance knowledge resources, generating a collaborative scheduling scheme, ensuring that critical tasks are processed first, resources are configured reasonably, and operation and maintenance response instructions that conform to the equipment operation scenario are generated.
It enables accurate identification and reasonable prioritization of operation and maintenance tasks, improves response speed and targeting, avoids resource waste, and enhances the quality and efficiency of operation and maintenance services.
Smart Images

Figure CN120729803B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of operation and maintenance technology, and more specifically, to a large-scale model reasoning optimization method and system for operation and maintenance technical services. Background Technology
[0002] In the current field of operation and maintenance (O&M) technical services, with the continuous increase in the types and complexity of equipment, the number and types of O&M service requests are also becoming increasingly diverse. Traditional O&M service processing methods often use relatively fixed processes to deal with various requests, lacking flexibility and adaptability to different O&M scenarios and needs.
[0003] Existing operation and maintenance systems typically do not adequately consider the device operating scenario identifier and specific operation and maintenance requirement descriptions associated with each request when processing operation and maintenance service requests. This results in the inability to reasonably prioritize requests based on their importance and urgency when faced with a large number of concurrent requests. Consequently, some critical and urgent operation and maintenance tasks may not be processed in a timely manner, while some relatively minor tasks consume a large amount of resources.
[0004] Meanwhile, in the application of large-scale model inference to operations and maintenance services, there is a lack of scientific and reasonable methods for allocating the computing resources and operations and maintenance knowledge resources required for inference. Often, a fixed resource allocation pattern is followed, without dynamic adjustments based on the characteristics and needs of different operations and maintenance inference tasks. This results in low resource utilization, failing to fully leverage the advantages of large-scale models in operations and maintenance technical services, and impacting the quality and efficiency of operations and maintenance services. Summary of the Invention
[0005] In view of this, the purpose of this application is to provide a large-scale model reasoning optimization method and system for operation and maintenance technical services.
[0006] According to a first aspect of this application, a large-scale model inference optimization method for operation and maintenance technical services is provided, the method comprising:
[0007] Obtain a set of operation and maintenance service requests under the operation and maintenance technical service scenario. The set of operation and maintenance service requests contains multiple operation and maintenance service requests to be processed. Each operation and maintenance service request carries operation and maintenance requirement description information and associated equipment operation scenario identifier.
[0008] Based on the operation and maintenance requirement description information and equipment operation scenario identifier of each operation and maintenance service request in the operation and maintenance service request set, the priority of the operation and maintenance inference task corresponding to each operation and maintenance service request is dynamically marked, and an operation and maintenance inference task priority list is generated.
[0009] Based on the priority list of operation and maintenance inference tasks and the inference link requirements corresponding to each operation and maintenance inference task, the computing resources and operation and maintenance knowledge resources required for large model inference are allocated to generate a collaborative scheduling scheme for computing resources and operation and maintenance knowledge resources.
[0010] Based on the aforementioned collaborative scheduling scheme for computing resources and operation and maintenance knowledge resources, an appropriate reasoning process is configured for each operation and maintenance reasoning task. The large model reasoning process is executed through the allocated computing resources to generate operation and maintenance reasoning results corresponding to the operation and maintenance service request.
[0011] Combining the device operation scenario identifier associated with the operation and maintenance service request, the operation and maintenance reasoning result is converted into an operation and maintenance response instruction that conforms to the operation specifications of the device operation scenario, and the operation and maintenance response instruction is sent to the corresponding operation and maintenance execution terminal.
[0012] According to a second aspect of this application, a large-scale model inference optimization system for operation and maintenance technical services is provided. The large-scale model inference optimization system for operation and maintenance technical services includes a machine-readable storage medium and a processor. The machine-readable storage medium stores machine-executable instructions. When the processor executes the machine-executable instructions, the large-scale model inference optimization system for operation and maintenance technical services implements the aforementioned large-scale model inference optimization method for operation and maintenance technical services.
[0013] According to a third aspect of this application, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, and when the computer-executable instructions are executed, the aforementioned large model inference optimization method for operation and maintenance technical services is implemented.
[0014] Based on any of the above aspects, the technical effect of this application is as follows:
[0015] By acquiring a set of operation and maintenance (O&M) service requests in O&M technical service scenarios and dynamically prioritizing each request based on its O&M requirement description and device operation scenario identifier, a priority list for O&M inference tasks is generated. This accurately identifies the importance and urgency of different O&M tasks, ensuring that critical tasks are processed first. This optimizes the processing order of O&M tasks from the source, improving the response speed and targeting of O&M services. Based on the priority list and inference chain requirements, the computing resources and O&M knowledge resources required for large-scale model inference are allocated, generating a collaborative scheduling scheme for computing resources and O&M knowledge resources. This achieves dynamic and reasonable resource allocation, avoiding resource waste and idleness, fully leveraging the inference capabilities of the large model, and improving resource utilization efficiency. Based on the scheduling scheme, an appropriate inference process is configured and executed for large-scale model inference processing, generating accurate O&M inference results. These results are then combined with device operation scenario identifiers to transform into O&M response instructions that conform to operational specifications and sent to the execution terminal, significantly improving the quality and efficiency of O&M technical services. Attached Figure Description
[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be introduced in a basic manner below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart illustrating the large-model inference optimization method for operation and maintenance technical services provided in an embodiment of this application is shown.
[0018] Figure 2 This illustration shows a schematic diagram of the component structure of a large-model inference optimization system for operation and maintenance technical services, provided in an embodiment of this application, for implementing the above-described large-model inference optimization method for operation and maintenance technical services. Detailed Implementation
[0019] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.
[0020] Figure 1 This document illustrates a flowchart of a large-model inference optimization method and system for operation and maintenance technical services provided in an embodiment of this application. It should be understood that in other embodiments, the order of some steps in the large-model inference optimization method for operation and maintenance technical services in this embodiment can be shared according to actual needs, or some steps can be omitted or maintained. The detailed steps of the large-model inference optimization method for operation and maintenance technical services include:
[0021] Step S110: Obtain the set of operation and maintenance service requests under the operation and maintenance technical service scenario. The set of operation and maintenance service requests contains multiple operation and maintenance service requests to be processed. Each operation and maintenance service request carries operation and maintenance requirement description information and associated equipment operation scenario identifier.
[0022] This embodiment uses the operation and maintenance scenario of intelligent equipment in chain stores as the application background, involving equipment such as smart air conditioners, store information display screens, and IoT control systems. The set of operation and maintenance service requests comes from technical support requests initiated by on-site engineers from various chain stores via mobile terminals, as well as automatic repair requests generated after the system automatically detects equipment anomalies. Each operation and maintenance service request contains complete operation and maintenance requirement description information and corresponding equipment operating scenario identifiers.
[0023] The description of maintenance requests varies depending on the request type. For requests initiated by on-site engineers, the description includes the time the problem occurred, the abnormal equipment symptoms, the initial troubleshooting operations performed, and the results. For system-generated requests, the description includes abnormal parameter monitoring values, the duration of the abnormal parameters, and the frequency of the abnormality. The equipment operating scenario identifier is a string of letters and numbers containing the store code, equipment type code, installation area code, and current operating mode code. For example, "STXYZ-KT-XX-C" indicates that an air conditioning unit installed in a specific area of a store is in cooling mode.
[0024] These maintenance service requests are aggregated to the request processing server in the technology center via an encrypted transmission protocol. The server performs format validation on each request to ensure the completeness of the maintenance requirement description information and the validity of the equipment operation scenario identification. After successful validation, all requests are temporarily stored in a request queue in order of receipt time, forming a set of maintenance service requests to be processed, awaiting subsequent priority assignment and processing.
[0025] Step S120: Based on the operation and maintenance requirement description information and equipment operation scenario identifier of each operation and maintenance service request in the operation and maintenance service request set, dynamically assign priority to the operation and maintenance inference task corresponding to each operation and maintenance service request, and generate an operation and maintenance inference task priority list.
[0026] For each request in the set of operation and maintenance service requests, the task labeling module of the technology center extracts its operation and maintenance requirement description information and equipment operation scenario identifier, and determines the priority of the corresponding operation and maintenance inference task through multi-dimensional analysis. First, it parses the core problem type in the operation and maintenance requirement description information to determine whether it is an emergency fault, parameter abnormality, or routine inspection requirement; at the same time, it decodes the equipment operation scenario identifier to obtain scenario information such as the importance of the store where the equipment is located and the type of equipment function.
[0027] Based on the urgency of the problem type and the importance of the equipment in the scenario information, each maintenance inference task is evaluated using a pre-defined priority assessment logic. For example, the inference task for an air conditioning cooling failure in a core store receives a higher evaluation result than the task for fine-tuning the display parameters on the information display screen of a regular store. All maintenance inference tasks are sorted from highest to lowest evaluation result to form a maintenance inference task priority list, which includes the task identifier, the corresponding requested equipment operation scenario identifier, and priority assessment details.
[0028] Step S121: Extract the operation and maintenance task type from the operation and maintenance requirement description information of each operation and maintenance service request. The operation and maintenance task type includes equipment fault handling task, equipment parameter configuration task, and equipment operation status monitoring task.
[0029] The task type extraction module performs semantic analysis on the description information of operation and maintenance requirements, and determines the type of operation and maintenance task through keyword matching and intent recognition. The description of equipment fault handling tasks usually contains keywords such as "fault", "abnormal", and "cannot run", such as "the air conditioner suddenly stopped during operation and the display screen has no display". The description of equipment parameter configuration tasks contains keywords such as "adjust", "set", and "calibrate", such as "the brightness parameter of the information display screen needs to be adjusted to the standard value". The description of equipment operation status monitoring tasks contains keywords such as "inspect", "monitor", and "patrol", such as "regularly monitor the communication status of the IoT controller".
[0030] After each maintenance service request is accurately categorized into its corresponding task type, the system adds a task type tag to it. The tag is stored in the form of a numeric code: "01" represents a device fault handling task, "02" represents a device parameter configuration task, and "03" represents a device operating status monitoring task. These tags will serve as one of the basic parameters for subsequent priority evaluation, ensuring that the priority evaluation logic for different types of tasks is differentiated.
[0031] Step S122: Parse the device operation scenario identifier associated with each operation and maintenance service request, and determine the system level to which the device belongs, the business link associated with the device, and the current operation stage of the device corresponding to the device operation scenario identifier.
[0032] The device operation scenario identifier parsing module decodes the identifier string segment by segment according to preset encoding rules. Taking the identifier "STXYZ-KT-XX-C" as an example, firstly, "STXYZ" is extracted as the store code. By querying the store information database, it is found that the store is a core store in the region, and the corresponding system level of the device is the first-level core layer. Next, "KT" is extracted as the device type code, which is determined to be an air conditioning device. The associated business link is the environmental control link, which directly affects the customer experience in the store. Then, "XX" is extracted as the installation area code, which corresponds to the store lobby area. Finally, "C" is extracted as the operation mode code, representing the cooling operation stage.
[0033] After parsing, structured scenario information is generated, including the system level to which the device belongs (Level 1 Core Layer / Level 2 Standard Layer / Level 3 Auxiliary Layer), the business links associated with the device (Environmental control / Information display / Security monitoring, etc.), and the current operating stage of the device (Startup / Stable operation / Standby, etc.). This information will be combined with the operation and maintenance task type to jointly determine the priority level.
[0034] Step S123: Determine the basic priority level based on the association between the operation and maintenance task type and the system level to which the device belongs. The basic priority level corresponding to the device fault handling task is higher than that of the device parameter configuration task, and the basic priority level corresponding to the device parameter configuration task is higher than that of the device operation status monitoring task.
[0035] The basic priority level classification follows a dual dimension of task type and system level. In the preset association rules, the basic priority of the equipment fault handling task is higher than the other two types of tasks at any system level; the basic priority of the equipment parameter configuration task is higher than the equipment operation status monitoring task, but lower than the equipment fault handling task; the basic priority of the equipment operation status monitoring task is the lowest.
[0036] Simultaneously, the system hierarchy fine-tunes the basic priority levels. Tasks of the same type on Level 1 core layer devices have a higher basic priority than those on Level 2 standard layer devices, and vice versa. For example, the basic priority of a Level 1 core layer device parameter configuration task is higher than that of a Level 2 standard layer task, and the basic priority of a Level 1 core layer device operation status monitoring task is higher than that of a Level 2 standard layer task. Through this dual-dimensional evaluation, an initial basic priority level is determined for each operation and maintenance inference task.
[0037] Step S124: Based on the business link attributes associated with the device, adjust the basic priority level. The priority level of the operation and maintenance inference task corresponding to the operation and maintenance service request of the core business link is adjusted upward.
[0038] The business link attribute assessment module determines whether a device is a core business link based on the type of business link associated with it. In the context of chain stores, environmental control links (air conditioning equipment) and main information display links (main display screen equipment) are core business links, and interruptions or abnormalities in these links will directly affect the normal operation of the store; while auxiliary lighting control, back-end data backup links, etc. are non-core business links.
[0039] For operation and maintenance inference tasks with the same basic priority level, if the associated business link is a core business link, its priority level will be adjusted up one level. For example, if both are device parameter configuration tasks and are at the same system level, the air conditioning parameter adjustment task in the core business link has a higher priority than the background device parameter adjustment task in the non-core business link. The adjusted priority level will be recorded, noting that the adjustment was based on the business link attribute.
[0040] Step S125: Based on the timeliness requirements of the operation and maintenance response at the current stage of equipment operation, further adjust the priority level. The priority level of the operation and maintenance inference task corresponding to the operation and maintenance service request that meets the immediate response requirement is adjusted upward again.
[0041] The current operating phase of the equipment determines the timeliness requirements for maintenance response. Equipment operating during peak hours (such as air conditioners in stores during peak business hours) has the highest timeliness requirements for maintenance response and needs to respond immediately; equipment operating during stable periods has the next lowest timeliness requirements; and equipment operating during standby or low-load periods has the lowest timeliness requirements.
[0042] The timeliness assessment module determines whether an immediate response is required based on the current operational phase of the equipment. For maintenance service requests requiring immediate response, the priority level is further increased from the previously adjusted level. For example, the task of handling air conditioning malfunctions in core stores during peak business hours, after adjustments at the system level and business chain, will have its priority level increased again due to the need for immediate response during peak operation, ensuring it receives the highest priority resource allocation.
[0043] Step S126: Integrate the final priority level corresponding to each operation and maintenance service request, arrange all operation and maintenance inference tasks corresponding to the operation and maintenance service requests in order of priority level from first to second priority, and generate an operation and maintenance inference task priority list.
[0044] The priority level integration module collects the final priority level of each operation and maintenance inference task after multi-dimensional adjustments. The levels are divided into four levels: "P0 (highest)", "P1", "P2" and "P3 (lowest)". All tasks are sorted from highest to lowest according to their final priority level, and tasks of the same level are arranged in order of receipt time.
[0045] The generated priority list of operation and maintenance inference tasks includes a unique task identifier, the corresponding device operating scenario identifier, the final priority level, priority adjustment details, and the estimated processing time. The list is stored in a structured data format and synchronized to the resource scheduling module in real time, serving as the core basis for allocating computing and knowledge resources and ensuring that high-priority tasks receive the necessary resource support first.
[0046] Step S130: Based on the priority list of operation and maintenance inference tasks and the inference link requirements corresponding to each operation and maintenance inference task, allocate the computing resources and operation and maintenance knowledge resources required for large model inference, and generate a collaborative scheduling scheme for computing resources and operation and maintenance knowledge resources.
[0047] After receiving the priority list of operation and maintenance inference tasks, the resource scheduling module first analyzes the inference chain requirements of each task to determine the type of computing resources (such as GPU nodes and CPU nodes), the quantity of computing resources, and the scope of operation and maintenance knowledge resources to be invoked. Based on the order of the priority list, resources are allocated to higher-priority tasks first.
[0048] When allocating computing resources, the load of currently available computing nodes must be considered to avoid resource overload. When allocating knowledge resources, the association between reasoning tasks and relevant knowledge sub-repositories must be established to ensure that the required knowledge data can be quickly obtained during the reasoning process. Finally, the resource allocation results are integrated into a collaborative scheduling scheme for computing resources and operational knowledge resources. The scheme clearly defines the allocation of computing nodes, the association with knowledge sub-repositories, and the resource usage time periods for each task.
[0049] Step S131: Analyze the inference chain requirements corresponding to each operation and maintenance inference task, and determine the type of inference computing node, the number of inference nodes, and the associated operation and maintenance knowledge domain required for each inference chain requirement.
[0050] The inference chain requirements analysis module breaks down each operation and maintenance inference task and determines the inference chain requirements based on the task's complexity and processing logic. For complex equipment fault diagnosis tasks, the inference chain is longer, requiring computational nodes with deep inference capabilities; for simple parameter query tasks, the inference chain is shorter, and basic inference nodes are sufficient.
[0051] By analyzing the technical fields and processing steps involved in the task, we determine the required inference computing node types (deep inference nodes or basic inference nodes), the number of inference nodes required to complete the task (single node or multi-node parallel operation), and the related operational knowledge domains (such as air conditioning fault handling, information dissemination equipment parameter domains, etc.). These analysis results will serve as a direct basis for resource allocation, ensuring the accuracy of resource allocation.
[0052] Step S1311: Decompose each operation and maintenance inference task into multiple inference subtasks. Each inference subtask corresponds to a link in the inference chain. Determine the processing logic and data interaction requirements of each inference subtask.
[0053] The task decomposition module breaks down the operation and maintenance reasoning task into multiple interconnected reasoning subtasks according to the reasoning process. Taking the air conditioning refrigeration fault diagnosis task as an example, it can be decomposed into stages such as "fault phenomenon identification subtask", "fault cause analysis subtask", "maintenance plan generation subtask", and "plan verification subtask".
[0054] Each inference subtask has a clearly defined processing logic. For example, the processing logic of the fault phenomenon identification subtask is "to extract key phenomenon features from the operation and maintenance requirement description and match them with the fault feature database." Simultaneously, data interaction requirements are determined, including input data types (such as phenomenon description text, equipment parameters), output data types (such as phenomenon feature vectors, matching degree scores), and data transfer relationships with other subtasks. This breakdown ensures that the resource requirements of each stage can be accurately assessed.
[0055] Step S1312: Based on the processing logic of each inference subtask, determine the type of inference computation node required for that inference subtask. Inference subtasks involving deep fault causal analysis correspond to the deep inference requirement node type, and inference subtasks involving basic parameter queries correspond to the basic inference requirement node type.
[0056] The computation node type determination module categorizes subtasks based on the processing logic complexity of the reasoning subtasks. Subtasks involving deep fault causal analysis, multi-factor correlation reasoning, and complex solution generation are classified as nodes requiring deep reasoning due to the need for multi-level feature extraction and logical reasoning. Subtasks involving basic parameter querying, simple rule matching, and standardized result output are classified as nodes requiring basic reasoning.
[0057] For example, the air conditioner fault cause analysis subtask requires analyzing the correlation between multiple parameters such as refrigerant pressure, compressor current, and ambient temperature, which falls under the category of deep reasoning; while the standard operating procedure query subtask in the repair plan falls under the category of basic reasoning. Different types of nodes are configured with different hardware resources and reasoning models to ensure processing efficiency and accuracy.
[0058] Step S1313: Count the number of inference subtasks included in each operation and maintenance inference task and the parallel processing possibility of each inference subtask, and determine the number of inference nodes required for the operation and maintenance inference task. The more inference subtasks that can be processed in parallel, the more inference nodes are required.
[0059] The inference node count module calculates the total number of inference subtasks for each operation and maintenance inference task and analyzes the dependencies between subtasks. Subtasks with no or weak dependencies are determined to be parallelizable; subtasks with strong dependencies (such as those that can only be executed after the completion of a preceding subtask) are determined to be serializable.
[0060] The required number of inference nodes is determined based on the number of subtasks that can be processed in parallel, following the principle of "matching the number of parallel subtasks with the number of nodes." For example, if a task contains multiple inference subtasks, some of which can be processed in parallel, then a corresponding number of inference nodes need to be allocated to achieve parallel processing and shorten the overall inference time. Node redundancy is also considered, with a certain number of spare nodes reserved to handle unforeseen circumstances.
[0061] Step S1314: Extract the operation and maintenance knowledge content that needs to be referenced during the processing of each inference subtask, and determine the associated operation and maintenance knowledge domain according to the category to which the operation and maintenance knowledge content belongs. Among them, the fault diagnosis type inference subtask is associated with the fault processing knowledge domain, and the parameter configuration type inference subtask is associated with the parameter standard knowledge domain.
[0062] The knowledge domain association module scans the processing logic of each inference subtask, extracting keywords of the required maintenance knowledge content, such as "air conditioner compressor fault code," "refrigerant charging standard pressure," and "display brightness adjustment parameter range." Based on the keyword matching knowledge domain classification system, it determines the maintenance knowledge domain associated with each subtask.
[0063] The fault diagnosis subtasks are primarily associated with the fault handling knowledge domain, which includes fault code libraries, fault cause-effect diagrams, and typical fault case sets. The parameter configuration subtasks are primarily associated with the parameter standards knowledge domain, which includes equipment parameter specification tables, parameter adjustment operation manuals, and parameter anomaly handling guidelines. Each subtask can be associated with multiple knowledge domains to ensure comprehensive knowledge coverage during the reasoning process.
[0064] Step S1315: Integrate the inference computing node type, number of inference nodes, and associated operation and maintenance knowledge domains corresponding to each operation and maintenance inference task to form an inference link requirement description for each operation and maintenance inference task.
[0065] The inference pipeline requirements integration module summarizes the various requirement parameters for each operations and maintenance inference task, forming a structured inference pipeline requirement description. The description includes the task identifier, a list of inference computing node types (arranged in subtask order), the total number of required inference nodes, a list of associated operations and maintenance knowledge domains, and the knowledge retrieval priority for each domain.
[0066] For example, in the inference chain requirement description for an air conditioner fault diagnosis task, it is clearly stated that deep inference requirement nodes are used for the fault cause analysis subtask, and basic inference requirement nodes are used for the fault code query subtask. A corresponding number of inference nodes need to be allocated, and they need to be associated with the fault handling knowledge domain (priority 1) and the parameter standard knowledge domain (priority 2). This inference chain requirement description serves as an input document for resource allocation, ensuring that resource allocation is accurately matched with task requirements.
[0067] Step S132: Organize the available computing resource pool under the operation and maintenance technical service scenario. The available computing resource pool includes multiple inference computing nodes. Each inference computing node is marked with its node service capabilities and the number of inference tasks it currently carries.
[0068] The computing resource management module periodically scans the computing resource cluster of the technology center, collecting status information of all available inference computing nodes. The annotation information for each inference computing node includes a unique node identifier, hardware configuration parameters (such as processor type, memory capacity, number of computing cores), supported inference model types, node service capability level (high / medium / low), and the number of inference tasks currently being carried.
[0069] Available computing resources are categorized and stored according to node service capability levels, forming a resource index table. This table is updated in real time; information is refreshed promptly when the number of tasks a node is handling changes or when a node's status changes (e.g., offline, maintenance). Through this process, the resource scheduling module can clearly understand the current resource supply situation.
[0070] Step S133: Organize the operation and maintenance knowledge resource library under the operation and maintenance technical service scenario. The operation and maintenance knowledge resource library is divided into multiple knowledge sub-libraries according to the operation and maintenance knowledge domain. Each knowledge sub-library stores operation and maintenance cases, operation and maintenance operation specifications and equipment characteristic data of the corresponding domain.
[0071] The knowledge resource organization module structures and categorizes the operation and maintenance knowledge resource base, dividing it into several specialized sub-bases based on the knowledge domain, such as the air conditioning equipment knowledge sub-base, the information publishing equipment knowledge sub-base, and the IoT control knowledge sub-base. Each knowledge sub-base is further subdivided according to knowledge type, including an operation and maintenance case library (storing historical fault handling records), an operation and maintenance specification library (storing standardized operating procedures), and an equipment characteristic database (storing equipment technical parameters and model difference data).
[0072] Each knowledge sub-repository is equipped with a knowledge indexing system and an update mechanism to ensure the accuracy and timeliness of the knowledge content. The knowledge indexing system supports multi-dimensional searching by keywords, device model, fault type, etc.; the update mechanism regularly extracts knowledge from new operation and maintenance cases and technical manuals and supplements it to the corresponding sub-repository.
[0073] Step S134: According to the operation and maintenance inference task priority list, prioritize the allocation of inference computing nodes with service capabilities matching the inference link requirements of core operation and maintenance inference tasks, so that the number of inference tasks currently carried by the inference computing nodes corresponding to the core operation and maintenance inference tasks does not exceed the preset ratio of the node's maximum capacity.
[0074] The core task resource allocation module identifies core operation and maintenance inference tasks (usually P0 and P1 level tasks) from the operation and maintenance inference task priority list and allocates resources in descending order of priority. For each core task, matching inference computing nodes are selected from the available computing resource pool based on the node type and service capability requirements in its inference chain.
[0075] During the selection process, the focus is on checking the number of inference tasks currently being handled by a node. This ensures that the node's load after allocation does not exceed a preset proportion of its maximum capacity, reserving resource margins to handle peak task loads. For example, if a deep inference node has a maximum capacity and is already handling some tasks, it's crucial to ensure that the newly allocated tasks do not exceed a preset proportion. If the current load is close to or exceeds this proportion, other nodes are selected. This rigorous load control ensures the inference efficiency of core tasks.
[0076] Step S1341: Extract core operation and maintenance reasoning tasks from the operation and maintenance reasoning task priority list, and process each core operation and maintenance reasoning task in order of priority from first to second priority.
[0077] The core task extraction module scans the priority list of operation and maintenance inference tasks and selects tasks with priority levels P0 and P1 as core operation and maintenance inference tasks. Tasks with priority level P0 are arranged in order of priority over tasks with priority level P1, and core tasks of the same priority level are sorted according to their original order in the list.
[0078] A core task processing queue is established, with each task item in the queue containing a task identifier, a summary of inference chain requirements, and a tag for the required resource type. The processing module sequentially calls the resource allocation interface according to the queue order, matching computing resources for each core task to ensure that high-priority tasks receive resource support first. During sequential processing, the system generates a dedicated resource allocation session for each core task, recording task processing progress and resource matching status to avoid resource allocation conflicts between tasks.
[0079] Step S1342: For the core operation and maintenance inference task currently being processed, select inference computing nodes from the available computing resource pool whose service capabilities match the inference link requirements of the operation and maintenance inference task, and form a candidate inference computing node set.
[0080] The resource filtering module analyzes the inference chain requirements of the current core operation and maintenance inference task, extracting the inference computing node type, service capability level, and hardware configuration requirements. Based on these parameters, it iterates through all nodes in the available computing resource pool and uses a node attribute matching algorithm to filter out inference computing nodes that meet the requirements.
[0081] During the screening process, the node type is first matched (nodes requiring deep inference or nodes requiring basic inference) to ensure consistency with task requirements. Secondly, the node's service capability level is verified to meet the processing performance requirements of the task. Finally, the node's hardware configuration is checked to ensure it supports the operation of the inference model involved in the task. All nodes meeting the criteria are included in a candidate inference computing node set, which contains information such as node identifier, current load status, and resource response speed for subsequent selection.
[0082] Step S1343: Query the number of inference tasks currently carried by each candidate inference computing node in the candidate inference computing node set, and calculate the ratio of the number of inference tasks currently carried by each candidate inference computing node to the maximum carrying capacity of the node.
[0083] The load query module obtains real-time load data for each node in the candidate inference computing node set through the resource monitoring interface, focusing on the number of inference tasks currently being carried. Simultaneously, it retrieves the maximum capacity parameter for each node. This maximum capacity parameter is pre-set based on the node's hardware configuration and historical performance test results, representing the maximum number of tasks the node can handle under stable operating conditions.
[0084] The calculation module calculates the node load rate by proportionally dividing the number of tasks currently being handled by each node by its maximum capacity. The load rate reflects the node's resource utilization; a lower ratio indicates that the node has sufficient remaining resources. The calculation results are stored in association with the node identifier to form a candidate node load list.
[0085] Step S1344: Select candidate inference computing nodes whose ratio of the number of inference tasks currently being carried to the maximum carrying capacity does not exceed a preset ratio threshold, and use them as target inference computing nodes.
[0086] The threshold filtering module invokes a preset load control strategy and extracts the load ratio threshold parameter. This load ratio threshold varies depending on the node type; the threshold for nodes requiring deep inference is lower than that for nodes requiring basic inference, to ensure that complex tasks receive more sufficient resource support.
[0087] The load rates of candidate nodes are compared with a threshold, and nodes whose load rates do not exceed the threshold are selected as target inference computing nodes. If multiple nodes meet the criteria, they are sorted by resource response speed from fastest to slowest, and the node with the fastest response speed is selected first. After the target node is determined, the system locks a portion of the node's resources to prevent it from being occupied by other tasks.
[0088] Step S1345: Assign the currently processed core operation and maintenance inference tasks to the target inference computing node, and update the number of inference tasks currently carried by the target inference computing node.
[0089] The task allocation module generates a task allocation instruction, which includes the core operation and maintenance inference task identifier, inference link parameters, and data transmission address. The instruction is sent to the target inference computing node via the resource scheduling interface. Upon receiving the instruction, the node returns a confirmation message, completing the task allocation process.
[0090] After task allocation is completed, the resource status update module immediately updates the current number of tasks carried by the target inference computing node, recalculates the node's load rate, and synchronizes it to the available computing resource pool. Simultaneously, it marks the task's allocation status as "allocated" in the core task processing queue and records the node identifier and allocation time.
[0091] Step S1346: Repeat the above steps to complete the allocation of inference computing nodes for all core operation and maintenance inference tasks, and then allocate inference computing nodes for ordinary operation and maintenance inference tasks in the same way.
[0092] The loop processing module monitors the completion status of the core task processing queue. Once all P0 and P1 level tasks have completed node allocation, it initiates the resource allocation process for ordinary operation and maintenance inference tasks (P2 and P3 levels). The allocation process for ordinary tasks is the same as that for core tasks, but it is lower than that for core tasks in terms of load threshold and resource selection priority.
[0093] During the normal task allocation process, the system prioritizes using the remaining available resources after the core tasks have been allocated, avoiding the use of spare resources that the core tasks might need. After all tasks have been allocated, a resource allocation summary report is generated, which includes the node allocation status, resource utilization, and load balancing status of tasks at each level.
[0094] Step S135: Match a knowledge sub-base corresponding to the operation and maintenance knowledge domain for each operation and maintenance inference task, and establish an association mapping between the inference computing node and the knowledge sub-base, so that the inference computing node can directly call the associated knowledge sub-base when performing inference processing.
[0095] The knowledge sub-base matching module searches for the corresponding knowledge sub-base in the operation and maintenance knowledge resource base based on the list of operation and maintenance knowledge domains associated with each operation and maintenance inference task. Through a domain tag matching algorithm, it ensures that each task can accurately match the relevant knowledge sub-base; for cross-domain tasks, it matches multiple relevant sub-bases.
[0096] The association mapping module establishes the association between inference computing nodes and matching knowledge sub-repositories, generating an association mapping table. This table records node identifiers, knowledge sub-repository identifiers, knowledge access permissions, and data transmission paths. When performing inference, the inference computing node can directly access the corresponding knowledge sub-repository through this mapping table to obtain the necessary operation and maintenance cases, operating procedures, and equipment characteristic data, without requiring an additional knowledge retrieval process.
[0097] Step S136: Record the allocation results of reasoning computing nodes and the matching results of knowledge sub-bases corresponding to each operation and maintenance reasoning task, and integrate them to form a collaborative scheduling scheme for computing resources and operation and maintenance knowledge resources.
[0098] The scheduling scheme generation module collects resource allocation data for all operation and maintenance inference tasks, including the inference computing node identifier, node configuration parameters, knowledge sub-base identifier, and associated mapping relationships for each task. This data is then categorized and organized according to task priority and resource type to form a structured scheduling scheme document.
[0099] The solution document includes a resource allocation list, a knowledge association list, a resource usage sequence table, and a conflict resolution mechanism. The resource usage sequence table clearly defines the resource usage periods for each task to avoid resource conflicts; the conflict resolution mechanism specifies the rules for adjusting task priorities when resources are insufficient. After the solution is generated, it undergoes validity verification to ensure that the resource requirements of all tasks are met. Once verification is successful, it is deployed to the inference execution module as the basis for inference process configuration.
[0100] Step S140: Based on the collaborative scheduling scheme of computing resources and operation and maintenance knowledge resources, configure an appropriate reasoning process for each operation and maintenance reasoning task, execute large model reasoning processing through the allocated computing resources, and generate operation and maintenance reasoning results corresponding to the operation and maintenance service request.
[0101] The inference process configuration module designs a dedicated inference process for each operation and maintenance inference task based on the collaborative scheduling scheme of computing resources and operation and maintenance knowledge resources. The process design determines the division of inference stages, the processing logic of each stage, and the timing of knowledge invocation based on the task's inference chain requirements, the types of allocated computing nodes, and the matching knowledge sub-repositories.
[0102] After configuration, the inference process parameters are sent to the corresponding inference computing nodes. The nodes load the large model inference engine and initialize according to the process parameters. Inference processing is executed using allocated computing resources. During inference, data from related knowledge sub-bases is invoked at preset times to assist in the development of the inference logic. After inference processing, an operations and maintenance inference result containing operation and maintenance suggestions, supporting data, and related data is generated. This result is stored in association with the corresponding operations and maintenance service request.
[0103] Step S141: Extract the allocation results of reasoning computing nodes and the matching results of knowledge sub-bases corresponding to each operation and maintenance reasoning task from the collaborative scheduling scheme of computing resources and operation and maintenance knowledge resources.
[0104] The solution parsing module connects to the collaborative scheduling solution of computing resources and operation and maintenance knowledge resources. It splits the solution content according to the task identifier and extracts the inference computing node allocation results for each operation and maintenance inference task, including node identifier, node type and resource configuration parameters. At the same time, it extracts the knowledge sub-base matching results, including knowledge sub-base identifier, knowledge domain and call priority.
[0105] The extracted results are organized into structured data according to tasks, with each task corresponding to a record containing node information and knowledge sub-base information. These records are transmitted to the inference process configuration module as the basic data for process design, ensuring that the configured inference process is compatible with the allocated resources and knowledge.
[0106] Step S142: Based on the matching results of the knowledge sub-base, determine the types of operation and maintenance knowledge data that each operation and maintenance reasoning task needs to call during the reasoning process and the timing of the call, and generate a knowledge call sequence plan.
[0107] The knowledge retrieval planning module analyzes the content of the knowledge sub-base in the knowledge sub-base matching results and determines the types of operation and maintenance knowledge data required by each operation and maintenance reasoning task at each stage of reasoning. For example, the fault diagnosis stage requires fault case data and causal relationship data, while the solution generation stage requires operation specification data and equipment parameter data.
[0108] Based on the phased division of the inference process, a call timing is set for each type of knowledge data to ensure timely access to knowledge support when needed for inference. The call timing is set according to the nodes of the inference phase, such as calling basic equipment parameter data at the start of inference, calling case reference data in the intermediate inference phase, and calling specification verification data in the result verification phase. The knowledge data type and call timing are integrated to generate a knowledge call sequence plan, which includes the call sequence number, knowledge type, sub-repository identifier, and triggering conditions.
[0109] Step S143: Based on the inference computing node type corresponding to the inference computing node allocation result, configure the inference module combination for each operation and maintenance inference task. The deep inference requirement node is configured with an inference module combination containing a multi-stage inference layer, and the basic inference requirement node is configured with an inference module combination containing a single-stage inference layer.
[0110] The inference module configuration module selects the corresponding inference module combination scheme based on the type of inference computing node (deep inference requirement node or basic inference requirement node). For deep inference requirement nodes, a multi-stage inference layer combination is configured to handle complex inference logic; for basic inference requirement nodes, a single-stage inference layer combination is configured to improve processing efficiency.
[0111] During module configuration, the functional responsibilities, input / output data formats, and inter-layer connections of each inference layer are clearly defined. Multi-stage inference layer configuration includes levels such as requirement analysis, knowledge association, multi-path inference, and result integration; single-stage inference layer configuration includes levels such as requirement mapping and direct inference. Configuration results are bound to inference computation nodes to ensure that nodes load the correct inference modules.
[0112] Step S1431: Identify the type of inference computing node in the inference computing node allocation result corresponding to each operation and maintenance inference task, and determine whether the inference computing node type is a deep inference requirement node or a basic inference requirement node.
[0113] The node type identification module parses the node attribute information in the inference computing node allocation results and determines the node type through the type identifier field. The identifier field of deep inference requirement nodes contains the keyword "deep inference", and the hardware configuration parameters include information on high-computing-power components that support the operation of complex models; the identifier field of basic inference requirement nodes contains the keyword "basic inference", and the hardware configuration is mainly for meeting the needs of lightweight model operation.
[0114] The identification results are stored in association with the operation and maintenance inference tasks, serving as the direct basis for the configuration of inference module combinations. For nodes with questionable identification results, the system automatically initiates configuration verification, using test tasks to verify the actual inference capabilities of the nodes and ensure accurate type judgment.
[0115] Step S1432: If the inference calculation node type is a deep inference requirement node, configure a multi-stage inference layer combination for the corresponding operation and maintenance inference task, including a requirement parsing layer, a knowledge association layer, a multi-path inference layer, and a result integration layer.
[0116] The multi-stage inference layer configuration module loads a combination of multi-stage inference layers for deep inference requirement nodes. The requirement parsing layer is responsible for structuring the operational requirement description information, converting unstructured text into structured requirement feature vectors; the knowledge association layer establishes the association between structured requirements and knowledge sub-base data, extracting relevant knowledge features; the multi-path inference layer generates multiple sets of intermediate inference results through different inference paths, with each path employing different inference logic; the result integration layer merges multiple sets of intermediate inference results, eliminates logical conflicts, and forms a unified inference conclusion.
[0117] The inference layers are connected sequentially, with the output of the previous layer serving as the input of the next. Inter-layer data transmission uses a standardized feature vector format. During configuration, the operating parameters of each layer are set, such as the feature extraction dimension of the requirement parsing layer and the number of paths in the multi-path inference layer, to ensure efficient and accurate inference.
[0118] Step S1433: The requirement parsing layer is used to structure the description information of operation and maintenance requirements, the knowledge association layer is used to establish the association between structured requirements and knowledge sub-base data, the multi-path reasoning layer is used to generate multiple sets of intermediate reasoning results through different reasoning paths, and the result integration layer is used to fuse multiple sets of intermediate reasoning results.
[0119] The requirement parsing layer employs natural language processing (NLP) technology to segment, identify entities, and classify intents in the description of operational requirements. It extracts key information and converts it into a structured requirement feature vector, which includes multiple dimensions such as requirement type, equipment features, and problem features. The knowledge association layer uses feature matching algorithms to compare the requirement feature vector with knowledge features in the knowledge sub-base, establishing associations and filtering out highly relevant knowledge data.
[0120] The multi-path reasoning layer sets up multiple parallel reasoning paths, each employing a different reasoning strategy, such as rule-based reasoning, case-based reasoning, and model-based reasoning. Each path independently generates intermediate reasoning results. The result integration layer uses a voting mechanism and logical verification methods to evaluate and merge multiple sets of intermediate reasoning results, retaining conclusions with high consistency, correcting conflicting content, and forming the final reasoning output.
[0121] Step S1434: If the inference computing node type is a basic inference requirement node, configure a single-stage inference layer combination that includes a requirement mapping layer and a direct inference layer for the corresponding operation and maintenance inference task.
[0122] The single-stage inference layer configuration module loads single-stage inference layer combinations based on the basic inference requirement nodes. The requirement mapping layer maps the operation and maintenance requirement description information into standardized inference inputs, and directly associates them with preset inference templates through keyword matching; the direct inference layer executes inference logic according to the inference template based on the standardized inference inputs and knowledge sub-base data, and generates a single inference result.
[0123] The two-layer structure employs a simple linear connection, with the output of the requirement mapping layer directly fed into the direct inference layer. During configuration, the matching accuracy of the inference templates is optimized to ensure that different types of basic requirements can be accurately mapped to their corresponding inference templates, thereby improving inference efficiency and accuracy.
[0124] Step S1435: The requirement mapping layer is used to map the operation and maintenance requirement description information into standardized reasoning input, and the direct reasoning layer is used to generate a single reasoning result based on the standardized reasoning input and knowledge sub-base data.
[0125] The requirement mapping layer uses preset mapping rules to convert keywords and core demands in the operation and maintenance requirement descriptions into a standardized inference input format. This maximum capacity format is consistent with the input requirements of the inference template. The mapping rules include the correspondence between keywords and inference templates, parameter extraction rules, and format conversion logic, ensuring that the mapped input data is accurate and complete.
[0126] The direct inference layer loads an inference template that matches the standardized input. This template contains fixed inference logic and knowledge reference paths. Following the guidance of the inference template, relevant data from the knowledge sub-base is invoked, and inference calculations are performed according to the logical steps to generate a single inference result. The result includes clear conclusions and supporting evidence, meeting the response requirements of basic operational needs.
[0127] Step S1436: Record the configuration results of the combination of inference modules corresponding to each operation and maintenance inference task, so that the configuration results are compatible with the inference computing node type.
[0128] The configuration result recording module stores the configuration details of the inference module combination for each operation and maintenance inference task in the configuration database. These details include the task identifier, node type, inference layer combination list, runtime parameters for each layer, and configuration time. An index linking tasks and configuration results is established for easy subsequent querying and tracing.
[0129] The system periodically verifies the configuration results, comparing whether the inference computing node types match the configured inference module combinations. If a mismatch is found, the system automatically triggers a reconfiguration process to ensure that the configuration results always match the node types and to guarantee the stability of the inference process.
[0130] Step S144: Embed the knowledge retrieval timing plan into the configured inference module combination to form an adapted inference process for each operation and maintenance inference task.
[0131] The timing embedding module integrates the knowledge retrieval timing plan with the inference module combination, inserting knowledge retrieval interfaces into the corresponding inference layer of the inference module combination according to the planned retrieval timing. An initial knowledge retrieval interface is inserted after the requirement parsing layer is completed to obtain basic device knowledge; a continuous knowledge retrieval interface is set in the knowledge association layer to dynamically obtain associated knowledge; and an on-demand knowledge retrieval interface is set in the multi-path inference layer and the direct inference layer to retrieve specific knowledge based on the inference progress.
[0132] The embedded, adapted inference process includes the complete sequence of inference logic and knowledge retrieval. Each knowledge retrieval interface is associated with a corresponding knowledge sub-repository and data type. After the process is configured, simulation tests are performed to verify the matching degree between the timing of knowledge retrieval and the inference progress, ensuring that knowledge can be accurately retrieved when needed for inference.
[0133] Step S145: Input the operation and maintenance requirement description information and equipment operation scenario identifier of each operation and maintenance service request into the corresponding adaptation inference process, and execute the large model inference processing through the allocated inference computing node. During the inference process, the matching knowledge sub-base data is called according to the knowledge call timing plan.
[0134] The inference execution module converts key information from the maintenance service request (maintenance requirement description and device operation scenario identifier) into inference input data and processes it according to the input requirements of the adapted inference process. The input data is then sent to the allocated inference computing node via a data transmission protocol. Upon receiving the data, the node initiates the adapted inference process.
[0135] The inference computing nodes execute large-scale model inference processing according to the workflow steps. When the trigger conditions set by the knowledge retrieval time sequence plan are reached, the nodes access the matching knowledge sub-repository through the association mapping table to obtain the required data. The knowledge data and inference input data are fused and processed to drive the inference logic to unfold layer by layer until all inference stages are completed. Intermediate data during the inference process is stored in real time in the node's temporary storage area for result verification and traceability.
[0136] Step S1451: Perform structured transformation on the operation and maintenance requirement description information of each operation and maintenance service request, decompose the unstructured operation and maintenance requirement description content into structured requirement data containing requirement type identifier, core demand points, and associated equipment parameters, and parse the equipment operation scenario identifier associated with the operation and maintenance service request, extract the equipment model information, operating environment parameters, and historical operation and maintenance record index contained in the equipment operation scenario identifier to form a standardized input package.
[0137] The structured transformation module uses natural language understanding technology to process unstructured operation and maintenance requirement descriptions. It extracts key information such as device names and fault symptoms through entity recognition and determines the requirement type identifier (such as fault handling or parameter configuration) through intent classification. The core requirements are broken down into quantifiable requirement indicators, and associated device parameters are extracted from the description or supplemented by querying through device identifiers to form structured requirement data.
[0138] The device operation scenario identifier parsing module splits the identifier string according to the encoding rules, extracting device model information (such as air conditioner model, display screen size), operating environment parameters (such as operating mode, ambient temperature range), and historical maintenance record index (pointing to the device's historical maintenance records). The structured requirement data is integrated with the parsed scenario information to form a standardized input package containing multiple data fields. The input package uses a unified data format to ensure that the inference calculation nodes can parse it correctly.
[0139] Step S1452: Transmit the standardized input package to the input interface of the corresponding adaptation inference process. According to the inference computing node type associated with the adaptation inference process, trigger the inference module initialization operation of the inference computing node. The adaptation inference process corresponding to the deep inference requirement node triggers the module loading of the multi-stage inference layer. The adaptation inference process corresponding to the basic inference requirement node triggers the module loading of the single-stage inference layer.
[0140] The data transmission module sends standardized input packets to the input interface adapted to the inference process via an encrypted channel. After verifying the integrity and format correctness of the input packets, the interface sends an initialization command to the inference computing node. The inference computing node initiates the corresponding module loading process based on its own type (deep or basic inference requirement node).
[0141] The deep reasoning requirement node loads all modules of the multi-stage reasoning layer, including the requirement parsing module, knowledge association module, multi-path reasoning engine, and result integration module. Each module is initialized according to preset parameters. The basic reasoning requirement node loads the requirement mapping module and the direct reasoning engine. During module initialization, the connection status with the knowledge sub-base is verified. After initialization, the node returns a ready signal, waiting for the reasoning start command.
[0142] Step S1453: Extract the knowledge call node and associated knowledge sub-base type corresponding to the current inference stage from the knowledge call timing plan, and send the knowledge call instruction to the inference computing node. The inference computing node establishes and matches the knowledge layer according to the knowledge call instruction to generate a single inference result based on the standardized inference input and knowledge sub-base data. The requirement mapping layer uses a pre-trained semantic mapping model to convert unstructured operation and maintenance requirement description information into standardized inference input vectors with a fixed format. The vector dimension matches the input dimension of the direct inference layer. The standardized inference input vector contains feature data of multiple dimensions such as requirement type features, equipment parameter features, and scenario features. Each dimension feature has been normalized to ensure that the data distribution meets the input requirements of the direct inference layer.
[0143] The direct inference layer loads a pre-trained lightweight inference model, which is generated based on a large amount of historical operation and maintenance case data. This model comprises three basic structures: an input layer, a hidden layer, and an output layer. The input layer receives standardized inference input vectors and maps the feature data to the output space through feature transformation and weight calculation in the hidden layer. The output layer uses the softmax activation function to generate probability distributions for different operation and maintenance solutions, selecting the solution with the highest probability as the single inference result. During inference, the direct inference layer continuously calls upon basic parameter data and rule bases in the knowledge sub-base to verify and correct the inference results, ensuring that the results comply with equipment operation specifications.
[0144] Step S1436: Record the configuration results of the inference module combination corresponding to each operation and maintenance inference task to ensure that the configuration results are compatible with the inference computing node type. The configuration result recording module stores detailed information about the inference module combination in the task configuration database, including task identifier, inference computing node type, inference module list, parameter settings of each module, and knowledge sub-base association. Each configuration result generates a unique configuration version number, which is bound to the inference task identifier for easy traceability and optimization later.
[0145] The database periodically performs statistical analysis on the configuration results to identify the matching efficiency of different inference computing node types and inference module combinations. When the hardware configuration of the inference computing node is updated or the inference model version is upgraded, the configuration result recording module automatically triggers a configuration adaptability check to ensure that the new configuration remains compatible with the node type.
[0146] Step S144: Embed the knowledge retrieval timing plan into the configured inference module combination to form an adapted inference flow for each operation and maintenance inference task. The timing embedding module parses each retrieval node in the knowledge retrieval timing plan and determines its insertion position in the inference module combination. For multi-stage inference layer combinations, knowledge retrieval nodes are embedded after the requirement parsing layer is completed, during the knowledge association layer, before the multi-path inference layer is started, and at the initial stage of the result integration layer, respectively; for single-stage inference layer combinations, knowledge retrieval nodes are embedded after the requirement mapping layer is completed and during the direct inference layer.
[0147] Each knowledge call node establishes a trigger association with its corresponding inference module. When the inference process reaches a specific stage of that module, the knowledge call instruction is automatically activated. The embedded, adapted inference process generates a visual flowchart, annotating the execution order of each inference module, the triggering timing of the knowledge call node, and the data interaction relationships between modules. The flowchart is stored in the process configuration library, serving as an operational guide for the inference computing nodes to execute inference tasks.
[0148] Step S145: Input the operation and maintenance requirement description information and equipment operation scenario identifier of each operation and maintenance service request into the corresponding adaptation inference process. The large model inference process is executed through the allocated inference computing nodes. During the inference process, matching knowledge sub-base data is called according to the knowledge call sequence plan. The data input module packages the operation and maintenance requirement description information and equipment operation scenario identifier into an input data set and transmits it to the starting node of the adaptation inference process through an interface. The inference computing nodes start each inference module according to the flowchart sequence of the adaptation inference process, gradually advancing the inference process.
[0149] During the inference process, when the process reaches the knowledge call node, a data request is automatically sent to the knowledge sub-repository to obtain the knowledge data required for the current inference stage. The knowledge data, after format conversion, is input into the inference module and fused with the feature data in the input dataset. The inference computation node records intermediate results, knowledge call records, and module running status in real time, ensuring the traceability of the inference process.
[0150] Step S1451: Perform structured transformation on the operation and maintenance requirement description information of each operation and maintenance service request. Decompose the unstructured operation and maintenance requirement description into structured requirement data containing requirement type identifiers, core demands, and associated equipment parameters. Simultaneously, parse the equipment operation scenario identifier associated with the operation and maintenance service request, extracting the equipment model information, operating environment parameters, and historical operation and maintenance record index contained within the equipment operation scenario identifier to form a standardized input package. The structured transformation module uses natural language processing technology to perform word segmentation and semantic parsing on the unstructured text, identifying requirement type identifiers (such as fault reporting, parameter adjustment, and status query).
[0151] The core requirement extraction module uses keyword weighting to filter out the core requirements in the maintenance needs description, such as "poor air conditioning cooling effect" and "insufficient display brightness." The associated equipment parameter extraction module identifies mentioned equipment parameters from the description, such as "set temperature 26℃" and "operating frequency 50Hz," forming a structured parameter list. The equipment operation scenario identifier parsing module splits the identifier string according to encoding rules, extracting equipment model information (e.g., "KFR-35GW / BP3"), operating environment parameters (e.g., "ambient temperature 32℃" and "humidity 60%)), and historical maintenance record indexes (e.g., "REC-20230615-002"). This data is then integrated with the structured requirements data into a standardized input package, which is encapsulated in JSON format to ensure data integrity and parsability.
[0152] Step S1452: The standardized input package is transmitted to the input interface of the corresponding adapted inference process. Based on the type of inference computing node associated with the adapted inference process, the inference module initialization operation of the inference computing node is triggered. For deep inference requirement nodes, the adapted inference process triggers multi-stage inference layer module loading; for basic inference requirement nodes, the adapted inference process triggers single-stage inference layer module loading. The data transmission module sends the standardized input package to the input interface of the adapted inference process through an encrypted channel. After verifying the integrity and format correctness of the input package, the interface sends an initialization command to the inference computing node. The inference computing node initiates the corresponding module loading process according to its own type (deep or basic inference requirement node).
[0153] The deep reasoning requirement node loads all modules of the multi-stage reasoning layer, including the requirement parsing module, knowledge association module, multi-path reasoning engine, and result integration module. Each module is initialized according to preset parameters. The basic reasoning requirement node loads the requirement mapping module and the direct reasoning engine. During module initialization, the connection status with the knowledge sub-base is verified. After initialization, the node returns a ready signal, waiting for the reasoning start command.
[0154] Step S1453: Extract the knowledge call node and associated knowledge sub-base type corresponding to the current inference stage from the knowledge call timing plan, and send the knowledge call instruction to the inference computing node. The inference computing node establishes a data transmission link between the knowledge call instruction and the matched knowledge sub-base. The timing extraction module parses the knowledge call timing plan in the order of the inference stages to determine the type of knowledge sub-base to be called in the current stage (such as fault handling knowledge sub-base, parameter standard knowledge sub-base) and the specific call node location.
[0155] The knowledge retrieval instruction includes the knowledge sub-repository identifier, the required knowledge type, data format requirements, and transmission encryption method. Upon receiving the instruction, the inference computing node sends a connection request to the knowledge management system through the knowledge resource scheduling interface. After verifying the request's validity, the knowledge management system allocates an access port for the knowledge sub-repository to the inference computing node and establishes a dedicated data transmission link. The link employs a two-way encryption mechanism to ensure the security and integrity of knowledge data transmission. A connection success confirmation signal is returned after the transmission link is established.
[0156] Step S1454: Through the established data transmission link, extract the operation and maintenance knowledge data required for the current inference stage from the matched knowledge sub-base. For the inference stage corresponding to the deep inference requirement node, prioritize extracting fault causal relationship data and complex operation and maintenance solution templates from the knowledge sub-base. For the inference stage corresponding to the basic inference requirement node, prioritize extracting basic parameter standard data and routine operation guidelines from the knowledge sub-base. The data extraction module performs a retrieval operation in the matched knowledge sub-base according to the knowledge type in the knowledge call instruction.
[0157] For the fault diagnosis phase of the deep inference requirement node, retrieve and extract causal correlation data (such as the correlation between compressor abnormal noise and insufficient refrigerant, and the impact curve of condenser blockage on heat dissipation efficiency) and complex operation and maintenance solution templates (such as the multi-component collaborative replacement operation process and the systemic fault troubleshooting step framework). For the parameter query phase of the basic inference requirement node, retrieve and extract basic parameter standard data (such as the air conditioner summer cooling temperature setting range and the display brightness adjustment parameter threshold) and routine operation guidelines (such as the standard steps for filter cleaning and the operation specifications for power restart). The extracted operation and maintenance knowledge data is packaged in a unified format and sent to the inference computing node through the data transmission link.
[0158] Step S1455: Using the inference computing node, the structured requirement data and extracted operation and maintenance knowledge data from the standardized input package are input into the initialized inference module to execute the first-stage inference operation and generate the first-stage inference intermediate result. The inference execution module converts the structured requirement data and operation and maintenance knowledge data into a feature vector form that the inference module can recognize. The structured requirement data vector includes requirement type features, equipment parameter features, and scenario features, while the operation and maintenance knowledge data vector includes knowledge association features, historical case features, and specification constraint features.
[0159] Two feature vectors are input into the initialized inference module. The module performs feature fusion and nonlinear transformation through a multi-layer neural network, calculating the correlation and influence weight between features. For the first stage (requirement parsing stage) of the deep inference requirement node, the output is a ranking of the importance of structured requirements and a knowledge matching score; for the first stage (requirement mapping stage) of the basic inference requirement node, the output is a standardized inference input vector and a knowledge association index. The intermediate results of the first stage inference are stored in the node's temporary cache, with a timestamp and module identifier.
[0160] Step S1456: Based on the next knowledge call node in the knowledge call sequence plan, repeat the above steps of knowledge sub-base data extraction and inference calculation until the inference stage corresponding to all nodes in the knowledge call sequence plan is completed. During this period, the intermediate inference results generated in each inference stage are temporarily stored in the temporary data storage area of the inference calculation node. The sequence advancement module monitors the completion status of the current inference stage. When the first stage of inference calculation is completed, it extracts the information of the next knowledge call node from the knowledge call sequence plan.
[0161] Following steps S1453 to S1455, the knowledge retrieval and inference operations of the next inference stage are executed sequentially until all inference stages corresponding to all knowledge retrieval nodes are completed. For example, a deep inference requirement node needs to complete four stages sequentially: requirement parsing, knowledge association, multi-path inference, and result integration. Each stage performs knowledge extraction and inference operations. The intermediate results of each stage's inference include feature data, operation parameters, and stage conclusions, which are stored in a temporary data storage area in stage order. The area adopts a partitioned storage strategy to ensure that data from different stages are not confused.
[0162] Step S1457: During the execution of each inference stage, the inference computation node compares the intermediate inference results of the current stage with the corresponding verification data in the knowledge sub-base to ensure that the core logic of the intermediate inference results and the verification data is consistent. If it is a deep inference requirement node, the multiple sets of intermediate inference results generated by the multi-path inference layer also need to be cross-validated to eliminate logical conflicts. The verification comparison module extracts verification data related to the current inference stage from the knowledge sub-base, such as the standard inference results of historical cases and the logical constraints of standardized operations.
[0163] The intermediate inference results are compared with the validation data using feature comparison and logical consistency checks. The deviation between the two is calculated, and if the deviation is within a preset threshold, the validation is considered successful. For multi-path inference layers requiring deep inference, different inference paths generate multiple sets of intermediate inference results. The cross-validation module calculates the logical similarity and conclusion consistency of each set of results, identifies and eliminates conflicting content, and weights and fuses divergent conclusions to form a unified intermediate result. Intermediate results that pass validation are marked as valid and allowed to proceed to the next inference stage; intermediate results that fail validation trigger a re-inference mechanism.
[0164] Step S1458: After completing all inference stages, the inference computing node integrates all intermediate inference results and organizes the data according to the output format requirements of the inference process to form an intermediate inference dataset containing a complete inference logic chain. This intermediate inference dataset is used to generate the operation and maintenance inference results corresponding to the operation and maintenance service requests. The result integration module reads all valid intermediate inference results in the temporary data storage area in the order of the inference stages and extracts the core conclusions, feature parameters, and knowledge reference records of each stage.
[0165] For nodes requiring deep reasoning, a result integration algorithm is used to weight and fuse the results of multi-path reasoning, constructing a reasoning logic chain based on logical connections. This chain includes the relationships between four stages: problem analysis, cause derivation, solution generation, and effect prediction. For nodes requiring basic reasoning, the results of each stage are arranged chronologically to form a linear reasoning logic chain. The intermediate reasoning dataset is organized in a structured table format, including fields such as reasoning stage number, core conclusion description, supporting evidence data, knowledge citation sources, and credibility score.
[0166] Step S146: Generate content containing the direction of operation and maintenance (O&M) operations, the basis for O&M operations, and related data of O&M operations through reasoning processing, and determine this content as the O&M reasoning result corresponding to the O&M service request. The result generation module parses the reasoning logic chain and core conclusions of the intermediate reasoning dataset, and extracts clear O&M operation directions, such as "replacing the air conditioner compressor", "adjusting the display brightness parameters", "cleaning the condenser heat sink", etc.
[0167] The operational and maintenance (O&M) operations are based on O&M knowledge data, historical case matching results, and logical reasoning conclusions referenced during the integrated reasoning process, demonstrating the rationality and scientific basis of the operational direction. The O&M operation-related data includes the equipment parameters, tool list, safety precautions, and expected performance indicators required to execute the operation. These three parts are combined according to a standardized format to generate a structured O&M reasoning result, which includes a reasoning process ID and a generation timestamp to ensure traceability.
[0168] Step S150: Combining the device operation scenario identifier associated with the operation and maintenance service request, the operation and maintenance inference result is converted into an operation and maintenance response instruction conforming to the operation specifications of the device operation scenario, and the operation and maintenance response instruction is sent to the corresponding operation and maintenance execution terminal. The instruction conversion module receives the operation and maintenance inference result and the device operation scenario identifier, performs format conversion and content adaptation on the inference result according to the operation specifications requirements corresponding to the scenario identifier, generates an operation and maintenance response instruction that can be directly executed by the operation and maintenance execution terminal, and sends it to the terminal device through the communication module.
[0169] Step S151: Parse the device operation scenario identifier associated with the maintenance service request to obtain the device operation specification version, device operation permission scope, and the operation interface type currently connected to the device corresponding to the device operation scenario identifier. The identifier parsing module splits the device operation scenario identifier string according to preset encoding rules and extracts the specification version field (e.g., "V2.3"), permission scope field (e.g., "ADMIN-01"), and interface type field (e.g., "RS485-02").
[0170] The device operation specification library is queried using the "Specification Version" field to retrieve the corresponding operation specification document, which includes terminology definitions, standard operating procedures, and instruction format requirements. The "Permission Scope" field is linked to the permission management system to determine the permitted functional scope and parameter adjustment permissions for the current maintenance operation, such as whether firmware upgrades or modification of core parameters are allowed. The "Interface Type" field is used to query the interface technical manual to clarify communication parameters such as data transmission protocol, baud rate, and verification method.
[0171] Step S152: Extract the operation direction, operation basis, and operation-related data from the operation and maintenance reasoning results, and provide a structured description of the operation direction according to the format requirements corresponding to the equipment operation specification version. The result extraction module separates three core parts from the operation and maintenance reasoning results, and the structured description module converts the operation direction into standardized operation terms by referring to the terminology system and expression requirements of the equipment operation specification version.
[0172] For example, "to improve the cooling effect of the air conditioner" can be transformed into "to adjust the operating parameters of the air conditioning system to improve cooling efficiency." The structured description includes three elements: the object of operation, the action of operation, and the goal of operation. Each element conforms to the definition standards in the specification version, ensuring consistent understanding of the operation direction across different terminal devices and operators.
[0173] Step S153: Based on the device operation permission range, filter the maintenance operation-related data to ensure that the content meets the permission requirements, and delete the related data that exceeds the device operation permission range. The permission filtering module compares the maintenance operation-related data with the device operation permission range item by item, retaining the content within the permission range, such as routine parameter adjustment data, basic operation tool lists, and general safety prompts.
[0174] Content exceeding the scope of permissions, such as core firmware upgrade packages, encryption parameter modification commands, and access permissions for advanced diagnostic tools, will be deleted and the filtering log will be recorded. The filtered related data will be labeled with permission identifiers, indicating the permission level and operation restrictions for each item, ensuring that the maintenance execution terminal can only access and execute operations within its authorized scope.
[0175] Step S154: Based on the type of the currently connected operation interface of the device, determine the instruction format of the maintenance response command. The instruction format includes the field structure corresponding to the interface protocol and the data transmission encoding method. The interface adaptation module queries the interface protocol specification according to the type of the currently connected operation interface of the device to determine the field structure of the command, such as the start identifier field, command type field, data length field, operation content field, verification field, and end identifier field.
[0176] The data transmission encoding method is selected based on the interface characteristics, such as ASCII encoding, hexadecimal encoding, or a custom binary encoding. The encoding method must ensure that the instruction content is unambiguous during transmission and that transmission efficiency is optimal. After the instruction format is determined, a format template is generated, which includes the length, data type, and padding rules for each field.
[0177] Step S155: Organize the structured description of the operation and maintenance direction, the basis for the operation and maintenance, and the filtered operation and maintenance related data according to the determined instruction format to form an operation and maintenance response instruction that conforms to the operation specifications of the equipment operation scenario. The instruction assembly module fills in the structured description of the operation and maintenance direction (up to the operation content field), the basis for the operation and maintenance (up to the additional description field), and the filtered operation and maintenance related data (up to the parameter data field) in the order of the fields in the format template.
[0178] The filled fields undergo format validation to ensure that the data type and length of each field conform to the template requirements. After successful validation, the value of the validated field is calculated and filled. The entire instruction is encoded using a defined data transmission encoding method to generate the final operation and maintenance response instruction. The instruction is stored in byte stream format for easy transmission through the operation interface.
[0179] Step S156: Confirm whether the field structure and data transmission encoding method in the maintenance response command fully match the type of the currently connected operation interface of the device, so that the maintenance response command can be correctly parsed by the maintenance execution terminal. The command verification module compares the field structure of the maintenance response command with the protocol specification corresponding to the operation interface type, and checks whether the field order, length and definition are consistent.
[0180] Simultaneously, the encoded instructions undergo decoding testing to verify whether the decoded content matches the original filled content, ensuring no data loss or distortion during the encoding process. If a mismatch or decoding anomaly is found during verification, the process returns to the instruction assembly module to readjust the format and encoding method; verified maintenance response instructions are marked as valid instructions and prepared to be sent to the maintenance execution terminal.
[0181] Step S157: The verified maintenance response command is sent to the corresponding maintenance execution terminal through the currently connected operation interface of the device. During the transmission process, the data transmission status is monitored to ensure that the command arrives completely at the terminal device. The communication module establishes a data transmission link according to the communication parameters (such as baud rate and timeout) corresponding to the operation interface type, and sends the maintenance response command in the form of a byte stream.
[0182] During transmission, the link status and data transmission progress are monitored in real time. If a transmission interruption or error occurs, a retransmission mechanism is automatically initiated until the command is successfully sent or the maximum number of retransmissions is reached. After the maintenance execution terminal receives the command, it returns an acknowledgment signal. Upon receiving the acknowledgment signal, the communication module records a command transmission success log, including the transmission time, terminal identifier, and command ID.
[0183] Figure 2 This application illustrates a large model inference optimization system 100 for operation and maintenance technical services, including a processor 1001, a memory 1003, and program code stored in the memory 1003. The processor 1001 executes the program code to implement the steps of the large model inference optimization method for operation and maintenance technical services.
[0184] Figure 2 The large-model inference optimization system 100 for operation and maintenance technical services shown includes a processor 1001 and a memory 1003. The processor 1001 and the memory 1003 are connected, for example, via a bus 1002. Optionally, the large-model inference optimization system 100 for operation and maintenance technical services may further include a transceiver 1004, which can be used for data interaction between this large-model inference optimization system for operation and maintenance technical services and other large-model inference optimization systems for operation and maintenance technical services, such as sending and / or receiving data. It should be noted that in actual scheduling, the transceiver 1004 is not limited to one, and the structure of this large-model inference optimization system 100 for operation and maintenance technical services does not constitute a limitation on the embodiments of this application.
[0185] The memory 1003 is used to store program code for executing the embodiments of this application, and its execution is controlled by the processor 1001. The processor 1001 is used to execute the program code stored in the memory 1003 to implement the steps shown in the foregoing method embodiments.
[0186] This application provides a computer-readable storage medium storing program code, which, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.
[0187] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application, without departing from the technical concept of this application, also fall within the protection scope of the embodiments of this application.
Claims
1. A large-scale model inference optimization method for operation and maintenance technical services, characterized in that, The method includes: Obtain a set of operation and maintenance service requests under the operation and maintenance technical service scenario. The set of operation and maintenance service requests contains multiple operation and maintenance service requests to be processed. Each operation and maintenance service request carries operation and maintenance requirement description information and associated equipment operation scenario identifier. Based on the operation and maintenance requirement description information and equipment operation scenario identifier of each operation and maintenance service request in the operation and maintenance service request set, the priority of the operation and maintenance inference task corresponding to each operation and maintenance service request is dynamically marked, and an operation and maintenance inference task priority list is generated. Based on the priority list of operation and maintenance inference tasks and the inference link requirements corresponding to each operation and maintenance inference task, the computing resources and operation and maintenance knowledge resources required for large model inference are allocated to generate a collaborative scheduling scheme for computing resources and operation and maintenance knowledge resources. Based on the aforementioned collaborative scheduling scheme for computing resources and operation and maintenance knowledge resources, an appropriate reasoning process is configured for each operation and maintenance reasoning task. The large model reasoning process is executed through the allocated computing resources to generate operation and maintenance reasoning results corresponding to the operation and maintenance service request. Combining the device operation scenario identifier associated with the operation and maintenance service request, the operation and maintenance reasoning result is transformed into an operation and maintenance response instruction that conforms to the operation specifications of the device operation scenario, and the operation and maintenance response instruction is sent to the corresponding operation and maintenance execution terminal; The step involves allocating computing resources and operational knowledge resources required for large-scale model inference based on the priority list of operational inference tasks and the inference chain requirements corresponding to each operational inference task, and generating a collaborative scheduling scheme for computing resources and operational knowledge resources, including: Analyze the inference chain requirements corresponding to each operation and maintenance inference task, and determine the type of inference computing node, the number of inference nodes, and the associated operation and maintenance knowledge domain required for each inference chain requirement; The available computing resource pool in the operation and maintenance technical service scenario is sorted out. The available computing resource pool contains multiple inference computing nodes. Each inference computing node is marked with its node service capabilities and the number of inference tasks it is currently carrying. Organize an operation and maintenance knowledge resource library for operation and maintenance technical service scenarios. The operation and maintenance knowledge resource library is divided into multiple knowledge sub-libraries according to the operation and maintenance knowledge domain. Each knowledge sub-library stores operation and maintenance cases, operation and maintenance specifications and equipment characteristic data of the corresponding domain. According to the operation and maintenance inference task priority list, inference computing nodes with service capabilities matching the inference link requirements are allocated to core operation and maintenance inference tasks in a priority manner, so that the number of inference tasks currently carried by the inference computing nodes corresponding to the core operation and maintenance inference tasks does not exceed the preset ratio of the node's maximum capacity. For each operation and maintenance inference task, a knowledge sub-base corresponding to the operation and maintenance knowledge domain is matched, and an association mapping between the inference computing node and the knowledge sub-base is established, so that the inference computing node can directly call the associated knowledge sub-base when performing inference processing; Record the allocation results of reasoning computing nodes and the matching results of knowledge sub-bases for each operation and maintenance reasoning task, and integrate them to form a collaborative scheduling scheme for computing resources and operation and maintenance knowledge resources.
2. The large-scale model inference optimization method for operation and maintenance technical services according to claim 1, characterized in that, Based on the operation and maintenance requirement description information and equipment operation scenario identifier of each operation and maintenance service request in the operation and maintenance service request set, the priority of the operation and maintenance inference task corresponding to each operation and maintenance service request is dynamically marked, and an operation and maintenance inference task priority list is generated, including: Extract the operation and maintenance task type from the operation and maintenance requirement description information of each operation and maintenance service request. The operation and maintenance task type includes equipment fault handling task, equipment parameter configuration task, and equipment operation status monitoring task. Analyze the device operation scenario identifier associated with each operation and maintenance service request to determine the system level to which the device belongs, the business link associated with the device, and the current operation stage of the device; Based on the association between the operation and maintenance task type and the system level to which the device belongs, the basic priority level is determined, wherein the basic priority level corresponding to the device fault handling task is higher than that of the device parameter configuration task, and the basic priority level corresponding to the device parameter configuration task is higher than that of the device operation status monitoring task. Based on the business link attributes associated with the device, the basic priority level is adjusted, and the priority level of the operation and maintenance inference task corresponding to the operation and maintenance service request of the core business link is adjusted upward. Based on the timeliness requirements of maintenance response during the current operating phase of the equipment, the priority level is further adjusted, and the priority level of maintenance inference tasks corresponding to maintenance service requests that meet the requirements of immediate response is adjusted upward again. Integrate the final priority level corresponding to each operation and maintenance service request, arrange all operation and maintenance inference tasks corresponding to the operation and maintenance service requests in order of priority level from first to second priority, and generate an operation and maintenance inference task priority list.
3. The large-scale model inference optimization method for operation and maintenance technical services according to claim 1, characterized in that, The analysis of each operational inference task corresponds to the inference chain requirements, determining the type of inference computing nodes, the number of inference nodes, and the associated operational knowledge domain required for each inference chain requirement, including: Each operation and maintenance inference task is broken down into multiple inference subtasks. Each inference subtask corresponds to a link in the inference chain, and the processing logic and data interaction requirements of each inference subtask are determined. Based on the processing logic of each inference subtask, determine the type of inference computing node required for that inference subtask. Inference subtasks involving deep fault causal analysis correspond to the deep inference requirement node type, while inference subtasks involving basic parameter queries correspond to the basic inference requirement node type. Count the number of inference subtasks included in each operation and maintenance inference task and the parallel processing possibility of each inference subtask to determine the number of inference nodes required for the operation and maintenance inference task. The more inference subtasks that can be processed in parallel, the more inference nodes are required. Extract the operation and maintenance knowledge content that needs to be referenced during the processing of each inference subtask, and determine the associated operation and maintenance knowledge domain according to the category to which the operation and maintenance knowledge content belongs. Among them, the fault diagnosis type inference subtask is associated with the fault processing knowledge domain, and the parameter configuration type inference subtask is associated with the parameter standard knowledge domain. The inference computing node type, number of inference nodes, and associated operation and maintenance knowledge domains corresponding to each operation and maintenance inference task are integrated to form an inference chain requirement description for each operation and maintenance inference task.
4. The large-scale model inference optimization method for operation and maintenance technical services according to claim 1, characterized in that, The step of prioritizing the allocation of inference computing nodes with service capabilities matching the inference chain requirements of core operation and maintenance inference tasks according to the operation and maintenance inference task priority list, so as to ensure that the number of inference tasks currently carried by the inference computing nodes corresponding to the core operation and maintenance inference tasks does not exceed a preset proportion of the node's maximum capacity, includes: Extract core operation and maintenance reasoning tasks from the operation and maintenance reasoning task priority list, and process each core operation and maintenance reasoning task in order of priority from first to second priority. For the core operation and maintenance inference task currently being processed, inference computing nodes whose service capabilities match the inference link requirements of the operation and maintenance inference task are selected from the available computing resource pool to form a candidate inference computing node set; Query the number of inference tasks currently carried by each candidate inference computing node in the candidate inference computing node set, and calculate the ratio of the number of inference tasks currently carried by each candidate inference computing node to the maximum capacity of that node; Candidate inference computing nodes that are currently carrying inference tasks and whose maximum capacity does not exceed a preset ratio threshold are selected as target inference computing nodes. Assign the currently processed core operation and maintenance inference tasks to the target inference computing node, and update the number of inference tasks currently carried by the target inference computing node; Repeat the above steps to complete the allocation of inference computing nodes for all core operation and maintenance inference tasks, and then allocate inference computing nodes for ordinary operation and maintenance inference tasks in the same way.
5. The large-scale model inference optimization method for operation and maintenance technical services according to claim 1, characterized in that, The method, based on the collaborative scheduling scheme of computing resources and operation and maintenance knowledge resources, configures an appropriate inference process for each operation and maintenance inference task, executes large model inference processing through the allocated computing resources, and generates operation and maintenance inference results corresponding to the operation and maintenance service request, including: Extract the allocation results of reasoning computing nodes and the matching results of knowledge sub-bases for each operation and maintenance reasoning task from the collaborative scheduling scheme of computing resources and operation and maintenance knowledge resources. Based on the matching results of the knowledge sub-base, determine the type of operation and maintenance knowledge data to be called and the timing of the call for each operation and maintenance reasoning task during the reasoning process, and generate a knowledge call sequence plan; Based on the inference computing node type corresponding to the inference computing node allocation result, configure the inference module combination for each operation and maintenance inference task. The deep inference requirement node is configured with an inference module combination containing a multi-stage inference layer, and the basic inference requirement node is configured with an inference module combination containing a single-stage inference layer. The knowledge retrieval timing plan is embedded into the configured inference module combination to form an adapted inference process for each operation and maintenance inference task; Input the operation and maintenance requirement description information and equipment operation scenario identifier of each operation and maintenance service request into the corresponding adaptation reasoning process, and execute the large model reasoning processing through the allocated reasoning computing nodes. During the reasoning process, the matching knowledge sub-base data is called according to the knowledge call sequence plan. The reasoning process generates content containing the direction of operation and maintenance, the basis for operation and maintenance, and related data of operation and maintenance. This content is then determined as the operation and maintenance reasoning result corresponding to the operation and maintenance service request.
6. The large-scale model inference optimization method for operation and maintenance technical services according to claim 5, characterized in that, The configuration of the inference module combination for each operation and maintenance inference task based on the inference computing node type corresponding to the inference computing node allocation result includes: Identify the type of inference computing node in the inference computing node allocation result corresponding to each operation and maintenance inference task, and determine whether the inference computing node type is a deep inference requirement node or a basic inference requirement node. If the inference computing node type is a deep inference requirement node, configure a multi-stage inference layer combination for the corresponding operation and maintenance inference task, including a requirement parsing layer, a knowledge association layer, a multi-path inference layer, and a result integration layer. The requirement parsing layer is used to structure the description information of operation and maintenance requirements; the knowledge association layer is used to establish the relationship between structured requirements and knowledge sub-base data; the multi-path reasoning layer is used to generate multiple sets of intermediate reasoning results through different reasoning paths; and the result integration layer is used to fuse multiple sets of intermediate reasoning results. If the inference computing node type is a basic inference requirement node, configure a single-stage inference layer combination that includes a requirement mapping layer and a direct inference layer for the corresponding operation and maintenance inference task. The requirement mapping layer is used to map the operation and maintenance requirement description information into standardized reasoning input, and the direct reasoning layer is used to generate a single reasoning result based on the standardized reasoning input and knowledge sub-base data. Record the configuration results of the inference module combination corresponding to each operation and maintenance inference task so that the configuration results are adapted to the inference computing node type.
7. The large-model inference optimization method for operation and maintenance technical services according to claim 5, characterized in that, The process involves inputting the operation and maintenance requirement description information and equipment operation scenario identifier of each operation and maintenance service request into the corresponding adaptation inference process. Large-scale model inference processing is then executed through the allocated inference computing nodes. During the inference process, matching knowledge sub-base data is invoked according to the knowledge invocation sequence plan, including: The operation and maintenance requirement description information of each operation and maintenance service request is processed by structure transformation. The unstructured operation and maintenance requirement description content is decomposed into structured requirement data containing requirement type identifier, core requirements, and associated equipment parameters. At the same time, the equipment operation scenario identifier associated with the operation and maintenance service request is parsed, and the equipment model information, operating environment parameters, and historical operation and maintenance record index contained in the equipment operation scenario identifier are extracted to form a standardized input package. The standardized input package is transmitted to the input interface of the corresponding adaptation inference process. According to the inference computing node type associated with the adaptation inference process, the inference module initialization operation of the inference computing node is triggered. The adaptation inference process corresponding to the deep inference requirement node triggers the module loading of the multi-stage inference layer, and the adaptation inference process corresponding to the basic inference requirement node triggers the module loading of the single-stage inference layer. Extract the knowledge call node and associated knowledge sub-base type corresponding to the current inference stage from the knowledge call timing plan, send the knowledge call instruction to the inference computing node, and the inference computing node establishes a data transmission link between the knowledge call instruction and the matching knowledge sub-base. Through the established data transmission link, the operation and maintenance knowledge data required for the current inference stage are extracted from the matching knowledge sub-base. The inference stage corresponding to the deep inference requirement node prioritizes the extraction of fault causal relationship data and complex operation and maintenance solution templates from the knowledge sub-base. The inference stage corresponding to the basic inference requirement node prioritizes the extraction of basic parameter standard data and routine operation guidelines from the knowledge sub-base. The structured requirement data and extracted operation and maintenance knowledge data in the standardized input package are input into the initialized inference module using the inference computing node to perform the first stage of inference operation and generate the first stage inference intermediate result. According to the next knowledge call node in the knowledge call sequence plan, repeat the above knowledge sub-base data extraction and reasoning operation steps until the reasoning stage corresponding to all nodes in the knowledge call sequence plan is completed. During this period, the intermediate reasoning results generated in each reasoning stage are temporarily stored in the temporary data storage area of the reasoning calculation node. During the execution of each inference stage, the inference computing node compares and correlates the intermediate inference results of the current stage with the corresponding verification data in the knowledge sub-base to ensure that the core logic of the intermediate inference results and the verification data is consistent. If it is a deep inference requirement node, it is also necessary to cross-verify the multiple sets of intermediate inference results generated by the multi-path inference layer to eliminate logical conflicts. After completing all inference stages, the inference computing node integrates all intermediate inference results and organizes the data according to the output format requirements of the inference process to form an intermediate inference dataset containing a complete inference logic chain. The intermediate inference dataset is used to generate operation and maintenance inference results corresponding to operation and maintenance service requests.
8. The large-scale model inference optimization method for operation and maintenance technical services according to claim 1, characterized in that, The step of combining the device operation scenario identifier associated with the operation and maintenance service request and converting the operation and maintenance inference result into an operation and maintenance response instruction that conforms to the operation specifications of the device operation scenario includes: Parse the device operation scenario identifier associated with the operation and maintenance service request, and obtain the device operation specification version, device operation permission scope and the operation interface type currently connected to the device corresponding to the device operation scenario identifier; Extract the operation and maintenance direction, operation and maintenance basis, and operation and maintenance related data from the operation and maintenance reasoning results, and describe the operation and maintenance direction in a structured manner according to the format requirements corresponding to the equipment operation specification version. Based on the device operation permission scope, filter the content in the operation and maintenance operation related data that meets the permission requirements, and delete the related data that exceeds the device operation permission scope; Based on the type of operation interface currently connected to the device, determine the instruction format of the maintenance response command. The instruction format includes the field structure corresponding to the interface protocol and the data transmission encoding method. The structured description of the operation and maintenance direction, the basis for operation and maintenance, and the filtered operation and maintenance related data are organized according to the determined instruction format to form operation and maintenance response instructions that conform to the operation specifications of the equipment operation scenario. Confirm that the field structure and data transmission encoding method in the maintenance response command are completely compatible with the type of the operation interface currently connected to the device, so that the maintenance response command can be correctly parsed by the maintenance execution terminal.
9. A large-scale model inference optimization system for operation and maintenance technical services, characterized in that, The method includes a processor and a computer-readable storage medium storing machine-executable instructions that, when executed by the processor, implement the large model inference optimization method for operation and maintenance technical services as described in any one of claims 1-8.
Citation Information
Patent Citations
Deployment method and device of reasoning service and processor readable storage medium
CN114745264A
Computing task scheduling method and system applied to numerical control machining system
CN120355196A