Operation and maintenance assisting method, program product, electronic equipment and storage medium

By combining multimodal recognition technology with visual and code structure, personalized operation and maintenance interface prompts and assistance modes are generated, which solves the problems of low efficiency and poor accuracy in operation and maintenance in existing technologies, achieves targeted assistance, and improves user experience.

CN122027465APending Publication Date: 2026-05-12BEIJING TOPSEC NETWORK SECURITY TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING TOPSEC NETWORK SECURITY TECH
Filing Date
2026-02-02
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies cannot provide targeted assistance information for operations and maintenance personnel, resulting in low efficiency and poor accuracy in operations and maintenance, especially prone to errors in complex tasks.

Method used

By using multimodal recognition technology, combined with visual and code structure recognition of the operation and maintenance interface, multimodal fusion features are generated. Combined with user context information and operation and maintenance knowledge base, personalized prompts and assistance modes are provided, including dynamic demonstrations and automated operations.

Benefits of technology

It improves the efficiency and accuracy of operation and maintenance, adapts to different user skill levels, reduces errors, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122027465A_ABST
    Figure CN122027465A_ABST
Patent Text Reader

Abstract

The invention provides an operation and maintenance assisting method, a program product, electronic equipment and a storage medium, and the method comprises the steps: obtaining an operation and maintenance interface accessed by a user, carrying out the multi-modal recognition of the operation and maintenance interface, and generating a multi-modal fusion feature; the multi-modal identification comprises visual identification and element structure identification; based on the multi-modal fusion features, the user context information and an operation and maintenance knowledge base, generating prompt content corresponding to the operation and maintenance interface, and displaying the prompt content to the user; the user context information comprises a user capability portrait determined based on user historical operation data; acquiring an assistance request of a user; the assistance request comprises an assistance mode; according to the multi-mode fusion feature and the user context information, generating guide information corresponding to the assistance mode; the guide information is used for assisting the user in operation and maintenance. Therefore, users with different experiences can obtain more appropriate support, the provided information is more targeted, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to an auxiliary operation and maintenance method, program product, electronic device, and storage medium. Background Technology

[0002] In modern network operations and maintenance environments, web management interfaces have become the core platform for operations and maintenance personnel to configure devices, manage services, and deploy applications. These interfaces typically contain numerous configuration options and multi-level menus. However, when users encounter problems during operations and maintenance, existing technologies can only provide fixed information and cannot offer more targeted assistance. Summary of the Invention

[0003] The purpose of this application is to provide an auxiliary operation and maintenance method, program product, electronic device and storage medium to improve the above-mentioned problems.

[0004] In a first aspect, embodiments of this application provide an assisted operation and maintenance method, comprising: acquiring an operation and maintenance interface accessed by a user; performing multimodal recognition on the operation and maintenance interface to generate multimodal fusion features; multimodal recognition including visual recognition and element structure recognition; generating prompt content corresponding to the operation and maintenance interface based on the multimodal fusion features, user context information, and an operation and maintenance knowledge base, and displaying the prompt content to the user; user context information including a user capability profile determined based on the user's historical operation data; acquiring the user's assistance request; the assistance request including an assistance mode; generating guidance information corresponding to the assistance mode according to the multimodal fusion features and user context information; the guidance information being used to assist the user in performing operation and maintenance operations.

[0005] In the above implementation process, by combining visual and code structure multimodal recognition, various styles of operation and maintenance interfaces can be accurately understood, and context-appropriate, conversational prompts can be generated, improving the efficiency and accuracy of operation and maintenance operations. Based on multimodal fusion features and user context information, guidance information corresponding to the assistance mode is generated, and the depth of prompts and assistance is dynamically adjusted according to the user's historical operation level, so that users with different experience can receive more appropriate support, the information provided is more targeted, and the user experience is improved.

[0006] Optionally, in this embodiment of the application, multimodal recognition is performed on the operation and maintenance interface to generate multimodal fusion features, including: using a trained convolutional neural network to identify visual elements in the operation and maintenance interface and extracting visual feature vectors; parsing the document object model tree of the operation and maintenance interface through the DOM parsing channel and extracting structural feature vectors of interface elements; obtaining the context feature vector of the operation and maintenance interface; the context feature vector is used to reflect the operating environment of the operation and maintenance interface; and fusing the visual feature vector, structural feature vector and context feature vector to obtain multimodal fusion features.

[0007] In the aforementioned implementation process, by combining visual recognition and DOM structure analysis, the system can not only "see" the appearance of interface elements but also understand their roles in page functionality. This overcomes the limitations of a single recognition method and provides excellent adaptability to various styles of operation and maintenance interfaces. Through contextual feature fusion, the system's understanding is not isolated but rather combines user identity and operational scenarios. The generated multimodal fusion features serve as a unified semantic representation, making subsequent prompt content generation and operational guidance more accurate and reliable.

[0008] Optionally, in this embodiment of the application, obtaining the context feature vector of the operation and maintenance interface includes: obtaining context information reflecting the operating environment of the operation and maintenance interface; the context information includes at least one of user role, current operation sequence and interface state; and generating a context feature vector based on the context information.

[0009] In the above implementation process, by generating contextual feature vectors, the specific environment in which the user operates can be dynamically perceived and quantified, making the assistive functions no longer rigid and unchanging. This allows the system to understand the potential different needs of users with different roles (such as novices and experts), to predict the user's subsequent intentions based on the steps the user has already completed, and to identify the interface area currently being focused on by the user. Based on the generated contextual feature vectors, subsequent prompts and assistance guidance can be closely integrated with the actual scenario.

[0010] Optionally, in this embodiment of the application, before generating the prompt content corresponding to the operation and maintenance interface based on multimodal fusion features, user context information, and operation and maintenance knowledge base, the method further includes: obtaining user historical operation data; the user historical operation data includes operation sequence, operation result, and operation time; and determining the user capability profile based on the user historical operation data and preset weights.

[0011] In the aforementioned implementation process, the system learns from users' past operations to automatically build personalized profiles reflecting their true skill levels. This allows for intelligent adjustments to the depth and method of subsequent assistance based on each user's historical proficiency. For users exhibiting high success rates and efficiency in their records, the system can reduce basic prompts and provide more efficient advanced techniques or automation options; for users showing difficulty or a tendency to make mistakes, the system provides more detailed and progressive guidance. This personalized adaptation based on objective historical data improves the speed and accuracy of various users in completing maintenance tasks, enhancing the user experience.

[0012] Optionally, in this embodiment of the application, the prompt content corresponding to the operation and maintenance interface is generated based on multimodal fusion features, user context information, and operation and maintenance knowledge base. This includes: using multimodal fusion features as a semantic index to query the operation and maintenance knowledge base and retrieve planning information corresponding to the operation and maintenance interface; the planning information includes concept explanations, parameter ranges, constraints, and risk warning information; determining the personalized style of the prompt content based on the user context information; and organizing the planning information according to the personalized style to generate the prompt content.

[0013] In the above implementation process, multimodal fusion features are used for knowledge retrieval, ensuring that the prompts accurately correspond to the interface elements currently seen by the user, providing accurate prompts and improving efficiency. Personalized styles are determined by combining user context information, allowing the generated prompts to vary in detail and expression depending on the individual and the scenario. This reduces redundant information interference for experienced users and minimizes errors for novice users due to insufficient guidance, improving the efficiency and accuracy of configuration and troubleshooting operations.

[0014] Optionally, in this embodiment, the assistance mode includes at least a first assistance mode and a second assistance mode; generating guidance information corresponding to the assistance mode includes: if it is the first assistance mode, then according to the multimodal fusion features and user context information, retrieving matching dynamic demonstration content, the dynamic demonstration content being an interactive micro-video generated based on the real-time interface state and reflecting the actual configuration process; if it is the second assistance mode, then analyzing the configuration form of the operation and maintenance interface, identifying the input parameter list and the deduced parameter list; generating a parameter collection interface to guide the user to input parameters based on the input parameter list; and automatically generating the parameters required in the deduced parameter list based on the parameters input by the user, the multimodal fusion features, and the operation and maintenance knowledge base.

[0015] In the above implementation process, the dynamic demonstration content provided by the first assistance mode transforms abstract operation steps into intuitive and visual animations, enabling users to quickly imitate and learn the correct operation methods. This is especially suitable for procedural and step-by-step tasks, reducing the learning cost. The second assistance mode goes a step further, intelligently analyzing forms and distinguishing parameter types, freeing users from tedious, repetitive, and error-prone manual filling. Users only need to focus on core decision parameters, and the system can automatically and accurately complete the filling and calculation of the remaining large number of parameters. These two methods shorten the time required to complete complex configuration tasks and improve the efficiency and reliability of operation and maintenance work.

[0016] Optionally, in this embodiment of the application, after automatically generating the required parameters in the derivation parameter list, the method further includes: verifying the filled parameters through a verification engine; wherein the verification engine integrates an operation and maintenance knowledge base; and generating a warning message and performing a rollback operation if the verification result fails to pass the characterization.

[0017] In the aforementioned implementation process, the verification engine utilizes an integrated knowledge base for real-time checks. It can detect formatting errors, out-of-bounds values, configuration conflicts, or potential security risks in parameter input before user submission or system application configuration, and immediately provide clear warnings. This allows errors to be corrected at the earliest stage, avoiding potential system failures, service interruptions, or security vulnerabilities caused by incorrect configurations. The automatic rollback operation quickly and cleanly undoes erroneous changes, restoring the interface to a safe state, eliminating the hassle of manual investigation and rollback.

[0018] Optionally, in this embodiment of the application, after guiding the user to input parameters based on the input parameter list on the parameter collection interface, the method further includes: calling a pre-recorded abstract process description, which contains semantic operation intent; and instantiating the abstract process description into an operation sequence based on the parameters input by the user and the parameter mapping graph, so as to generate the parameters required in the derivation parameter list.

[0019] In the above implementation process, calling pre-recorded abstract process descriptions allows the system to reuse validated standard operating procedures, improving the correctness and reliability of the operation methods, while reducing the tediousness of repeatedly recording fixed scripts for each different interface style. Semantic operation intent recording makes the process template independent of specific interface implementation details, enhancing adaptability. This reduces the workload for users when filling out complex forms and improves processing efficiency.

[0020] Optionally, in this embodiment of the application, the method further includes: acquiring interactive behavior data of the user during the operation based on the guidance information; and using successful operations in the interactive behavior data as positive feedback samples to optimize the fusion model used to generate multimodal fusion features.

[0021] In the above implementation process, real user operation data after receiving guidance is obtained, providing a basis for analyzing the system's assistance effectiveness and user behavior patterns. Successful operation cases are used as positive feedback to optimize the core model for generating multimodal fusion features, enabling the system to become increasingly accurate over time.

[0022] Secondly, embodiments of this application also provide an assisted operation and maintenance device, comprising: a fusion module, used to acquire the operation and maintenance interface accessed by the user, perform multimodal recognition on the operation and maintenance interface, and generate multimodal fusion features; the multimodal recognition includes visual recognition and element structure recognition; a prompt information module, used to generate prompt content corresponding to the operation and maintenance interface based on the multimodal fusion features, user context information, and an operation and maintenance knowledge base, and display the prompt content to the user; the user context information includes a user capability profile determined based on the user's historical operation data; an acquisition request module, used to acquire the user's assistance request; the assistance request includes an assistance mode; and an assistance module, used to generate guidance information corresponding to the assistance mode based on the multimodal fusion features and user context information; the guidance information is used to assist the user in performing operation and maintenance operations.

[0023] Thirdly, embodiments of this application also provide a computer program product, including computer program instructions, which are executed by a processor to perform the method provided in the first aspect or any implementation thereof.

[0024] Fourthly, embodiments of this application also provide an electronic device, including: a processor and a memory, the memory storing computer program instructions, which are executed by the processor to perform the method provided in the first aspect or any implementation thereof.

[0025] Fifthly, embodiments of this application also provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, perform the method provided in the first aspect or any implementation thereof.

[0026] The operation and maintenance assistance method, program product, electronic device, and storage medium provided in this application, through multimodal recognition combining visual and code structure, can accurately understand various styles of operation and maintenance interfaces and generate context-appropriate, conversational prompts, improving the efficiency and accuracy of operation and maintenance operations. Based on multimodal fusion features and user context information, guidance information corresponding to the assistance mode is generated. The depth of prompts and assistance is dynamically adjusted according to the user's historical operational level, ensuring that users with varying experience receive more appropriate support, providing more targeted information, and enhancing the user experience. Attached Figure Description

[0027] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0028] Figure 1 A flowchart illustrating an assisted operation and maintenance method provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of the maintenance assistance device provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0029] The embodiments of the technical solution of this application will now be described in detail with reference to the accompanying drawings. These embodiments are only used to more clearly illustrate the technical solution of this application and are therefore merely examples, and should not be used to limit the scope of protection of this application.

[0030] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this application.

[0031] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.

[0032] In modern network operations and maintenance (O&M) environments, web management interfaces have become the core platform for O&M personnel to configure devices, manage services, and deploy applications. However, with the increasing complexity of IT systems, these O&M interfaces typically contain numerous configuration options and multi-level menus, posing significant challenges for users, especially non-O&M personnel. Existing O&M platforms generally suffer from high learning costs, complex operating procedures, and high error risks, requiring users to repeatedly consult technical documentation or seek professional support, severely reducing O&M work efficiency.

[0033] Currently, solutions on the market mainly fall into two categories: one is basic prompting systems, such as simple tooltips and help documentation links. These systems can only provide static information and cannot provide dynamic guidance based on context. The other is video demonstration systems, which guide users through pre-recorded operation videos, but lack interactivity and real-time capabilities. These solutions cannot adapt to different user skill levels and real-time operating environments. Especially when handling complex operation and maintenance tasks, users still need to manually complete all operation steps, which can easily lead to system failures or security vulnerabilities due to operational errors.

[0034] This application's embodiments, by combining visual and code structure multimodal recognition, can accurately understand various styles of operation and maintenance interfaces and generate context-appropriate, conversational prompts, improving the efficiency and accuracy of operation and maintenance operations. Based on multimodal fusion features and user context information, guidance information corresponding to the assistance mode is generated. The depth of prompts and assistance is dynamically adjusted according to the user's historical operational level, ensuring that users with varying experience receive more suitable support, providing more targeted information, and enhancing the user experience.

[0035] Please see Figure 1 The illustrated diagram shows a flowchart of an assisted operation and maintenance method provided in an embodiment of this application. The assisted operation and maintenance method provided in this application can be applied to electronic devices, which may include physical devices such as servers, PCs, tablets, or smartphones, or virtual devices such as virtual machines or containers. The electronic device can be a single device, a combination of multiple devices, or a cluster of a large number of devices. The assisted operation and maintenance method may include: Step S110: Obtain the operation and maintenance interface accessed by the user, perform multimodal recognition on the operation and maintenance interface, and generate multimodal fusion features; multimodal recognition includes visual recognition and element structure recognition.

[0036] Step S120: Based on multimodal fusion features, user context information, and operation and maintenance knowledge base, generate prompt content corresponding to the operation and maintenance interface and display the prompt content to the user; user context information includes user capability profile determined based on user historical operation data.

[0037] Step S130: Obtain the user's assistance request; the assistance request includes the assistance mode.

[0038] Step S140: Generate guidance information corresponding to the assistance mode based on the multimodal fusion features and user context information; the guidance information is used to assist users in performing operation and maintenance operations.

[0039] In step S110, visual recognition specifically processes pixel-level information of the interface. Real-time screenshots of the user-accessed operations and maintenance interface are taken via browser extensions or injected scripts. A lightweight convolutional neural network (CNN) model pre-trained for the operations and maintenance interface is used to analyze the screenshots. The model detects and locates various visual elements in the interface, such as buttons, input boxes, drop-down menus, topology map icons, dashboard charts, etc., and extracts a high-dimensional visual feature vector for each detected region. The visual feature vector encodes the visual appearance of the elements, such as color, shape, texture, and embedded text information extracted through optical character recognition (OCR).

[0040] Element structure recognition is the identification of the internal code structure of an interface. This is achieved by accessing and parsing the Document Object Model (DOM) tree of the webpage. For example, this involves listening to dynamic changes in the DOM to obtain the tag type, HTML attributes, hierarchical position in the DOM tree, and parent-child relationships of each interface element. Based on the information obtained, a structural feature vector is generated, which describes the role of each element in the interface's functional logic.

[0041] The extracted visual and structural feature vectors, along with a contextual feature vector reflecting the operating environment, are input into a cross-modal attention fusion network. The contextual feature vector encodes information such as the currently logged-in user's role, the user's recent action sequence, and the currently active tab or pop-up state. The attention network dynamically evaluates the importance of the three features in the current scene and assigns them appropriate weights. For example, for graphical topology maps, the weight of visual features is increased; for standardized forms, the weight of structural features is dominant. A unified multimodal fusion feature can be output through weighted summation. The multimodal fusion feature integrates visual, structural, and contextual information, forming the most comprehensive and accurate semantic representation of the target interface elements.

[0042] In step S120, the multimodal fusion features are used as semantic fingerprints to query the operation and maintenance knowledge base. The operation and maintenance knowledge base can be organized in a graph format, containing entities such as device models, configuration parameters, service protocols, and common risks, as well as their relationships. The operation and maintenance knowledge base can also be in the form of tables or documents. By querying the operation and maintenance knowledge base, content related to elements in the operation and maintenance interface is retrieved, such as the conceptual definition of the configuration item, its legal value range, its constraint relationship with other parameters, and the faults that may be caused by improper configuration.

[0043] The organization of this content is then determined by the user's skill profile within the user's context information. This skill profile is calculated by continuously analyzing historical user operation data, such as success rate, efficiency, and error types, and applying an algorithm model with time decay. The user skill profile dynamically reflects the user's proficiency level for different operational tasks. For novice users, detailed prompts with explanations of basic concepts can be generated; for expert users, only advanced options, risk warnings, or efficiency tips can be displayed.

[0044] One implementation method is to display prompts to users via smart floating boxes, or through tables and other graphical representations. By monitoring the user's mouse hover or focus behavior, when the system detects that the user intends to seek help on a certain interface element, it immediately triggers the rendering and display of prompts. The generated prompts can be simple text or they can be assembled from logical knowledge fragments, such as the sequence of steps or causal relationships, using natural language and organized in a multi-layered structure for better guidance.

[0045] In step S130, the assistance request is initiated by the user by clicking a specific button on the interface, selecting a right-click menu option, or directly selecting the prompt content in the floating box displayed in S120. The assistance request includes the assistance mode selected by the user. For example, the user may click the "Watch Demo" button (corresponding to the first assistance mode) or click the "Auto Fill" button (corresponding to the second assistance mode). By listening to these preset page design (UI) component interaction events, the user's assistance intent and mode selection are captured, thereby activating the corresponding deep assistance process.

[0046] In step S140, different guidance is executed according to the mode selected by the user. For example, if the user selects the first assistance mode, the multimodal fusion features are used as context to retrieve or generate a dynamic demonstration content in real time. The dynamic demonstration content can be a screen operation micro-video or animation focusing on the current operation area, dynamically showing the complete key steps such as mouse clicks and keyboard inputs to complete a similar configuration task from the current state, providing the user with an intuitive operation example that can be imitated.

[0047] If the user selects the second assistance mode, an automated assistance process can be initiated. The current configuration form is analyzed to automatically distinguish which parameters require user input and which can be automatically derived based on rules, a knowledge base, or previously entered parameters. A parameter collection interface is generated, guiding the user to fill in only key information. A predefined semantic operation flow is then invoked. This flow, recorded during multimodal fusion features, captures the intent sequence of "what operation to perform on what semantic element." Combined with the key parameters just entered by the user, the abstract flow is instantiated into a set of automated operation sequences that can be precisely executed on the current interface. During this process, a validation engine integrated with a knowledge base simultaneously checks the validity of the parameters and the compliance of the configuration. By automatically executing this sequence, the remaining fields are filled in, options are selected, and the configuration is submitted, thus transforming the guidance information into tangible automated actions that reduce operational burden and errors.

[0048] In the implementation of the above embodiments: by combining visual and code structure multimodal recognition, various styles of operation and maintenance interfaces can be accurately understood, and context-appropriate, conversational prompts can be generated, improving the efficiency and accuracy of operation and maintenance operations. Based on multimodal fusion features and user context information, guidance information corresponding to the assistance mode is generated, and the depth of prompts and assistance is dynamically adjusted according to the user's historical operation level, so that users with different experience can obtain more appropriate support, the information provided is more targeted, and the user experience is improved.

[0049] Optionally, in this embodiment of the application, multimodal recognition is performed on the operation and maintenance interface to generate multimodal fusion features, including: A trained convolutional neural network (CNN) is used to identify visual elements in the operations and maintenance (O&M) interface and extract visual feature vectors. The trained CNN is a deep learning model pre-trained using a large number of labeled screenshots of the O&M interface. When a user opens the O&M interface, a real-time screenshot of the current viewport is captured via a browser plugin and input into the network. The network's front-end backbone (e.g., MobileNet) is responsible for extracting general features from the image, while subsequent detection heads are specifically used to identify and locate typical visual elements in the O&M scenario, such as buttons, input boxes, charts, and topology icons. For each identified element region, the network extracts a fixed-dimensional numerical array, i.e., a visual feature vector, from its deep feature map. This vector encodes the visual appearance semantics of the element, such as color distribution, shape contour, and the visual form of embedded text, providing a visual basis for subsequent semantic understanding.

[0050] The document object model (DOM) tree of the operations and maintenance (O&M) interface is parsed through a DOM parsing channel to extract structural feature vectors of the interface elements. The DOM parsing channel refers to the processing path through which a script injected into the O&M page accesses and analyzes the underlying DOM structure. This script dynamically monitors and parses the DOM tree of the O&M interface, accurately obtaining information about each HTML node. The process of extracting structural feature vectors from interface elements involves collecting structured information such as the target element's tag name, type, various attributes, depth and position in the DOM tree, and its relationships with parent, child, and sibling nodes. This discrete structured information is then mapped into a continuous, dense structural feature vector through an embedding layer or encoder (such as a self-attention-based Transformer encoder). The structural feature vector characterizes the element's role in the interface's functional logic, such as whether it is a submittable form button, a required text input box, or a cell displaying read-only information.

[0051] The context feature vector of the operations and maintenance (O&M) interface is obtained; this vector reflects the operational environment of the O&M interface. It is used to obtain the instantaneous state of the user's current operational environment. By collecting various signals reflecting the operational environment of the O&M interface, these signals constitute contextual information, mainly including: the currently logged-in user role (e.g., administrator, auditor), the sequence of operations executed in this session, the currently active browser tab or pop-up, and the stage of the task flow. This multi-dimensional, discrete, or serialized contextual information is transformed into a unified digital identifier, then projected onto the same vector space through an independent embedding layer, and concatenated or weighted to generate a comprehensive contextual feature vector. The contextual feature vector provides the system with a background framework for understanding "who operates under what circumstances."

[0052] Visual feature vectors, structural feature vectors, and contextual feature vectors are fused to obtain a multimodal fusion feature. These three vectors are then input into a cross-modal attention fusion network. The core of this network is a learnable attention mechanism. First, the three vectors are linearly transformed, and then their mutual attention weights are calculated. These weights dynamically reflect the importance of each feature for the final semantic determination in a specific scenario. For example, for a graph-rich topological map, visual features may have higher weights; for a standard form, structural features are more critical. Finally, the network performs a weighted sum of the three transformed features based on the calculated weights, outputting a novel, unified feature representation that integrates visual, structural, and contextual information—the multimodal fusion feature. This feature forms the foundation for subsequent intelligent prompts and assisted decision-making.

[0053] As one implementation method, to overcome the challenge of non-strict alignment between visual elements and DOM nodes, this patent proposes a cross-modal attention fusion network. This network first establishes a soft correspondence between visual and structural features through a learnable alignment matrix. Its core is the formula F_fusion=Attn(V,S,C)=α·V+β·S+γ·C: F_fusion: Multimodal fusion feature. The unified semantic feature representation generated after multimodal information fusion is obtained by weighted summation of three sub-feature vectors. Function: As the final output of the intelligent perception engine, it is used to match the types of interface elements in the knowledge base.

[0054] V: Visual Feature Vector. The visual information extracted from the interface screenshot is encoded and forward-propagated through a lightweight CNN (such as a simplified version of MobileNet or ResNet), taking the output of the penultimate layer. Purpose: To identify the shape, color, texture, and other appearance features of visual elements such as buttons, input boxes, and selectors.

[0055] S: Structural Feature Vector. Encodes the element's structural information obtained from parsing the DOM tree. The DOM parser extracts the element's tag name, attributes, hierarchical position, and adjacency relationships, then converts it into a vector using an attention-based encoder. Purpose: To understand the structural role (e.g., form fields, navigation menus, submit buttons) and logical relationships of elements within the interface.

[0056] C: Context Feature Vector. Encodes the contextual information of the current operating environment, combining user role, historical operation sequence, current task stage, system state, etc., through an embedding layer. Purpose: To provide a foundation for personalized, context-specific semantic understanding, enabling the system to distinguish the different meanings of the same element in different contexts.

[0057] α: Visual weight parameter. Controls the contribution of visual feature vectors to the final decision. Effect: For graphical, icon-rich interfaces (such as dashboards), the α value is higher; for text-intensive configuration pages, the α value is lower.

[0058] β: Structural weight parameter. Controls the contribution of structural feature vectors to the final decision. Effect: For standardized form elements (such as input and select), the β value is higher; for custom components, the β value is relatively lower.

[0059] γ: Context weight parameter. Controls the contribution of the context feature vector to the final decision, γ = 1 - (α + β), ensuring weight normalization. Purpose: When the user operates frequently or is in a complex task flow, the importance of the context increases, and the value of γ increases accordingly.

[0060] Dynamic attention weight generation: The weights α, β, and γ are not fixed values, nor are they calculated independently according to simple rules. Instead, they are dynamically generated based on the overall context of the current fusion task. The system can also concatenate the projected visual feature vector V, structural feature vector S, and contextual feature vector C to form a comprehensive scene representation. This representation is then input into a lightweight attention module, which calculates initial attention scores for the three feature sources using its internally learnable parameter matrix and biases. These scores are then normalized using a Softmax function to obtain a set of dynamic weights that sum to 1. These dynamically generated weights, rather than fixed values, are ultimately used to perform weighted fusion of the three feature vectors, generating multimodal fusion features that accurately reflect the semantics of the current scene.

[0061] In the implementation of the above embodiments: by combining visual recognition and DOM structure analysis, the system can not only "see" the appearance of interface elements, but also understand their roles in page functions, thereby overcoming the limitations of a single recognition method and possessing good adaptability to various styles of operation and maintenance interfaces. Through contextual feature fusion, the system's understanding is not isolated, but rather combines user identity and operation scenario. The generated multimodal fusion features serve as a unified semantic representation, making subsequent prompt content generation and operational guidance more accurate and reliable.

[0062] Optionally, in this embodiment of the application, obtaining the context feature vector of the operation and maintenance interface includes: Acquiring contextual information reflecting the operational environment of the operations and maintenance interface; this contextual information includes at least one of user role, current operation sequence, and interface state. Acquiring contextual information reflecting the operational environment of the operations and maintenance interface refers to the system's real-time collection and integration of multi-dimensional state data about the user's current scenario during operations. Contextual information includes, but is not limited to, the following: 1) User role, obtained directly from login credentials or the permission management system, such as distinguishing between system administrators, regular operators, or auditors; 2) Current operation sequence, recorded by the front-end event listener, referring to the time sequence record of a series of operations performed by the user in this session (such as clicking a menu item or filling in a specific field); 3) Interface state, referring to the instantaneous interface status such as the active view in the current browser tab, open modal dialog boxes, selected tabs, and page scroll position. This information collectively constitutes an immediate description of "who, what they are doing, and which part of the interface they are in," providing an environmental basis for subsequent personalized decision-making.

[0063] Generating contextual feature vectors based on contextual information is a process of encoding discrete, multi-source contextual states into a unified numerical representation that can be processed by a computer. Each type of contextual information is processed separately: for categorical data such as user roles, a learnable embedding lookup table maps it to a fixed-length vector; for sequential data such as the current operation sequence, a lightweight recurrent neural network or Transformer encoder is used to capture its sequential dependencies and output a vector summarizing the entire sequence; for interface states, their components are similarly embedded and encoded before being concatenated or averaged. Finally, the vectors generated above are fused and dimensionality reduced through a fully connected layer, outputting a comprehensive, fixed-dimensional contextual feature vector. This vector compactly represents the overall semantics of the current operating environment, facilitating subsequent fusion with visual and structural features.

[0064] In the implementation of the above embodiments: by generating context feature vectors, the specific environment in which the user operates can be dynamically perceived and quantified, making the assistive functions no longer rigid and unchanging. This enables the system to understand the potential different needs of users with different roles (such as novices and experts), to predict their subsequent intentions based on the user's completed operation steps, and to identify the interface area currently being focused on by the user. Based on the context feature vectors generated in this way, the subsequent generation of prompts and assistance operation guidance can be closely integrated with the actual scenario.

[0065] Optionally, in this embodiment of the application, before generating the prompt content corresponding to the operation and maintenance interface based on multimodal fusion features, user context information, and operation and maintenance knowledge base, the method further includes: The system acquires historical user operation data, which includes operation sequences, operation results, and operation time. Operation sequences are the sequential records of UI events triggered by a user to complete a specific task, such as configuring the server, including events like clicks, input, and navigation. Operation results indicate whether each operation achieved its intended goal, such as whether configuration was successfully verified or command execution was error-free, typically recorded as success / failure flags or detailed error codes. Operation time is the duration of each independent step or the entire task. Historical user operation data is captured by a front-end event listener and stored in a structured format in a back-end database via a log service, providing raw factual evidence for building user profiles.

[0066] Based on users' historical operation data and preset weights, a user capability profile is determined. Periodic statistical analysis is performed on the acquired user historical operation data. The preset weights, defined by domain experts or data analysts, are used to assign different levels of importance to different operation types; for example, high-risk configurations have different weights than ordinary viewing. Different scores are also awarded or penalized for successful and incorrect operation results.

[0067] The implementation method for determining user capability profiles includes: the system uses a specific mathematical model, for example, subtracting the weighted number of erroneous operations from the weighted number of successful operations, and then dividing by the total operation time or number of operations, to comprehensively calculate one or more numerical capability indices. Capability indices can be subdivided by operational domain and may incorporate time decay factors to more accurately reflect the user's current true skill level, ultimately forming a dynamically updated, structured user capability profile.

[0068] As one implementation method, clustering algorithms can be used to classify user behavior and define different levels of capability indices. The formula for calculating the capability index is as follows:

[0069] Where Ci is the ability index of user i, n is the number of operations, Sj is the weight of successful operation j, Wj is the operation weight, Ej is the erroneous operation, Pj is the error penalty coefficient, and Ti is the total operation time.

[0070] Dynamic decay mechanism: The capability index is not permanently accumulated; the system adopts an exponential decay model, where λ is the decay coefficient (usually set to 0.01 / day). This means that if a user does not perform a certain type of task for a long time, their capability index will slowly decrease, prompting the system to provide appropriate prompts when the user resumes operation.

[0071] In the implementation of the above embodiments: the system can learn from users' past operations and automatically construct personalized profiles reflecting their true skill levels. This allows for intelligent adjustment of the depth and method of subsequent assistance based on different users' historical proficiency. For users exhibiting high success rates and efficiency in their records, the system can reduce basic prompts and provide more efficient advanced techniques or automation options; for users showing difficulty or a tendency to make mistakes, the system will provide more detailed and progressive guidance. Personalized adaptation based on objective historical data improves the speed and accuracy of various users in completing maintenance tasks, enhancing the user experience.

[0072] Optionally, in this embodiment, based on multimodal fusion features, user context information, and an operation and maintenance knowledge base, the prompt content corresponding to the operation and maintenance interface is generated, including: Using multimodal fusion features as semantic indexes, we can query the operation and maintenance knowledge base and retrieve the planning information corresponding to the operation and maintenance interface. The planning information includes concept explanations, parameter ranges, constraints, and risk warnings.

[0073] Querying the operations and maintenance knowledge base is achieved by calculating the similarity between the feature vector and the pre-stored entity vectors in the knowledge base. For example, the operations and maintenance knowledge base is constructed in the form of a graph database, where nodes represent entities such as devices, parameters, and concepts, and edges represent relationships such as constraints and dependencies. The system finds the most relevant knowledge nodes through vector retrieval, and then traverses its associated edges to retrieve planning information directly related to the current interface element. This information consists of structured knowledge fragments, including the element's conceptual explanation, parameter range, constraints, and risk warnings that may be caused by improper configuration, preparing the raw materials for generating prompts.

[0074] The personalized style of the prompts is determined based on user context information. This context information includes comprehensive characteristics such as the user's role, skill profile, and current task stage. Based on this context information, a predefined set of rules or a lightweight classification model maps the current scenario to several preset style templates. For example, for a user whose skill profile indicates a novice and whose role is operator, the system might choose a "detailed guidance" style, emphasizing basic concepts and step-by-step explanations; for an expert administrator, it might choose a "concise reminder" style, focusing on key parameter ranges and advanced risk warnings. The style determines the level of detail, technical depth, and expression of subsequent content.

[0075] The system organizes planning information according to a personalized style to generate prompts. This organization includes using templates corresponding to the chosen personalized style, logically sorting, filtering, and linguistically organizing the retrieved planning information fragments. For example, in a "detailed guidance" style, the system organizes information in a specific order, using more common vocabulary and complete sentences; in a "concise reminder" style, it may only list key parameter ranges and major risk items, using short phrases. Generating prompts typically utilizes template-based natural language generation technology or lightweight neural text generation models to smoothly convert the organized structured information into coherent paragraphs or multi-level segments, ultimately forming personalized prompts that can be displayed in a floating window and tailored to the user's needs.

[0076] In the implementation of the above embodiments: Multimodal fusion features are used for knowledge retrieval, ensuring that the prompts accurately correspond to the interface elements currently seen by the user, providing accurate prompts and improving efficiency. Personalized styles are determined by combining user context information, allowing the generated prompts to vary in detail and expression depending on the individual and the scenario. This reduces redundant information interference for experienced users and minimizes errors for novice users due to insufficient guidance, effectively improving the efficiency and accuracy of configuration and troubleshooting operations.

[0077] Optionally, in this embodiment of the application, the assistance mode includes at least a first assistance mode and a second assistance mode; generating guidance information corresponding to the assistance mode includes: If it is the first assistance mode, then based on the multimodal fusion features and user context information, the matching dynamic demonstration content is retrieved. The dynamic demonstration content is an interactive micro-video that reflects the real configuration process and is generated based on the real interface status.

[0078] When a user selects the first assistance mode, the multimodal fusion features are combined with the user's contextual information to form a comprehensive query condition. This query condition is then used to search a pre-stored case or material library for the most relevant operation examples. The retrieved materials are not fixed video files, but rather interactive micro-videos generated based on the real-time interface state, reflecting the actual configuration process. For example, a lightweight screen operation recording engine can be launched in the background. Based on the retrieved operation example script, it automatically simulates key operation steps, such as clicking and inputting, in the current real-world operation and maintenance interface environment, and records or renders this process in real time as a short, dynamic demonstration, such as an GIF or video. This GIF or video focuses on the user's current interface context, improving the relevance and operability of the demonstration.

[0079] If the second assistance mode is selected, the configuration form of the operation and maintenance interface is analyzed to identify the input parameter list and the derivation parameter list; a parameter collection interface is generated to guide the user to input parameters based on the input parameter list; based on the user-input parameters, multimodal fusion features, and the operation and maintenance knowledge base, the required parameters in the derivation parameter list are automatically generated.

[0080] When the user selects the second assistance mode, the system first analyzes the configuration form in the operations and maintenance interface. It identifies the semantics of each form item by parsing the DOM structure and combining multimodal fusion features. Based on the rules defined in the operations and maintenance knowledge base, all parameters are categorized into a list of input parameters provided by the user and a list of derived parameters that can be automatically calculated or queried by the system based on rules, formulas, or existing parameters.

[0081] Next, a parameter collection interface is generated, which can be a floating window or a sidebar, clearly listing the items in the input parameter list and guiding the user to fill in the required parameters. Based on the user-input parameters, multimodal fusion features, and the operation and maintenance knowledge base, the parameters required in the derivation parameter list are automatically generated. Implementation methods include using key user-input parameters as input, calling predefined parameter mapping logic or scripts, and combining this with the operation and maintenance knowledge base for verification and value retrieval. This automatically calculates the specific values ​​of all parameters in the derivation parameter list and prepares them for filling in the corresponding positions on the form.

[0082] In the implementation of the above embodiments: the dynamic demonstration content provided by the first assistance mode transforms abstract operation steps into intuitive and visual animations, enabling users to quickly imitate and learn the correct operation methods. This is especially suitable for procedural and step-by-step tasks, reducing the learning cost. The second assistance mode goes a step further, intelligently analyzing forms and distinguishing parameter types, freeing users from tedious, repetitive, and error-prone manual filling. Users only need to focus on core decision parameters, and the system can automatically and accurately complete the filling and calculation of the remaining large number of parameters. These two methods shorten the time required to complete complex configuration tasks and improve the efficiency and reliability of operation and maintenance work.

[0083] Optionally, in this embodiment of the application, after automatically generating the parameters required in the derivation parameter list, the method further includes: The verification engine validates the entered parameters; it integrates an operations and maintenance knowledge base. The verification engine is a standalone software module with this knowledge base. The knowledge base stores rich verification rules in a graph format, including parameter format specifications, value ranges, device compatibility constraints, service dependencies, and security policies. The verification engine is triggered when a user or the system automatically enters parameters.

[0084] The real-time validation method involves the engine extracting the parameters to be validated and their context, querying all relevant rules in the operations and maintenance knowledge base, and performing logical matching and calculations for each rule. For example, it checks the IP address format, port conflicts, and whether the selected service supports the current device model. The validation process is instantaneous, providing preliminary results immediately after user input.

[0085] If the verification result fails, a warning message is generated and a rollback operation is performed.

[0086] A failed validation result indicates that the validation engine, based on knowledge base rules, determines that one or more parameters have format errors, out-of-bounds errors, conflicts, or security risks. In this case, the system will first generate a warning message, typically displayed next to the corresponding parameter on the interface with a red border or a pop-up window, showing the specific error reason and improvement suggestions. Simultaneously, a rollback operation will be automatically triggered.

[0087] For example, a temporary state snapshot or original value can be created or recorded before performing any configuration operation or autofill that might change the system state. When verification fails, based on this record, all changes made to the relevant configuration items in the current session can be automatically rolled back, and the data in the interface form can be restored to the state before the operation. This ensures that incorrect configurations are not partially submitted or left behind, maintaining the stability and consistency of the system.

[0088] In the implementation of the above embodiments: the verification engine utilizes an integrated knowledge base for real-time checks, enabling it to promptly detect format errors, out-of-bounds values, configuration conflicts, or potential security risks in parameter input before user submission or system application configuration, and immediately provide clear warnings. This allows errors to be corrected at the earliest stage, avoiding system failures, service interruptions, or security vulnerabilities that may be caused by incorrect configurations taking effect. The automatic rollback operation quickly and cleanly undoes erroneous changes, restoring the interface to a safe state, eliminating the hassle of manual investigation and rollback.

[0089] Optionally, in this embodiment of the application, after guiding the user to input parameters based on the input parameter list on the parameter collection interface, the method further includes: The system invokes a pre-recorded abstract process description, which contains semantic operational intents. In the second assistance mode, invoking a pre-recorded abstract process description refers to the system reading a pre-recorded and stored process template related to the current task. During the recording phase, these pre-recorded abstract process descriptions record a sequence of semantic operational intents for interface elements by analyzing the multimodal fusion characteristics at the time, such as "enter a value in the IP address input field." This description is decoupled from specific interface styles and layouts, ensuring that it can be reused for different styles of but functionally similar operation and maintenance interfaces. Upon invocation, the system matches and loads the appropriate abstract process description based on the contextual characteristics of the current task.

[0090] Based on the user-input parameters and parameter mapping graph, the abstract process description is instantiated into an operation sequence to generate the parameters required for deriving the parameter list. The parameter mapping graph is constructed synchronously during the process recording phase, defining the dependencies between all parameters in the process and distinguishing between user-input parameters and internal parameters derived by the system based on rules, formulas, or knowledge bases.

[0091] The instantiation process involves substituting the actual parameter values ​​input by the user into an abstract process description and parameter mapping graph. Following the logic in the graph, the engine calculates specific values ​​for all internally derived parameters and binds the abstract semantic operation intent (such as "fill in the value") with the actual target elements identified through multimodal features on the current operations interface and the specific parameter values. This generates a concrete sequence of operations that can be executed step-by-step on the current real interface. The final execution result of this sequence is the generation of the parameters needed in the deduced parameter list, which are then automatically filled in.

[0092] In the implementation of the above embodiments: calling pre-recorded abstract process descriptions allows the system to reuse validated standard operating procedures, improving the correctness and reliability of the operation methods, while reducing the tediousness of repeatedly recording fixed scripts for each different interface style. Semantic operation intent recording makes the process template independent of specific interface implementation details, enhancing adaptability. This reduces the workload for users when filling out complex forms and improves processing efficiency.

[0093] Optionally, in this embodiment of the application, the method further includes: This involves acquiring user interaction data during the process of operating based on guided information. For example, it can structurally capture and record all user interactions after receiving various system-provided guides, including hover prompts, demo videos, and auto-filled interfaces. The interaction data is a comprehensive dataset, specifically including: the user's specific actions on interface elements, the order and timestamps of these actions, the duration of pauses before specific steps or prompts, acceptance or modification of auto-filled content, and the final result of the operation. This data is captured in real-time by a listening script deployed in the browser and associated with the specific guided information context and user session identifier. It is then sent to the backend server in a structured log format for storage, forming a raw data pool for analyzing user behavior and the effectiveness of system guidance.

[0094] Successful actions in interaction behavior data are used as positive feedback samples to optimize the fusion model used to generate multimodal fusion features. Samples that have been verified as successful actions can be selected from the interaction behavior data, such as configuration submissions that ultimately pass the verification engine without rollback. These successful samples, along with their complete context at the time of occurrence, including the UI screenshot and DOM state at that time, are marked as positive feedback samples.

[0095] During optimization, samples are used as training data for incremental learning. For example, the original multimodal inputs from the samples are fed back into the fusion model to be optimized (i.e., a cross-modal attention fusion network), and the successfully completed task is used as the target signal. The model's internal parameters, such as the attention weight matrix, are fine-tuned through backpropagation. This enables the model to more accurately generate multimodal fusion features representing the semantics of elements when encountering similar interfaces or scenarios, thereby achieving self-improvement in understanding specific operation and maintenance interfaces or user operation patterns.

[0096] In the implementation of the above embodiments: real operation data of users after receiving guidance is obtained, providing a basis for analyzing the system's assistance effect and user behavior patterns. Successful operation cases are used as positive feedback to optimize the core model for generating multimodal fusion features, which enables the system to become increasingly accurate with the increase of usage time.

[0097] This application provides an assisted operation and maintenance method based on intelligent floating box prompts: The system first captures the visual features of the user's current operation and maintenance interface, and identifies interface elements by analyzing pixel information; simultaneously, it parses the DOM (Document Object Model) structure of the page to understand the logical composition and hierarchical relationship of the elements; and combines this with the user's real-time context information, such as role and task stage. These three types of information are then sent to a multimodal feature fusion module for integration and deep semantic understanding.

[0098] After integration, user behavior is continuously monitored to determine if any user hovering over specific interface elements is detected. Once a hovering intent is detected, the system immediately generates personalized prompts based on the integrated context information and displays them to the user in a layered manner (such as concise prompts and detailed explanations) via a floating window. Users can select an assistance mode from the floating window or relevant menus according to their needs.

[0099] If the user chooses simple assistance, the system will display a dynamic operation video matching the context of the operation for the user to watch and learn. If the user chooses detailed assistance, the system will guide the user to input only a few key parameters, and then automatically complete the intelligent filling of the remaining form items. If the user chooses learning assistance, the system will launch an interactive step-by-step guide, using visual focusing technologies such as highlighting and masking to gradually guide the user to complete the operation manually. Regardless of the path chosen, the final filled-in or modified configuration will be checked by a built-in configuration verification engine, which verifies the legality and security of the parameters according to knowledge base rules. If the verification passes, the configuration record is submitted and the user's capability profile is updated, and the process ends successfully. If the verification fails, the process may prompt errors and guide corrections, thereby ensuring the correctness of operation and maintenance operations and system stability.

[0100] Please see Figure 2 The diagram shown is a structural schematic of the maintenance assistance device provided in this application embodiment; this application embodiment provides a maintenance assistance device 200, including: The fusion module 210 is used to acquire the operation and maintenance interface accessed by the user, perform multimodal recognition on the operation and maintenance interface, and generate multimodal fusion features; the multimodal recognition includes visual recognition and element structure recognition; The prompt information module 220 is used to generate prompt content corresponding to the operation and maintenance interface based on multimodal fusion features, user context information and operation and maintenance knowledge base, and display the prompt content to the user; the user context information includes the user capability profile determined based on the user's historical operation data. The request acquisition module 230 is used to acquire the user's assistance request; the assistance request includes the assistance mode. The assistance module 240 is used to generate guidance information corresponding to the assistance mode based on the multimodal fusion features and user context information; the guidance information is used to assist users in performing operation and maintenance operations.

[0101] Optionally, in this embodiment of the application, the assisted operation and maintenance device 200 and the fusion module 210 are specifically used to identify visual elements in the operation and maintenance interface using a trained convolutional neural network and extract visual feature vectors; parse the document object model tree of the operation and maintenance interface through the DOM parsing channel and extract structural feature vectors of the interface elements; obtain the context feature vector of the operation and maintenance interface; the context feature vector is used to reflect the operating environment of the operation and maintenance interface; and fuse the visual feature vector, structural feature vector and context feature vector to obtain multimodal fusion features.

[0102] Optionally, in this embodiment of the application, the assisted operation and maintenance device 200 and the fusion module 210 are specifically used to obtain context information reflecting the operating environment of the operation and maintenance interface; the context information includes at least one of user role, current operation sequence and interface state; and generate a context feature vector based on the context information.

[0103] Optionally, in this embodiment of the application, the maintenance assistance device 200 further includes a capability profile module for acquiring user historical operation data; the user historical operation data includes operation sequence, operation result and operation time; and the user capability profile is determined based on the user historical operation data and preset weights.

[0104] Optionally, in this embodiment of the application, the operation and maintenance assistance device 200 and the prompt information module 220 are specifically used to use multimodal fusion features as semantic indexes to query the operation and maintenance knowledge base and retrieve the planning information corresponding to the operation and maintenance interface; the planning information includes concept explanations, parameter ranges, constraints and risk warning information; determine the personalized style of the prompt content according to the user context information; organize the planning information according to the personalized style and generate prompt content.

[0105] Optionally, in this embodiment, the maintenance assistance device 200 has at least a first assistance mode and a second assistance mode. The assistance module 240 is specifically used to: in the first assistance mode, retrieve matching dynamic demonstration content based on multimodal fusion features and user context information, wherein the dynamic demonstration content is an interactive micro-video generated based on the real-time interface status and reflecting the actual configuration process; in the second assistance mode, analyze the configuration form of the maintenance interface, identify the input parameter list and the deduced parameter list; generate a parameter collection interface to guide the user to input parameters based on the input parameter list; and automatically generate the required parameters in the deduced parameter list based on the user-input parameters, multimodal fusion features, and the maintenance knowledge base.

[0106] Optionally, in this embodiment of the application, the operation and maintenance assistance device 200 further includes a verification module, which is used to verify the entered parameters through a verification engine; wherein, the verification engine integrates an operation and maintenance knowledge base; if the verification result fails, a warning message is generated and a rollback operation is performed.

[0107] Optionally, in this embodiment of the application, the maintenance assistance device 200 further includes a parameter derivation module, which is used to call a pre-recorded abstract process description, the abstract process description containing semantic operation intent; and to instantiate the abstract process description into an operation sequence based on the parameters input by the user and the parameter mapping map, so as to generate the parameters required in the parameter derivation list.

[0108] Optionally, in this embodiment of the application, the operation and maintenance assistance device 200 further includes a model optimization module, which is used to acquire interactive behavior data of the user during the operation based on the guidance information; and to optimize the fusion model used to generate multimodal fusion features by using successful operations in the interactive behavior data as positive feedback samples.

[0109] It should be understood that this device corresponds to the above-described assisted operation and maintenance method embodiments and is capable of performing the various steps involved in the above method embodiments. The specific functions of this device can be found in the description above, and detailed descriptions are omitted here to avoid repetition. The device includes at least one software functional module that can be stored in memory or embedded in the device's operating system (OS) in the form of software or firmware.

[0110] Please see Figure 3 The diagram shows a structural schematic of an electronic device provided in an embodiment of this application. An electronic device 300 provided in this application includes a processor 310 and a memory 320. The memory 320 stores machine-readable instructions executable by the processor 310. When the machine-readable instructions are executed by the processor 310, the method described above is performed.

[0111] Figure 3 The components shown can be implemented using hardware, software, or a combination thereof. Electronic device 300 may be a physical device, such as a server or PC, or a virtual device, such as a virtual machine or virtualization container. Furthermore, electronic device 300 is not limited to a single device; it can be a combination of multiple devices or a cluster of numerous devices.

[0112] This application also provides a storage medium storing a computer program, which is executed by a processor to perform the above-described method.

[0113] The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0114] This application also provides a computer program product, including computer program instructions, which are executed by a processor to perform the method described above.

[0115] It should be understood that the disclosed apparatus and methods can also be implemented in other ways, given the several embodiments provided in this application. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0116] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0117] The above description is only an optional implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application.

Claims

1. A method for assisting in operation and maintenance, characterized in that, include: The system obtains the operation and maintenance interface accessed by the user, performs multimodal recognition on the operation and maintenance interface, and generates multimodal fusion features; the multimodal recognition includes visual recognition and element structure recognition; Based on the multimodal fusion features, user context information, and operation and maintenance knowledge base, the corresponding prompt content for the operation and maintenance interface is generated and displayed to the user; The user context information includes a user capability profile determined based on the user's historical operation data; Obtain the user's assistance request; the assistance request includes an assistance mode. Based on the multimodal fusion features and the user context information, guide information corresponding to the assistance mode is generated; The guidance information is used to assist users in performing operation and maintenance tasks.

2. The method according to claim 1, characterized in that, The operation and maintenance interface is subjected to multimodal recognition to generate multimodal fusion features, including: A trained convolutional neural network is used to identify visual elements in the operation and maintenance interface and extract visual feature vectors. The document object model tree of the operation and maintenance interface is parsed through the DOM parsing channel to extract the structural feature vectors of the interface elements; Obtain the context feature vector of the operation and maintenance interface; the context feature vector is used to reflect the operating environment of the operation and maintenance interface. The visual feature vector, the structural feature vector, and the context feature vector are fused to obtain the multimodal fusion feature.

3. The method according to claim 2, characterized in that, Obtaining the context feature vector of the operation and maintenance interface includes: Obtain context information reflecting the operating environment of the operation and maintenance interface; the context information includes at least one of user role, current operation sequence, and interface state; Based on the context information, the context feature vector is generated.

4. The method according to claim 1, characterized in that, Before generating the prompt content corresponding to the operation and maintenance interface based on the multimodal fusion features, user context information, and operation and maintenance knowledge base, the method further includes: Acquire the user's historical operation data; the user's historical operation data includes operation sequence, operation result, and operation time. The user capability profile is determined based on the user's historical operation data and preset weights.

5. The method according to claim 1, characterized in that, Based on the multimodal fusion features, user context information, and the operation and maintenance knowledge base, the corresponding prompt content for the operation and maintenance interface is generated, including: The multimodal fusion features are used as semantic indexes to query the operation and maintenance knowledge base and retrieve the planning information corresponding to the operation and maintenance interface; the planning information includes concept explanations, parameter ranges, constraints, and risk warning information. The personalized style of the prompt content is determined based on the user context information; The planning information is organized according to the personalized style to generate the prompt content.

6. The method according to claim 1, characterized in that, The assistance mode includes at least a first assistance mode and a second assistance mode; Generate the guidance information corresponding to the assistance mode, including: If it is the first assistance mode, then according to the multimodal fusion features and the user context information, the matching dynamic demonstration content is retrieved. The dynamic demonstration content is an interactive micro-video that reflects the real configuration process and is generated according to the real-time interface status. If it is the second assistance mode, the configuration form of the operation and maintenance interface is analyzed to identify the input parameter list and the derivation parameter list; a parameter collection interface is generated to guide the user to input parameters based on the input parameter list; based on the parameters input by the user, the multimodal fusion features, and the operation and maintenance knowledge base, the required parameters in the derivation parameter list are automatically generated.

7. The method according to claim 6, characterized in that, After automatically generating the required parameters in the derivation parameter list, the method further includes: The entered parameters are validated through a verification engine; the verification engine integrates the operation and maintenance knowledge base. If the verification result fails, a warning message is generated and a rollback operation is performed.

8. The method according to claim 6, characterized in that, After guiding the user to input parameters based on the input parameter list on the parameter collection interface, the method further includes: Invoke a pre-recorded abstract process description, which contains semantic operation intent; Based on the parameters input by the user and the parameter mapping graph, the abstract process description is instantiated into an operation sequence to generate the parameters required in the derivation parameter list.

9. The method according to claim 1, characterized in that, The method further includes: Acquire user interaction data during the operation based on the guidance information; Successful operations from the interaction behavior data are used as positive feedback samples to optimize the fusion model used to generate the multimodal fusion features.

10. A computer program product, characterized in that, It includes computer program instructions that are executed by a processor to perform the method as described in any one of claims 1 to 9.

11. An electronic device, characterized in that, include: A processor and a memory, the memory storing computer program instructions that, when executed by the processor, perform the method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, perform the method as described in any one of claims 1 to 9.