Instruction processing method and device based on multiple modes
Through multimodal instruction processing methods, the problem of user operation complexity under traditional interactive methods is solved, precise operation in complex scenarios is achieved, and user experience and system adaptation efficiency are improved.
Patent Information
- Application Number
- CN202510746514.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-08-26
AI Technical Summary
Traditional interaction methods increase users' cognitive burden in complex scenarios, resulting in non-professional users being unable to complete the operation smoothly.
Through multimodal instruction processing methods, including ambiguity detection, ambiguity elimination, parsing intentions, splitting atomic operation sequences and generating operation logic using low-code orchestration to ensure the accuracy of the operation.
In complex scenarios, the accuracy and efficiency of user operations are achieved, the multi-step disassembly requirement is reduced, and the efficiency of calling high-frequency functions and user satisfaction are improved.
Smart Images

Figure CN120540564A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a multimodal instruction processing method and device. Background Art
[0002] With the deepening of enterprise digital transformation, modern management systems have gradually integrated complex business processes and data interaction functions.
[0003] Traditional interaction methods rely on fixed forms, buttons, and templated inputs, requiring users to memorize every step and click sequence. However, these fixed forms and multi-level operation paths impose a significant cognitive burden on non-expert users. Especially when users are unsure of what to do next, they often become lost in the complex interface hierarchy, making it difficult to successfully complete operations in complex scenarios.
[0004] Therefore, how to achieve precise operations in complex scenarios has become an urgent problem that needs to be solved. Summary of the Invention
[0005] The present application provides a multimodal instruction processing method and device, the purpose of which is to achieve precise operation in complex scenarios.
[0006] In order to achieve the above objectives, this application provides the following technical solutions:
[0007] A multimodal instruction processing method, comprising:
[0008] When receiving an operation instruction, performing ambiguity detection on the operation instruction;
[0009] When the operation instruction is ambiguous, the operation instruction is ambiguous to obtain an ambiguous operation instruction;
[0010] Parsing the eliminated operation instruction to obtain the parsed intention;
[0011] Splitting the parsed intent to obtain an atomic operation sequence; the atomic operation sequence includes multiple atomic operations; the atomic operations indicate small tasks that are executed separately;
[0012] Process the atomic operation sequence using a low-code orchestration method to obtain the operation logic;
[0013] Call the interface to execute the operation logic, obtain the operation result, and display the operation result.
[0014] Optionally, when the operation instruction is ambiguous, performing ambiguity elimination on the operation instruction to obtain an operation instruction after ambiguity elimination includes:
[0015] Obtain interface status information, user historical behavior information and business rules;
[0016] Constructing context information based on the interface state information, the user historical behavior information and the business rules;
[0017] Filtering information corresponding to the operation instruction from the context information and marking it as new information;
[0018] generating a new operation instruction based on the new information and the operation instruction;
[0019] The new operation instruction is ambiguous-eliminated using a preset elimination scheme to obtain an ambiguous-eliminated operation instruction.
[0020] Optionally, parsing the eliminated operation instruction to obtain the parsed intent includes:
[0021] performing a cleaning process on the eliminated operation instructions to obtain a cleaned operation instruction;
[0022] Converting pre-built context feature information into a context feature vector; the context feature information is pre-built based on interface state information, user historical behavior information and business rules;
[0023] The context feature vector and the cleaned operation instruction are input into a large language model to obtain the parsed intent.
[0024] Optionally, also include;
[0025] When feedback information sent based on the operation result is received, determining the feedback information as a reward function;
[0026] The model parameters of the large language model are adjusted according to the reward function to obtain an optimized large language model.
[0027] Optionally, also include:
[0028] Get the correction information within the preset historical time period;
[0029] A high-frequency problem report is generated based on the correction information, and the high-frequency problem report is fed back.
[0030] A multimodal instruction processing device, comprising:
[0031] a detection unit, configured to perform ambiguity detection on an operation instruction when the operation instruction is received;
[0032] an elimination unit, configured to eliminate ambiguity in the operation instruction when the operation instruction is ambiguous, and obtain an operation instruction after the ambiguity is eliminated;
[0033] a parsing unit, configured to parse the eliminated operation instruction to obtain a parsed intention;
[0034] A splitting unit, configured to split the parsed intent to obtain an atomic operation sequence; the atomic operation sequence includes a plurality of atomic operations; the atomic operations indicate small tasks that are executed separately;
[0035] A processing unit, configured to process the atomic operation sequence using a low-code orchestration method to obtain operation logic;
[0036] The execution unit is used to call the interface to execute the operation logic, obtain the operation result, and display the operation result.
[0037] Optionally, the elimination unit is specifically configured to:
[0038] Obtain interface status information, user historical behavior information and business rules;
[0039] Constructing context information based on the interface state information, the user historical behavior information and the business rules;
[0040] Filtering information corresponding to the operation instruction from the context information and marking it as new information;
[0041] generating a new operation instruction based on the new information and the operation instruction;
[0042] The new operation instruction is ambiguous-eliminated using a preset elimination scheme to obtain an ambiguous-eliminated operation instruction.
[0043] Optionally, the parsing unit is specifically configured to:
[0044] performing a cleaning process on the eliminated operation instructions to obtain a cleaned operation instruction;
[0045] Converting pre-built context feature information into a context feature vector; the context feature information is pre-built based on interface state information, user historical behavior information and business rules;
[0046] The context feature vector and the cleaned operation instruction are input into a large language model to obtain the parsed intent.
[0047] Optionally, also include:
[0048] a determining unit, configured to, upon receiving feedback information sent based on the operation result, determine the feedback information as a reward function;
[0049] An adjustment unit is used to adjust the model parameters of the large language model according to the reward function to obtain an optimized large language model.
[0050] Optionally, also include:
[0051] An acquisition unit, configured to acquire correction information within a preset historical time period;
[0052] A generating unit is configured to generate a high-frequency problem report based on the correction information and to feed back the high-frequency problem report.
[0053] The technical solution provided by this application is to perform ambiguity detection on the operation instruction when it receives the operation instruction; when the operation instruction is ambiguous, the operation instruction is disambiguated to obtain the eliminated operation instruction; the eliminated operation instruction is parsed to obtain the parsed intent; the parsed intent is split to obtain an atomic operation sequence; the atomic operation sequence is processed using a low-code orchestration method to obtain the operation logic; the interface is called to execute the operation logic to obtain the operation result, and the operation result is displayed. By disambiguating and parsing the operation instructions input by the user, clear and unambiguous instructions are generated to ensure accurate operation during subsequent execution. It can be seen that the user only needs to enter the operation instruction to achieve accurate operation even in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0055] Figure 1 A flowchart of a multimodal instruction processing method provided in an embodiment of the present application;
[0056] Figure 2 A flowchart of a method for eliminating ambiguity in an operation instruction provided in an embodiment of the present application;
[0057] Figure 3 A flowchart of a method for parsing an operation instruction provided in an embodiment of the present application;
[0058] Figure 4 A schematic diagram of the architecture of a multimodal instruction processing system provided in an embodiment of the present application;
[0059] Figure 5 A schematic diagram of the architecture of a multimodal instruction processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0060] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0061] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0062] like Figure 1 FIG. 1 is a flowchart of a multimodal instruction processing method provided in an embodiment of the present application, comprising the following steps:
[0063] S101: When an operation instruction is received, an ambiguity detection is performed on the operation instruction.
[0064] Operation instructions include but are not limited to fuzzy instructions. Fuzzy instructions refer to instructions issued by users through natural language, which have incomplete constraints and uncertainty compared to traditional interaction methods. For example, the operation instruction is to delete an order.
[0065] Optionally, the operation instruction includes one or more combinations of text, voice or image.
[0066] It is understandable that different operation instructions can be identified as ambiguous through different detection mechanisms. For details, see Table 1.
[0067] Table 1
[0068] Ambiguous category Trigger scenario example Detection Mechanism Object missing type "Delete Record" (no target specified) Check the interface focus object + historical operation correlation Parameter conflict type "Export orders from Beijing and Shanghai" (the system only supports single-region filtering) Verify API parameter constraints Permission violation "Approve Manager Zhang's leave request" (user does not have approval authority) Match the RBAC permission matrix in the business rules Timing-dependent "Undo the previous step" (the history operation stack is empty) Analyze the operation sequence state machine
[0069] It should be noted that the content of Table 1 above is only for illustration.
[0070] S102: When the operation instruction is ambiguous, the operation instruction is ambiguous to obtain an ambiguous operation instruction.
[0071] When an operation instruction is ambiguous, the operation instruction is disambiguated. Specifically, context information is constructed, and the operation instruction is disambiguated according to the context information to obtain the disambiguated operation instruction.
[0072] Optionally, in another embodiment of the present application, the specific implementation of step S102 is as follows: Figure 2 As shown, the following steps are included:
[0073] S201: Obtaining interface status information, user historical behavior information and business rules.
[0074] Interface status information includes capturing the current page's DOM elements (e.g., field names, button states, and focused object attributes) and parsing the UI layout structure (e.g., table selected rows and form input values). For example, when a user is on an order management page, the current filter condition ("Status = Overdue") and the selected order ID are recorded.
[0075] Specifically, user historical behavior information includes: extracting the last N operation sequences (for example: [create new customer → save failed → modify fields → save successfully]), counting high-frequency operation paths (for example, a user repeatedly executed "export Beijing customers" 5 times within 3 days), and recording user error correction behavior (for example, a second correction to the "delete record" command).
[0076] Optional business rules include: loading permission policies (such as the database scope accessible by user roles), injecting data validation logic (such as the order amount must be greater than 0), and associating cross-system constraints (such as customer ID mapping rules between CRM and ERP).
[0077] S202: Construct context information based on interface state information, user historical behavior information and business rules.
[0078] Among them, contextual information can more accurately understand the user's intentions, needs or behaviors.
[0079] S203: Filter out information corresponding to the operation instruction from the context information and mark it as new information.
[0080] The information corresponding to the operation instruction filtered out from the context information includes but is not limited to order information, interface status information, and business rules.
[0081] S204: Generate a new operation instruction based on the new information and the operation instruction.
[0082] It is understandable that a new operating instruction is generated based on the new information and the operating instruction, so that the new operating instruction is a clearer instruction than the operating instruction.
[0083] S205: Eliminate ambiguity on the new operation instruction using a preset elimination solution to obtain an operation instruction after ambiguity elimination.
[0084] Among them, the preset elimination solutions include but are not limited to active questioning, automatic correction and authority downgrade.
[0085] It is understandable that the new operation instructions are clearer than the original instructions, but in order to better execute user operations, we still need to continue to use the preset elimination solution to eliminate ambiguity in the new operation instructions to ensure that the user's needs can be accurately met.
[0086] For example, by proactively asking questions and providing interactive responses, the system can resolve ambiguity in new operation instructions based on the user's answer (e.g., "Please confirm the customer ID you want to delete"). Automatic correction can be used to replace conflicting parameters in operation instructions (e.g., "Shanghai" and "Beijing" can be replaced with "Beijing"). Alternatively, an alternative operation can be performed through permission downgrade (e.g., if the user does not have deletion permission, they must submit a deletion request).
[0087] S103: Analyze the eliminated operation instruction to obtain the analyzed intention.
[0088] It's understandable that parsing the disambiguated instruction primarily aims to extract the instruction's core intent, i.e., the specific action the user wishes to perform. For example, deleting a specific customer ID (if the user has permission) and determining whether additional actions (such as submitting an application) are required based on the permission.
[0089] Optionally, in another embodiment of the present application, the specific implementation of step S103 is as follows: Figure 3 As shown, the following steps are included:
[0090] S301: Cleaning the eliminated operation instructions to obtain cleaned operation instructions.
[0091] Among them, instruction cleaning includes the processing of different types of instructions: for text instructions, key information is extracted through word segmentation and entity recognition (such as converting "Manager Zhang" into {type:Person,value:Zhang_123}); for voice instructions, end-to-end speech recognition (ASR) is used to convert voice into text, and the user identity is bound through voiceprint authentication; for image instructions, OCR technology is used to extract text information, and target detection technology is combined for area annotation and UI coordinate mapping to ensure the accuracy and executability of the instructions.
[0092] S302: Convert pre-constructed context feature information into a context feature vector.
[0093] Among them, context feature information is pre-built based on interface state information, user historical behavior information and business rules.
[0094] It should be noted that the specific implementation method of pre-building context feature information based on interface state information, user historical behavior information and business rules can be found in step S201 and step S202.
[0095] It is understandable that feature vectors are the standard input form that machine learning models can process. By converting contextual information into vectors, the model can better understand the relationships and structures between data, thereby improving the accuracy of predictions and inferences. For example, the specific form of a contextual feature vector is:
[0096] ---prompts
[0097] [System Status] Current page: Order Management | Selected order ID: ORD_20240501_001 | User role: Sales Manager
[0098] [Business Rules] Actionable: Query / Export | Prohibited: Delete
[0099] [Historical behavior] Recent operation: Create a new Beijing customer (successful)
[0100] Please parse the user instruction into JSON: {"action":"","params":{}}
[0101] Command: "Show all projects that Manager Zhang is responsible for"
[0102] Output constraints: force the generation of structured JSON (avoid free text risks)
[0103] ---.
[0104] S303: Input the context feature vector and the cleaned operation instruction into the large language model to obtain the parsed intent.
[0105] As you can understand, the context feature vector and the cleaned operation instructions are input into the large language model. The large language model analyzes the context feature vector and understands the instruction to parse the user's true intent. This includes identifying the task type, extracting keywords, and understanding the context. Thus, the parsed intent is obtained.
[0106] In addition, the parsed intent includes explicit parameters and implicit parameters. Explicit parameters: These are the parameter values explicitly given and directly appear in the instruction. For example: "Beijing customer" can be converted to {"location": "Beijing"}. Implicit parameters: These are the default parameters deduced according to certain business rules or context. They are not explicitly given in the instruction, but are inferred by the system based on the context. For example: "recently" can be converted to {"time_range": "last 7 days"}, which is the time range inferred according to the system's default rules (such as the past 7 days).
[0107] The examples are as follows:
[0108] ---json
[0109] {
[0110] "primary_action": "Export overdue orders",
[0111] "sub_actions":
[0112] {"type": "Filter", "params": {"status": "overdue"}}, <00ooo253>{"type": "Export", "params": {"format": "Excel"}}
[0114] ,
[0115] "context": {
[0116] "user_role": "Sales supervisor",
[0117] "current_page": "Order management"
[0118] }}
[0119] ---。
[0120] S104: Split the parsed intent to obtain an atomic operation sequence.
[0121] Among them, the atomic operation sequence includes multiple atomic operations; an atomic operation indicates a small task to be executed separately (such as filtering, exporting, approving).
[0122] Specifically, the specific manifestation form of the atomic operation library is shown in Table 2.
[0123] Table 2
[0124] Atomic operation types System Mapping Example Data Query SQL query / API call SELECT * FROM orders Data Filtering WHERE condition injection WHERE status='overdue' Data Export File generation interface export(format='CSV') Notification sent Message queue push notify(channel='email') Approval trigger Workflow engine call start_approval_flow() Record modification UPDATE / PATCH requests PATCH / orders / {id} <00002s98>It should be noted that the contents shown in Table 2 are only for illustration.
[0126] It can be understood that, through the corresponding recognition algorithm, the parsed intent is mapped into atomic operations according to the atomic operation library, and then the atomic operations are split to obtain an atomic operation sequence.
[0127] For example, if the parsed intent is "export overdue orders and contact the person in charge", the atomic operation sequence obtained by splitting the parsed intent is:
[0128] ---json
[0129] {
[0130] "actions": [
[0131] { "type": "FILTER", "params": { "status": "overdue"}},
[0132] { "type": "EXPORT", "params": { "format": "CSV"}},
[0133] { "type": "NOTIFY", "params": { "recipient": "owner"}}
[0134] ]}
[0135] --.
[0136] For example, through the corresponding recognition algorithm, the user intention is mapped to an atomic operation. The example is as follows:
[0137] -- PYTHON
[0138] def map_to_atomic_actions(intent: dict) -> list:
[0139] atomic_ops = []
[0140] # Main operation decomposition (e.g. "Export overdue orders" → filter + export)
[0141] if intent["primary_action"] in COMPLEX_ACTION_MAP:
[0142] atomic_ops.extend(COMPLEX_ACTION_MAP[intent["primary_action"]])
[0143] # Sub-operation conversion (direct mapping atomic operation)
[0144] for sub_action in intent.get("sub_actions", []):
[0145] atomic_ops.append({
[0146] "type": ATOMIC_MAP[sub_action["type"]],
[0147] "params": sub_action["params"]
[0148] })
[0149] return atomic_ops
[0150] # Predefined complex operation mapping library
[0151] COMPLEX_ACTION_MAP = {
[0152] "Export overdue orders": [
[0153] {"type": "DATA_FILTER", "params": {"field": "status", "value": "overdue"}},
[0154] {"type": "DATA_EXPORT", "params": {"format": "CSV"}}
[0155] ],
[0156] "Contact Person": [
[0157] {"type": "USER_LOOKUP", "params": {"role": "owner"}},
[0158] {"type": "NOTIFY", "params": {"channel": "email"}} ]
[0160] }
[0161] ---.
[0162] In addition, the parameters in the user intent can be static (such as a fixed "overdue" status) or dynamic (for example, obtaining the current time range). In order to flexibly adapt to different situations, it is necessary to handle static parameter injection and dynamic parameter parsing.
[0163] Static parameters refer to parameters provided directly by the user in the intent. For example, in the "Export overdue orders" operation, it is clearly stated that the field to be filtered is "status" and its value is "overdue". The specific implementation logic is:
[0164] ---json
[0165] {
[0166] "type": "DATA_FILTER",
[0167] "params": {
[0168] "field": "status",
[0169] "value": "overdue" / / explicit parameter from intent
[0170] }
[0171] }
[0172] --.
[0173] If the user does not explicitly provide certain parameters, the system can dynamically obtain these parameter values based on the context. For example, when filtering orders, if the user does not provide a time range (time_range), the system will obtain a default time range from the context. The specific implementation logic is as follows:
[0174] ---PYTHON
[0175] if op["type"] == "DATA_FILTER":
[0176] # If the parameter is not clear (such as "recent orders")
[0177] if "time_range" not in op["params"]:
[0178] # Get default values from dynamic context
[0179] op["params"]["time_range"] = context.get_default("filter_time_range")
[0180] ---.
[0181] Some parameters may change dynamically at runtime. For example, the "Person in charge" may be the person in charge of the currently selected customer. In this case, variable replacement is required at runtime. The specific implementation logic is as follows:
[0182] ---json
[0183] {
[0184] "type": "USER_LOOKUP",
[0185] "params": {
[0186] "name": "{Currently selected customer. Responsible person}" / / Replaced with actual value at runtime
[0187] }}
[0188] ---.
[0189] S105: Use low-code orchestration to process the atomic operation sequence to obtain the operation logic.
[0190] Among them, the dependency relationship in the atomic operation sequence is first analyzed to obtain the analysis results (some atomic operations may require the output of other atomic operations as input); the atomic operation sequence is orchestrated according to the analysis results using the orchestration rules to obtain the operation logic.
[0191] For example, the dependency relationship is: filter expected orders → export to Excel → send email notification → create operation log.
[0192] Specifically, the rules for defining the order of operations are given, and a Python function sequence_ops is given for sorting operations.
[0193] Rule 1: Data operations must be performed in the order of "query → filter → modify". For example, before modifying data, you must first query the data and then filter it. The specific implementation logic is:
[0194] fany(op["type"] in ("DATA_UPDATE", "DATA_DELETE") for op in atomic_ops):
[0195] ordered_ops.append(get_op_by_type(atomic_ops, "DATA_QUERY"))
[0196] ordered_ops.append(get_op_by_type(atomic_ops, "DATA_FILTER")).
[0197] Rule 2: Notification operations must be performed after all other operations (usually after all data operations are completed). The specific implementation logic is as follows:
[0198] ordered_ops += [op for op in atomic_ops if op["type"].startswith("NOTIFY_")]
[0199] return ordered_ops
[0200] ---.
[0201] In addition, a visual interface for operation orchestration can be implemented through low-code tools (such as drag-and-drop components): Drag-and-drop components: Users can design workflows by dragging operation components (such as filters and export buttons). Wiring logic: Users use wiring to define the input and output relationships between operations, indicating the dependencies between operations. Automatically generate DSL: The system automatically generates a corresponding DSL (domain-specific language) based on the user's operations to describe the workflow. Examples are:
[0202] ---yaml
[0203] - action: filter
[0204] params:
[0205] field: status
[0206] value: overdue
[0207] output: filtered_orders # output variable name
[0208] - action: export
[0209] params:
[0210] source: filtered_orders
[0211] format: excel
[0212] - action: notify
[0213] params:
[0214] template: order_export_success
[0215] recipients: [user.email]
[0216] ---.
[0217] S106: Call the interface to execute the operation logic, obtain the operation result, and display the operation result.
[0218] Among them, the API or database interface of the target system can be called to execute the operation logic, obtain the operation results, and return the operation results to the user in natural language and visual form.
[0219] In addition, after the operation results are displayed, the user's behavior data will also be recorded.
[0220] Optionally, after S106, the large language model is adjusted based on the correction information provided by the user based on the operation results, which can help the model identify potential biases or sources of error, thereby optimizing and reducing misleading or inaccurate answers. Therefore, another embodiment of the present application provides a method for optimizing a large language model, including:
[0221] When feedback information sent based on the operation result is received, the feedback information is determined as a reward function.
[0222] The feedback information includes but is not limited to correction information and satisfaction scores.
[0223] As you can see, the reward function is the criterion used to evaluate the quality of the model's output. Feedback (such as corrections or satisfaction ratings) is used to inform the reward function. If the model's output meets expectations or better, the reward function gives a higher score; if the output does not meet expectations, the reward function gives a lower score.
[0224] The model parameters of the large language model are adjusted according to the reward function to obtain an optimized large language model.
[0225] Among them, after the large language model is adjusted, the model parameters will change, thereby improving the performance and effect of the model, enabling it to better complete the task.
[0226] Optionally, after S106, when the user is satisfied with the feedback on the operation result, correction information within a preset historical time period can be collected, and a high-frequency report can be generated based on this correction information. Through the high-frequency report, the large language model can be optimized and adjusted later. Therefore, another embodiment of the present application provides a method for generating a high-frequency problem report, including:
[0227] Get the correction information within the preset historical time period.
[0228] Among them, the correction information within the preset historical time period is obtained, that is, the data of the failed collection scenario is collected.
[0229] Generate a high-frequency problem report based on the correction information and provide feedback on the high-frequency problem report.
[0230] It's understandable that acquiring historical correction information, identifying failure scenarios, generating frequent problem reports, generating optimization suggestions based on these reports, and providing feedback on these reports and suggestions can help the system continuously optimize, promptly resolve frequent user-reported issues, and improve overall performance and user experience.
[0231] It should be noted that, based on the above process shown in S101-S106, this embodiment can achieve the following beneficial effects:
[0232] 1. Through the natural language interaction engine, users can directly trigger operations through spoken instructions (such as "export overdue orders and contact the person in charge"), thereby reducing the need for multi-step disassembly and significantly improving the efficiency of calling high-frequency functions.
[0233] 2. Through the dynamic context-aware module, interface status information, historical behavior data, and business rules are associated in real time, effectively resolving ambiguity issues caused by large language models being out of context (for example, "deleting records" without a clear target object), thereby achieving accurate operation analysis in complex scenarios.
[0234] 3. By collecting user correction behaviors and failure scenarios (such as repeated correction instructions) from interaction logs, we dynamically adjust model strategies using a reinforcement learning framework, forming a closed loop of "use-feedback-evolution." Pilot data shows that the speed of correcting high-frequency issues has increased by 40%, while also improving user satisfaction.
[0235] 4. By introducing low-code instruction orchestration technology and using domain-specific language (DSL) to automatically generate operation logic, it not only greatly reduces the workload of custom development, but also supports plug-in adaptation across systems (such as ERP, CRM).
[0236] 5. This not only optimizes individual user experiences but also supports enterprise decision-making through data value mining. For example, high-frequency requirements (such as automated approval processes) can be extracted from interaction logs and fed back into product feature iterations, thus forming a positive ecological cycle.
[0237] like Figure 4 , which is a schematic diagram of the architecture of a multimodal instruction processing system provided in an embodiment of the present application, the instruction processing system includes: a basic operation module and a feedback iteration module.
[0238] The basic operation module is used to receive operation instructions input through text or voice, and use the natural language interaction engine to eliminate ambiguity and parse the intent of the operation instructions based on context information to obtain the parsed intent; the operation instruction mapper converts the parsed intent into an atomic operation sequence, and the atomic operation sequence is processed using a low-code orchestration method to obtain the operation logic, and the interface is called to execute the operation logic to obtain the operation result.
[0239] The feedback iteration module is used to record the feedback information sent by users based on the operation results; and adjust the large language model based on the feedback information.
[0240] For example, a user enters the command "Show all projects for which Manager Zhang is responsible." By combining contextual information (e.g., the current interface status (e.g., department head identity) and the employee database), an API is called to retrieve Manager Zhang's ID. This command is then mapped to a "Filter project list" operation, and the results are displayed as a Gantt chart. If the user provides satisfactory feedback, the relevant data is incorporated into the model training set, improving the efficiency of responding to subsequent similar requests.
[0241] In summary, by disambiguating and parsing the user's input commands, clear and unambiguous instructions are generated, ensuring accurate operation during subsequent execution. This shows that users only need to enter the operation command to achieve accurate operation even in complex scenarios.
[0242] like Figure 5 As shown, it is a schematic diagram of the architecture of a multimodal instruction processing device provided in an embodiment of the present application, and the instruction processing device includes: a detection unit 100, an elimination unit 200, a parsing unit 300, a splitting unit 400, a processing unit 500 and an execution unit 600.
[0243] The detection unit 100 is configured to perform ambiguity detection on an operation instruction when the operation instruction is received.
[0244] The elimination unit 200 is used to eliminate ambiguity in an operation instruction when the operation instruction is ambiguous, and obtain an operation instruction after the ambiguity is eliminated.
[0245] The elimination unit 200 is specifically used to: obtain interface status information, user historical behavior information and business rules; construct context information based on the interface status information, user historical behavior information and business rules; filter out information corresponding to the operation instruction from the context information and identify it as new information; generate a new operation instruction based on the new information and the operation instruction; use a preset elimination scheme to eliminate ambiguity on the new operation instruction to obtain the eliminated operation instruction.
[0246] The parsing unit 300 is used to parse the eliminated operation instruction to obtain the parsed intention.
[0247] The parsing unit 300 is specifically used to: clean the eliminated operation instructions to obtain the cleaned operation instructions; convert the pre-constructed context feature information into a context feature vector; the context feature information is pre-constructed based on the interface status information, the user's historical behavior information and the business rules; the context feature vector and the cleaned operation instructions are input into the large language model to obtain the parsed intention.
[0248] The splitting unit 400 is used to split the parsed intent to obtain an atomic operation sequence; the atomic operation sequence includes multiple atomic operations; the atomic operation indicates a small task that is executed separately.
[0249] The processing unit 500 is used to process the atomic operation sequence using a low-code orchestration method to obtain the operation logic.
[0250] The execution unit 600 is used to call the interface to execute the operation logic, obtain the operation result, and display the operation result.
[0251] In summary, by disambiguating and parsing the user's input commands, clear and unambiguous instructions are generated, ensuring accurate operation during subsequent execution. This shows that users only need to enter the operation command to achieve accurate operation even in complex scenarios.
[0252] Combine Figure 5 According to the content shown, the instruction processing device also includes: a determination unit and an adjustment unit.
[0253] A determination unit is configured to, when receiving feedback information sent based on the operation result, determine the feedback information as a reward function.
[0254] The adjustment unit is used to adjust the model parameters of the large language model according to the reward function to obtain an optimized large language model.
[0255] Combine Figure 5 The content shown in FIG. 4 shows, the instruction processing device further includes: an acquisition unit and a generation unit.
[0256] The acquisition unit is used to obtain the correction information within a preset historical time period.
[0257] The generating unit is used to generate a high-frequency problem report based on the correction information and feed back the high-frequency problem report.
[0258] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Ordinary technicians in this field can understand and implement it without expending creative work.
[0259] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.
[0260] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multimodal instruction processing method, characterized in that: include: When receiving an operation instruction, performing ambiguity detection on the operation instruction; When the operation instruction is ambiguous, the operation instruction is ambiguous to obtain an ambiguous operation instruction; Parsing the eliminated operation instruction to obtain the parsed intention; Splitting the parsed intent to obtain an atomic operation sequence; the atomic operation sequence includes multiple atomic operations; the atomic operations indicate small tasks that are executed separately; Process the atomic operation sequence using a low-code orchestration method to obtain the operation logic; Call the interface to execute the operation logic, obtain the operation result, and display the operation result.
2. The method according to claim 1, characterized in that When the operation instruction is ambiguous, performing ambiguity elimination on the operation instruction to obtain an operation instruction after ambiguity elimination includes: Obtain interface status information, user historical behavior information and business rules; Constructing context information based on the interface state information, the user historical behavior information and the business rules; Filtering information corresponding to the operation instruction from the context information and marking it as new information; generating a new operation instruction based on the new information and the operation instruction; The new operation instruction is ambiguous-eliminated using a preset elimination scheme to obtain an ambiguous-eliminated operation instruction.
3. The method according to claim 1, characterized in that Parsing the eliminated operation instruction to obtain the parsed intent includes: performing a cleaning process on the eliminated operation instructions to obtain a cleaned operation instruction; Converting pre-built context feature information into a context feature vector; the context feature information is pre-built based on interface state information, user historical behavior information and business rules; The context feature vector and the cleaned operation instruction are input into a large language model to obtain the parsed intent.
4. The method according to claim 3, characterized in that Also includes; When feedback information sent based on the operation result is received, determining the feedback information as a reward function; The model parameters of the large language model are adjusted according to the reward function to obtain an optimized large language model.
5. The method according to claim 1, characterized in that Also includes: Get the correction information within the preset historical time period; A high-frequency problem report is generated based on the correction information, and the high-frequency problem report is fed back.
6. A multimodal instruction processing device, characterized in that: include: a detection unit, configured to perform ambiguity detection on an operation instruction when the operation instruction is received; an elimination unit, configured to eliminate ambiguity in the operation instruction when the operation instruction is ambiguous, and obtain an operation instruction after the ambiguity is eliminated; a parsing unit, configured to parse the eliminated operation instruction to obtain a parsed intention; A splitting unit, configured to split the parsed intent to obtain an atomic operation sequence; the atomic operation sequence includes a plurality of atomic operations; the atomic operations indicate small tasks that are executed separately; A processing unit, configured to process the atomic operation sequence using a low-code orchestration method to obtain operation logic; The execution unit is used to call the interface to execute the operation logic, obtain the operation result, and display the operation result.
7. The device according to claim 6, characterized in that The elimination unit is specifically used for: Obtain interface status information, user historical behavior information and business rules; Constructing context information based on the interface state information, the user historical behavior information and the business rules; Filtering information corresponding to the operation instruction from the context information and marking it as new information; generating a new operation instruction based on the new information and the operation instruction; The new operation instruction is ambiguous-eliminated using a preset elimination scheme to obtain an ambiguous-eliminated operation instruction.
8. The device according to claim 6, characterized in that The parsing unit is specifically used for: performing a cleaning process on the eliminated operation instructions to obtain a cleaned operation instruction; Converting pre-built context feature information into a context feature vector; the context feature information is pre-built based on interface state information, user historical behavior information and business rules; The context feature vector and the cleaned operation instruction are input into a large language model to obtain the parsed intent.
9. The device according to claim 8, characterized in that Also includes: a determining unit, configured to, upon receiving feedback information sent based on the operation result, determine the feedback information as a reward function; An adjustment unit is used to adjust the model parameters of the large language model according to the reward function to obtain an optimized large language model.
10. The device according to claim 6, characterized in that Also includes: An acquisition unit, configured to acquire correction information within a preset historical time period; A generating unit is configured to generate a high-frequency problem report based on the correction information and to feed back the high-frequency problem report.
Citation Information
Cited By
Car business SaaS voice control system based on multi-mode instruction analysis
CN121306124A