Intelligent interactive collaborative operation system and method based on AI large model technology
The intelligent interactive collaborative work system based on AI large model technology provides a unified task entry point and real-time status synchronization, solving the problem of users frequently switching between different applications and realizing a proactive and intelligent interactive experience and efficient task processing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-04-03
AI Technical Summary
In existing collaboration tools, users need to frequently switch between independent applications to perform tasks, resulting in redundant operation paths and broken context information. Traditional AI assistants cannot proactively interpret ambiguous task intentions, and the degree of automation for complex tasks is low.
The intelligent interactive collaborative operation system based on AI big data model technology includes an AI interaction framework module, a dynamic task decomposition module, and an automatic status synchronization module. It provides a unified task entry point, actively infers user intent, parses natural language or multimedia instructions into structured sub-tasks, and synchronizes status in real time through a persistent connection channel.
It achieves a highly integrated and proactive intelligent interactive experience, enhances the ability to process complex and multimedia commands, ensures the real-time performance and transparency of the task execution process, and significantly improves operational efficiency and convenience.
Smart Images

Figure CN121785696A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent collaborative office system technology, specifically to an intelligent interactive collaborative work system and method based on AI large model technology. Background Technology
[0002] Existing collaboration tools generally suffer from technical flaws: Interaction fragmentation problem: Users need to frequently switch between independent applications to perform tasks, resulting in redundant operation paths and broken context information; for example, task creation, progress tracking and file retrieval are scattered across different interfaces, reducing collaboration efficiency; Limitations of AI's passive response: Traditional AI assistants only support the execution of simple commands and cannot proactively interpret ambiguous task intentions. They lack systematic support for task decomposition, status tracking, and closed-loop human-machine interaction, resulting in a low degree of automation for complex tasks. Summary of the Invention
[0003] This invention aims to at least partially address one of the technical problems in related technologies. To this end, one objective of this invention is to propose an intelligent interactive collaborative work system based on AI large-scale model technology, comprising: The AI interaction framework module provides a unified task entry point. This entry point supports receiving user input and actively inferring user intent to provide dynamic suggestions. The suggestions are derived from the input text, locally stored documents, recently visited web pages, frequently used business systems, and command-line tool history. It also supports directly executing content startup and status operations within the window via shortcut keys. The dynamic task decomposition module parses natural language or multimedia instructions into structured subtasks based on large language models and multimodal models; The automatic status synchronization module receives and stores the execution status of subtasks in real time through a persistent connection channel.
[0004] Preferably, the AI interaction framework module includes: A semi-structured confirmation window is used to parse ambiguous instructions and provide a secondary modification interface; Agent task progress indicator, visually displaying the execution status and processing history of subtasks; The task details sub-window displays the task context and dialogue history in response to click actions.
[0005] Preferably, the operation of the semi-structured confirmation window includes: An initial task decomposition scheme for receiving outputs from large language models and multimodal models; It supports users in modifying and splitting task objects and execution steps; The confirmed structured instructions are then sent to the dynamic task breakdown module.
[0006] Preferably, the execution process of the dynamic task decomposition module includes: Identify the intent of action, the target of the task, and the time frame in natural language or multimedia instructions; Generate a structured list of subtasks containing their execution order; Subtasks are dispatched to the automatic state synchronization module and external function modules.
[0007] Preferably, the automatic status synchronization module includes: Streamable long-lived connection channels enable real-time capture of Agent execution state changes; A state database persistently stores task execution steps, intermediate outputs, and final results; The log recording unit associates the task status with the operation timestamp.
[0008] Preferred options also include: The file preview unit in the secondary window responds to file clicks in the sidebar and renders the file content directly in the floating window. The dialogue history retrieval unit supports natural language retrieval and locating historical dialogue fragments.
[0009] Preferably, the dynamic task breakdown module is linked with the smart calendar module: Parse time constraints in natural language; Automatically generate calendar events and associate them with subtask time nodes; Monitor the overdue status of tasks through the status synchronization module.
[0010] The intelligent interactive collaborative work method based on AI large model technology includes the following steps: S1. Launch the AI dialogue window using a global hotkey; S2. Receive natural language or multimedia instructions and break them down into a chain of structured subtasks; S3. Synchronize the execution status of subtasks in real time through a persistent connection channel; S4. Integrate task progress indicators, file previews, and historical search functions into a unified interface.
[0011] Preferably, the decomposition into a structured subtask chain includes: The large language model and multimodal model are invoked to identify action entities and object entities in the instructions; Generate a modifiable semi-structured task confirmation window; The final subtask dispatch instruction is generated based on the user's adjustments.
[0012] Preferably, it also includes a closed-loop coding process: Inject UI component context in response to drag-and-drop operations; Generate business adaptation code in the embedded workbench; The generated code will be deployed directly to the associated application environment for execution.
[0013] The above-described solution of the present invention has at least the following beneficial effects: It achieves a highly integrated and proactively intelligent interactive experience: by providing a unified task entry point and supporting proactive inference of user intent, the system has completely changed the operating mode of users having to frequently switch between different applications and passively input data; it can integrate the user's entire working context (documents, web pages, business systems, etc.) to provide dynamic suggestions, which greatly improves the intelligence and convenience of interaction and reduces the operational burden; The system has improved its ability to understand and process complex and multimedia instructions: By combining a large language model with a multimodal model, the system can not only process text instructions, but also understand and process multimedia instructions such as images and audio, which greatly expands the application scenarios and scope of the system and enables it to adapt to more diverse task requirements. It ensures the real-time performance and transparency of the task execution process: the status of subtasks is synchronized in real time through persistent connection channels, ensuring the timeliness and accuracy of task progress information; users and managers can clearly grasp the real-time execution status of each task, breaking the "black box" operation and improving the visibility and controllability of collaborative work; Significantly improves operational efficiency and convenience: It supports the function of directly executing content initiation and status operations within the window via shortcut keys, allowing users to quickly complete various operations without leaving the current working interface, greatly shortening the path of task initiation and management, and improving overall work efficiency.
[0014] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0016] Figure 1 This is a flowchart of an intelligent interactive collaborative work system based on AI large model technology provided in an embodiment of the present invention; Figure 2 This is a flowchart of an intelligent interactive collaborative operation method based on AI large model technology provided in an embodiment of the present invention.
[0017] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0019] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "circumferential," and "radial," etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the accompanying drawings and are only for the convenience of describing this invention and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.
[0020] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0021] In this invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0022] In this invention, unless otherwise explicitly specified and limited, "above" or "below" the second feature can include direct contact between the first and second features, or contact between the first and second features through another feature between them. Furthermore, "above," "over," and "on top" of the second feature includes the first feature directly above or diagonally above the second feature, or simply indicates that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature includes the first feature directly below or diagonally below the second feature, or simply indicates that the first feature is at a lower horizontal level than the second feature.
[0023] The following describes in detail, with reference to the accompanying drawings, an intelligent interactive collaborative work system based on AI large model technology according to an embodiment of the present invention.
[0024] Please see Figure 1 In this embodiment, the system includes: an AI interaction framework module that provides a unified task entry point. This entry point supports receiving user input and actively inferring user intent to provide dynamic suggestions. The suggestions are derived from input text, locally stored documents, recently visited web pages, frequently used business systems, and command-line tool history. It also supports direct execution of content initiation and status operations within the window via shortcut keys; a dynamic task decomposition module that parses natural language or multimedia instructions into structured subtasks based on a large language model and a multimodal model; and an automatic status synchronization module that receives and stores the execution status of subtasks in real time through a persistent connection channel. Unified Access and Intelligent Sensing: Users initiate interactions through a unified task entry point provided by the AI interaction framework module. This module not only receives natural language or multimedia commands directly input by the user, but its core feature is its ability to proactively infer the user's intent. By analyzing multi-source contextual information such as input text, locally stored documents, recently visited web pages, frequently used business systems, and command-line tool history, the system dynamically generates task suggestions highly relevant to the current scenario, thereby proactively assisting the user in decision-making. Users can efficiently invoke this entry point via shortcut keys and directly execute content initiation and status operations, achieving seamless interaction. Intelligent parsing and task structuring: User instructions (whether in natural language or multimedia form) are passed to the dynamic task decomposition module. This module utilizes the core capabilities of Large Language Model (LLM) and multimodal model to perform deep semantic parsing and understanding of the instructions, identifying the core intent, entities, and constraints. Subsequently, this module automatically decomposes and transforms the fuzzy and complex initial instructions into a series of well-defined, executable, structured subtasks, laying the foundation for automated processing. Real-time state synchronization and closed loop: The structured subtasks are dispatched to the corresponding execution units (which can be internal system components or external functional modules); the automatic state synchronization module listens for and captures all subtask execution state changes in real time through its persistent connection channel (such as WebSocket long connection); any state update (such as start, in progress, completed, failed) will be captured and persisted in real time to ensure that the progress information on the system overview and user interface is always up-to-date, forming a real-time closed loop from task dispatch to state feedback; The three modules mentioned above work together to form a complete intelligent interactive collaborative operation process with a unified entry point, intelligent parsing as the core, and state synchronization as the guarantee.
[0025] This embodiment has the following effects: It achieves a highly integrated and proactively intelligent interactive experience: by providing a unified task entry point and supporting proactive inference of user intent, the system has completely changed the operating mode of users having to frequently switch between different applications and passively input data; it can integrate the user's entire working context (documents, web pages, business systems, etc.) to provide dynamic suggestions, which greatly improves the intelligence and convenience of interaction and reduces the operational burden; The system has improved its ability to understand and process complex and multimedia instructions: By combining a large language model with a multimodal model, the system can not only process text instructions, but also understand and process multimedia instructions such as images and audio, which greatly expands the application scenarios and scope of the system and enables it to adapt to more diverse task requirements. It ensures the real-time performance and transparency of the task execution process: the status of subtasks is synchronized in real time through persistent connection channels, ensuring the timeliness and accuracy of task progress information; users and managers can clearly grasp the real-time execution status of each task, breaking the "black box" operation and improving the visibility and controllability of collaborative work; Significantly improves operational efficiency and convenience: It supports the function of directly executing content initiation and status operations within the window via shortcut keys, allowing users to quickly complete various operations without leaving the current working interface, greatly shortening the path of task initiation and management, and improving overall work efficiency.
[0026] In this embodiment, the AI interaction framework module includes: a semi-structured confirmation window for parsing ambiguous commands and providing a secondary modification interface; an Agent task progress indicator for visually displaying the execution status and processing history of subtasks; and a task details sub-window that displays the task context and dialogue history in response to click operations. Example 1: Automatic generation of suggestions for project weekly reports based on multi-source context I. Scenario-based approach and multi-source information perception User behavior: The user is writing code when an email notification about the "Project Alpha" pops up in the lower right corner of the screen. The user then presses the global hotkey Ctrl+ / / to bring up the AI dialogue window. The system actively infers: The AI interaction framework module does not remain blank and wait for input; instead, it immediately and proactively analyzes the current user context. Recently visited webpage: The project management system was opened in the browser tab, displaying the iteration dashboard of "Project Alpha"; Local storage documents: Several document modifications related to "Project Alpha" were found in the "This Week's Work Log" folder; Commonly used business systems: The code IDE detected that all recently committed code comments contain the #proj-alpha tag; Command line tool history: We discovered that the user has been frequently executing the command `git pull origin feature / alpha-api` recently.
[0027] Based on the above multi-source information, the system proactively infers that the user's current core task may be related to the weekly report summary of "Project Alpha".
[0028] II. Generation and Presentation of Dynamic Suggestions The system proactively provides dynamic suggestions in the dialog window: "You are detected working on Project Alpha. Do you need my help?" 1. Automatically generate this week's project progress report? (This will integrate code commit history, document updates, and Kanban status.) 2. Could you please schedule a time for your team's weekly meeting on Friday? 3. Other operations...; Please select or enter your command directly. It is recommended to display the corresponding shortcut key (such as F1, F2) next to the option.
[0029] III. Execute directly using shortcut keys Users do not need to type; they can simply press the F1 key. The system directly executes the content startup: immediately starts the "Generate Weekly Report" task, and pops up a semi-structured confirmation window for the user to finally confirm and fine-tune the scope of the weekly report content.
[0030] In this embodiment, the operation of the semi-structured confirmation window includes: receiving the initial task decomposition scheme output by the large language model and the multimodal model; supporting users to modify and split the task object and execution steps; and sending the confirmed structured instructions to the dynamic task decomposition module. Example 2: Generating a competitive analysis report combining images and text I. Multimodal command input User actions: In the dialogue window with the AI, the user performs the following operations; Text input, such as: "Please help me analyze the main functional differences between this new product and the competing products shown in the image below." Drag and drop to upload: Simultaneously drag a high-resolution screenshot of a competitor's product launch event into the input box.
[0031] II. Multimodal Model Collaborative Analysis and Scheme Generation The dynamic task decomposition module synchronously calls the Large Language Model (LLM) and the Multimodal Model (MLLM): LLM parses text instructions and identifies core actions (analysis, comparison) and objects (new products, competitors, functions).
[0032] MLLM parses the uploaded image, performs image recognition, and extracts: The competitor's brand logo (identified as "Brand C"); Key features and keywords on the presentation slides (such as "ultra-long battery life" and "AI vision").
[0033] The outputs of the two models are fused to generate an initial multimodal task decomposition scheme, which is then presented in a semi-structured confirmation window: "[Main Task] Comparative Analysis of New Product and Competitor Features" ├── [Subtask 1] Information Extraction │ ├── Target: Our new product specification (automatically linked to local documents) │ └── Target: Brand C (derived from image recognition) Competitor features (derived from image recognition: ultra-long battery life, AI-powered vision...) ├── [Subtask 2] Function Comparison │ └── Method: Compare the parameters and descriptions of our new product and Brand C's competitor's product item by item in terms of "battery life" and "AI function". └── [Subtask 3] Report Generation Format: Markdown comparison table; In the confirmation window, the information identified from the image (such as "Brand C" and "ultra-long battery life") is clearly marked with its source (such as a label with "[Image Recognition]").
[0034] III. User Adjustments and Confirmation of the Multimodal Solution Users have noticed a slight deviation in the multimodal model recognition: The "AI Vision" feature in the image was accurately identified, but the user wanted to compare it with the "Computational Photography" feature. The user clicks on the word "AI Insight" in subtask 2 and changes it to "Computational Photography".
[0035] After the user confirms that the plan is correct, they click the "Execute" button. The system will then issue the final adjusted structured instructions.
[0036] In this embodiment, the execution process of the dynamic task decomposition module includes: identifying the action intent, task object, and time range in natural language or multimedia instructions; generating a structured subtask list containing the execution order; and dispatching the subtasks to the automatic state synchronization module and external function modules.
[0037] Example 3: Supply Chain Query Task Based on Product Sketch I. Multimedia Command Input User interaction: In the dialogue window with the AI, the user did not enter any text, but directly uploaded a clear photo of a hand-drawn sketch of a product part.
[0038] II. Multimedia Command Parsing and Intent Recognition After receiving a multimedia instruction (image), the dynamic task decomposition module calls the Multimodal Model (MLLM) for parsing. Image recognition: The model identified the sketch as a ring-shaped metal fastener, accompanied by a handwritten label "Material: 304 stainless steel"; Intent inference: Based on common business processes, the module infers that the user's intent is likely to be to query the supplier or inventory information of the part (rather than to design or manufacture it).
[0039] III. Generating a Structured Subtask List Based on the analysis results of the multimodal model, the module automatically generates a structured chain of subtasks: "[Task 1] Part Information Confirmation" Object: Ring fastener (derived from image recognition) │Attribute: Material = 304 Stainless Steel (Source: Image Recognition) │ Operation: Search for matching specifications and models in the parts library. [Task 2] Supplier Inquiry │Operation: Based on the part model number output from Task 1, query the list of qualified suppliers in the ERP system. [Task 3] Inventory and Delivery Check │ Operation: In the SCM system, query the real-time inventory and shortest lead time of the preferred supplier.
[0040] IV. Dispatch and Execution The structured task list is immediately dispatched to the corresponding external functional modules: Task 1 → Enterprise Internal Parts Inventory Query System; Task 2 → ERP System Interface; Task 3 → Supply Chain Management (SCM) System Interface.
[0041] In this embodiment, the automatic state synchronization module includes: a Streamable long-connection channel for real-time capture of Agent execution state changes; a state database for persistently storing task execution steps, intermediate outputs, and final results; and a log recording unit for associated storage of task state and operation timestamps. Example 4: Multi-level Task State Synchronization and Persistence I. Establishment of Long Connection Channel When the dynamic task decomposition module dispatches subtasks to the email sending agent: The automatic state synchronization module creates a Streamable long-lived connection channel; Bind Task ID: Task#2025-0810-1732; Subscribe to the message topic: / agent / email / status.
[0042] II. Real-time Status Capture Process When the Agent executes a task, it pushes state changes via a persistent connection: / / Initial state {"task_id": "Task#2025-0810-1732", "step": "Email template rendering", "progress": 30%, "output": null, "timestamp": "2025-08-10T17:33:21Z"}"; The status synchronization module receives and decodes messages in real time; Trigger a state database write operation.
[0043] III. Structured Data Storage The state database stores data hierarchically:
[0044] The database tables are designed with a versioned chain structure, which supports state backtracking.
[0045] IV. Log Recording Mechanism Each state change triggers a logging unit: "[2025-08-10T17:35:47Z] TASK_UPD Task#2025-0810-1732 STATE: progress=80% ACTION: Add encrypted attachment HASH: a1b2c3d4e5 (Data Integrity Check Code) Log and database records are aligned using dual timestamps: event_timestamp: The time when the state occurred (taken from the Agent message); store_timestamp: The time the data was written to the database.
[0046] V. Abnormal Status Handling When the Agent returns an error status: {"step": "Recipient Verification", "error_code": "ERR_ADDR_INVALID", "detail": "Email format incorrect: zhang@comany"}"; The state synchronization module executes: Mark the database task status as FAILED; The log records error details and the time when it occurred; The alarm event is triggered to the AI interaction framework module.
[0047] In this embodiment, it also includes: a file sub-window preview unit, which responds to the sidebar file click operation and directly renders the file content in the floating window; and a dialogue history retrieval unit, which supports natural language retrieval and locates historical dialogue fragments. Example 5: Real-time preview and related queries of multimodal design drafts I. Multimodal File Preview User interaction: In the context of the UI design task review dialog, a sidebar displays a video file named "New App Interface Model.mp4" and an image file named "UI_Sketch.png"; the user clicks on either file link.
[0048] System response: For UI_sketch.png: A floating preview window renders and displays the content of the image in real time, and provides basic operation controls such as scaling and rotation; For the new APP interface model.mp4: A lightweight video player is embedded in the floating preview window, which supports play, pause, and volume control without launching an external video application.
[0049] II. Implementation of Multimodal Content Rendering Technology The system automatically selects different built-in rendering engines based on the file extension (e.g., .png, .mp4). Image rendering engine: processes formats such as JPG, PNG, and GIF, performing fast decoding and display; Video / audio rendering engine: processes formats such as mp4, mov, and mp3, and performs streaming media decoding and playback in a sandbox environment; Preview window adaptive layout: The preview window size is dynamically adjusted according to the original aspect ratio of the media file to ensure that the content is displayed without distortion.
[0050] III. Dialogue Retrieval Combined with Multimodal Preview User interaction: While watching a video preview, the user noticed an interesting UI interaction animation and then entered a natural language query in the conversation history search box: "Have we discussed the implementation of this sliding menu animation before?"
[0051] System response: Multimodal retrieval: The retrieval unit not only scans historical dialogue text, but also associates dialogue records containing "side menu" design drafts (images / videos); Results Presentation: In the search results, one historical record is highlighted: "[2025-08-10] Side-sliding menu scheme confirmed, see attached menu_demo.mp4 for details." A thumbnail preview is displayed next to this result; Seamless transition: When the user clicks on the result, the dialog window automatically jumps to the current dialog context, and at the same time, the preview window of the menu_demo.mp4 file is opened to play the relevant animation video.
[0052] In this embodiment, the dynamic task decomposition module works in conjunction with the intelligent calendar module to: parse time constraints in natural language; automatically generate calendar events and associate them with sub-task time nodes; and monitor task overdue status through the status synchronization module. Example 6: Automatic Schedule Creation Based on Video Conferencing Screen Recording I. Multimedia Command Input and Time Information Extraction User action: The user receives a video conference recording in which the host verbally announces, "Our next project review meeting is scheduled for 2 PM on the 5th of next month." The user drags the video file into the AI dialogue window.
[0053] System response: The dynamic task decomposition module calls a multimodal model to process the video: First, we performed Automatic Speech Recognition (ASR) to convert the audio into text: "Our next project review meeting is scheduled for 2 PM on the 5th of next month." Subsequently, the Large Language Model (LLM) parses the converted text and accurately identifies the time entity: 2 PM on the 5th of next month.
[0054] The module converts the identified time string into a structured timestamp: 2025-09-05T14:00:00Z.
[0055] II. Calendar Event Generation and Task Association The system automatically generates calendar events: Event ID: CAL#Proj-Review Title: Project Review Meeting (Source: Video Recording) Time: 2025-09-05 14:00-15:30 Related task: Task#Prepare-Review-Materials; At the same time, it creates a new preparation task, Task#Prepare-Review-Materials; And set the deadline for this task to the same time the day before the meeting begins: (2025-09-04T14:00:00Z), and associated with that calendar event.
[0056] III. Status Monitoring of Multimedia Tasks The automatic status synchronization module continuously monitors the progress of Task#Prepare-Review-Materials; If a task is detected as incomplete at 2025-09-04T14:00:00Z, the system will not only mark the task status as overdue, but also: Send the user a reminder containing a summary of the original video clip: "You have not yet completed the materials prepared for the 'Project Review Meeting' (from the video conference recording you received on August 10). Please see Figure 2 In this embodiment 7, the intelligent interactive collaborative operation method based on AI large model technology includes the following steps: S1. Invoking the AI dialogue window through a global hotkey; S2. Receiving natural language or multimedia instructions and breaking them down into a structured sub-task chain; S3. Synchronizing the execution status of sub-tasks in real time through a persistent connection channel; S4. Integrating task progress indicators, file previews, and historical retrieval functions into a unified interface; Multimedia Marketing Material Generation Task Complete Process I. Hotkey activation and multimedia command input (corresponding to S1) When a user presses the global hotkey Ctrl+ / / in the design software, an AI dialogue window pops up in the lower right corner of the screen. Instead of typing text, users dragged in a poster image of a new product and added a voice command: "Please apply this poster style to all social media posts for next week's promotion, and get it done by next Wednesday." II. Instruction Decomposition and Structured Subtask Chain Generation (corresponding to S2) After the system receives a multimedia command (image + audio), the dynamic task breakdown module is activated: Multimodal model analysis of images to extract core visual elements: main color tone, font style, layout, and logo position.
[0057] The large language model parses speech and identifies action intent ("application style", "generate text and images"), object ("promotional activities"), and key time constraints ("before next Wednesday").
[0058] The module merges the parsed results to generate a structured chain of subtasks: [Task 1] Define visual specifications: Extract style elements such as color and font from the images → [Task 2] Size adaptation: Generate templates of different sizes for Weibo long images, Xiaohongshu images, WeChat Moments images, etc. → [Task 3] Content generation: Fill in promotional copy for each template → [Task 4] Review and delivery: Complete by next Wednesday.
[0059] III. State Synchronization and Functional Integration (corresponding to S3, S4) The persistent connection channel begins real-time synchronization of subtask states. Users can view this on the same interface: Task progress indicator: Displays "Style extraction complete, generating Weibo template...".
[0060] File Preview: Click on the output file of Task 2, and a preview window will slide out on the right to directly render the generated image templates of different sizes.
[0061] Historical search: When a user enters "the style of the last promotion", the search function uses natural language to locate similar marketing task dialogues and their output files in the past.
[0062] In this embodiment, the decomposition into a structured subtask chain includes: calling the large language model and multimodal model to identify the action entities and object entities in the instructions; generating a modifiable semi-structured task confirmation window; and generating the final subtask dispatch instruction based on the user's adjustment results. Example 8: Decomposition and Confirmation of Multimodal Design Requirements I. Multimodal Command Input and Entity Recognition User interaction: In the dialogue window with the AI, the user enters the text command: "Make this UI concept diagram into an interactive prototype", and at the same time drags and drops an image named Future Home APP Concept Sketch.jpg.
[0063] System calls for multi-model collaborative work: Large Language Model (LLM) parses text instructions, identifies action entities: make (i.e., create, implement), and outputs object entities: interactive prototypes.
[0064] Multimodal modeling (MLLM) parses uploaded images, performs visual understanding, identifies objects in the images such as navigation bars, card lists, bottom tab bars, and floating buttons, and infers possible interactive actions such as clicking and swiping.
[0065] The outputs of the two models are merged to jointly complete the comprehensive identification of "actions" and "objects" in the instructions.
[0066] II. Generate a modifiable semi-structured task confirmation window Based on the identified entities, the system generates an initial task plan and displays it in the confirmation window: "[Main Task] Create an interactive prototype based on the concept art." ├── [Subtask 1] Page Structure Implementation │ ├── Object: Navigation bar (derived from image recognition) │ ├── Object: Card list (derived from image recognition) │ └── Object: Bottom tab bar (derived from image recognition) ├── [Subtask 2] Implementation of Interaction Logic │ ├── Action: Click the "floating button" (derived from image recognition) → Pop-up menu │ └── Action: Slide the "card list" (derived from image recognition) → Load more └── [Subtask 3] Prototype Preview and Export Format: "Figma editable file"; In the confirmation window, all elements and interactions identified from the image are clearly labeled with the source of the [image recognition].
[0067] III. User Adjustment and Final Command Generation After the user reviews the proposal, modifications will be made: Users discovered a detail missing from the multimodal model: a delete option should appear when a card is long-pressed; The user adds a line to subtask 2: "Action: Long press on a card list item → Show delete button".
[0068] After the user clicks "confirm," the system generates a detailed structured instruction (such as in JSON format) based on the final adjustment result and dispatches it to an external UI design tool or front-end coding assistance module for execution.
[0069] In this embodiment, a closed-loop coding process is also included: injecting UI component context in response to drag-and-drop operations; generating business adaptation code in the embedded workbench; and directly deploying the generated code to the associated application environment for execution. Example 9: Closed Loop of UI Code Generation for Order Management System 1. Drag and drop to inject context Users drag and drop the [Order Inquiry Panel] UI component from the VSCode workbench to the code generator; The system automatically captures component metadata: "{"Component Type": "DataGrid", "Binding fields": ["Order ID", "Customer Name", "Amount"], "Action Buttons": ["View Details", "Export"]}”; Context analysis engine output: Requirement: Generate an order data table with query conditions, supporting details navigation and export.
[0070] II. Business Adaptation Code Generation Embedded AI encoder execution: # Generating pseudocode logic 1. Parse UI components → Extract field mapping rules 2. Matching business template: CRM_DataGrid_WithAction 3. Inject the current project's technology stack: React + Ant Design; Output the complete code file, for example: / / OrderTable.jsx import { Table, Button} from 'antd'; const columns = [ { title: 'Order ID', dataIndex: 'orderId'}, { title: 'Customer', dataIndex: 'customer'}, { title: 'Amount', dataIndex: 'amount'}, { title: 'Operation', render: (_, record) =>( <> <button onclick="{()" => showDetail(record)}>View Details< / button> <button onclick="{()" => exportOrder(record)}> Export< / button> ) } ]; const OrderTable = ({ data}) =>( );
[0071] III. Immediate Deployment and Execution The system will execute automatically: Hot update injection: Insert OrderTable.jsx into the current React component tree; Simulated data population: Load 50 order records from the test database; Runtime environment activation: The running results are rendered on the right side of the embedded workbench; User operation verification: 1. Click "View Details" → A pop-up order details drawer component will appear; 2. Click "Export" → Download the CSV file (containing the data of the current page).
[0072] IV. Closed-loop correction mechanism The user discovered that the amount was not formatted: Add the following comment: "Amount fields need to be formatted with thousands separators"; The system regenerates the code and performs hot updates, for example: / / Modified render: (amount) =>( {new Intl.NumberFormat().format(amount)} ) ".
[0073] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0074] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. All equivalent structural transformations made under the inventive concept of the present invention using the contents of the present invention specification and drawings, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present invention.
Claims
1. An intelligent interactive collaborative work system based on AI large-scale model technology, characterized in that: include: The AI interaction framework module provides a unified task entry point. This entry point supports receiving user input and actively inferring user intent to provide dynamic suggestions. The suggestions are derived from the input text, locally stored documents, recently visited web pages, frequently used business systems, and command-line tool history. It also supports directly executing content startup and status operations within the window via shortcut keys. The dynamic task decomposition module parses natural language or multimedia instructions into structured subtasks based on large language models and multimodal models; The automatic status synchronization module receives and stores the execution status of subtasks in real time through a persistent connection channel.
2. The intelligent interactive collaborative work system based on AI large model technology according to claim 1, characterized in that, The AI interaction framework module includes: A semi-structured confirmation window is used to parse ambiguous instructions and provide a secondary modification interface; Agent task progress indicator, visually displaying the execution status and processing history of subtasks; The task details sub-window displays the task context and dialogue history in response to click actions.
3. The intelligent interactive collaborative work system based on AI large model technology according to claim 2, characterized in that, The operation of the semi-structured confirmation window includes: An initial task decomposition scheme for receiving outputs from large language models and multimodal models; It supports users in modifying and splitting task objects and execution steps; The confirmed structured instructions are then sent to the dynamic task breakdown module.
4. The intelligent interactive collaborative work system based on AI large model technology according to claim 1, characterized in that, The execution process of the dynamic task decomposition module includes: Identify the intent, target, and time frame in natural language or multimedia instructions; Generate a structured list of subtasks containing their execution order; Subtasks are dispatched to the automatic state synchronization module and external function modules.
5. The intelligent interactive collaborative work system based on AI large model technology according to claim 1, characterized in that, The automatic status synchronization module includes: Streamable long-lived connection channels enable real-time capture of Agent execution state changes; A state database persistently stores task execution steps, intermediate outputs, and final results; The log recording unit associates the task status with the operation timestamp.
6. The intelligent interactive collaborative work system based on AI large model technology according to claim 1, characterized in that, Also includes: The file preview unit in the secondary window responds to file clicks in the sidebar and renders the file content directly in the floating window. The dialogue history retrieval unit supports natural language retrieval and locating historical dialogue fragments.
7. The intelligent interactive collaborative work system based on AI large model technology according to claim 1, characterized in that, The dynamic task breakdown module works in conjunction with the smart calendar module: Parse time constraints in natural language; Automatically generate calendar events and associate them with subtask time nodes; Monitor the overdue status of tasks through the status synchronization module.
8. An intelligent interactive collaborative operation method based on AI large-scale model technology, applied to the intelligent interactive collaborative operation system based on AI large-scale model technology as described in claims 1 to 9, characterized in that, Includes the following steps: S1. Launch the AI dialogue window using a global hotkey; S2. Receive natural language or multimedia instructions and break them down into a chain of structured subtasks; S3. Synchronize the execution status of subtasks in real time through a persistent connection channel; S4. Integrate task progress indicators, file previews, and historical search functions into a unified interface.
9. The intelligent interactive collaborative operation method based on AI large model technology according to claim 8, characterized in that, The decomposition into a structured subtask chain includes: The large language model and multimodal model are invoked to identify action entities and object entities in the instructions; Generate a modifiable semi-structured task confirmation window; The final subtask dispatch instruction is generated based on the user's adjustments.
10. The intelligent interactive collaborative operation method based on AI large model technology according to claim 8, characterized in that, It also includes the coding closed-loop process: Inject UI component context in response to drag-and-drop operations; Generate business adaptation code in the embedded workbench; The generated code will be deployed directly to the associated application environment for execution.