AI desktop workstation system based on multi-modal interaction and working method

By integrating multimodal signal acquisition and AI decision-making, the AI ​​desktop workstation system solves the problems of limited functionality and data silos in traditional desktop workstations, and achieves efficient and intelligent task execution and improved user experience in complex office scenarios.

CN122019096APending Publication Date: 2026-05-12HANGZHOU LINGFENG INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU LINGFENG INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-02-02
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Traditional desktop workstations have limited functionality and cannot meet the diverse needs of complex office scenarios. They suffer from severe data silos, and intelligent tools cannot deeply understand the context of office work, resulting in high operational complexity and low efficiency.

Method used

The system employs an AI desktop workstation based on multimodal interaction, integrating a high-performance computing unit, a multimodal signal acquisition module, and a dedicated AI processing chip. It collects user voice, gestures, facial expressions, and eye movement data in real time, and performs feature fusion and intent recognition through an AI decision layer to automatically execute office tasks.

Benefits of technology

It enables efficient processing of multimodal data and intelligent linkage of office tasks, reducing the need for users to switch between different devices and improving office efficiency and user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122019096A_ABST
    Figure CN122019096A_ABST
Patent Text Reader

Abstract

The invention discloses an AI desktop work station system based on multi-modal interaction and a working method. The AI desktop work station system comprises a hardware layer, an interface layer and a control layer, wherein the hardware layer is composed of an integrated high-performance computing unit, a multi-modal signal acquisition module, a special AI processing chip (NPU) and at least two display output interfaces; the sensing layer is used for collecting four types of multi-modal data including voice instructions, gesture actions, facial expressions and eye movement tracks of a user in real time; the AI decision-making layer serves as a system core control unit, and a central task scheduling engine is arranged in the AI decision-making layer; and the application layer is in instruction connection with the AI decision layer. According to the invention, the special AI processing chip and the multi-mode acquisition module are integrated through the hardware layer, office context data are generated in real time in combination with the sensing layer, and the AI decision-making layer dynamically schedules multi-mode data processing, so that the problems of single interaction mode, rigid hardware resource allocation and lack of context sensing capability of a traditional workstation are solved; the method has the advantages that the multi-modal data processing efficiency is improved, intelligent office task linkage execution is realized, and the user interaction experience is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent office equipment technology, and in particular to an AI desktop workstation system and its working method based on multimodal interaction. Background Technology

[0002] In today's office environment, traditional desktop workstations have revealed numerous drawbacks. Functionally, they are often limited to single-task processing and cannot meet the diverse needs of complex office scenarios. Software applications operate independently, making data sharing and circulation difficult, thus creating data silos. For example, when conducting data analysis, users may need to export data from one data processing software and then manually import it into report generation software, a cumbersome and error-prone process. When users face complex office tasks, such as simultaneously performing data analysis, report writing, and participating in collaborative meetings, they need to frequently switch between different devices such as computers, mobile phones, and conference tablets, and jump back and forth between various office applications. This not only increases the complexity of operations but also leads to a significant waste of time, seriously affecting work efficiency. While some voice assistants and intelligent software have emerged in the market, their integration with hardware remains relatively loose. In actual office work, these intelligent tools struggle to deeply understand the context of the work environment and provide coherent, intelligent services. For example, when users are using document editing software, voice assistants cannot accurately understand user commands based on the current document content and provide targeted assistance. These issues highlight the urgency and necessity of developing a completely new desktop workstation system that deeply integrates AI capabilities with hardware. Summary of the Invention

[0003] This application aims to provide an AI desktop workstation system and its working method based on multimodal interaction, which has the advantages of improving multimodal data processing efficiency, realizing intelligent office task linkage execution, and enhancing user interaction experience.

[0004] This application provides an AI desktop workstation system based on multimodal interaction, comprising: a hardware layer: consisting of an integrated high-performance computing unit, a multimodal signal acquisition module, a dedicated AI processing chip (NPU), and at least two display output interfaces; a perception layer: connected to the hardware layer, used to acquire four types of multimodal data in real time: user voice commands, gestures, facial expressions, and eye movements, and extracting software interface elements and document type information currently displayed on the screen through image recognition algorithms to generate office context data packages; an AI decision-making layer: serving as the core control unit of the system, with a built-in central task scheduling engine, which connects to AI capability modules via a data bus, including a natural language processing (NLP) module, a computer vision (CV) module, and a user behavior analysis module; and an application layer: connected to the AI ​​decision-making layer's instructions, including a real-time meeting minutes generation module, an intelligent data insight visualization module, and a cross-document knowledge base retrieval and recommendation module, capable of automatically calling the corresponding modules to execute office tasks based on the intent commands output by the AI ​​decision-making layer.

[0005] Furthermore, the multimodal signal acquisition module includes a high-definition camera, a microphone array, and an eye-tracking sensor; the dedicated AI processing chip (NPU) has a computing power of no less than 10 TOPS to support INT8 / FP16 mixed-precision computing.

[0006] Furthermore, the central task scheduling engine can receive multimodal data and office context data packets output by the perception layer, analyze the correlation of user behavior through feature fusion algorithms, and dynamically identify the user's office intentions.

[0007] Furthermore, the high-performance computing unit at the hardware layer adopts a multi-core processor with a main frequency of no less than 3.0GHz and a memory capacity of no less than 32GB, supporting the DDR5 memory protocol; the display output interface includes HDMI 2.1 and DisplayPort 1.4 interfaces, supporting 4K@60Hz dual-screen extended display.

[0008] Furthermore, the image recognition algorithm in the perception layer adopts a lightweight CNN model with no more than 5M model parameters and an inference time of no more than 100ms on a dedicated AI processing chip (NPU); the office context data package also includes the current system time, the editing time of opened documents, and the software operation history.

[0009] Furthermore, the feature fusion algorithm of the AI ​​decision layer adopts an attention mechanism, assigning weight coefficients to voice, image, and eye-tracking data respectively. The weight coefficients are dynamically adjusted according to the user's historical operation preferences, with an adjustment cycle of no more than 24 hours.

[0010] Furthermore, the application layer's real-time meeting minutes generation module supports multilingual transcription, including Chinese, English, and Japanese, with a transcription accuracy of no less than 95%. The intelligent data insight visualization module supports three chart types: line charts, bar charts, and heatmaps, and automatically recommends the optimal chart style based on the data dimensions.

[0011] Furthermore, a working method for an AI desktop workstation based on multimodal interaction is provided. This method is applied to the aforementioned AI desktop workstation system based on multimodal interaction and includes the following steps: S1: Hardware layer initialization, the multimodal signal acquisition module enters real-time monitoring state, the dedicated AI processing chip (NPU) loads model parameters, and the high-performance computing unit establishes communication links with each layer; S2: The perception layer collects user operation data, including a high-definition camera that collects gesture and facial expression images at a sampling frame rate of no less than 30fps; a microphone array that collects voice commands, which are then converted into 16kHz mono audio data after noise reduction; an eye-tracking sensor that records eye movement trajectory coordinates at a sampling frequency of no less than 120Hz; and simultaneously, screen capture technology is used to extract the currently running software process name and document format information to generate an office context dataset containing timestamps; S3: The AI ​​decision layer receives the dataset output from S2, and the central task scheduling engine inputs the multimodal data into the corresponding AI modules: the Natural Language Processing (NLP) module performs semantic parsing on the voice commands, and the Computer Vision (CV) module processes the gestures and facial expressions. The system extracts features from the context image, and the user behavior analysis module builds a user operation sequence model by combining the office context dataset; S4: The central task scheduling engine integrates the output results of various AI modules and compares them with the preset office task library through the intent matching algorithm to determine the user's target task; if the target task is a meeting record, it sends a meeting minutes generation instruction to the application layer; if it is a data analysis task, it sends a data visualization instruction; if it is an information retrieval task, it sends a knowledge base retrieval instruction; S5: After receiving the instruction, the real-time meeting minutes generation module automatically transcribes the audio content and extracts key issues; the intelligent data insight visualization module generates trend charts based on the user's currently opened Excel / CSV file; the cross-document knowledge base retrieval and recommendation module selects the top 5 most relevant documents from the local document library and cloud knowledge base based on context keywords and pushes them to the display output interface; S6: After the task is completed, the AI ​​decision layer monitors the user's feedback on the results through the user behavior analysis module. If the user modifies the minutes content or adjusts the chart style, the task execution model is automatically updated to optimize the accuracy of subsequent intent recognition.

[0012] Furthermore, in step S2, the eye-tracking sensor locates the eye coordinates using corneal reflection, with a positioning accuracy error of no more than 0.5°; the screen capture technology uses the system's underlying API call method to avoid causing lag in software operation, with a capture delay of no more than 50ms.

[0013] Furthermore, the intent matching algorithm in step S4 adopts a deep learning model. The model training dataset contains more than 100,000 user operation samples in office scenarios, and the intent recognition accuracy is no less than 92%. The preset office task library includes four primary tasks: meeting collaboration, data processing, document management, and schedule reminders. Each primary task is further subdivided into no less than five secondary sub-tasks.

[0014] Furthermore, in step S6, user feedback operations include mouse click confirmation, voice command correction, and keyboard editing. The AI ​​decision layer updates the task execution model through reinforcement learning algorithms, uses user feedback results as reward signals, and optimizes the weight allocation logic of the feature fusion algorithm. Beneficial effects

[0015] This invention integrates a dedicated AI processing chip and a multimodal acquisition module at the hardware layer, combines real-time generation of office context data at the perception layer, dynamically schedules multimodal data processing at the AI ​​decision layer, and intelligently executes office tasks at the application layer. This solves the problems of traditional workstations, such as single interaction methods, rigid hardware resource allocation, and lack of context awareness. It has the advantages of improving multimodal data processing efficiency, realizing intelligent office task linkage execution, and enhancing user interaction experience. Attached Figure Description

[0016] The invention will now be further described with reference to the accompanying drawings; Figure 1 This is a system framework diagram proposed in this invention; Figure 2 This is a schematic diagram of the method flow proposed in this invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] The technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of this application described and shown in the accompanying drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0019] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0020] Reference Figure 1 - Figure 2 This application discloses an AI desktop workstation system based on multimodal interaction, comprising a hardware layer, a perception layer, an AI decision-making layer, and an application layer. The hardware layer consists of an integrated high-performance computing unit, a multimodal signal acquisition module, a dedicated AI processing chip, and a display output interface. The perception layer collects voice commands, gestures, facial expressions, and eye-tracking data in real time, and extracts interface elements using image recognition algorithms to generate office context data packages. A central task scheduling engine connects to AI capability modules via a data bus, including a natural language processing module, a computer vision module, and a user behavior analysis module. The application layer includes a meeting minutes generation module, a data visualization module, and a knowledge base retrieval module, executing office tasks based on intent commands.

[0021] The hardware layer includes a multimodal signal acquisition module that integrates multiple sensors, such as a high-definition camera, microphone array, and eye-tracking sensor, to synchronously acquire user behavior data. A dedicated AI processing chip is a neural network processor with parallel computing capabilities, typically implemented using chips supporting mixed-precision computing, to accelerate real-time processing of multimodal data. The perception layer's office context data package contains a set of data including operational environment characteristics, obtained through screen capture technology to acquire software process information and document formats, providing multidimensional environmental parameters for intent recognition. The AI ​​decision-making layer's central task scheduling engine coordinates multi-module collaboration, connected via a data bus to unify the parsing of cross-modal information. The application layer's cross-document knowledge base retrieval and recommendation module is a related information query system, implemented by establishing local and cloud document indexes to automatically filter highly relevant references.

[0022] The hardware layer provides the system with basic computing power and physical interface support, while the multimodal signal acquisition module continuously captures user operation data. The perception layer transforms raw data into structured office context data packets, where image recognition algorithms identify current software interface elements and combine them with timestamps to form operational environment features. The AI ​​decision-making layer parses voice commands through a natural language processing module, extracts gesture and facial expression features through a computer vision module, and establishes an operation sequence model through a user behavior analysis module. The central task scheduling engine integrates multi-dimensional analysis results and matches them with a preset task library to generate intent commands. The application layer calls corresponding functional modules based on the command type: the meeting minutes generation module automatically transcribes voice content, the data visualization module generates charts, and the knowledge base module pushes related documents. Each layer forms a closed-loop workflow from data acquisition to task execution, enabling intelligent operation across software and devices.

[0023] Traditional workstations rely on manual operation of multiple independent software programs, while this solution automatically identifies user intent through multimodal data fusion. Existing systems require manual switching between devices to handle different tasks, while this solution achieves automatic task scheduling through context awareness. Traditional intelligent tools cannot connect with operating environment information, while this solution establishes multi-dimensional environmental parameters through office context data packages, making AI decisions more aligned with real-world scenario needs.

[0024] This application achieves unified acquisition and parsing of multimodal data, reducing the number of times users switch between different devices through context awareness and intent recognition technologies. Deep integration of hardware and AI modules improves the efficiency of complex office tasks. Ultimately, it forms an intelligent office support system capable of automatically identifying user needs and allocating resources. Furthermore, it proposes a multimodal signal acquisition module including a high-definition camera, microphone array, and eye-tracking sensor, with a dedicated AI processing chip boasting a computing power of no less than 10 TOPS and supporting INT8 / FP16 mixed-precision computing.

[0025] A high-definition camera refers to a video capture device with high-resolution imaging capabilities, specifically using an 8-megapixel CMOS sensor, to capture the displacement trajectory of user hand gestures and subtle muscle changes in facial expressions. A microphone array is an audio capture system consisting of multiple microphones arranged in a specific geometric structure, specifically a six-microphone ring array, using beamforming technology to eliminate environmental noise interference. An eye-tracking sensor is a visual focus detection device based on the principle of optical reflection, specifically using an infrared light source in conjunction with a high-speed image sensor, to capture the trajectory of the user's eye movements.

[0026] A high-definition camera continuously captures the user's upper body dynamics at a sampling rate of no less than 30 frames per second, extracting gesture contour features through edge computing. The microphone array employs an adaptive beamforming algorithm, and an eye-tracking sensor records the pupil center coordinates at a sampling frequency of 120Hz, establishing a visual focus mapping relationship in conjunction with the screen display area. A dedicated AI processing chip, through hardware-level memory optimization, maintains a 1.2x throughput improvement when running image recognition algorithms. Multimodal data streams, after passing through a timestamp alignment module, enter the mixed-precision computing unit in a pipelined manner, enabling parallel processing of speech, vision, and eye-tracking data.

[0027] Traditional workstations typically use a combination of a standalone camera and a unidirectional microphone, which cannot achieve coordinated spatial sound source localization and visual focus tracking. Existing general-purpose processors suffer from computational resource contention when processing multimodal data, resulting in speech recognition latency exceeding 200 milliseconds. This solution optimizes the spatial layout of multiple sensors, improving gesture recognition accuracy to 98% while keeping multimodal data processing latency below 80 milliseconds. Compared to a pure FP16 computing mode, the hybrid precision architecture of the dedicated AI processing chip reduces memory bandwidth usage by 40%.

[0028] This application achieves synchronous acquisition and efficient processing of multi-dimensional user interaction data, ensuring precise alignment of voice commands, gesture operations, and visual focus in the temporal dimension. Hardware computing resource utilization is increased to over 85%, supporting continuous multimodal interaction task processing for 8 hours without performance degradation. The real-time matching error between eye-tracking data and screen interface elements is controlled within 3 pixels, providing reliable input for subsequent intent recognition.

[0029] This application further proposes a technical solution for the central task scheduling engine to receive multimodal data and office context data packets output by the perception layer, and to analyze the correlation of user behavior and dynamically identify the user's office intentions through feature fusion algorithms.

[0030] The central task scheduling engine is the core control unit responsible for coordinating multi-source data processing. It can be implemented using a distributed message queue architecture to decouple data reception from task distribution. Multimodal data refers to heterogeneous data sets including voice, gestures, and eye-tracking trajectories. Synchronous transmission can be achieved through timestamp alignment technology to ensure the temporal consistency of data from different sensors. Office context data packages are metadata datasets containing software interface elements, document types, and operation history. These can be encapsulated using JSON structured data format to provide environmental awareness information. Feature fusion algorithms are mathematical models that integrate features from multiple data sources. They can be implemented using a graph neural network architecture to mine cross-modal associations by constructing user behavior relationship graphs. User behavior correlation refers to the strength of logical connections between different operational behaviors. State transition probabilities can be calculated using a Hidden Markov Model to quantify the coherence of user operation sequences. Dynamic recognition of user office intentions refers to a decision-making mechanism that analyzes user needs in real time. This can be achieved using online learning algorithms to update model parameters and adapt to constantly changing office scenarios.

[0031] The raw multimodal data collected by the perception layer is preprocessed to form a standardized data stream, which is then transmitted to the central task scheduling engine along with the office context data package. The feature fusion algorithm first encodes the speech commands into semantic vectors, converts gestures into spatial coordinate sequences, and maps eye-tracking data to a screen hotspot distribution map. Document type information in the office context data package is converted into domain knowledge vectors, which, together with the current system time and operation history, constitute the environmental feature matrix. After receiving the above multi-dimensional feature inputs, the graph neural network model calculates the contribution weights of each modality feature through an attention mechanism, generating a user behavior relevance scoring matrix. The online learning module adjusts model parameters based on real-time feedback data. When a document type switch or software interface change is detected, incremental training of the intent recognition model is automatically triggered to ensure dynamic adaptation to the new office context environment.

[0032] Traditional desktop workstations typically only process single-modal input data and lack the ability to continuously perceive the office context. Existing intent recognition systems mostly use static rule engines, which cannot effectively handle the non-linear relationships between multimodal data and are even less able to adapt to dynamically changing office task scenarios. This solution overcomes the technical bottlenecks of traditional systems in data integration and adaptive capabilities by constructing a multimodal feature fusion and dynamic model update mechanism.

[0033] Through the above technical solution, this application can accurately capture users' real needs in complex office scenarios. For example, when processing data analysis tasks simultaneously during a video conference, the system can combine keywords in voice commands, chart areas pointed to by gestures, and document paragraphs focused on by eye movements to automatically identify the user's combined intent to generate meeting minutes and update data visualizations. When the user switches from document editing to email processing, the feature fusion algorithm can quickly adjust the weight allocation based on the operation history, avoiding misjudgment of intent caused by scene switching.

[0034] This application further proposes that the high-performance computing unit at the hardware layer adopts a multi-core processor. The parallel computing architecture of the multi-core processor is configured to simultaneously process multimodal data streams from the perception layer and algorithmic computation tasks from the AI ​​decision layer. For example, in the parallel execution of voice command parsing and image feature extraction, the multi-core architecture can avoid task blocking. The DDR5 memory protocol improves data throughput through a dual-channel data transmission mechanism. For example, when loading multiple large documents simultaneously, the memory subsystem can quickly respond to data read and write requests. The combination of HDMI 2.1 and DisplayPort 1.4 interfaces is designed to support independent dual-screen display control. For example, in data analysis scenarios, one screen can display raw data tables, while the other screen simultaneously presents visualized charts. The bandwidth allocation mechanism of the two interfaces ensures zero-latency screen refresh.

[0035] Traditional workstations typically use a single-core processor paired with DDR4 memory, which can easily lead to computational bottlenecks when running multimodal data processing tasks. Furthermore, the single-display interface design cannot meet the needs of multi-window collaborative operation. This solution, through a collaborative design of a multi-core processor and DDR5 memory, improves data processing speed by approximately 40%. Simultaneously, the resolution and refresh rate supported by the dual-display interface are more than 50% higher than traditional solutions, effectively solving the performance degradation problem during multi-task parallel processing.

[0036] This application enables efficient data processing and multi-screen collaborative operation in complex office scenarios. The combination of a multi-core processor and DDR5 memory significantly reduces the waiting time for loading large documents and data analysis. For example, when handling video conferencing and data modeling tasks simultaneously, the system response time is reduced to 60% of the original level. The high bandwidth of the dual display interfaces supports the simultaneous display of multiple application windows at 4K resolution. For example, when performing cross-document comparison, the two screens can display original files in different formats respectively, eliminating operation interruptions caused by window switching.

[0037] This application further proposes that the image recognition algorithm of the perception layer adopts a lightweight CNN model with no more than 5M model parameters and an inference time of no more than 100ms on a dedicated AI processing chip; the office context data package also includes the current system time, the editing time of the opened document, and the software operation history.

[0038] Lightweight CNN models refer to convolutional neural networks built by reducing the number of network layers and channels. Specifically, they can be implemented using MobileNet or ShuffleNet architectures, reducing computational complexity through depthwise separable convolutions. Dedicated AI processing chips refer to embedded neural network processors with parallel computing capabilities. Specifically, they can be implemented using chip architectures with tensor computing cores, improving inference efficiency through hardware acceleration. Office context data packages refer to datasets containing user operating environment characteristics. Specifically, they can be obtained through system API calls to retrieve timestamps, document metadata, and operation logs, enhancing context awareness through multi-dimensional data fusion.

[0039] Image recognition algorithms reduce the number of parameters by optimizing the network structure, enabling the model to adapt to the storage capacity and computing resources of embedded processing chips, avoiding the dependence of traditional large models on discrete graphics cards. The parallel computing architecture of dedicated processing chips accelerates the model inference process, ensuring that image feature extraction is completed within a limited time. New time-dimensional data added to the office context data package is used to analyze user operation rhythm, document editing time data is used to evaluate task priority, and operation history data is used to build a user behavior sequence model. These three elements together form a contextual feature system with temporal correlation. This combination of technologies, while maintaining real-time performance, compensates for the cognitive limitations of traditional systems that rely solely on current interface elements through multi-dimensional data analysis.

[0040] Compared to existing technologies, traditional intelligent tools typically use general-purpose computing units to run large image recognition models, resulting in processing latency that fails to meet real-time interaction requirements. Existing context-aware systems only collect information about current interface elements, lacking the ability to track user operating habits over long periods. This solution achieves efficient computation in an embedded environment through the co-design of dedicated hardware and lightweight models. Simultaneously, it constructs dynamically updated user profiles through multi-dimensional data fusion, significantly improving the accuracy and response speed of intent recognition.

[0041] Through the above technical solution, this application effectively solves the interaction latency problem caused by insufficient computing resources in traditional systems, achieving real-time data processing through the combination of low-power hardware and optimized models. The newly added contextual dimension data enables the system to accurately capture user operation patterns and predict behavioral intentions by combining time series analysis, avoiding misunderstandings of instructions due to missing information. This combination of technologies enhances the contextual understanding capability in complex office scenarios while maintaining system response speed.

[0042] This application further proposes an AI decision-making layer feature fusion algorithm that adopts an attention mechanism, assigning weight coefficients to speech, image, and eye-tracking data respectively. The weight coefficients are dynamically adjusted according to the user's historical operation preferences, with an adjustment cycle of no more than 24 hours.

[0043] Attention mechanisms refer to techniques that allocate processing priorities by calculating the relevance of different modalities to the current task. Specifically, this can be implemented using a multi-head self-attention module combined with a gated recurrent unit to capture the semantic relationships between multimodal data. The weight coefficients are numerical parameters reflecting the importance of different modalities in the decision-making process. These can be implemented by modeling the probability distribution of historical operation data using a normalized exponential function to balance the matching relationship between user habits and real-time data. Dynamic adjustment refers to the process of automatically updating the weight allocation strategy based on changes in user behavior. This can be achieved by using a sliding time window to statistically analyze user operation frequency and modal usage preferences, ensuring the system's adaptability to changes in user behavior.

[0044] When the system receives voice, image, and eye-tracking data, the attention mechanism first extracts the feature vectors of each modality and calculates their relevance score to the current work task. Based on the score, the system assigns a first weight coefficient to voice commands, a second weight coefficient to gesture images, and a third weight coefficient to eye-tracking trajectories. The user behavior analysis module continuously records user operations in meeting minutes, data analysis, and document retrieval scenarios. By statistically analyzing the number of times each modality data was actively corrected by the user and the final adoption rate over 24 hours, dynamic adjustment coefficients are generated. The central task scheduling engine weights and fuses these dynamic adjustment coefficients with the real-time calculated attention weights to form the final comprehensive weight value used for intent recognition.

[0045] Compared to existing technologies, traditional multimodal fusion methods employ fixed weight allocation strategies, which cannot adapt to the differences in operating habits among different users. For example, users accustomed to voice commands may misjudge operations due to excessive weighting of gesture data. This solution introduces an attention mechanism and dynamic preference learning, enabling the system to automatically optimize the contribution of each modality based on the user's actual behavior patterns. While maintaining stability over a 24-hour period, it effectively reduces intent recognition bias caused by differences in user habits.

[0046] Through the above technical solution, this application solves the problem of insufficient intent recognition accuracy caused by differences in user operating habits during multimodal data fusion. In meeting recording scenarios, the system can automatically increase the weighting coefficient of voice data based on the user's frequent use of voice to correct meeting minutes; in data analysis scenarios, when a user continuously focuses on a specific chart through eye tracking, the system will increase the decision-making weight of eye tracking data. This dynamic adaptation mechanism significantly improves interaction accuracy and reduces the frequency of manual correction operations by users.

[0047] This application further proposes an application-layer real-time meeting minutes generation module that supports multilingual transcription, including Chinese, English, and Japanese, and an intelligent data insight visualization module that supports three chart types: line chart, bar chart, and heatmap, automatically recommending the optimal chart style based on the data dimension.

[0048] Multilingual transcription refers to the cross-language text conversion of speech content. Specifically, it can be achieved using a speech recognition framework based on a pre-trained acoustic model. By constructing a multilingual phoneme mapping table and language-specific decoders, it addresses recognition errors caused by differences in pronunciation between different languages. Automatic recommendation of optimal chart styles based on data dimensions refers to matching visualization formats according to data structure features. Specifically, feature engineering can be used to extract time-series attributes, classification label distribution, and spatial correlation indicators from the data. A pre-defined rule engine then triggers the corresponding chart generation logic, eliminating subjective bias from manual chart selection.

[0049] The real-time meeting minutes generation module, upon receiving multilingual voice input, automatically distinguishes the current language type using a language detection model, calls the corresponding language's acoustic model and dictionary resources for decoding, and generates timestamped text records. The intelligent data insight visualization module, after parsing the user-opened spreadsheet file, performs feature analysis on the data columns. When a date field is detected with continuously changing values, it automatically triggers a line chart generation command; when a correspondence between discrete category labels and numerical data is identified, a bar chart is used for comparison; and when the data contains latitude and longitude coordinates or matrix distribution characteristics, a heatmap is prioritized to present spatial density information.

[0050] Compared to existing technologies, traditional meeting recording systems only support transcription in a single language, making them unsuitable for multinational team collaboration scenarios. Furthermore, data visualization tools require users to manually select chart types, lacking intelligent analysis capabilities based on data structure characteristics. This solution achieves synchronous recording of cross-language content through a multilingual parallel decoding mechanism. Combined with an automatic data feature recognition algorithm, it can accurately recommend chart types based on common data structure characteristics found in office scenarios.

[0051] Through the above technical solutions, this application effectively solves the language barrier problem in cross-border conference scenarios, ensures the complete recording of multilingual conference content, and reduces the complexity of user operations through an automated chart recommendation mechanism, avoids data display deviations caused by improper manual selection, and improves information processing efficiency and data display adaptability in office scenarios.

[0052] This application further proposes a working method including the following steps: hardware layer initialization, multimodal signal acquisition module enters real-time monitoring state, dedicated AI processing chip loads model parameters, and high-performance computing unit establishes communication links with each layer; perception layer collects user operation data, captures gesture and facial expression images through high-definition camera, captures voice commands through microphone array, records eye movement trajectory coordinates through eye-tracking sensor, and extracts the currently running software process name and document format information to generate an office context dataset; AI decision layer inputs multimodal data into the corresponding AI module for processing, and determines the user's target task through intent matching algorithm; application layer triggers different functional modules to execute office tasks according to the instruction type; monitors user feedback operations and updates the task execution model to optimize recognition accuracy.

[0053] The multimodal signal acquisition module refers to a composite sensor system integrating vision, hearing, and eye tracking. Specifically, it can use a high-definition camera with an infrared illumination module for gesture capture, a ring microphone array with beamforming algorithms for voice acquisition, and an eye-tracking sensor for gaze localization based on pupil-corneal reflection. These devices work together to simultaneously capture multidimensional user behavior data. The office context dataset refers to operational environment information containing timestamps. Specifically, it can use window handle capture technology to obtain the currently active software process, identify document format types through file header parsing, and combine this with the system clock to generate a time-stamped dataset for constructing a spatiotemporal correlation model of user operation scenarios. The intent matching algorithm is a task recognition mechanism based on multimodal features. Specifically, it can use a multi-label classification model to predict the task type from the fused feature vectors, and match them with typical scenarios in a preset task library using cosine similarity calculation to achieve a mapping transformation from raw data to office intent. The feedback optimization mechanism is a dynamic update system that adjusts model parameters based on user operations. Specifically, it can use an online learning framework to receive user correction behavior data in real time, and update the neural network weight parameters through a backpropagation algorithm, allowing the system to continuously adapt to individual user operating habits.

[0054] This method establishes a basic environment for multimodal data acquisition and processing through hardware initialization, ensuring low-latency communication between sensors and computing units. During data acquisition, visual sensors capture subtle changes in body movements in frame sequences, audio acquisition devices extract clean speech signals through noise suppression, and eye-tracking devices continuously record the path of gaze focus shifts. These elements, combined with screen content capture, form a complete operational context record. The data processing stage employs a modular division of labor mechanism, inputting speech, image, and behavioral data into dedicated analysis units to extract feature vectors. Feature fusion is then used to construct a multi-dimensional representation of user intent. In the task determination stage, the fused features are matched against a pre-defined task library, triggering corresponding application modules to execute automated processing flows based on the matching results. The execution result feedback stage monitors user modifications to the system output, generating reward signals for reinforcement learning and continuously optimizing the decision logic of the intent recognition model.

[0055] Compared to existing technologies, traditional office systems employ a single-modal input method and lack context awareness; for example, voice assistants can only process isolated voice commands and cannot relate them to the current document content. This method uses multi-sensor collaborative data acquisition to create a three-dimensional model of the operational scenario, constructing a multi-dimensional contextual dataset including time, space, and application environment, enabling intent recognition to possess scene relevance. Existing technologies often drive task execution based on preset rules, failing to adapt to dynamically changing office needs. This method employs feature fusion and dynamic matching mechanisms, continuously optimizing model parameters through online learning to form a closed-loop processing system, effectively improving task recognition accuracy in complex scenarios.

[0056] Through the above technical solution, this application achieves synchronous acquisition and fusion analysis of multimodal data across application scenarios, and can automatically identify complex office intentions and trigger corresponding functional modules. In meeting scenarios, the system can simultaneously process speech transcription, speaker gesture recognition, and document content association to generate intelligent meeting minutes with key points marked. In data analysis scenarios, by combining user gaze focus distribution and historical operating habits, the system automatically recommends the optimal visualization scheme and generates interactive charts. This method reduces the manual switching between multiple software interfaces and achieves automated pipeline processing of office tasks through context awareness and intent prediction.

[0057] This application further proposes that the eye-tracking sensor locates eye coordinates using corneal reflection, with a positioning accuracy error of no more than 0.5 degrees; the screen capture technology adopts the system's underlying API call method to avoid affecting software operation, with a capture delay of no more than 50 milliseconds.

[0058] The corneal reflectance method is a technique that tracks eye movements by measuring changes in the position of the reflected light spot on the surface of the eyeball. Specifically, it can be achieved using an infrared light source and a high-speed camera working together. By analyzing the relative displacement between the corneal reflectance point and the center of the pupil, high-precision positioning can be achieved. The system's low-level API call method refers to a technique that bypasses the operating system's graphical interface layer to directly access the display buffer data. This can be implemented using the graphics device interface functions provided by the operating system. By reducing the number of data transfer layers, resource consumption can be reduced.

[0059] During user interaction, an eye-tracking sensor illuminates the surface of the eyeball with an infrared light source to create a reflected light spot. A high-speed camera captures the movement of this light spot at a sampling frequency of at least 120 Hz, and the gaze point is calculated by combining this with the pupil center coordinates. Screen capture technology directly reads pixel data from the video memory by calling the operating system kernel-level graphics interface, avoiding interference from traditional screenshot software on the graphics rendering pipeline. These two technologies work together to ensure both the accuracy of user intent recognition and the real-time performance of the software.

[0060] Traditional eye-tracking often uses pupil-center tracking, whose accuracy is greatly affected by ambient light and head movement. Corneal reflection tracking, however, significantly improves tracking stability by establishing a stable reflection reference point. Conventional screen capture technologies rely on graphical interface layer screenshot commands, which can easily cause interface refresh delays. In contrast, the underlying API call method, by directly accessing video memory data, effectively avoids interface lag.

[0061] Through the above technical solution, this application can accurately capture the user's gaze point position, ensuring that the subsequent intent analysis module obtains reliable visual attention data, while maintaining the smooth operation of office software and avoiding work interruptions caused by data collection. The improved eye-tracking coordinate positioning accuracy ensures the accuracy of user interface element click prediction, while the low-latency screen capture technology enables real-time synchronization of the operation interface state, providing time-aligned basic data for multimodal data analysis.

[0062] This application further proposes that the intent matching algorithm in step S4 adopts a deep learning model, the model training dataset contains more than 100,000 user operation samples in office scenarios, and the intent recognition accuracy is not less than 92%; the preset office task library includes four primary tasks: meeting collaboration, data processing, document management, and schedule reminders, and each primary task is further subdivided into no less than 5 secondary sub-tasks.

[0063] Deep learning models refer to algorithmic frameworks that analyze multimodal user behavior data through multi-layered neural network structures. Specifically, they can be implemented using Transformer architecture or graph neural networks, capturing potential correlation patterns in user operation sequences through end-to-end training. The training dataset contains over 100,000 user operation samples from various office scenarios, meaning it constructs a behavioral data set covering typical scenarios such as meeting minutes, report creation, and document retrieval. This is achieved through a combination of enterprise office system log collection and manual annotation, ensuring that the samples cover the operational characteristics of different positions and workflows. The intent recognition accuracy rate of no less than 92% refers to the model's classification accuracy on the test set, which is verified using cross-validation methods, with F1-score used as the evaluation metric to ensure a balance between precision and recall. The four primary task categories in the pre-defined office task library are basic task types categorized based on office scenario characteristics. Specifically, core scenarios can be determined through cluster analysis of office software usage logs. The secondary subtasks under each primary task category are implemented through business process decomposition; for example, document management can be further subdivided into version control, permission allocation, and format conversion.

[0064] The deep learning model processes multimodal input data, including voice command text, gesture feature vectors, and eye-tracking coordinates, through an encoder structure. After generating a fused feature representation, it calculates similarity with standard task templates in a task library. During model training, a contrastive learning strategy is employed, using positive and negative sample pairs to enhance the model's ability to distinguish similar intentions. A pre-defined office task library stores task definitions in a hierarchical tree structure. After the model outputs preliminary intention classification results, the system traverses the task tree downwards to match the most suitable secondary sub-tasks for the current office context. For example, after recognizing an intention related to "data processing," it further distinguishes between data cleaning, trend prediction, and visualization requirements.

[0065] This application can accurately analyze users' complex operational intentions in multi-task scenarios such as collaborative meetings and report creation, and automatically match them to refined office sub-tasks such as document version comparison and pivot table generation. The tree-structured design of the task library allows the system to maintain coverage of mainstream scenarios while supporting the rapid expansion of service scope by adding second-level sub-task nodes, adapting to the personalized business process needs of different enterprises.

[0066] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

[0067] The above formulas are all derived from software simulation using a large amount of data, and are selected to be close to the actual values. The coefficients in the formulas are set by those skilled in the art based on the actual situation. The above are only preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or changes made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. An AI desktop workstation system based on multimodal interaction, characterized in that, include: Hardware layer: Composed of an integrated high-performance computing unit, a multimodal signal acquisition module, a dedicated AI processing chip (NPU), and at least two display output interfaces; Perception layer: Connected to the hardware layer, it is used to collect four types of multimodal data in real time: user voice commands, gestures, facial expressions, and eye movements. It also uses image recognition algorithms to extract the software interface elements and document type information currently displayed on the screen and generate office context data packages. AI Decision Layer: As the core control unit of the system, it has a built-in central task scheduling engine. The central task scheduling engine is connected to the AI ​​capability modules through a data bus. The AI ​​capability modules include a natural language processing (NLP) module, a computer vision (CV) module, and a user behavior analysis module. Application Layer: Connects to AI decision-making layer commands and includes a real-time meeting minutes generation module, an intelligent data insight visualization module, and a cross-document knowledge base retrieval and recommendation module.

2. The AI ​​desktop workstation system based on multimodal interaction according to claim 1, characterized in that, The multimodal signal acquisition module includes a high-definition camera, a microphone array, and an eye-tracking sensor; the dedicated AI processing chip (NPU) has a computing power of no less than 10 TOPS and is used to support INT8 / FP16 mixed-precision calculations.

3. The AI ​​desktop workstation system based on multimodal interaction according to claim 2, characterized in that, The central task scheduling engine can receive multimodal data and office context data packets output by the perception layer, analyze the correlation of user behavior through feature fusion algorithms, and dynamically identify the user's office intentions.

4. The AI ​​desktop workstation system based on multimodal interaction according to claim 1, characterized in that, The high-performance computing unit of the hardware layer adopts a multi-core processor with a main frequency of no less than 3.0GHz and a memory capacity of no less than 32GB, supporting the DDR5 memory protocol; the display output interface includes HDMI2.1 and DisplayPort1.4 interfaces, supporting 4K@60Hz dual-screen extended display.

5. The AI ​​desktop workstation system based on multimodal interaction according to claim 1, characterized in that, The image recognition algorithm of the perception layer adopts a lightweight CNN model with no more than 5M model parameters and an inference time of no more than 100ms on a dedicated AI processing chip (NPU); the office context data packet also includes the current system time, the editing time of the opened documents, and the software operation history.

6. The AI ​​desktop workstation system based on multimodal interaction according to claim 3, characterized in that, The feature fusion algorithm of the AI ​​decision layer adopts an attention mechanism, assigning weight coefficients to voice, image, and eye-tracking data respectively. The weight coefficients are dynamically adjusted according to the user's historical operation preferences.

7. The AI ​​desktop workstation system based on multimodal interaction according to claim 1, characterized in that, The application layer's real-time meeting minutes generation module supports multilingual transcription, including Chinese, English, and Japanese, with a transcription accuracy of no less than 95%. The intelligent data insight visualization module supports three chart types: line charts, bar charts, and heatmaps, and automatically recommends the optimal chart style based on the data dimensions.

8. A method for operating an AI desktop workstation based on multimodal interaction, wherein the method is applied to the AI ​​desktop workstation system based on multimodal interaction as described in any one of claims 1 to 7, characterized in that, Includes the following steps: S1: Hardware layer initialization, multimodal signal acquisition module enters real-time monitoring state, dedicated AI processing chip (NPU) loads model parameters, high-performance computing unit establishes communication links with each layer; S2: The perception layer collects user operation data, including a high-definition camera that captures gestures and facial expression images with a sampling frame rate of no less than 30fps; a microphone array that collects voice commands, which are then converted into 16kHz mono audio data after noise reduction; an eye-tracking sensor that records eye movement trajectory coordinates with a sampling frequency of no less than 120Hz; and screen capture technology that extracts the process name and document format information of the currently running software to generate an office context dataset containing timestamps. S3: The AI ​​decision layer receives the dataset output by S2. The central task scheduling engine inputs the multimodal data into the corresponding AI modules: the Natural Language Processing (NLP) module performs semantic parsing of voice commands, the Computer Vision (CV) module extracts features from gesture and facial expression images, and the User Behavior Analysis module combines the office context dataset to build a user operation sequence model. S4: The central task scheduling engine integrates the output results of each AI module, compares them with the preset office task library through the intent matching algorithm, determines the user's target task, and sends a meeting minutes generation instruction to the application layer if the target task is a meeting record, a data visualization instruction if it is a data analysis, and a knowledge base retrieval instruction if it is an information retrieval. S5: After receiving the instruction, the real-time meeting minutes generation module automatically transcribes the audio content and extracts key topics. The intelligent data insight visualization module generates trend charts based on the user's currently open Excel / CSV file. The cross-document knowledge base retrieval and recommendation module selects the top 5 most relevant documents from the local document library and cloud knowledge base based on context keywords and pushes them to the display output interface. S6: After the task is completed, the AI ​​decision-making layer monitors the user's feedback on the results through the user behavior analysis module. If the user modifies the minutes or adjusts the chart style, the task execution model is automatically updated to optimize the accuracy of subsequent intent recognition.

9. The AI ​​desktop workstation working method based on multimodal interaction according to claim 8, characterized in that, In step S2, the eye-tracking sensor locates the eye coordinates using corneal reflection, and the screen capture technology uses the system's underlying API call method to avoid causing lag in software operation.

10. The AI ​​desktop workstation working method based on multimodal interaction according to claim 8, characterized in that, The intent matching algorithm in step S4 uses a deep learning model. The model training dataset contains more than 100,000 user operation samples in office scenarios. The preset office task library includes four primary tasks: meeting collaboration, data processing, document management, and schedule reminders. Each primary task is further subdivided into no less than five secondary sub-tasks. In step S6, user feedback operations include mouse click confirmation, voice command correction, and keyboard editing. The AI ​​decision layer updates the task execution model through reinforcement learning algorithms, uses user feedback results as reward signals, and optimizes the weight allocation logic of the feature fusion algorithm.