Multi-modal intelligent input method system and implementation method

Through the multi-modal intelligent input method system, the cloud disk knowledge base is integrated, the input data is automatically captured and classified management, and the cross-device knowledge system synchronization and real-time intelligent assistance are supported. The pain points of file search and cross-platform operation are solved, the input efficiency and convenience are improved, and an intelligent and personalized input environment is built.

CN120353348APending Publication Date: 2025-07-22张翼飞
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510432744.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing input method has problems such as scattered file storage, complex retrieval, and cumbersome operation in file search and cross-platform operation, resulting in inefficient input and unable to meet users' efficient input needs.

Method used

It provides a multimodal intelligent input method system, which realizes automatic data crawling, classified storage and cross-device synchronization through cloud disk knowledge base integration and management, input data crawling and classification management, cross-device knowledge system synchronization, and real-time intelligent auxiliary functions, and provides real-time operation suggestions.

Benefits of technology

It significantly improves the user's input efficiency and convenience, reduces file search time, improves the convenience and consistency of data management, enhances the seamless experience across devices, and improves input accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120353348A_ABST
    Figure CN120353348A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of information input interaction of software application, and particularly relates to a multi-mode intelligent input method system and an implementation method. According to the multi-modal intelligent input method system, by innovatively integrating the cloud disk knowledge base, automatic grabbing and classified management of input data are achieved, synchronization of a cross-device knowledge system is supported, a real-time intelligent auxiliary function is provided, the problem that a current input method has pain points in the aspects of file searching, cross-platform operation and the like is solved, and the user experience is improved. According to the system, the text searching efficiency in the interaction process of the user and the application software is improved, the data management of a client is optimized, cross-device seamless experience can be achieved, intelligent real-time assistance is provided, the input efficiency and convenience of the user are remarkably improved, and an intelligent and personalized input environment can be created for the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information input and interaction of software applications, and particularly relates to a multi-modal intelligent input method system and an implementation method thereof. Background Art

[0002] In modern digital office and information interaction scenarios, users face many troubles when using existing input methods, which seriously affect the input efficiency and experience. The main problems are reflected in the following two key aspects:

[0003] (1) The contradiction between file search and input process: When performing input operations in various application software (including APPs and computer software), users' needs are not limited to simple text entry. For example, in copywriting creation, project planning and other work, it is often necessary to refer to existing files stored locally or in the cloud to obtain inspiration, materials, and paste for reference. However, the current operation process has the following significant drawbacks: 1. The dispersion of file storage: Files may be stored dispersedly on different local devices, such as USB drives, home computers, company computers, etc., and may also be distributed in cloud disks provided by multiple different service providers, such as Huawei Cloud Disk, Baidu Cloud Disk, etc. This dispersed storage method makes it extremely difficult for users to find specific files, consuming a large amount of time and energy. 2. The complexity of file opening and retrieval: Even if the storage location of the file is determined, opening the file requires the use of specialized application software, such as using Word to open documents, specific graphic software to open pictures, etc. In addition, retrieving information within the file is also quite cumbersome. This series of operations frequently interrupt the user's input thinking, seriously reducing the input efficiency.

[0004] (2) The complexity of cross-platform operations: During the input process, users often need to frequently switch between multiple application software. Taking the creation of a complete copy as an example, it may be necessary to first form a draft on an AI platform that is good at copywriting creation, and then switch to other software for subsequent processing such as polishing, typesetting, and format adjustment. In this process, users need to constantly perform copy and paste operations. The operation process is cumbersome, not only increasing the possibility of errors, but also greatly reducing the work efficiency.

[0005] It can be seen that the existing input methods can no longer meet the growing demand of users for efficient input. There is an urgent need for an innovative solution to optimize the input process, improve the user experience, and meet the personalized usage needs of users. Summary of the Invention

[0006] The present invention aims to provide a multi-modal intelligent input method system and an implementation method thereof, so as to realize personalized intelligent auxiliary input information through the multi-modal intelligent input method system, and improve the efficiency of information interaction and the quality of input information of users in various application software.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] Provide a multimodal intelligent input method system, including:

[0009] A cloud disk knowledge base integration and management module, which is used to log in to the cloud disk knowledge base and retrieve and organize the data stored in the cloud disk knowledge base according to the user input text;

[0010] An input data capture and classification management module: used to capture and classify and store the multimodal interaction data between the user and the application software on the current device in real time;

[0011] A cross-device knowledge system synchronization module: used to build a unified user account system, associate the data generated by the user on each device with the user account, unify the data in each device of the user to the cloud server and merge it into the knowledge base, and supply each device to synchronize the data of the knowledge base to the local;

[0012] A real-time intelligent assistance module: used to parse the interaction information input by the user, screen and match relevant data from the knowledge base based on the parsing result, and generate real-time operation suggestions for the user based on the screened data.

[0013] Preferably, the cloud disk knowledge base integration and management module includes:

[0014] A multi-cloud disk association sub-module: used for the user to log in and authorize multiple cloud disks to realize the unified association of various cloud disks with the input method system;

[0015] An intelligent retrieval and organization sub-module: used to analyze the user input content in real time, retrieve relevant information in the associated cloud disk files, and refine and organize it for presentation to the user.

[0016] Preferably, the input data capture and classification management module includes:

[0017] A data capture sub-module: used to capture the interaction data between the user and the application software in real time based on the data capture method and perform local encryption processing. The data capture method includes hook functions, process monitoring and clipboard listening;

[0018] A classification storage sub-module: used to classify and store the captured data in the database on the current device according to multi-dimensional information including application software type, data format and operation time, add metadata tags and build indexes.

[0019] Preferably, the cross-device knowledge system synchronization module includes:

[0020] Account system and data identification sub-module: used to build a unified user account system, assign a unique ID to the user, associate the data generated by the user on each device with this ID, and clarify the data ownership;

[0021] Data synchronization and version control sub-module: used to realize the data interaction between the user's each device and the cloud server, unify the data of each device to the cloud server for data synchronization, and push the data in the latest version of the knowledge base after synchronization to each device of the user.

[0022] Preferably, the real-time intelligent assistance module includes:

[0023] Real-time analysis and matching sub-module: used to parse the user interaction information based on the technologies of natural language processing and knowledge graph, combined with the application software environment information, and screen and match relevant data from the knowledge base;

[0024] Intelligent suggestion generation sub-module: used to generate real-time operation suggestions based on the screened data through a deep learning model, and present the real-time operation suggestions to the user in the form of a pop-up window or quick input after quality evaluation.

[0025] Preferably, the intelligent retrieval and sorting sub-module retrieves relevant information in the associated cloud disk files in the following way: for text files, calculate the similarity using the TF-IDF algorithm combined with the BM25 algorithm, and for picture files, extract feature vectors using an image recognition algorithm for comparison and retrieval.

[0026] Preferably, the data capture sub-module locally encrypts the captured data using the AES encryption algorithm, and the classification storage sub-module stores the data using a MySQL or PostgreSQL database and constructs a B+ tree index.

[0027] Preferably, the data synchronization and version control sub-module transmits data through a network channel based on the SSL / TLS protocol, and uses a conflict resolution algorithm based on timestamps to handle data conflicts.

[0028] The present invention also provides a multi-modal intelligent input implementation method, which includes:

[0029] Cloud disk knowledge base integration and management: log in to the cloud disk knowledge base, retrieve and sort the data stored in the cloud disk knowledge base according to the user input text;

[0030] Input data capture and classification management: capture the multi-modal interaction data of the user with the application software on the current device in real time and classify and store it;

[0031] Cross-device knowledge system synchronization: Build a unified user account system, associate the data generated by the user on each device with the user account, unify the data on each user device to the cloud server and merge it into the knowledge base, and each device synchronizes the data in the knowledge base to the local;

[0032] Real-time intelligent assistance: Analyze the interactive information input by the user, screen and match relevant data from the knowledge base based on the analysis results, and generate real-time operation suggestions for the user based on the screened data.

[0033] Preferably, the cloud disk knowledge base integration and management includes:

[0034] Multi-cloud disk association: The user authorizes the login of multiple cloud disks to achieve the unified association of various cloud disks with the input method system;

[0035] Intelligent retrieval and collation: Analyze the content input by the user in real time, retrieve relevant information in the associated cloud disk files, and refine and present it to the user;

[0036] The input data capture and classification management includes:

[0037] Data capture: Based on the data capture method, capture the user's interaction data with the application software in real time and perform local encryption processing. The data capture methods include hook functions, process monitoring, and clipboard listening;

[0038] Classification storage: Classify and store the captured data in the database on the current device according to multi-dimensional information including application software type, data format, and operation time, add metadata tags and build indexes;

[0039] The cross-device knowledge system synchronization includes:

[0040] Account system and data identification: Build a unified user account system, assign a unique ID to the user, associate the data generated by the user on each device with this ID, and clarify the data ownership;

[0041] Data synchronization and version control: Implement data interaction between each user device and the cloud server, unify the data of each device to the cloud server for data synchronization, and push the data in the latest version of the knowledge base after synchronization to each user device;

[0042] The real-time intelligent assistance includes:

[0043] Real-time analysis and matching: Based on the technologies of natural language processing and knowledge graph, combined with the application software environment information, analyze the user's interactive information and screen and match relevant data from the knowledge base;

[0044] Intelligent suggestion generation: Based on the filtered data, real-time operation suggestions are generated through a deep learning model, and after quality assessment of the real-time operation suggestions, they are presented to the user in the form of pop-up windows or quick inputs.

[0045] Compared with the prior art, the beneficial effects of the present invention are as follows: The multi-modal intelligent input method system solves the pain points existing in current input methods in aspects such as file searching and cross-platform operations by innovatively integrating the cloud disk knowledge base, realizing automatic capture and classification management of input data, supporting cross-device knowledge system synchronization, and providing real-time intelligent assistance functions. This system improves the document search efficiency during the interaction between the user and application software, optimizes the customer's data management, enables seamless cross-device experience, and provides intelligent real-time assistance, significantly improving the user's input efficiency and convenience, and can create an intelligent and personalized input environment for the user. Among them, the multi-modal intelligent input method system can uniformly store and manage the user's data on different cloud disks through the cloud disk knowledge base integration and management module. Moreover, by storing data in the cloud disk, the user does not need to frequently switch between different storage devices and software to search for files, can quickly obtain the required materials, and the data in the cloud disk is intelligently associated with the input content, greatly saving the search time and improving the input efficiency. The input data capture and classification management module automatically collects multi-modal data of the user's interaction with application software and stores them classified, facilitating the user to query and use at any time, effectively integrating fragmented personal information, and improving the convenience and orderliness of data management. The cross-device knowledge system synchronization module can unify the data on each cloud disk to different devices of the user, ensuring that the user can obtain a consistent knowledge system when logging in to the input method on any device. The real-time intelligent assistance module generates operation suggestions in real time according to the user's input scenario and intention, provides intelligent support for the user, reduces the time for the user to actively think and search, and significantly improves the accuracy and efficiency of input. Brief Description of the Drawings

[0046] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification, and are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0047] Figure 1 It is the overall architecture diagram of an embodiment of the multi-modal intelligent input method system of the present invention.

[0048] Figure 2 It is the block diagram of the cloud disk knowledge base integration and management module in an embodiment of the multi-modal intelligent input method system of the present invention.

[0049] Figure 3 It is the block diagram of the input data capture and classification management module in an embodiment of the multi-modal intelligent input method system of the present invention.

[0050] Figure 4Block diagram of the cross-device knowledge system synchronization module in an embodiment of the multi-modal intelligent input method system of the present invention.

[0051] Figure 5 Block diagram of the real-time intelligent assistance module in an embodiment of the multi-modal intelligent input method system of the present invention.

[0052] Figure 6 Flowchart of an embodiment of the multi-modal intelligent input implementation method of the present invention. Detailed implementation manner

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0054] In one embodiment, a multi-modal intelligent input method system is provided. As Figure 1 shown, the multi-modal intelligent input method system includes a cloud disk knowledge base integration and management module 100, an input data capture and classification management module 200, a cross-device knowledge system synchronization module 300, and a real-time intelligent assistance module 400. Among them, the cloud disk knowledge base integration and management module 100 is used to log in to the cloud disk knowledge base and retrieve and organize the data stored in the cloud disk knowledge base according to the user input text; the input data capture and classification management module 200 is used to capture the multi-modal interaction data between the user and the application software on the current device in real time and classify and store it; the cross-device knowledge system synchronization module 300 is used to build a unified user account system, associate the data generated by the user on each device with the user account, unify the data in each device of the user to the cloud server and merge it into the knowledge base, and supply each device to synchronize the data in the knowledge base to the local; the real-time intelligent assistance module 400 is used to parse the interaction information input by the user, screen and match relevant data from the knowledge base based on the parsing result, and generate real-time operation suggestions for the user based on the screened data.

[0055] The multi-modal intelligent input method system can uniformly store and manage the user's data on different cloud disks through the cloud disk knowledge base integration and management module 100. And by storing data on the cloud disk, users do not need to frequently switch between different storage devices and software to search for files, and can quickly obtain the required materials. The data in the cloud disk is intelligently associated with the input content, greatly saving the search time and improving the input efficiency. Through actual tests, in the input scenario involving file reference, the average operation time of users is shortened by 40%. This is because in the traditional method, it takes about 5 minutes on average for users to search for files and locate relevant content in the files, while after using this system, this time is shortened to about 3 minutes. The input data capture and classification management module 200 automatically collects the multi-modal data of the user's interaction with application software and classifies and stores it, which is convenient for users to query and use at any time, effectively integrating fragmented personal information, and improving the convenience and orderliness of data management. Here, the multi-modal data includes different types and formats of data, such as text data in different formats such as word and pdf, and picture data in different formats. In the research and development test, the user feedback shows that the efficiency of data search and use has increased by 50%. In the past, users needed to flip through 3-5 different files or records to search for specific interaction data on average. Now, through the classified storage and efficient indexing of this system, it only takes 1-2 operations on average to find the required data. The cross-device knowledge system synchronization module 300 can unify the data on each cloud disk to different devices of the user, ensuring that when the user logs in to the input method on any device, they can obtain a consistent knowledge system, realizing seamless switching across devices and enhancing the user's stickiness to the input method. In the research and development test, when users switch to use the input method between different devices, the data synchronization success rate reaches more than 98%. After a large number of device and network environment tests, in every 100 device-to-device data synchronization operations, there are less than 2 times that may fail to synchronize due to extreme network failures and other reasons, and the system has an automatic retry mechanism and can complete the synchronization in a short time. The real-time intelligent assistance module 400 generates operation suggestions in real time according to the user's input scenario and intention, provides intelligent support for users, reduces the time for users to actively think and search, and especially significantly improves the accuracy and efficiency of input in complex input scenarios. In the test, in scenarios such as business communication and copywriting creation, after users use the intelligent suggestion function, the quality and completion speed of the input content are increased by 30% on average. For example, in the scenario of writing a business email, it originally took about 20 minutes to write an email on average. After using the intelligent suggestion function, it can be shortened to about 14 minutes on average, and the professionalism and logic of the email content are also enhanced.

[0056] Based on the above embodiments, the multimodal intelligent input method system innovatively integrates the cloud disk knowledge base, realizes the automatic capture and classification management of input data, supports cross-device knowledge system synchronization, and provides real-time intelligent assistance functions, solving the pain points existing in current input methods in aspects such as file search and cross-platform operation. This system improves the document search efficiency during the interaction between users and application software, optimizes the customer's data management, can achieve a seamless cross-device experience, and provides intelligent real-time assistance, significantly improving the user's input efficiency and convenience, and can create an intelligent and personalized input environment for users.

[0057] Further, as Figure 2 shown, the cloud disk knowledge base integration and management module 100 in the multimodal intelligent input method system includes a multi-cloud disk association sub-module 101 and an intelligent retrieval and collation sub-module 102. The multi-cloud disk association sub-module 101 is used for the user to log in and authorize multiple cloud disks, realizing the unified association of various cloud disks with the input method system; the intelligent retrieval and collation sub-module 102 is used for real-time analysis of the user's input content, retrieving relevant information in the associated cloud disk files, and refining and presenting it to the user.

[0058] In the input method software architecture of the multimodal intelligent input method system, a dedicated multi-cloud disk association sub-module 101 is designed in the cloud disk knowledge base integration and management module 100. The multi-cloud disk association sub-module 101 supports the user to input the cloud disk account and password for login authorization through the standard API interface protocol, or adopts third-party authorization methods such as OAuth to realize the unified association of various cloud disks (such as Huawei Cloud Disk, Baidu Cloud Disk, Tencent Weiyun, etc.) used by the user to the input method system, and has the ability to dock with multiple cloud disk services. To ensure data security, an encrypted transmission protocol, such as the SSL / TLS protocol, is used during the authorization login process to encrypt the user's account information and authorization requests.

[0059] The intelligent retrieval and sorting sub-module 102 in the cloud disk knowledge base integration and management module 100 is constructed by means of natural language processing (NLP) and machine learning algorithms. After the user establishes a data acquisition channel with different cloud disks through the multi-cloud disk association sub-module 101, when the user inputs text in the application software, the intelligent retrieval and sorting sub-module 102 analyzes the content input by the user in real time, extracts key information such as keywords and themes, and then performs in-depth retrieval in the associated cloud disk files based on these key information. For text files, text matching algorithms such as the TF-IDF algorithm combined with the BM25 algorithm are used to calculate the similarity between the input content and the file content to quickly locate relevant files. For picture files, image recognition algorithms such as the resnet algorithm are used to extract the feature vectors of the pictures and compare them with the pre-stored picture feature library to find pictures related to the input content. After retrieving relevant files, the intelligent retrieval and sorting sub-module 102 uses information extraction algorithms such as text summarization algorithms based on deep learning (such as BERT-Sum) to extract key paragraphs and sentences from text files; for pictures, it generates text descriptions of the pictures. The extracted information is presented to the user in a concise and clear manner to realize the intelligent association between the input content and the cloud disk data for the user to call when inputting, reducing the time for the user to search for files and improving the input efficiency of the user in the application software interaction process.

[0060] Furthermore, as Figure 3 shown, the input data capture and classification management module 200 in the multi-modal intelligent input method system includes a data capture sub-module 201 and a classification storage sub-module 202. The data capture sub-module 201 is used to capture user-application software interaction data in real time based on data capture methods and perform local encryption processing. The data capture methods include hook functions, process monitoring, and clipboard listening; the classification storage sub-module 202 is used to classify and store the captured data in the database on the current device according to multi-dimensional information including application software type, data format, and operation time, add metadata tags, and build indexes.

[0061] On the premise of obtaining the explicit authorization of the application software used by the user, the data scraping sub-module 201 of the input data scraping and classification management module 200 uses technical means such as hook functions and process monitoring to monitor the interaction operations between the user and various application software in real time. For example, for the input content of the user in the application software, the input box monitoring technology is adopted to capture the text information input by the user in the input boxes of various application software; for the modification records, by monitoring the document editing events of the application software, the operations of the user adding, deleting, and modifying the document content are recorded; for the copied information, the system clipboard monitoring technology is used to obtain the data copied by the user to the clipboard. All the scraped data is initially encrypted locally, and a symmetric encryption algorithm (such as AES) can be used to ensure the security of the data during transmission and storage.

[0062] The classification and storage sub-module 202 of the input data scraping and classification management module 200 is used to classify and store the scraped data according to multi-dimensional information such as the type of application software, the format of the data, and the operation time. For example, according to the type of application software, the data interacting with Word is stored under the category of "Office Software - Word", and the data interacting with Photoshop is stored under the category of "Graphic Processing Software - Photoshop". For each category, it is further subdivided according to the data format, such as text files, picture files, etc. At the same time, detailed metadata tags are added to each piece of data, including the operation time, operation type (input, modification, copy, etc.), relevant application software version information, etc. A database management system (such as MySQL or PostgreSQL) is used to store the classified data, and an efficient index structure, such as a B+ tree index, is constructed so that users can quickly query and retrieve relevant data.

[0063] Furthermore, as Figure 4 shown, the cross-device knowledge system synchronization module 300 in this multi-modal intelligent input method system includes an account system and data identification sub-module 301 and a data synchronization and version control sub-module 302. The account system and data identification sub-module 301 is used to construct a unified user account system, assign a unique ID to the user, associate the data generated by the user on each device with this ID, and clarify the data ownership; the data synchronization and version control sub-module 302 is used to realize the data interaction between the user's each device and the cloud server, unify the data of each device to the cloud server for data synchronization, and push the data in the latest version of the knowledge base after synchronization to each device of the user.

[0064] Build a unified user account system for the input method system through the account system and data identification sub-module 301 in the cross-device knowledge system synchronization module 300, enabling users to register using unique identifiers such as mobile phone numbers and email addresses, and authenticate their identities by setting passwords. When a user registers, the system assigns a unique user ID to each user, which serves as the user's identity identifier in the system. For various types of data generated by the user on different devices, whether it is cloud disk associated data or data obtained by the data scraping sub-module 201, it is associated with the user ID. When storing data, the user ID is used as the key identification field and stored in the database table to ensure clear data ownership.

[0065] The data synchronization and version control sub-module 302 uses a distributed version control method (such as Git) to achieve cross-device data synchronization. When a user operates on data on a certain device (such as adding file associations, modifying input data, etc.), the system automatically records operation logs, including operation time, operation type, data content involved, etc. The operation logs and the updated data are transmitted to the cloud server through an encrypted network channel (such as an HTTP / HTTPS channel based on the SSL / TLS protocol). After receiving the data, the cloud server first performs identity verification and data integrity verification. After successful verification, it finds the corresponding user data storage location based on the user ID and merges the new data with the original data. During the merging process, a version control algorithm is used, such as a conflict resolution algorithm based on timestamps. When data conflicts occur, the version with the latest operation time is used as the standard. At the same time, the cloud server pushes the updated synchronization information to other devices of the user. After receiving the synchronization information, other devices automatically download the updated data to ensure that users can obtain the latest and consistent knowledge system when logging in to the input method on any device.

[0066] Furthermore, as Figure 5 shown, the real-time intelligent assistance module 400 in this multi-modal intelligent input method system includes a real-time analysis and matching sub-module 401 and an intelligent suggestion generation sub-module 402. The real-time analysis and matching sub-module 401 is used to parse user interaction information based on natural language processing and knowledge graph technologies, combined with application software environment information, and screen and match relevant data from the knowledge base; the intelligent suggestion generation sub-module 402 is used to generate real-time operation suggestions through a deep learning model based on the screened data, and present the real-time operation suggestions to the user in the form of a pop-up window or quick input after quality evaluation.

[0067] When the user interacts with the application software for input, the real-time analysis and matching sub-module 401 in the real-time intelligent assistance module 400 is activated. The real-time analysis and matching sub-module 401 uses techniques such as syntactic analysis and semantic understanding in natural language processing to deeply analyze the interaction information between the user and the software. For example, for the input text content, analyze its grammatical structure and extract key information such as the theme and intention. At the same time, combined with the application software environment information where the user is currently located, such as the software type, the current operation interface, etc., relevant data is screened from the knowledge base integrated by the input method. For example, if the user enters content in the WeChat chat interface, the system determines that the current is a social scenario and preferentially retrieves data related to social communication from the knowledge base, such as common greetings, reply templates, etc. Using knowledge graph technology, establish the association relationship between the application software scenario and the knowledge base data to improve the accuracy and efficiency of matching. Based on the results of the real-time analysis and matching sub-module 401, the intelligent suggestion generation sub-module 402 uses a deep learning model (such as a generation model based on the Transformer architecture) to generate real-time response solutions or operation suggestions. For different application scenarios and user input intentions, the model generates corresponding text suggestions. For example, in the scenario of writing a business email, according to the theme and previous content entered by the user, generate appropriate email beginnings, endings, and key paragraphs of the body. For the generated suggestions, quality evaluation is carried out through language generation quality evaluation indicators (such as BLEU scores, ROUGE scores, etc.) to ensure the accuracy and reasonableness of the suggestions. The generated suggestions are presented to the user in the form of pop-up windows or quick inputs. The user can directly select to use the suggested content or make minor modifications according to actual needs to quickly complete the input task.

[0068] It should be noted that in this multi-modal intelligent input method system, the intelligent retrieval and collation sub-module 102 and the intelligent suggestion generation sub-module 402 are two sub-modules with different functions. The intelligent retrieval and collation sub-module 102 focuses on retrieval feedback, and the intelligent suggestion generation sub-module 402 focuses on content generation, and are fed back to the user through two different windows.

[0069] The usage method of this multi-modal intelligent input method system is as follows:

[0070] (I) Usage of the cloud disk knowledge base integration and management module

[0071] 1. Association of multiple cloud disks

[0072] (1) The user opens the input method settings interface and selects the cloud disk service provider to be associated, such as Huawei Cloud Disk, in the options of the cloud disk association sub-module 101.

[0073] (2) The input method system pops up an authorization interface, prompting the user to enter the account and password of Huawei Cloud Disk, or select to use the third-party authorization method (such as authorizing through the Huawei account).

[0074] (3) After the user enters the account password or completes the third-party authorization, the input method system sends an authorization request to the Huawei Cloud Disk server through an SSL / TLS encrypted channel. The request contains the account password or authorization token entered by the user.

[0075] (4) The Huawei Cloud Disk server verifies the legitimacy of the user identity and the authorization request. If the verification passes, it returns an authorization success message and the relevant API access token.

[0076] (5) The input method system saves the authorization information and the API access token, establishes a connection with the Huawei Cloud Disk, and completes the cloud disk association operation. Repeating the above steps, other cloud disks can be associated. By associating multiple cloud disks of the user with this input method system, personalized data related to the user can be obtained.

[0077] 2. Intelligent Retrieval and Sorting

[0078] (1) The user enters text in the application software, and the intelligent retrieval and sorting sub-module 102 of the input method monitors the input content in real time. For example, the user enters "Marketing Plan Materials".

[0079] (2) The intelligent retrieval and sorting sub-module 102 uses natural language processing technology to perform word segmentation on the input content, and extracts keywords such as "Marketing", "Plan", and "Materials".

[0080] (3) According to the extracted keywords, the intelligent retrieval and sorting sub-module 102 sends a retrieval request to the associated cloud disk. Taking the Huawei Cloud Disk as an example, by calling the API provided by the Huawei Cloud Disk and passing the keyword parameters, a request is made to retrieve relevant files.

[0081] (4) The cloud disk server searches in the cloud disk files according to the retrieval request, and returns a list of matching files and file metadata (such as file name, file type, file size, etc.).

[0082] (5) The intelligent retrieval and sorting sub-module 102 further filters and sorts the returned file list. For text files, the TF-IDF algorithm combined with the BM25 algorithm is used to calculate the similarity score between the file and the input content, and the files are sorted according to the score. For picture files, image recognition technology is used to extract the picture feature vectors, which are compared with the pre-trained picture feature library to screen out pictures related to the input content, and a text description of the pictures is generated.

[0083] (6) The intelligent retrieval and sorting sub-module 102 presents the sorted file information (including the key paragraph summary of text files, the text description of pictures, etc.) to the user in the form of a pop-up window or a sidebar. The user can directly click on the file link to open the file to view the detailed content.

[0084] (2) Usage of the Input Data Scraping and Classification Management Module 200

[0085] 1. Data Scraping

[0086] (1) After obtaining the authorization of the application software, the input method system starts the data scraping sub-module 201 in the background to perform data scraping services. For example, when the user opens the Word software for document editing.

[0087] (2) For the input content, the data scraping sub-module 201 listens to the input box events of the Word document through the hook function, captures the text information input by the user in the input box, and records the input time.

[0088] (3) For the modification records, the data scraping sub-module 201 listens to the document editing events of Word, such as paragraph insertion, deletion, text modification, etc., and records the operation type, operation location, and the content before and after the modification.

[0089] (4) For the copied information, the data scraping sub-module 201 uses the system clipboard monitoring technology. When the user performs a copy operation, it obtains the data content in the clipboard and records the copy time.

[0090] (5) Encrypt all the scraped data locally. Use the AES encryption algorithm, set a 256-bit encryption key, and encrypt the data. The encrypted data is temporarily stored in the local cache and waits for further processing.

[0091] 2. Classification and Storage

[0092] (1) The data scraping service transfers the encrypted data in the local cache to the classification and storage sub-module 202 at regular intervals (such as every 1 minute).

[0093] (2) The classification and storage sub-module 202 first decrypts the encrypted data, using the pre-set AES decryption key for the decryption operation.

[0094] (3) Classify and store the decrypted data according to the type of the source application software. For example, store the data from Word in the "Office Software - Word" table in the database.

[0095] (4) When storing the data, the classification and storage sub-module 202 adds detailed metadata tags to each piece of data, such as operation time, operation type, relevant application software version information, etc. For text data, store the complete text content; for picture data, store the path or binary data of the picture, and generate a thumbnail of the picture to store in the database for quick preview.

[0096] (5) The classification storage sub-module 202 uses the MySQL database management system to build a B+ tree index for each data table to improve the efficiency of data query and retrieval. For example, indexes are built for the "operation time" field and the "operation type" field of the "office software - Word" table respectively, facilitating users to quickly query relevant data according to the time range and operation type.

[0097] (III) Use of the cross-device knowledge system synchronization module 300

[0098] 1. Account system and data identification

[0099] (1) When the user first uses the input method system, the user opens the registration interface provided by the account system and data identification sub-module 301, enters the mobile phone number or email address as the registered account, and sets a high-strength password containing letters, numbers, and special characters.

[0100] (2) The account system and data identification sub-module 301 of the input method system sends a registration request to the server that provides the registration service for it. The request contains the registration information input by the user. The server verifies the uniqueness of the registered account. If the account has not been registered, a unique user ID is generated for the user, and the user registration information (including the encrypted hash values of the account and password, user ID, etc.) is stored in the user information table.

[0101] (3) When the user performs data operations on the device, whether it is associating cloud disk files or generating input data, the account system and data identification sub-module 301 of the input method system adds the user ID as an identifier to the data. For example, when the user associates a Baidu cloud disk file on the mobile phone, the account system and data identification sub-module 301 stores information such as the user ID, cloud disk association information (such as the associated cloud disk name, file path, etc.), and operation time in the cloud disk association data table.

[0102] 2. The data synchronization and version control sub-module 302 is arranged in the cloud server where the cloud disk is located and the user's device to perform data synchronization and version control, specifically as follows:

[0103] (1) The user performs operations on data on a certain device (such as a computer), such as entering a paragraph of text in a Word document. The input data capture and classification management module 200 captures this operation, records the operation log, including the operation time, operation type (input), the content of the input text, etc., and stores it for classification and later use by the data synchronization and version control sub-module 302.

[0104] (2) The data synchronization and version control sub-module 302 transmits the operation log and the updated data (such as the document content containing the newly entered text) together to the cloud server through an encrypted network channel based on the SSL / TLS protocol.

[0105] (3) After the cloud server receives the data, the data synchronization and version control sub-module 302 first verifies the user identity. It queries the user's authentication information in the user information table through the user ID to verify the legal source of the data. Then it performs data integrity verification by comparing the calculated hash value of the data with the hash value attached during the transmission process to ensure that the data has not been tampered with during the transmission.

[0106] (4) After the verification passes, the cloud server, based on the data synchronization and version control sub-module 302, finds the corresponding user data storage location according to the user ID. For example, it finds the records related to the user in the "Office Software - Word" data table. It merges the new data with the original data and adopts a conflict resolution algorithm based on timestamps. For example, if the user modifies the same document on another device at the same time, the cloud server compares the timestamps of the two operations and updates the data based on the operation version with the newer timestamp.

[0107] (5) The cloud server, through the data synchronization and version control sub-module 302, pushes the updated synchronization information (such as the updated data list, operation logs, etc.) to other devices of the user (such as mobile phones, tablets). After receiving the synchronization information, the other devices automatically download the updated data and update the data records in the local database to ensure that the data on each device is consistent with the cloud server.

[0108] (4) Usage of the real-time intelligent assistance module 400

[0109] 1. Real-time analysis and matching

[0110] (1) The user enters content in an application software (such as WeChat), and the real-time analysis and matching sub-module 401 obtains the input text content "There is an important meeting tomorrow. What should I prepare?".

[0111] (2) The real-time analysis and matching sub-module 401 uses a syntactic analysis tool (such as StanfordCoreNLP) in natural language processing technology to analyze the input text, determine the grammatical structure of the sentence, and extract key information such as "tomorrow", "important meeting", "prepare", etc.

[0112] (3) The real-time analysis and matching sub-module 401 combines the application software environment information where the user is currently located (WeChat chat interface, judged as a social communication scenario), and screens relevant data from the knowledge base integrated with the input method. Using knowledge graph technology, it searches for knowledge nodes related to "preparing for an important meeting", such as "preparing meeting materials", "key points of meeting speeches", etc.

[0113] (4) The real-time analysis and matching sub-module 401 retrieves relevant data records from the knowledge base according to the association relationships in the knowledge graph, such as chat records, preparation lists, etc. in similar scenarios in the past. The cosine similarity algorithm is used to calculate the similarity between the retrieved data and the input content, and the data is sorted according to the similarity level.

[0114] 2. Intelligent suggestion generation

[0115] (1) The intelligent suggestion generation sub-module 402 generates real-time response solutions or operation suggestions based on the relevant data screened by the real-time analysis and matching sub-module, using a generation model based on the Transformer architecture. The model is trained according to the input content and relevant data to learn the language patterns and response strategies in different scenarios. For example, the model generates a suggestion: "You can first prepare the materials related to the meeting, such as project reports, data charts, etc., and at the same time sort out your speaking points to clarify the core content to be conveyed."

[0116] (2) For the generated suggestions, the intelligent suggestion generation sub-module 402 conducts quality evaluation through language generation quality evaluation metrics (such as BLEU score, ROUGE score, etc.). If the evaluation score does not reach the set threshold, the model adjusts and optimizes the suggestions until the quality requirements are met.

[0117] (3) The intelligent suggestion generation sub-module 402 displays the generated suggestions in the form of pop-ups near the application software interface. The pop-up window is designed to be simple and clear, highlighting the key content. Users can directly click on the suggestion content and insert it into the input box, or make minor modifications according to the actual situation and then use it.

[0118] In one embodiment, a multi-modal intelligent input implementation method is provided, in combination with Figure 6 As shown, the method includes the following steps:

[0119] Step S100: Cloud disk knowledge base integration and management: Log in to the cloud disk knowledge base, and retrieve and organize the data stored in the cloud disk knowledge base according to the user input text.

[0120] Step S200: Input data capture and classification management: Real-time capture the multi-modal interaction data between the user and the application software on the current device and classify and store it.

[0121] Step S300: Cross-device knowledge system synchronization: Build a unified user account system, associate the data generated by the user on each device with the user account, unify the data on each device of the user to the cloud server and merge it into the knowledge base, and each device synchronizes the data in the knowledge base to the local.

[0122] Step S400: Real-time intelligent assistance: Parse the interaction information input by the user, screen and match relevant data from the knowledge base based on the parsing result, and generate real-time operation suggestions for the user based on the screened data.

[0123] Among them, in step S100, the integration and management of the cloud disk knowledge base includes: (1) Multi-cloud disk association sub-step: The user logs in and authorizes multiple cloud disks to achieve the unified association of various cloud disks with the input method system; (2) Intelligent retrieval and sorting sub-step: Analyze the user input content in real time, retrieve relevant information in the associated cloud disk files, and refine and present it to the user. In step S200, the input data capture and classification management includes: (1) Data capture sub-step: Based on the data capture method, real-time capture the user's interaction data with the application software and perform local encryption processing. The data capture methods include hook functions, process monitoring, and clipboard listening; (2) Classification storage sub-step: Classify and store the captured data in the database on the current device according to multi-dimensional information including application software type, data format, and operation time, add metadata tags and build indexes. In step S300, the cross-device knowledge system synchronization includes: (1) Account system and data identification sub-step: Build a unified user account system, assign a unique ID to the user, associate the data generated by the user on each device with this ID, and clarify the data ownership; (2) Data synchronization and version control sub-step: Realize the data interaction between the user's devices and the cloud server, unify the data of each device to the cloud server for data synchronization, and push the data in the latest version of the knowledge base after synchronization to each device of the user. In step S400, the real-time intelligent assistance includes: (1) Real-time analysis and matching sub-step: Based on the technologies of natural language processing and knowledge graph, combined with the application software environment information, parse the user interaction information, and screen and match relevant data from the knowledge base; (2) Intelligent suggestion generation sub-step: Based on the screened data, generate real-time operation suggestions through a deep learning model, and present the real-time operation suggestions to the user in the form of a pop-up window or quick input after quality evaluation.

[0124] In the above embodiments of the implementation method of the multi-modal intelligent input method system, for the multi-cloud disk association sub-step in the cloud disk knowledge base integration and management step, when the user operates, select the cloud disk service provider in the input method settings interface, enter the account password or use the third-party authorization method, send an authorization request through the encrypted channel, and establish a connection to complete the association after the cloud disk server verifies and passes. For the intelligent retrieval and sorting sub-step, by monitoring the user input content, extract keywords, send a retrieval request to the associated cloud disk, screen and sort the returned files, and present relevant information in the form of a pop-up window or sidebar.

[0125] For the data scraping sub-step in the input data scraping and classification management step, after obtaining the authorization of the application software, the background starts the data scraping service, listens for data obtained from the input box, edit events, and clipboard, and locally encrypts and temporarily stores it. In the classification storage sub-step, the encrypted data is regularly transmitted to the classification storage sub-module, decrypted, classified and stored according to the application software type, metadata tags are added, and database and indexing technologies are used to improve the query efficiency.

[0126] In the account system and data identification sub-step of the cross-device knowledge system synchronization step, the user registers to obtain a unique ID, and this ID is added to all data operations on the device. In the data synchronization and version control sub-step, the device encrypts and transmits the operation logs and updated data to the cloud server. After authentication and integrity verification, the data is updated according to the version control algorithm and synchronization information is pushed to other devices.

[0127] In the real-time analysis and matching sub-step of the real-time intelligent assistance step, by obtaining the user input content, combined with the application software environment, natural language processing and knowledge graph technologies are used to screen relevant data from the knowledge base. In the intelligent suggestion generation sub-step, based on the screened data, a deep learning model is used to generate suggestions, which are presented to the user in the form of pop-ups or quick inputs after quality evaluation.

[0128] Through the innovative design of the multimodal intelligent input method system, significant improvements have been achieved in input efficiency, data management, cross-device collaboration, and intelligent experience. The specific beneficial effects are as follows: 1. Improved input efficiency: Through the integration of the cloud disk knowledge base and intelligent retrieval function, the average time taken by users to search for files in scenarios such as copywriting creation and project planning has been reduced from 5 minutes in the traditional method to 3 minutes, with a 40% efficiency improvement. In steps of real-time intelligent assistance, in scenarios such as business communication and copywriting writing, operation suggestions are generated through deep learning, resulting in an average 30% increase in the completion speed of input content and a significant improvement in input accuracy. 2. Optimized data management: In the steps of input data capture and classification management, it can automatically capture users' multimodal interaction data (such as text, pictures, operation logs), and classify and store them based on application software type, data format, and time dimension, constructing a B+ tree index, which improves the data search efficiency by 50%. Tests show that the number of operations for users to search for specific interaction data has been reduced from 3 - 5 times to 1 - 2 times. 3. Seamless cross-device collaboration: In the steps of cross-device knowledge system synchronization, through unified user IDs and cloud version control, the success rate of multi-device data synchronization reaches over 98%. When users switch between devices such as mobile phones and computers, they can obtain the latest knowledge base data in real-time, avoiding information breaks caused by device differences and enhancing work continuity. 4. Intelligent interaction experience: In the steps of real-time intelligent assistance, based on NLP and knowledge graph technologies, combined with application scenarios (such as WeChat social interaction, Word office work), personalized suggestions are generated. In the business email writing test, the adoption rate of suggestions reaches 70%, the content professionalism score is increased by 25%, effectively reducing the user's cognitive burden. 5. Enhanced data security: AES encryption is used to store users' interaction data, SSL / TLS protocols are used to ensure the security of cloud transmission, and at the same time, a timestamp conflict resolution algorithm is used to ensure data consistency. The system security level reaches the financial level standard, and the risk of user data leakage is reduced by more than 90%.

[0129] In summary, this method systematically solves the pain points of traditional input methods in file retrieval, cross-platform operation, data dispersion, etc. The comprehensive input efficiency is improved by 40% - 50%, creating an efficient, intelligent, and secure input ecosystem for users. In this multimodal intelligent input implementation method, by innovatively integrating the cloud disk knowledge base, realizing the automatic capture and classification management of input data, supporting cross-device knowledge system synchronization, and providing real-time intelligent assistance functions, it solves the pain points existing in current input methods in aspects such as file search and cross-platform operation, significantly improving the input efficiency and convenience for users, and can create an intelligent and personalized input environment for users.

[0130] It should be noted that in this text, terms such as "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such a process, method, article or device.

[0131] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multimodal intelligent input method system, characterized in that, Including: A cloud disk knowledge base integration and management module, which is used to log in to the cloud disk knowledge base and retrieve and organize the data stored in the cloud disk knowledge base according to the user input text; An input data capture and classification management module: which is used to capture and classify and store the multi-modal interaction data between the user and the application software on the current device in real time; A cross-device knowledge system synchronization module: which is used to build a unified user account system, associate the data generated by the user on each device with the user account, unify the data in each device of the user to the cloud server and merge it into the knowledge base, and enable each device to synchronize the data in the knowledge base to the local; A real-time intelligent assistance module: which is used to parse the interaction information input by the user, screen and match relevant data from the knowledge base based on the parsing result, and generate real-time operation suggestions for the user based on the screened data.

2. The multimodal intelligent input method system according to claim 1, wherein The cloud disk knowledge base integration and management module includes: A multi-cloud disk association sub-module: which is used for the user to log in and authorize multiple cloud disks to realize the unified association of various cloud disks with the input method system; An intelligent retrieval and organization sub-module: which is used to analyze the user input content in real time, retrieve relevant information in the associated cloud disk files, and refine and present it to the user.

3. The multimodal intelligent input method system according to claim 1, wherein The input data capture and classification management module includes: A data capture sub-module: which is used to capture the interaction data between the user and the application software in real time based on the data capture method and perform local encryption processing. The data capture method includes the methods of hook function, process monitoring and clipboard listening; A classification storage sub-module: which is used to classify and store the captured data in the database on the current device according to multi-dimensional information including application software type, data format and operation time, add metadata tags and build indexes.

4. The multimodal intelligent input method system according to claim 1, wherein, The cross-device knowledge system synchronization module includes: An account system and data identification sub-module: which is used to build a unified user account system, assign a unique ID to the user, associate the data generated by the user on each device with this ID, and clarify the data ownership; A data synchronization and version control sub-module: which is used to realize the data interaction between each device of the user and the cloud server, unify the data of each device to the cloud server for data synchronization, and push the data in the latest version of the knowledge base after synchronization to each device of the user.

5. The multimodal intelligent input method system according to claim 1, wherein The real-time intelligent assistance module includes: A real-time analysis and matching sub-module: which is used to parse the user interaction information based on the technologies of natural language processing and knowledge graph, combined with the application software environment information, and screen and match relevant data from the knowledge base; An intelligent suggestion generation sub-module: which is used to generate real-time operation suggestions through a deep learning model based on the screened data, and present the real-time operation suggestions to the user in the form of a pop-up window or quick input after quality evaluation.

6. The multimodal intelligent input method system according to claim 2, wherein: The method for the intelligent retrieval and organization sub-module to retrieve relevant information in the associated cloud disk files is as follows: for text files, the TF-IDF algorithm is combined with the BM25 algorithm to calculate the similarity, and for picture files, the image recognition algorithm is used to extract feature vectors for comparison and retrieval.

7. The multimodal intelligent input method system according to claim 3, wherein: The data scraping sub-module locally encrypts the captured data using the AES encryption algorithm. The classification storage sub-module stores data in a MySQL or PostgreSQL database and constructs a B+ tree index.

8. The multimodal intelligent input method system according to claim 4, wherein: The data synchronization and version control sub-module transmits data through a network channel based on the SSL / TLS protocol and uses a conflict resolution algorithm based on timestamps to handle data conflicts.

9. A method for implementing multimodal intelligent input, characterized in that, The method includes: Cloud disk knowledge base integration and management: Log in to the cloud disk knowledge base, and retrieve and organize the data stored in the cloud disk knowledge base according to the user input text. Input data scraping and classification management: Real-time capture the interaction data between the user and the application software on the current device and classify and store it. Cross-device knowledge system synchronization: Build a unified user account system, associate the data generated by the user on each device with the user account, unify the data on each device of the user to the cloud server and merge it into the knowledge base, and each device synchronizes the data in the knowledge base to the local. Real-time intelligent assistance: Parse the interaction information input by the user, screen and match relevant data from the knowledge base based on the parsing result, and generate real-time operation suggestions for the user based on the screened data.

10. The multi-modal intelligent input implementation method according to claim 9, wherein: The cloud disk knowledge base integration and management includes: Multi-cloud disk association: The user authorizes the login of multiple cloud disks to achieve the unified association of various cloud disks and the input method system. Intelligent retrieval and organization: Analyze the user input content in real time, retrieve relevant information in the associated cloud disk files, and refine and present it to the user. The input data scraping and classification management includes: Data scraping: Real-time capture the interaction data between the user and the application software based on the data scraping method and perform local encryption processing. The data scraping method includes the methods of hook function, process monitoring, and clipboard listening. Classification storage: Classify and store the scraped data in the database on the current device according to multi-dimensional information including application software type, data format, and operation time, add metadata tags and construct an index. The cross-device knowledge system synchronization includes: Account system and data identification: Build a unified user account system, assign a unique ID to the user, associate the data generated by the user on each device with this ID, and clarify the data ownership. Data synchronization and version control: Realize the data interaction between each device of the user and the cloud server, unify the data of each device to the cloud server for data synchronization, and push the data in the latest version of the knowledge base after synchronization to each device of the user. The real-time intelligent assistance includes: Real-time analysis and matching: Based on the technologies of natural language processing and knowledge graph, combined with the application software environment information, parse the user interaction information and screen and match relevant data from the knowledge base. Intelligent suggestion generation: Based on the screened data, generate real-time operation suggestions through a deep learning model, and present the real-time operation suggestions to the user in the form of a pop-up window or quick input after quality evaluation.