Method and system for mixed interaction help based on real scene across software operations

By constructing a real-scene semantic understanding system based on deep learning, the problem of users having difficulty finding functions in the operation of electronic device software has been solved, and efficient interactive assistance across software and devices has been achieved, improving user experience and the real-scene understanding capabilities of large models.

CN118964566BActive Publication Date: 2026-06-26XIDIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2024-07-30
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

In the operation of existing electronic device software, users have difficulty finding the functions they need, which is inconvenient and inefficient when operating across software. Existing large models are also insufficient in understanding continuous user operations and real-world perception.

Method used

By combining deep learning-based natural language processing and image recognition with real-world semantic understanding, a cross-software hybrid interactive assistance system for software operation is constructed. This system accurately perceives the user's operating environment and continuous actions, providing precise operation guidance.

Benefits of technology

It improves the efficiency of software operation and user experience, enhances cross-software and cross-device interactive assistance capabilities, reduces training and usage costs, and improves the ability of large models to handle real-world problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118964566B_ABST
    Figure CN118964566B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on real scene can cross software's software operation mixed interactive help method and system, including extracting electronic equipment software information, button information, generate corresponding training data;And pre-constructing using deep learning based on real scene semantic understanding can cross software's software operation mixed interactive help model, according to question information, input software operation mixed interactive help model, output each semantic analysis result's ranking score, select one or more semantic analysis results as user problem understanding result, provide interactive operation help for user, including step-by-step prompt or automatic completion button click operation.The application greatly widens the problem solving range of software operation interactive help system based on real scene semantic understanding, improves the convenience and problem solving efficiency, solves the problem that user uses general help instruction is not intuitive, enhances the good experience of user using product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a hybrid interactive assistance method and system for software operation based on real-scene semantic understanding, which is cross-software and applicable to various types of computers, mobile terminals such as mobile phones, and electronic devices such as various household appliances and industrial appliances. Background Technology

[0002] In current electronic devices, the ease of use of the software determines the overall usability. Users typically encounter three main problems when using electronic device software: first, they don't know if a particular software has the function they need, or if any software has already developed the function they require; second, they can't find the required function buttons within the software; and third, even if they know a software has the function they need and where it is, they still "don't know how to use it" or can't use it effectively. Although software comes with user manuals, understanding the manual and finding functions is still far from the experience of having a professional guide you. These three problems significantly impact software usability and hinder its widespread adoption.

[0003] Meanwhile, some electrical appliances, such as household appliances, have many operation buttons and the operation interface is not very user-friendly. The instruction manual is inconvenient to use. In addition, each appliance is equipped with a user APP, which increases the difficulty of user management and searching. It is imperative to change these problems by using artificial intelligence technology to realize a unified electronic instruction manual that can ask questions in natural language and intuitively guide users to operate in real-world scenarios.

[0004] Because all software has its limitations, users often encounter situations where the functions of the software they are using cannot meet their requirements well, or even not at all. For example, drawing a flowchart in Word is not as convenient as drawing it in Visio and inserting it into Word. Sometimes, when operating a mobile phone, it is necessary to complete the work in computer software first, and then complete the work through the mobile phone. At this time, it is necessary to solve the problem across software. Since the operating system is also software, cross-software operation can not only complete cross-software interaction on the local machine, but also cross-platform and electronic device interaction. Of course, the premise is that different platforms and devices can communicate with each other.

[0005] With the development of deep learning-based natural language processing, image recognition, and multimodal technologies, it has become possible to understand users' natural language questions and then guide them in operating software. There are two modes in this guidance process: one is to use automation technology to minimize user intervention, and the other is to guide users through specific operations, such as clicking buttons. Each mode caters to different needs and has its own advantages and disadvantages. Automation is sometimes more efficient, while guiding users step-by-step through specific operations helps them learn the techniques. When encountering similar problems later, they can directly perform the corresponding actions, saving time spent asking questions and thus improving efficiency.

[0006] In reality, natural language conversations in real-world scenarios often involve numerous omissions and ambiguities. For example, "Where is the translator?" might mean finding the person who translates in real life, while in Word it might mean finding the translate button. Without contextual information, accurate understanding can only be achieved by increasing the number of dialogue turns. However, when combined with a real-world context, such questions are sometimes unambiguous, or even if they are, the number of dialogue turns or the complexity of the conversation can be reduced. In today's rapidly developing artificial intelligence landscape, the ability to perceive the real-world context of software operation and correctly handle the multimodal problem of language and continuous button clicks is crucial.

[0007] Currently, large model technology has evolved from simple large language (text) models (LLM) to image-language multimodal models (VLM). However, in the field of language-video multimodal models, it is necessary to combine language to perceive and understand continuous actions. Current large models are not strong in this regard. Even OpenAI's Sora, which generates one-minute videos, has many flaws. Software operation and real-world interactive assistance need to perceive continuous user operations, sometimes with operation time far exceeding one minute. In addition, the guidance information returned by large models is text and images, which is still far from accurate and intuitive.

[0008] In fact, continuous motion perception and understanding is the most important aspect of understanding the real world and is currently a major factor restricting the further improvement of large-scale model intelligence. Selecting real-world scenarios that can be accurately perceived and understood and organically combining them with existing large-scale models is the only way to improve the capabilities of large-scale models under the current conditions. Real-world software operation is such a scenario, and cross-software operation interaction can further significantly improve the ability of large-scale models to handle real-world problems.

[0009] Meanwhile, the training and use costs in the language-video multimodal domain are high, and the success rate of accurately understanding videos is not ideal. By adopting deep learning-based natural language processing technology combined with computer technology to capture users' continuous software operation actions, it is imperative to significantly improve the success rate and reduce training and use costs, thereby achieving accurate perception of users' software operation process, accurate understanding of users' natural language problems in real-world software operation scenarios, and ultimately achieving precise guidance for users' software operation technology. Summary of the Invention

[0010] To address the aforementioned deficiencies in existing technologies, the present invention aims to provide a software operation hybrid interactive assistance method and system based on real-scene and cross-software capabilities. This invention significantly expands the scope of problems that can be solved by software operation interactive assistance systems based on real-scene semantic understanding, improves ease of use and problem-solving efficiency, solves the problem of unintuitive user manuals, improves the efficiency of solving problems encountered in the use of electronic device software, and enhances the user experience quality of the product.

[0011] The present invention is achieved through the following technical solution.

[0012] One aspect of the present invention provides a real-world, cross-software-based, hybrid interactive assistance method for software operation, comprising the following steps:

[0013] Step 111: Extract software information from electronic devices, button information of interactive software operation buttons, collect help information on usage problems and other problems, and generate training data for a cross-software hybrid interactive help system based on real-scene semantic understanding.

[0014] Step 112: Extract the hardware and software information from the user's electronic device and construct a data model of the user's electronic device;

[0015] Step 113: Based on the training data, pre-construct a software operation hybrid interaction help model that uses deep learning and is based on real-scene semantic understanding and can be used across software.

[0016] Step 114: Extract the natural language question information from the user's question;

[0017] Step 115: Based on the constructed user electronic device data model, extract the current real-world information of the software, and generate real-world natural language question information based on the user's questions and operations;

[0018] Step 116: Input the real-scene natural language question information or user question information into the software operation hybrid interactive help model that is cross-software based on real-scene semantic understanding, and output the ranking score of each semantic parsing result;

[0019] Step 117: Select one or more semantic parsing results as the understanding results of the user's question based on the ranking score, and provide interactive operation assistance to the user based on the cross-software hybrid interactive assistance model.

[0020] Preferably, the software information in the electronic device includes:

[0021] The operating system and its version name in electronic devices; software name, software version and its specific functional information.

[0022] Button information includes:

[0023] Button name, button range information, relationship information between buttons, and button function information;

[0024] Button range information refers to the range contained in the button's border or vertex information, and the transformation relationship of the button range between different screen sizes and resolutions;

[0025] Button relationships refer to the relationship between the current button and the sub-buttons that appear when the current button is clicked;

[0026] Button function information refers to the descriptive information about the button's function;

[0027] Cross-software refers to software on different electronic devices or different software on the same device, including operating systems.

[0028] As a preferred option, step 111 specifically involves:

[0029] Step 1-1: Extract button information of interactive software operation buttons based on the real-scene cross-software software operation hybrid interactive help system;

[0030] Steps 1-2: Based on the knowledge of button usage issues, use rule-based methods or if-then-based methods to generate training data that combines real-world information with real-world natural language question information, user intent, and help information to generate a software button relationship tree.

[0031] The button relationship tree refers to the tree-like hierarchical relationship formed by the parent and child buttons of the interactive software.

[0032] The real-world natural language question information, user intent, and help information that combine real-world information are used in a way that combines real-world information as part of the question information in the form of natural language abbreviation recovery, or in the form of real-world information as part of the question information in the form of an attention mechanism.

[0033] Steps 1-3: Further generate training data based on the thesaurus and the real-world natural language question information, user intent, and help information that incorporates real-world information;

[0034] Steps 1-4: Collect the usage problems and help information of the interactive software and help information of other problems, including problems that can be solved across software and their help information; collect text, voice and image data of other types of problems and solutions from the Internet and books, and label them as training data.

[0035] Help information for other issues refers to the types of problems that can be solved by existing generative AI-based large models, but does not include software operation and usage issues;

[0036] Steps 1-5: Use the training data as training samples.

[0037] As a preferred option, step 112 specifically involves:

[0038] Extract software and hardware information from users' electronic devices, including software and hardware information from users' computers, mobile phones, tablets, iPads and home appliances, to build a data model of users' electronic devices;

[0039] Hardware and software information includes: CPU, memory size, hard disk size, and user screen resolution and scaling ratio in the user's electronic device; operating system and its version name; software name, software version and its storage location information.

[0040] As a preferred option, step 113 specifically involves:

[0041] Using the training samples and the help information, a software operation hybrid interactive help model based on real-scene semantic understanding that can be used across software is trained. The software operation hybrid interactive help model adopts a dialogue form based on deep learning, and the software real-scene information corresponding to the user's software operation questions is added to the training model in a weighted manner during training.

[0042] The dialogue format, which employs natural language processing techniques in deep learning, incorporates real-world software information into the training model. This can be either a Transformer model using an attention mechanism or a generative artificial intelligence model using Transformer.

[0043] As a preferred approach, when using the Transformer model, the real-scene encoding layer is fused with the embedding layer to obtain the real-scene embedding vector, which is then input into the encoding component of the Transformer for feature learning and weight modification.

[0044] Based on the embedding vector of the real scene fused into the coding layer, multi-head self-attention calculation is performed and the resulting attention vector is used as the intermediate state of the coding layer;

[0045] After passing through a layer-by-layer feedforward neural network, the hidden layer feature vectors corresponding to all labels in the input sequence are obtained;

[0046] By selecting the cross-entropy loss function, we obtain the cross-entropy loss of the module after real-scene omission restoration.

[0047] As a preferred option, step 115 specifically involves:

[0048] Step 5-1: Extract natural language questions asked by users when using interactive software;

[0049] Extract the current real-time information of the software, including the software name, the actual state of the menu bar, button information, button click information, and information on the screen size and resolution of the pop-up menus and buttons, as well as whether the software usage problem button has been clicked;

[0050] The acquisition of real-world information includes "button information" provided by the operating system and software interfaces, or button and menu operation information obtained by using hook methods, system monitoring or polling scheduling methods, or deep learning-based image processing technology.

[0051] Software usage issues, including button location finding, button function querying, and help on how to use buttons, as well as issues that need to be resolved across different software programs;

[0052] Step 5-2: If the user clicks the software usage question button, the real-scene information and natural language question information will be combined to generate real-scene natural language question information. The combination method includes using the real-scene information as part of the question information in the form of language omission restoration, or using the real-scene information as part of the question information in the form of attention mechanism.

[0053] If the user does not click the "Software Usage Questions" button, the real-world natural language question information will only include the user's question information.

[0054] Users can improve system response speed by selecting either general questions or software operation questions.

[0055] As a preferred option, step 117 specifically involves:

[0056] Step 7-1: Select one or more semantic parsing results based on the ranking score as the understanding results of the user's question;

[0057] Step 7-2: Provide interactive operation assistance to users based on the cross-software hybrid interactive assistance model. The cross-software hybrid interactive assistance model includes:

[0058] Button location search help: Based on the current menu state and button relationship tree of the interactive software, it provides sequential prompts for clicking one or more buttons. The prompts are arrows pointing to the buttons, or highlights or flashes at the button locations.

[0059] Button Function Help: Provides descriptions of button functions, including text, images, and videos;

[0060] How to use the buttons: This section introduces the help functions for combined buttons and provides real-world examples of how to find the button locations.

[0061] Click-button help: Includes step-by-step prompts and automatic button clicks, which can be configured by the user;

[0062] The step-by-step prompts provide operation instructions and guide the user to click the buttons sequentially to complete the task.

[0063] The automatic button click function: automatically completes the button click or calls the corresponding method to complete the button click, allowing the user to complete the required operation;

[0064] Cross-software recommendation: For user problems that require cross-software or cross-platform solutions, the system provides recommended usage information based on the user's electronic device information model, including prioritizing the recommendation of software suitable for the local device, and recommending other suitable software if no suitable software is available on the local device.

[0065] Cross-device step-by-step guidance or automatic clicks: For user problems that require cross-device solutions, interactive help information can be transmitted to the corresponding device via network communication, with the user's permission, and the corresponding software and functions can be activated to achieve cross-platform step-by-step prompts and automatic button clicks;

[0066] Software usage issues: button location finding, button function query and button usage help, as well as issues that need to be resolved across different software programs;

[0067] Step 7-3: If the user does not click the software usage problem button, extract the results that match the real-scene information and the solution to the current software usage problem based on the real-scene information, and have a dialogue with the user. Initiate the next round of dialogue based on the real-scene information to let the user confirm the returned result.

[0068] If the user confirms that the problem is being addressed using the software or a cross-software solution, then guide them to resolve the corresponding issue; if the user confirms that the problem is not being addressed using the software, then return all the content from the system feedback; if the returned results do not contain any content related to the real-world information, then return all the content from the system feedback.

[0069] Step 7-4: When the user's purpose is not clearly understood, initiate a dialogue to further clarify the user's needs.

[0070] Preferably, the natural language question information is voice information or text information;

[0071] For speech information, before semantic parsing of natural language information, speech recognition is used to convert the natural language information into text information.

[0072] Another aspect of the present invention provides a reality-based, cross-software hybrid interactive assistance system for software operation, comprising:

[0073] The interactive software button information acquisition module is used to help system developers obtain button information through mixed interaction of software operation. Button function information, button border or vertex information, current button and its parent button are filled in manually in the menu provided by the development platform, or automatically filled in after object detection or other image recognition in deep learning using OCR.

[0074] The real-scene natural language question-answer pair generation module is used to generate a software button relationship tree based on button information and button usage questions; further generate training data; use the training data as training samples; generate text, voice, and image data based on a thesaurus and real-scene natural language question information combined with real-scene information, user intent and help information, and other types of questions and solutions from the Internet and books, and further generate training data.

[0075] The software operation interaction help model construction module is used to train a cross-software hybrid interaction help model based on real-scene semantic understanding according to the training samples.

[0076] The real-scene acquisition module is used to acquire real-scene information, including software name and version, current menu status information, operating system and its version, and "button information" provided by the software interface, or button and menu operation information obtained by using hook methods, system monitoring methods, polling scheduling methods, or deep learning-based image processing technology.

[0077] The real-scene natural language question information generation module extracts the natural language question information asked by the user while using the interactive software, and combines the real-scene information with the natural language question information to generate real-scene natural language question information or only uses the user's question information.

[0078] The cross-software hybrid interactive help module based on real-scene semantic understanding adopts a deep learning-based dialogue system. It is used to understand users' natural language questions containing real-scene information. Users can improve the system's response speed by selecting general questions or software operation questions, and interactive help is provided based on the questions and the above selections, including help to query button functions, help to find button locations, or help on how to use buttons.

[0079] Hints are arrows pointing to the button, or highlighted or flashing at the button's location; button click help methods include step-by-step hints and automatic button clicks, which can be configured by the user.

[0080] The system also includes a speech recognition module, used to convert the natural language information into text information through speech recognition before the semantic parsing module performs semantic parsing on the natural language information.

[0081] The present invention, by adopting the above technical solution, has the following beneficial effects:

[0082] This invention provides precise guidance to users on solving problems by accurately sensing their software operating environment and continuous operation actions, combined with natural language questions posed during software operation. This solves the comprehension difficulties and even errors caused by the lack of contextual semantic information in general deep learning methods, and improves the human-computer interaction effect and efficiency of cross-software operation. Furthermore, the integration of step-by-step problem-solving prompts, automatic click-based problem-solving, and problem-solving in specific software operation scenarios with other types of problem-solving further improves the efficiency of software interaction. Attached Figure Description

[0083] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, do not constitute an undue limitation of the invention. In the drawings:

[0084] Figure 1 This is a flowchart of a software operation hybrid interactive assistance method based on real-world scenarios and compatible with various software, according to an embodiment of the present invention.

[0085] Figure 2 This is a schematic diagram of a development module for a real-world, cross-software software operation hybrid interactive assistance system according to an embodiment of the present invention;

[0086] Figure 3 This is a schematic diagram of a service module of a software operation hybrid interactive help system based on real-world scenarios and capable of operating across software, according to an embodiment of the present invention.

[0087] Figure 4 This is a schematic diagram of a cross-device step-by-step guidance or automatic click structure. Detailed Implementation

[0088] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions of the present invention are used to explain the present invention, but are not intended to limit the present invention.

[0089] like Figure 1 The diagram shown is a flowchart of a natural language understanding method provided by an embodiment of the present invention. The present invention provides a real-world, cross-software-enabled, hybrid interactive assistance method for software operation, comprising the following steps:

[0090] Step 111: Extract software information from electronic devices, button information of interactive software operation buttons, collect help information on usage problems and other issues, and generate training data for a cross-software hybrid interactive help system based on real-scene semantic understanding.

[0091] Specifically, the following steps are included:

[0092] Step 1-1: Extract button information of interactive software operation buttons based on the real-scene cross-software software operation hybrid interactive help system.

[0093] Cross-software refers to software on different electronic devices or different software on the same device; the software here includes the operating system.

[0094] Software information in electronic devices includes: the operating system and its version name, software name, software version and its specific function information.

[0095] Button information includes button name, button range information, button relationship information, and button function information. Button range information refers to the range contained by the button's border or vertex information and the variation relationship of the button range across different screen sizes and resolutions; button relationship refers to the relationship between the current button and the sub-buttons that appear when the current button is clicked; button function information refers to the description of the button's function; cross-software refers to software on different electronic devices or different software on the same device, and software includes the operating system.

[0096] In one embodiment: To obtain the button information of the "flowchart arrows" in Word, first, you need to launch Word and the interactive help system that supports cross-software operation based on a real-world scenario. Place the mouse over the upper left and lower right corners of the "Arrow Collection" button and follow the system prompts to obtain the button's range information. Then, manually enter the button name "Arrow Collection," the parent node "Shape" button, and the function information and usage steps. For functions that require cross-software integration, the Word function information should include the advantages of the corresponding Visio software, such as more convenient drawing and adjustment tools.

[0097] Steps 1-2: Based on the knowledge of button usage issues, use rule-based methods or if-then-based methods to generate training data that combines real-world natural language questions about usage issues, user intent, and help information, and generate a software button relationship tree.

[0098] A button relationship tree refers to the hierarchical tree-like relationship between parent and child buttons in interactive software.

[0099] This approach combines real-world information usage questions with real-world natural language question information, user intent, and help information. The method integrates real-world information as part of the question information in the form of natural language abbreviation recovery, or as part of the question information through an attention mechanism.

[0100] Steps 1-3 involve generating training data based on a thesaurus and by incorporating real-world information, usage questions, user intent, and help information.

[0101] Steps 1-4: Collect usage problems and help information of the interactive software and help information of other problems, including problems that can be solved across software and their help information; collect text, voice and image data of other types of problems and solutions from the Internet and books, and label them as training data; use the above training data as training samples.

[0102] Other types of problems refer to the types of problems that can be solved by existing generative artificial intelligence models, but do not include software operation and usage problems.

[0103] In one embodiment, a synonym for "save" could be: save a file.

[0104] Collect and label other user questions and help information about the interactive software, and use them as training data; in one embodiment, "saving" can also be described as: saving the file, etc.

[0105] The input sequence INPUT1 for the user's query question is:

[0106] INPUT1 = <x1,x2,…,x m >;

[0107] Then, the input sequence INPUT2 based on the real-world scene should be:

[0108] INPUT2 = <y1,y2,…,y n ,x1,x2,…,x m >;

[0109] Where y1, y2, ..., y n For real-world environment information and user operation information

[0110] The processed combined real-scene input sequence INPUT3 is:

[0111] INPUT3 = <z,x1,x2,…,x m >;

[0112] Where z represents the content corresponding to the real scene, which can be the dialogue content restored by omitting the real scene, or the content corresponding to the real scene after being processed by the attention mechanism.

[0113] Steps 1-5: Use the above training data as training samples.

[0114] Step 112: Extract the hardware and software information from the user's electronic device and construct a data model of the user's electronic device, specifically including:

[0115] Extract software and hardware information from users' electronic devices, including information on users' computers, mobile phones, tablets (including iPads), and home appliances, to construct a data model of users' electronic devices.

[0116] Hardware and software information includes: hardware information such as CPU, memory size, hard disk size, and user screen resolution and scaling ratio in the user's electronic device; operating system and its version name; software name, software version and its storage location.

[0117] Step 113: Based on the training data, pre-construct a deep learning-based, real-scene semantic understanding-based, cross-software hybrid interactive assistance model for software operation, specifically including:

[0118] A software operation hybrid interactive help model based on real-scene semantic understanding and capable of operating across software is obtained by training training samples and help information. The software operation hybrid interactive help model adopts a dialogue form based on deep learning, and the software real-scene information corresponding to the user's software operation questions is added to the training model in a weighted manner during training.

[0119] The dialogue format, which employs natural language processing techniques in deep learning, incorporates real-world software information into the training model. This can be either a Transformer model using an attention mechanism or a generative artificial intelligence model using Transformer.

[0120] When using the Transformer model, the reality encoding layer is fused with the embedding layer to obtain the reality embedding vector E. sense The input is then fed into the encoding component of the Transformer for feature learning and weight modification. The encoding component consists of multiple identical encoding layers stacked on top of each other. Each encoding layer contains a multi-head self-attention sublayer and a layer-by-layer feedforward network sublayer. Each sublayer uses residual connections and layer normalization operations.

[0121] After the embedded vector of the fused real-world scene is input into the coding layer, multi-head self-attention is first calculated, and the resulting attention vector is used as the intermediate state of the coding layer, as shown in the following formula:

[0122]

[0123] in, Let represent the input vector of the i-th encoding layer, which is used as the query vector, key vector, and value vector of the attention machine function. Let d represent the attention vector obtained from the i-th coding layer. k The dimension of the input key vector.

[0124] After passing through a layer-by-layer feedforward neural network, the feature vector of the current encoding layer, which incorporates real-world information, is obtained. Finally, after passing through the last encoding layer, the hidden layer feature vectors corresponding to all labels in the input sequence are obtained, which can be represented as:

[0125] H sense =FFN(Z sense )

[0126] H sense =(h1,h2,…,h l ),h i ∈R M

[0127] Wherein, FFN is a feedforward neural network, H sense Z is the real-world feature vector representation output by the last encoder. sense h is the attention vector output by the last encoder. i Let represent the i-th element of the input vector, l be the length of the input sequence containing special labels, and M represent the vector dimension of the hidden layer.

[0128] For the loss function, we choose the cross-entropy loss function, and the specific calculation formula is shown in the following formula:

[0129]

[0130] y max =max(P(y) I ,y S |I * ))

[0131] L Sense =∑ i (y i ln(y max )+(1-y i ln(1-y max )))

[0132] Among them, y i The output y represents the probability of all intent categories and slot names corresponding to the input sequence that incorporates real-world information. max L represents the maximum probability of intent category and slot name in the probability output. Sense This represents the cross-entropy loss of the module after real-scene omission restoration.

[0133] Step 114: Extract the user's natural language question information, including: if the user did not click the software usage question button, the user's question will be treated as a non-software usage question and resolved.

[0134] Step 115: Based on the constructed user electronic device data model, extract the current real-world information of the software, and generate real-world natural language question information based on the user's questions and operations, including:

[0135] Step 5-1: Extract the user's natural language questions and current real-world information from the interactive software.

[0136] Extracting the current real-world information of the software refers to: the user's experience while using the software, including the software name, the actual state of the menu bar, button information, button click information, and information on the screen size and resolution of pop-up menus and buttons, as well as whether the software usage problem button has been clicked; the real-world information is obtained using "button information" provided by the operating system and software interface, or by using hook methods, system monitoring methods, polling scheduling methods, or deep learning-based image processing technology to identify button and menu operation information; software usage problems include button location search, button function query and button usage help, as well as problems that need to be solved across software.

[0137] Step 5-2: If the user clicks the "Software Usage Questions" button, the real-world information and the natural language question information will be combined to generate real-world natural language question information. The combination method includes incorporating real-world information as part of the question information through language omission restoration, or using an attention mechanism to incorporate real-world information as part of the question information. If the user does not click the "Software Usage Questions" button, the real-world-based natural language question information will only include the user's question information.

[0138] Users can improve system response speed by selecting either general questions or software operation questions.

[0139] Step 116: Input the real-scene-based natural language question information or user question information into the real-scene-based semantic understanding cross-software hybrid interactive help model, and output the ranking scores of each semantic parsing result, including:

[0140] Input the user's natural language questions or real-scene-based natural language questions into a deep learning-based real-scene semantic understanding cross-software hybrid interactive assistance model; output the ranking score of each semantic parsing result, providing the semantic parsing results combined with the real-scene and their ranking scores.

[0141] Step 117: Select one or more semantic parsing results based on the ranking score as the understanding result of the user's question, and then provide interactive operation assistance to the user based on the cross-software hybrid interactive assistance model, including:

[0142] Step 7-1: Select one or more semantic parsing results based on the ranking score as the understanding results of the user's question;

[0143] Step 7-2: Provide interactive operation assistance to users based on the cross-software hybrid interactive assistance model. The cross-software hybrid interactive assistance model includes:

[0144] Button location search help: Based on the current menu status and button relationship tree of the interactive software, it provides sequential prompts for clicking one or more buttons. The prompts can be arrows pointing to the buttons, or highlights or flashes at the button locations.

[0145] Button Function Help: Provides descriptions of button functions, including text, images, and videos;

[0146] How to use the buttons: This section introduces the practical results of using the combination button function to query help and the button location to find help.

[0147] Help methods for clicking buttons include step-by-step prompts and automatic button clicks, which can be configured by the user.

[0148] Step-by-step prompts: Provide operation instructions and guide users to click the buttons in sequence to complete the process.

[0149] Automatic button click: Automatically completes button clicks or calls the corresponding method to complete button clicks, allowing the user to complete the required operation.

[0150] Cross-software recommendation: For user problems that require cross-software or cross-platform solutions, the system provides recommended usage information based on the user's electronic device information model. This includes prioritizing software suitable for the local device, and recommending other suitable software if no suitable software is available on the local device.

[0151] Cross-device step-by-step guidance or automatic clicks: For user problems that require cross-device solutions, interactive help information can be transmitted to the corresponding device via network communication, with the user's permission, and the corresponding software and functions can be activated to achieve cross-platform step-by-step prompts and automatic button clicks.

[0152] Software usage issues: button location finding, button function query, help on how to use buttons, and issues that need to be resolved across different software.

[0153] Step 7-3: If the user does not click the software usage problem button, extract the results that match the real-world information from the returned results, as well as the solutions to the current software usage problem, and engage in dialogue with the user. Initiate the next round of dialogue based on the real-world information to allow the user to confirm the returned results.

[0154] If the user confirms that the problem is being addressed using the software or a cross-software solution, then guide them to resolve the issue. If the user confirms that the problem is not being addressed using the software, then return all the information returned by the system. If the returned results do not contain any information related to the real-world scenario, then return all the information returned by the system.

[0155] After processing the task-oriented dialogue using deep learning, the output is:

[0156] OUTPUT1 = <y1,y2,…,y n >

[0157] The output sequence includes button operation commands, cross-software recommendations, and cross-device operation commands.

[0158] This solution addresses the shortcomings of general user manuals being unintuitive. By combining intuitive help with real-world scenarios, it greatly enhances the user experience and improves efficiency. Furthermore, cross-software recommendations and cross-device operation further improve the quality of problem-solving.

[0159] The probability of predicting a solution to the problem, based on a combination of real-world information and user-generated questions, is as follows:

[0160] p1 = (y|x1, x2, ..., x) i-1 x i x i+1 , ..., x n )

[0161]

[0162] x1,x2,…,x i-1 This represents real-world information about user actions, including sequences of software inputs. i ,x i+1 ,…,x n Represents the user's natural language description of the problem;

[0163] The probability p2 of predicting the solution to the problem is given by combining real-world information with natural language question information to generate real-world natural language question information.

[0164] p2=q(y|x0,x1,x2,…,x m )

[0165]

[0166] x1,x2,…,x m Represents the user's natural language description of the problem;

[0167] p1>=p2

[0168] Clearly, given that computer technology can accurately capture continuous user actions, the probability of understanding a user's natural language questions is greater than or equal to the probability of understanding a user's natural language questions through deep learning-video multimodal technology, especially considering that the success rate of existing video multimodal technologies is not ideal.

[0169] In practice, in one embodiment 1, when a user asks how to draw a flowchart while using Word, if it is found that the user has Visio software on their machine, it can be recommended that they use Visio to draw the flowchart with a reason for the recommendation. After the user agrees, the system will automatically start or guide them to start their Visio software. The system will automatically switch to Visio interactive help, and after the flowchart is drawn, it will guide them to insert it into the Word text.

[0170] In one embodiment 2, when a user is using Word and asks how to extract large blocks of text from images in a Word document, if no corresponding image text extraction software is installed, and the user frequently uses WeChat on their mobile phone, they can be prompted to select the image and right-click to select the "Save image as" button. After saving the image, they can open WeChat on their computer, select a friend or tool such as a file transfer assistant, send the image to WeChat on their mobile phone, open the image again, long-press the image, and click the "Extract Text" button in WeChat to extract the information. Then, they can select all the information and send it in WeChat. After extracting the text on their computer, they can insert it into the appropriate position in Word. The button prompts can be implemented by gradually prompting the user to click the corresponding button, or the operation can be completed automatically by clicking the button after completing the corresponding action.

[0171] Step 7-4: When the user's purpose is not clearly understood, initiate a dialogue to further clarify the user's needs.

[0172] Furthermore, natural language questions can be in the form of voice or text.

[0173] For speech information, before semantic parsing of natural language information, speech recognition is used to convert natural language information into text information.

[0174] Accordingly, this invention also provides a real-scene-based, cross-software hybrid interactive help system for software operation, and the structural diagram of the development module is shown below. Figure 2 As shown; a structural diagram of the service module, as follows. Figure 3 As shown; and a schematic diagram of a step-by-step guide or automatic click structure across devices, such as Figure 4 As shown.

[0175] This reality-based, cross-software-compatible, hybrid interactive help system includes:

[0176] Interactive software button information acquisition module 201, such as Figure 2As shown, this software is used to help system developers obtain button information through a combination of software operation and interaction. Button function information, button border or vertex information, current button and its parent button can be manually filled in the menu provided by the development platform, or automatically filled in after object detection or other image recognition in deep learning using OCR.

[0177] The real-scene natural language question-answer pair generation module 202 is used to generate a software button relationship tree based on button information and button usage questions; further generate training data; use the training data as training samples; generate text, voice and image data based on a thesaurus and real-scene natural language question information combined with real-scene information, user intent and help information and other types of questions and solutions from the Internet and books, and further generate training data.

[0178] The software operation interaction help model construction module 203 is used to train a software operation hybrid interaction help model that can be used across software based on real-scene semantic understanding according to training samples.

[0179] Reality acquisition module 301, such as Figure 3 As shown, the information used to obtain real-world information includes the software name and version, the current status of the menu, the operating system and its version, the "button information" provided by the software interface, or the button and menu operation information obtained by using hook methods, system monitoring methods, polling scheduling methods, or image processing technology based on deep learning.

[0180] The real-scene natural language question information generation module 302 extracts the natural language question information of the user using the interactive software according to the user's selection, and combines the real-scene information with the natural language question information to generate real-scene natural language question information or only uses the user's question information.

[0181] The cross-software hybrid interactive help module 303, based on real-scene semantic understanding, adopts a deep learning-based dialogue system. It is used to understand the user's natural language questions containing real-scene information. The user can choose between general questions or software operation questions to improve the system's response speed. Interactive help is provided based on the question and the above selection. The combination button function queries help and the button location finds help in real-scene results.

[0182] Hints can be arrows pointing to the button, or highlights or flashes at the button's location; button click help methods include step-by-step hints and automatic button clicks, which can be configured by the user.

[0183] Step-by-step prompts: Provide operation instructions and guide users to click the buttons in sequence.

[0184] Automatic button click: Automatically completes button clicks or calls the corresponding method to complete button clicks, allowing the user to complete the required operation.

[0185] Cross-software recommendation: For user problems that require cross-software or cross-platform solutions, recommendations are provided based on the user's electronic device information model, including prioritizing suitable software for the local device and recommending other suitable software if the local device is not available.

[0186] Step-by-step guidance or automatic clicks across devices: such as Figure 4 For user problems that require cross-device solutions, with the user's permission, electronic device I 401 sends information to electronic device II 402 through step-by-step prompts or automatic button clicks. Electronic device II 402 then transmits the interactive help information to the corresponding electronic device I 401 via network communication, activating the corresponding software and functions to achieve cross-platform step-by-step prompts and automatic button clicks.

[0187] Software usage issues: button location finding, button function query, help on how to use buttons, and issues that need to be resolved across different software.

[0188] If the user does not click the "Software Usage Issues" button, the returned results will extract the matching results based on the real-world information, along with solutions for the current software usage issues. A dialogue will then be initiated with the user, using the real-world information to allow them to confirm the returned results. If the user confirms that the solution is for the current software or a cross-software issue, guidance will be provided to resolve the corresponding problem. If the user confirms that the solution is not for the current software, all system feedback will be returned. If the returned results do not contain any content related to the real-world information, all system feedback will be returned.

[0189] The system also includes a speech recognition module, which converts natural language information into text information through speech recognition before the semantic parsing module performs semantic parsing on the natural language information.

[0190] This invention employs a real-scene-based, cross-software hybrid interactive assistance method for software operation, solving problems related to software usage efficiency and difficulty in finding solutions. This invention can accurately perceive continuous actions, accurately understand them, and accurately guide users, achieving high-quality interactive assistance for software operation. By integrating its training data with existing general-purpose large-scale model training data, it can expand the application scenarios of large-scale models from simple large language (text) and image-language multimodal models (VLM) to continuous action-language multimodal models for software operation. Due to the wide user base of real-scene interactive assistance systems, it not only enables the application of large-scale models in the field of software operation problem solving but also promotes the widespread adoption of large-scale models.

[0191] This invention is not limited to the above embodiments. Based on the technical solutions disclosed in this invention, those skilled in the art can make some substitutions and modifications to some of the technical features without creative effort, and all such substitutions and modifications are within the protection scope of this invention.

Claims

1. A method for providing hybrid interactive assistance for software operation across multiple software applications based on real-world scenarios, characterized in that: Includes the following steps: Step 111: Extract software information from electronic devices, button information of interactive software operation buttons, collect help information on usage problems and other problems, and generate training data for a cross-software hybrid interactive help system based on real-scene semantic understanding. Step 112: Extract the hardware and software information from the user's electronic device and construct a data model of the user's electronic device; Step 113: Based on the training data, pre-construct a software operation hybrid interaction help model that uses deep learning and is based on real-scene semantic understanding and can be used across software. Step 114: Extract the natural language question information from the user's question; Step 115: Based on the constructed user electronic device data model, extract the current real-world information of the software, and generate real-world natural language question information based on the user's questions and operations; Step 116: Input the real-scene natural language question information or user question information into the software operation hybrid interactive help model that is cross-software based on real-scene semantic understanding, and output the ranking score of each semantic parsing result; Step 117: Select one or more semantic parsing results as the understanding results of the user's question based on the ranking score, and provide interactive operation assistance to the user based on the cross-software hybrid interactive assistance model; Step 111 specifically involves: Step 1-1: Extract button information of interactive software operation buttons based on the real-scene cross-software software operation hybrid interactive help system; Steps 1-2: Based on the knowledge of button usage issues, use rule-based methods or if-then-based methods to generate training data that combines real-world information with real-world natural language question information, user intent, and help information to generate a software button relationship tree. Steps 1-3: Further generate training data based on the thesaurus and the real-world natural language question information, user intent, and help information that incorporates real-world information; Steps 1-4: Collect the usage problems and help information of the interactive software and help information of other problems, including problems that can be solved across software and their help information; collect text, voice and image data of other types of problems and solutions from the Internet and books, and label them as training data. Steps 1-5: Use the training data as training samples; The button relationship tree refers to the tree-like hierarchical relationship formed by the parent and child buttons of the interactive software. The real-world natural language question information, user intent, and help information that combine real-world information are used in a way that combines real-world information as part of the question information in the form of natural language abbreviation recovery, or in the form of real-world information as part of the question information in the form of an attention mechanism. The help information for other issues refers to the types of problems that can be solved by existing large generative AI models, but does not include the part about software operation and usage.

2. The method according to claim 1, wherein the software information in the electronic device includes: The operating system and its version name in electronic devices; software name, software version and its specific functional information. Button information includes: Button name, button range information, relationship information between buttons, and button function information; Button range information refers to the range contained in the button's border or vertex information, and the transformation relationship of the button range between different screen sizes and resolutions; Button relationships refer to the relationship between the current button and the sub-buttons that appear when the current button is clicked; Button function information refers to the descriptive information about the button's function; Cross-software refers to software on different electronic devices or different software on the same device, including operating systems.

3. The method according to claim 1, characterized in that, Step 112 specifically involves: Extract software and hardware information from users' electronic devices, including software and hardware information from users' computers, mobile phones, tablets, iPads and home appliances, to build a data model of users' electronic devices; Hardware and software information includes: CPU, memory size, hard disk size, and user screen resolution and scaling ratio in the user's electronic device; operating system and its version name; software name, software version and its storage location information.

4. The method according to claim 1, characterized in that, Step 113 specifically refers to: Using the training samples and the help information, a software operation hybrid interactive help model based on real-scene semantic understanding that can be used across software is trained. The software operation hybrid interactive help model adopts a dialogue form based on deep learning, and the software real-scene information corresponding to the user's software operation questions is added to the training model in a weighted manner during training. The dialogue format, which employs natural language processing techniques in deep learning, incorporates real-world software information into the training model. This can be either a Transformer model using an attention mechanism or a generative artificial intelligence model using Transformer.

5. The method according to claim 1, characterized in that, Step 115 specifically refers to: Step 5-1: Extract natural language questions asked by users when using interactive software; Extract the current real-time information of the software, including the software name, the actual state of the menu bar, button information, button click information, and information on the screen size and resolution of the pop-up menus and buttons, as well as whether the software usage problem button has been clicked; The acquisition of real-world information includes "button information" provided by the operating system and software interfaces, or button and menu operation information obtained by using hook methods, system monitoring or polling scheduling methods, or deep learning-based image processing technology. Software usage issues, including button location finding, button function querying, and help on how to use buttons, as well as issues that need to be resolved across different software programs; Step 5-2: If the user clicks the software usage question button, the real-scene information and natural language question information will be combined to generate real-scene natural language question information. The combination method includes using the real-scene information as part of the question information in the form of language omission restoration, or using the real-scene information as part of the question information in the form of attention mechanism. If the user does not click the "Software Usage Questions" button, the real-world natural language question information will only include the user's question information. Users can improve system response speed by selecting either general questions or software operation questions.

6. The method according to claim 1, characterized in that, Step 117 specifically refers to: Step 7-1: Select one or more semantic parsing results based on the ranking score as the understanding results of the user's question; Step 7-2: Provide interactive operation assistance to users based on the cross-software hybrid interactive assistance model. The cross-software hybrid interactive assistance model includes: Button location search help: Based on the current menu state and button relationship tree of the interactive software, it provides sequential prompts for clicking one or more buttons. The prompts are arrows pointing to the buttons, or highlights or flashes at the button locations. Button Function Help: Provides descriptions of button functions, including text, images, and videos; How to use the buttons: This section introduces the help functions for combined buttons and provides real-world examples of how to find the button locations. Click-button help: Includes step-by-step prompts and automatic button clicks, which can be configured by the user; The step-by-step prompts provide operation instructions and guide the user to click the buttons sequentially to complete the task. The automatic button click function: automatically completes the button click or calls the corresponding method to complete the button click, allowing the user to complete the required operation; Cross-software recommendation: For user problems that require cross-software or cross-platform solutions, the system provides recommended usage information based on the user's electronic device information model, including prioritizing the recommendation of software suitable for the local device, and recommending other suitable software if no suitable software is available on the local device. Cross-device step-by-step guidance or automatic clicks: For user problems that require cross-device solutions, interactive help information can be transmitted to the corresponding device via network communication, with the user's permission, and the corresponding software and functions can be activated to achieve cross-platform step-by-step prompts and automatic button clicks; Software usage issues: button location finding, button function query and button usage help, as well as issues that need to be resolved across different software programs; Step 7-3: If the user does not click the software usage problem button, extract the results that match the real-scene information and the solution to the current software usage problem based on the real-scene information, and have a dialogue with the user. Initiate the next round of dialogue based on the real-scene information to let the user confirm the returned result. If the user confirms that the problem is being addressed using the software or a cross-software solution, then guide them to resolve the corresponding issue; if the user confirms that the problem is not being addressed using the software, then return all the content from the system feedback; if the returned results do not contain any content related to the real-world information, then return all the content from the system feedback. Step 7-4: When the user's purpose is not clearly understood, initiate a dialogue to further clarify the user's needs.

7. The method according to any one of claims 1 to 6, characterized in that, The natural language query information can be either voice or text. For speech information, before semantic parsing of natural language information, speech recognition is used to convert the natural language information into text information.

8. A hybrid interactive assistance system for software operation based on real-world scenarios and compatible with various software, as described in claim 7, characterized in that... include: The interactive software button information acquisition module is used to help system developers obtain button information through mixed interaction of software operation. Button function information, button border or vertex information, current button and its parent button are filled in manually in the menu provided by the development platform, or automatically filled in after object detection or other image recognition in deep learning using OCR. The real-scene natural language question-answer pair generation module is used to generate a software button relationship tree based on button information and button usage questions; Further training data is generated; Use the training data as training samples; Generate text, voice, and image data based on a thesaurus and real-world natural language questions that combine real-world information with usage questions, user intent and help information, and other types of questions and solutions from the Internet and books, and further generate training data; The software operation interaction help model construction module is used to train a software operation hybrid interaction help model that is cross-software based on real-scene semantic understanding according to the training samples. The real-scene acquisition module is used to acquire real-scene information, including software name and version, current menu status information, operating system and its version, and "button information" provided by the software interface, or button and menu operation information obtained by using hook methods, system monitoring methods, polling scheduling methods, or deep learning-based image processing technology. The real-scene natural language question information generation module extracts the natural language question information asked by the user while using the interactive software, and combines the real-scene information with the natural language question information to generate real-scene natural language question information or only uses the user's question information. The cross-software hybrid interactive help module based on real-scene semantic understanding adopts a dialogue system based on deep learning. It is used to understand the user's natural language questions containing real-scene information. Users can improve the system's response speed by selecting general questions or software operation questions. Based on the questions and the above selections, interactive help is provided, including help to query button functions, help to find button locations, or help on how to use buttons. The hint is an arrow pointing to the button, or a highlight or flashing icon at the button's location; Click-button help options include step-by-step suggestions and automatic button clicks, which can be configured by the user. The system also includes a speech recognition module, used to convert the natural language information into text information through speech recognition before the semantic parsing module performs semantic parsing on the natural language information.

Citation Information

Patent Citations

  • Voice interaction method, system and platform for software operation live-action semantic understanding

    CN113223520A