Intelligent equipment operation assisting method and system based on large model and related equipment

By obtaining the operation interface images of the smart device, using large-scale models to analyze the working software and content information, and generating operation prompt information, the complex operation of the smart device is solved and the user experience is improved.

CN120447811AInactive Publication Date: 2025-08-08SHENZHEN XUNFANG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510518963.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The lack of real-time operation assistance methods for the use of smart devices in the prior art, resulting in high difficulty in user operation and poor user experience.

Method used

By obtaining the operation interface image of the smart device, using the preset work content analysis model to determine the work software information and content information, obtain the target work area image matched by the operation cursor, and generate operation prompt information based on the multimodal model.

Benefits of technology

Real-time operation assistance of smart devices is realized, reducing operation difficulty and improving user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120447811A_ABST
    Figure CN120447811A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent equipment operation assisting method and system based on a large model and related equipment, and relates to the technical field of software, and the method comprises the following steps: obtaining an operation interface image corresponding to intelligent equipment; according to the operation interface image, working software information and working content information corresponding to the intelligent equipment are determined through a preset working content analysis large model; if it is determined that working software corresponding to the intelligent device belongs to preset supported software according to the working software information, a target working area image matched with an operation cursor corresponding to the intelligent device is obtained, and the target working area image is used for representing part of information, corresponding to the position where the operation cursor belongs, in the working content information; and according to the target working area image, through a preset multi-mode large model, generating operation prompt information matched with the working content information for the intelligent device. Therefore, the operation assistance of the intelligent equipment can be realized, the operation difficulty of the intelligent equipment is reduced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of software technology, and in particular to a large model-based intelligent device operation assistance method, system and related equipment. Background Art

[0002] With the advancement of science and technology, the application of smart devices is becoming more and more extensive, and the operation process of various smart devices is becoming more and more complicated.

[0003] In the existing technology, there is a lack of real-time operation assistance methods for the use process of smart devices. Users can only rely entirely on their own experience to operate smart devices, which is not conducive to reducing the difficulty of operating smart devices and thus not conducive to improving the user experience.

[0004] Therefore, relevant technologies still need to be improved and developed. Summary of the Invention

[0005] The main purpose of this application is to provide a large-scale model-based smart device operation assistance method, system and related equipment, aiming to solve the technical problem that the related technology lacks a real-time operation assistance method for the use process of smart devices. When using smart devices, users can only rely entirely on their own experience to operate, which is not conducive to reducing the difficulty of operating smart devices and thus not conducive to improving the user experience.

[0006] In order to achieve the above-mentioned objectives, the present application provides, in a first aspect, a method for assisting smart device operation based on a large model, wherein the method for assisting smart device operation based on a large model comprises:

[0007] Obtain the operation interface image corresponding to the smart device;

[0008] According to the operation interface image, the working software information and working content information corresponding to the smart device are determined by using a preset working content analysis model;

[0009] If it is determined according to the working software information that the working software corresponding to the smart device belongs to the preset supported software, then obtaining a target working area image that matches the operation cursor corresponding to the smart device, wherein the target working area image is used to represent the portion of the work content information corresponding to the position of the operation cursor;

[0010] According to the target work area image, operation prompt information matching the work content information is generated for the smart device through a preset multimodal large model.

[0011] Optionally, obtaining an operation interface image corresponding to the smart device includes:

[0012] The screen image corresponding to the smart device is captured as the operation interface image.

[0013] Optionally, determining the work software information and work content information corresponding to the smart device based on the operation interface image and using a preset work content analysis model includes:

[0014] Get content analysis prompt words;

[0015] The operation interface image is input into the work content analysis model, and according to the content analysis prompt word, the operation interface image is subjected to content analysis by the work content analysis model to obtain the work software information and work content information corresponding to the smart device.

[0016] Optionally, the working software information includes software name, software version and software type.

[0017] Optionally, acquiring a target work area image that matches an operation cursor corresponding to the smart device includes:

[0018] Obtaining position information corresponding to an operation cursor of the smart device;

[0019] Determining a target working area corresponding to the smart device according to the location information and preset area size limiting parameters;

[0020] A screenshot is taken of the target working area to obtain an image of the target working area.

[0021] Optionally, generating operation prompt information matching the work content information for the smart device based on the target work area image through a preset multimodal large model includes:

[0022] Inputting the target work area image into a preset multimodal macromodel, and determining work content details matching the target work area image through a work detail analysis agent in the multimodal macromodel;

[0023] According to the work content details, operation prompt information matching the work content information is generated for the smart device through the work auxiliary agent in the multimodal large model.

[0024] Optionally, the smart device includes a computer.

[0025] A second aspect of the present application provides a large-scale model-based intelligent device operation assistance system, wherein the large-scale model-based intelligent device operation assistance system includes:

[0026] A first image acquisition module is used to acquire an operation interface image corresponding to the smart device;

[0027] An information determination module is used to determine the working software information and working content information corresponding to the smart device based on the operation interface image and a preset working content analysis model;

[0028] a second image acquisition module, configured to acquire a target work area image that matches an operation cursor corresponding to the smart device if it is determined based on the work software information that the work software corresponding to the smart device is preset supported software, wherein the target work area image is used to represent a portion of the work content information corresponding to a position of the operation cursor;

[0029] An operation prompt module is used to generate operation prompt information matching the work content information for the smart device based on the target work area image through a preset multimodal large model.

[0030] The third aspect of the present application provides an intelligent terminal, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of any one of the large model-based intelligent device operation assistance methods are implemented.

[0031] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any one of the large model-based smart device operation assistance methods.

[0032] As can be seen from the above, in the present application scheme, an operation interface image corresponding to the smart device is obtained; based on the operation interface image, the work software information and work content information corresponding to the smart device are determined through a preset work content analysis model; if it is determined according to the work software information that the work software corresponding to the smart device belongs to the preset supported software, then a target work area image matching the operation cursor corresponding to the smart device is obtained, wherein the target work area image is used to represent part of the work content information corresponding to the position of the operation cursor; based on the target work area image, operation prompt information matching the work content information is generated for the smart device through a preset multimodal large model.

[0033] Compared to the prior art, the solution corresponding to the large-scale model-based smart device operation assistance method provided in this application uses a large-scale work content analysis model to analyze the operation interface image of the smart device to determine the corresponding work software information and work content information of the smart device. After determining that the work software corresponding to the smart device belongs to the preset supported software based on the work software information, the target work area image corresponding to the smart device is further obtained, and based on the preset multimodal large-scale model, operation prompt information matching the work content information is generated for the smart device. This can automatically provide real-time operation prompts during the use of the smart device to achieve smart device operation assistance, which is conducive to reducing the difficulty of operating the smart device and thus helping to improve the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0035] Figure 1 This is a flow chart of a method for assisting smart device operation based on a large model provided in an embodiment of the present application;

[0036] Figure 2 This is a schematic diagram of the components of a large-scale model-based intelligent device operation assistance system provided in an embodiment of the present application;

[0037] Figure 3 This is a block diagram of the internal structure principle of a smart terminal provided in an embodiment of the present application. DETAILED DESCRIPTION

[0038] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it should be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0039] It will be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.

[0040] It should also be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0041] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0042] As used in this specification and the appended claims, the term "if" can be interpreted as meaning "when" or "upon" or "in response to determining" or "in response to being classified into," depending on the context. Similarly, the phrase "if it is determined" or "if it is classified into [described condition or event]" can be interpreted as meaning "upon determination" or "in response to determining" or "upon classification into [described condition or event]" or "in response to being classified into [described condition or event]," depending on the context.

[0043] The following is a clear and complete description of the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0044] In the following description, many specific details are set forth to facilitate a full understanding of the present application. However, the present application may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.

[0045] At present, the application of smart devices is becoming more and more extensive, and the functions of smart devices are becoming more and more diverse. Correspondingly, the operation process of various smart devices is becoming more and more complicated.

[0046] In the field of information technology, the development of large multimodal models marks a new height in artificial intelligence technology. These models can not only recognize and generate a variety of information forms such as text, voice, images, and video, but also show great potential in assisting computer operations. In this application, large multimodal models are used to efficiently assist in the operation of computers and other smart devices, thereby improving work efficiency and user experience.

[0047] In order to solve at least one of the above-mentioned technical problems, in the solution of the present application, an operation interface image corresponding to the smart device is obtained; based on the operation interface image, the work software information and work content information corresponding to the smart device are determined through a preset work content analysis model; if it is determined according to the work software information that the work software corresponding to the smart device belongs to the preset supported software, a target work area image matching the operation cursor corresponding to the smart device is obtained, wherein the target work area image is used to represent part of the work content information corresponding to the position of the operation cursor; based on the target work area image, operation prompt information matching the work content information is generated for the smart device through a preset multimodal large model.

[0048] Compared to the prior art, the solution corresponding to the large-scale model-based smart device operation assistance method provided in this application uses a large-scale work content analysis model to analyze the operation interface image of the smart device to determine the corresponding work software information and work content information of the smart device. After determining that the work software corresponding to the smart device belongs to the preset supported software based on the work software information, the target work area image corresponding to the smart device is further obtained, and based on the preset multimodal large-scale model, operation prompt information matching the work content information is generated for the smart device. This can automatically provide real-time operation prompts during the use of the smart device to achieve smart device operation assistance, which is conducive to reducing the difficulty of operating the smart device and thus helping to improve the user experience.

[0049] like Figure 1 As shown, the embodiment of the present application provides a method for assisting operation of a smart device based on a large model. Specifically, the method includes the following steps:

[0050] Step S100, obtaining an operation interface image corresponding to the smart device;

[0051] Step S200: determining the working software information and working content information corresponding to the smart device based on the operation interface image and a preset working content analysis model;

[0052] Step S300: If it is determined based on the working software information that the working software corresponding to the smart device is preset supported software, a target working area image that matches the operation cursor corresponding to the smart device is obtained, wherein the target working area image is used to represent the portion of the work content information corresponding to the position of the operation cursor;

[0053] Step S400: Based on the target work area image, operation prompt information matching the work content information is generated for the smart device through a preset multimodal large model.

[0054] The smart device is a device that requires operation assistance. In an embodiment of the present application, the software usage of the smart device is automatically analyzed and processed, and corresponding operation prompt information is generated to provide the user with guidance on the next operation. The operation prompt information is used to instruct the user on the next operation, and the operation prompt information is generated based on the current work content information and matches the user's current work content. For example, if the user is currently editing code, the operation prompt information can be prompt information for completing the code.

[0055] Furthermore, the operation prompt information is directly output (displayed) by the smart device to assist the user in operating the smart device.

[0056] In this way, the work content analysis macromodel is used to analyze the operating interface image of the smart device to determine the corresponding work software and work content information of the smart device. After determining that the work software corresponding to the smart device is pre-set supported software based on the work software information, the target work area image corresponding to the smart device is further obtained. Based on the pre-set multimodal macromodel, operation prompt information matching the work content information is generated for the smart device. This can automatically provide real-time operation prompts during the use of the smart device, providing operation assistance for the smart device, reducing the difficulty of operating the smart device and thus improving the user experience.

[0057] It should be noted that the smart device includes a computer. In the embodiments of the present application, the smart device is a computer (i.e., a computer) as an example for specific description, but in actual use, the smart device may also include other devices such as mobile phones and tablet computers, which is not specifically limited here.

[0058] In the embodiment of the present application, computer operation assistance is provided based on a large model to reduce the difficulty of users in using smart devices and improve processing efficiency and operation effects.

[0059] Specifically, obtaining the operation interface image corresponding to the smart device includes: capturing a screen image corresponding to the smart device as the operation interface image.

[0060] Furthermore, the determining of the working software information and working content information corresponding to the smart device based on the operation interface image and through a preset working content analysis model includes:

[0061] Get content analysis prompt words;

[0062] The operation interface image is input into the work content analysis model, and according to the content analysis prompt word, the operation interface image is subjected to content analysis by the work content analysis model to obtain the work software information and work content information corresponding to the smart device.

[0063] The content analysis prompt is used to optimize the preset work content analysis model. The specific content of the content analysis prompt can be set or adjusted according to actual needs and is not specifically limited here.

[0064] In a specific application scenario, a screenshot of a smart device's screen is captured as an operational interface image and submitted to a large-scale work content analysis model. The large-scale work content model is optimized based on content analysis prompts and can analyze work software for the Windows interface. Specifically, it analyzes the current work software and work content, clarifying the software used by the current user and the current user's work content.

[0065] A content analysis prompt word is as follows:

[0066] "

Character Limited

[0067] You are a professional screen interface analysis expert, specializing in identifying and analyzing various software interface features. Please perform the following systematic processing based on visual information:

[0068]

Processing Flow

[0069] 1. Global Observation

[0070] -Analyze the overall layout of the interface (menu bar structure / toolbar arrangement / color scheme)

[0071] - Identify prominent brand logos (logos / exclusive icons / featured controls)

[0072] 2. Element analysis

[0073] -Extract title bar text (including version number)

[0074] -Mark core functional areas (identify at least 3 characteristic areas)

[0075] - Capture status bar information

[0076] 3. Feature Matching

[0077] -Comparative software database (including 5000+ common software features)

[0078] - Identify unique interface components (such as the PS layer panel / WinRAR compression package icon)

[0079] 4. Logical Reasoning

[0080] - Eliminate interference from the browser web interface

[0081] - Differentiate between products from the same company (such as Adobe series)

[0082] -Handle multi-window overlay

[0083] Output format

[0084]

[0085]

Special instructions

[0086] 1. Web applications are marked as "Browser App: [application name]"

[0087] 2. The game interface needs to be differentiated by engine (Unity / Unreal)

[0088] 3. Development tools need to identify programming language support

[0089] 4. If the confirmation fails, it means that the feature item is not matched.

[0090] ”);

[0091] ”.

[0092] In the embodiment of the present application, the working software information includes software name, software version and software type.

[0093] Furthermore, the working software information is compared with preset supported software information to determine whether the working software is supported. Supported software is pre-configured software that supports operation assistance. If the working software is supported, a target work area image that matches the operation cursor corresponding to the smart device is obtained. Otherwise, if the working software is not supported, a prompt message is output to inform the user that operation assistance is unavailable.

[0094] Specifically, obtaining a target working area image that matches an operation cursor corresponding to the smart device includes:

[0095] Obtaining position information corresponding to an operation cursor of the smart device;

[0096] Determining a target working area corresponding to the smart device according to the location information and preset area size limiting parameters;

[0097] A screenshot is taken of the target working area to obtain an image of the target working area.

[0098] The area size limit parameter is a pre-set parameter used to limit the size of the target working area to be captured. Its specific value can be set and adjusted according to actual needs. For example, it can be set to capture the area 5 lines above and below the cursor position as the target working area, but it is not a specific limit.

[0099] In a specific application scenario, if it is a supported work software, a screenshot is taken based on the user's current cursor position, and the screenshot containing the cursor information and work context (i.e., the target work area image) is submitted to the work details analysis intelligent agent, and the user's work content context is analyzed using a large model.

[0100] Specifically, generating operation prompt information matching the work content information for the smart device based on the target work area image through a preset multimodal large model includes:

[0101] Inputting the target work area image into a preset multimodal macromodel, and determining work content details matching the target work area image through a work detail analysis agent in the multimodal macromodel;

[0102] According to the work content details, operation prompt information matching the work content information is generated for the smart device through the work auxiliary agent in the multimodal large model.

[0103] Among them, the work details analysis agent (work details analysis agent) and the work assistance agent (work assistance agent) are respective agents set in the multimodal large model, which are used to perform different specific data processing operations.

[0104] It should be noted that different work detail analysis agents and work auxiliary agents can be set for different software respectively, and the corresponding work detail analysis agents and work auxiliary agents can be retrieved according to the work software information during use.

[0105] In the embodiment of this application, the user's working software is code editing software, and the work content is code writing. Specifically, taking code completion as an example, a prompt word for implementing a large model is as follows:

[0106] {

[0107] Prompt word template: screenshot analysis instructions

[0108] Please describe the screenshot using the following structured prompts, and I will explain the programming context for you:

[0109] 1. Programming language identification

[0110] Key code features:

[0111] Keywords / symbols of the code in the screenshot (such as import, def, {}, ->, ;, etc.).

[0112] Code structure: function / class definition, indentation style (space vs. tab), end-of-line symbols (colon / semicolon).

[0113] Example description:

[0114] "The code contains from torch.utils.data import Dataset and class CustomDataset(Dataset):, use 4 spaces for indentation."

[0115] 2. Programming tool identification

[0116] Interface features:

[0117] Editor / IDE logo: The logo in the top menu bar (such as the blue square of VS Code and the orange snake icon of PyCharm).

[0118] Terminal style: terminal background color, command line prompt (such as $, >>>), or debugger interface.

[0119] Special function buttons: such as "Run", "Debug", "Git Commit", etc.

[0120] Example description:

[0121] "There is a green 'Run and Debug' button at the top of the editor, and the terminal area has green text on a black background, showing (venv)$python train.py."

[0122] 3. Coding Context Analysis

[0123] Code snippet near the cursor:

[0124] Provide at least 5 lines of code before and after the cursor location (hide sensitive information).

[0125] Code function description: such as data preprocessing, API calls, class method definitions, etc.

[0126] Example description:

[0127] "The cursor is located inside line 12 def forward(self, x):, the context is the PyTorch model class, containing self.layers = nn.Sequential(...)."

[0128] 4. Cursor position details

[0129] Position Description:

[0130] Inline position (e.g., “after the colon in if x>0:”).

[0131] Auto-completion prompt: If there is an IDE code completion pop-up window at the cursor, the keywords are listed (for example, after entering model., fit() and predict() are prompted).

[0132] Example description:

[0133] "The cursor is after lab in plt.plot(x,y,lab, and the completion prompt shows label= and linestyle=."

[0134] 5. Error / Warning Messages

[0135] Error text: The error message highlighted in the terminal or editor (such as SyntaxError, ModuleNotFoundError).

[0136] Example description:

[0137] "The terminal displays a red error: ImportError: cannot import name'Dataset'from'torch.utils.data'."

[0138] Output example

[0139] Based on your description, I would generate an analysis similar to the following:

[0140] Programming language: Python (using the PyTorch framework).

[0141] Programming tools: VS Code (with Python extension and Jupyter support).

[0142] Coding context: Define the forward method of the neural network model, missing import or dependency conflict.

[0143] Problem location: The Dataset class import failed due to incompatible PyTorch versions.

[0144] Please click on this template to provide the textual information in the screenshot to get an accurate analysis!

[0145] }.

[0146] The multimodal big model is used to obtain the programming language, programming tools, current coding context, and cursor position. The user's work content context is submitted to the work assistant agent (various software assistant models developed based on the multimodal big model, such as the Office completion big model and the code completion big model, are routed through the software obtained in the first step. The work agent can support the mixed use of multiple agents). Based on the user's work context, the assistant agent corresponding to the current work software generates the next step prompt.

[0147] Take code completion as an example, based on the context and content obtained, use the following prompt words:

[0148]

[0149]

[0150] It should be noted that the method provided in the embodiment of the present application can be used to complete code completion work, and can also be used to assist other operations, such as completing formulas in Excel worksheets, etc., which is not specifically limited here. Based on operations such as screenshots and cursor acquisition, multi-agent collaboration is used to assist user operations.

[0151] Furthermore, the large model can be integrated into an assistant software, installed on the user's desktop, and implemented as non-invasive software. Instead of using a plug-in model, it uses GUI + multimodal recognition technology to complete screen work content analysis and embed prompts for operating software.

[0152] It should be further explained that the multi-scenario Agent required by other scenarios can be adaptively set according to actual needs to expand the application scenarios of the large model.

[0153] like Figure 2 As shown in , corresponding to the large model-based intelligent device operation assistance method, the embodiment of the present application further provides a large model-based intelligent device operation assistance system, and the large model-based intelligent device operation assistance system includes:

[0154] The first image acquisition module 210 is used to acquire an operation interface image corresponding to the smart device;

[0155] An information determination module 220 is configured to determine the working software information and working content information corresponding to the smart device based on the operation interface image and a preset working content analysis model;

[0156] The second image acquisition module 230 is configured to acquire a target work area image that matches the operation cursor corresponding to the smart device if it is determined based on the work software information that the work software corresponding to the smart device is a preset supported software, wherein the target work area image is used to represent the portion of the work content information corresponding to the position of the operation cursor;

[0157] The operation prompt module 240 is used to generate operation prompt information matching the work content information for the smart device based on the target work area image through a preset multimodal large model.

[0158] Thus, in the solution corresponding to the large-scale model-based intelligent device operation assistance system provided by this application, the operating interface image of the intelligent device is analyzed through the work content analysis large-scale model to determine the work software information and work content information corresponding to the intelligent device. After determining that the work software corresponding to the intelligent device belongs to the preset supported software based on the work software information, the target work area image corresponding to the intelligent device is further obtained, and based on the preset multimodal large-scale model, operation prompt information matching the work content information is generated for the intelligent device. In this way, it is possible to automatically provide real-time operation prompts for the use process of the intelligent device to realize intelligent device operation assistance, which is conducive to reducing the difficulty of operating the intelligent device, thereby helping to improve the user experience.

[0159] It should be noted that the specific structure and implementation of the large model-based intelligent device operation assistance system and its various modules or units can refer to the corresponding description in the method embodiment and will not be repeated here.

[0160] It should be noted that the division method of each module of the large model-based intelligent device operation assistance system is not unique and is not specifically limited here.

[0161] Based on the above embodiment, the present application also provides a smart terminal, whose principle block diagram can be as follows: Figure 3 As shown. The above-mentioned intelligent terminal includes a processor, a memory, a network interface and a display screen connected via a system bus. Among them, the processor of the intelligent terminal is used to provide computing and control capabilities. The memory of the intelligent terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the intelligent terminal is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the steps of any one of the above-mentioned intelligent device operation assistance methods based on a large model are implemented. The display screen of the intelligent terminal can be a liquid crystal display screen or an electronic ink display screen.

[0162] Those skilled in the art will understand that Figure 3 The principle block diagram shown in the figure is only a block diagram of a partial structure related to the solution of the present application, and does not constitute a limitation on the smart terminal to which the solution of the present application is applied. The specific smart terminal may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0163] In one embodiment, a smart terminal is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of any one of the large model-based smart device operation assistance methods provided in the embodiments of the present application are implemented.

[0164] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the large model-based smart device operation assistance methods provided in the embodiment of the present application are implemented.

[0165] It should be understood that the serial numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0166] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the above-mentioned device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned device can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0167] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.

[0168] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0169] In the embodiments provided herein, it should be understood that the disclosed systems / terminal devices and methods can be implemented in other ways. For example, the system / terminal device embodiments described above are merely illustrative. For example, the division of the modules or units described above is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another system, or omitting or not implementing certain features.

[0170] If the above-mentioned integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The above-mentioned computer program can be stored in a computer-readable storage medium, and the computer program, when executed by the processor, can implement the steps of the above-mentioned various method embodiments. Among them, the above-mentioned computer program includes computer program code, and the above-mentioned computer program code can be in source code form, object code form, executable file or some intermediate form. The above-mentioned computer-readable medium may include: any entity or device capable of carrying the above-mentioned computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, electric signal and software distribution medium, etc. It should be noted that the content contained in the above-mentioned computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction.

[0171] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for assisting intelligent device operation based on a large model, characterized in that: The method comprises: Obtain the operation interface image corresponding to the smart device; According to the operation interface image, the working software information and working content information corresponding to the smart device are determined by using a preset working content analysis model; If it is determined according to the working software information that the working software corresponding to the smart device belongs to the preset supported software, then obtaining a target working area image that matches the operation cursor corresponding to the smart device, wherein the target working area image is used to represent the portion of the work content information corresponding to the position of the operation cursor; According to the target work area image, operation prompt information matching the work content information is generated for the smart device through a preset multimodal large model.

2. The method for assisting operation of intelligent devices based on a large model according to claim 1, characterized in that: The obtaining of the operation interface image corresponding to the smart device includes: The screen image corresponding to the smart device is captured as the operation interface image.

3. The method for assisting operation of intelligent devices based on a large model according to claim 1, characterized in that: The determining, based on the operation interface image and through a preset work content analysis model, the work software information and work content information corresponding to the smart device includes: Get content analysis prompt words; The operation interface image is input into the work content analysis model, and according to the content analysis prompt word, the operation interface image is subjected to content analysis by the work content analysis model to obtain the work software information and work content information corresponding to the smart device.

4. The method for assisting intelligent device operation based on a large model according to claim 3, characterized in that: The working software information includes software name, software version and software type.

5. The method for assisting operation of intelligent devices based on a large model according to claim 1, characterized in that: The acquiring of a target working area image that matches an operation cursor corresponding to the smart device includes: Obtaining position information corresponding to an operation cursor of the smart device; Determining a target working area corresponding to the smart device according to the location information and preset area size limiting parameters; A screenshot is taken of the target working area to obtain an image of the target working area.

6. The method for assisting intelligent device operation based on a large model according to claim 1, characterized in that: The step of generating, based on the target work area image and using a preset multimodal large model, operation prompt information matching the work content information for the smart device includes: Inputting the target work area image into a preset multimodal macromodel, and determining work content details matching the target work area image through a work detail analysis agent in the multimodal macromodel; According to the work content details, operation prompt information matching the work content information is generated for the smart device through the work auxiliary agent in the multimodal large model.

7. The method for assisting operation of an intelligent device based on a large model according to any one of claims 1 to 6, characterized in that: The smart device includes a computer.

8. A large model-based intelligent device operation assistance system, characterized in that: The system comprises: A first image acquisition module is used to acquire an operation interface image corresponding to the smart device; An information determination module is used to determine the working software information and working content information corresponding to the smart device based on the operation interface image and a preset working content analysis model; a second image acquisition module, configured to acquire a target work area image that matches an operation cursor corresponding to the smart device if it is determined based on the work software information that the work software corresponding to the smart device is preset supported software, wherein the target work area image is used to represent a portion of the work content information corresponding to a position of the operation cursor; An operation prompt module is used to generate operation prompt information matching the work content information for the smart device based on the target work area image through a preset multimodal large model.

9. An intelligent terminal, characterized in that: The smart terminal includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the steps of the large model-based smart device operation assistance method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the large model-based intelligent device operation assistance method according to any one of claims 1 to 7.