Human-computer interaction processing method, device, equipment and storage medium
By introducing image processing and multimodal large model technology into human-computer interaction, the problem of insufficient multimodal information integration and reasoning decision-making capabilities in the existing technology is solved, and efficient human-computer interaction in complex industrial scenarios is achieved.
Patent Information
- Application Number
- CN202411427795.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-10-14
AI Technical Summary
The human-computer interaction schemes in the prior art lack the ability to integrate multimodal information and reason and decision-making, and their application in complex industrial scenarios is immature.
By obtaining the images of the user interface and user's operation instructions, the images are input into the pre-trained image processing model for content extraction, structured information is generated, and the images, structured information and operation instructions are input into the multimodal large model, the task statements corresponding to the operation instructions are generated, and task statements are executed to realize the integration of multimodal information and inference decisions.
It realizes accurate integration and reasoning decisions for multimodal information, improves the automation level and adaptability of human-computer interaction, and is suitable for complex industrial scenarios.
Smart Images

Figure CN118963623B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of industrial interaction technology, and in particular to a human-computer interaction processing method, device, equipment and storage medium. Background Art
[0002] Human-computer interaction refers to the information exchange process between people and computers using a certain conversational language and a certain interactive method to complete a certain task. In the field of industrial software, traditional human-computer interaction methods are limited to a single linear communication channel, such as text input or mouse clicks, which not only reduces work efficiency, but also easily leads to misunderstandings when handling complex business processes. In recent years, with the rapid development of deep learning and natural language processing technologies, deep learning and natural language processing technologies can be introduced into human-computer interaction in existing technologies, thereby improving the automation level and adaptability of industrial software.
[0003] However, most existing solutions focus on single-modal interactions. For example, some intelligent Robotic Process Automation (RPA) solutions lack the ability to integrate multi-modal information and make inference decisions, and their applications in complex industrial scenarios are still immature. Summary of the invention
[0004] The purpose of this application is to provide a human-computer interaction processing method, device, equipment and storage medium to address the deficiencies in the above-mentioned prior art, so as to solve the problems that the human-computer interaction solutions in the prior art lack integration of multimodal information and reasoning decisions, and are immature in application in complex industrial scenarios.
[0005] To achieve the above purpose, the technical solution adopted in the embodiment of the present application is as follows:
[0006] In a first aspect, an embodiment of the present application provides a human-computer interaction processing method, the method comprising:
[0007] Acquire an image of the user interface at a preset frequency;
[0008] Obtain user operation instructions;
[0009] Inputting the image into a pre-trained image processing model, extracting content from the image, and generating structured information corresponding to the image, wherein the structured information is used to characterize the type and position of each element in the image;
[0010] Inputting the image, the structured information and the operation instruction into a pre-trained multimodal large model to generate a task statement corresponding to the operation instruction, wherein the task statement at least includes an area to be operated and an operation to be performed corresponding to the operation instruction;
[0011] The task statement is executed on the user interface, and an execution result is output to the user.
[0012] In a possible implementation, before inputting the image, the structured information, and the operation instruction into a pre-trained multimodal large model to generate a task statement corresponding to the operation instruction, the method further includes:
[0013] The operation instructions are standardized according to the pre-trained instruction processing model to obtain standardized operation instructions.
[0014] In a possible implementation, the image processing model includes: a visual positioning model and a text recognition model; the step of inputting the image into the pre-trained image processing model, extracting content from the image, and generating structured information corresponding to the image includes:
[0015] Inputting the image into the visual positioning model to generate the type and position of each element in the image;
[0016] Inputting the image into the text recognition model to generate labels and text descriptions of each element in the image;
[0017] According to the type, position, label and text description of each element, structured information corresponding to the image is generated.
[0018] In a possible implementation, the multimodal large model includes: a preprocessing module, a splicing module and a processing module; the inputting the image, the structured information and the operation instruction into the pre-trained multimodal large model to generate a task statement corresponding to the operation instruction includes:
[0019] Inputting the image into the preprocessing module for preprocessing to obtain a preprocessed image sequence;
[0020] Inputting the image sequence, the structural information and the operation instruction into the stitching module to generate a sequence to be processed;
[0021] The sequence to be processed is input into the processing module to generate a task statement corresponding to the operation instruction.
[0022] In a possible implementation, the preprocessing module includes: a segmentation module and a linear mapping layer; the inputting the image into the preprocessing module for preprocessing to obtain a preprocessed image sequence includes:
[0023] Inputting the image into the segmentation module, performing image segmentation processing on the image, and generating a plurality of sub-image blocks corresponding to the image;
[0024] Each of the sub-image blocks is input into the linear mapping layer for projection to generate the preprocessed image sequence.
[0025] In a possible implementation, executing the task statement on the user interface and outputting the execution result to the user includes:
[0026] Determine the operation area to be operated and the operation to be performed corresponding to the task statement in the user interface;
[0027] The operation to be performed is performed in the operation area, and after the execution is completed, the execution result is output to the user.
[0028] In a possible implementation, determining the area to be operated and the operation to be performed corresponding to the task statement in the user interface includes:
[0029] Extracting the task statement to obtain a region to be operated field and a field to be performed operation in the task statement;
[0030] Determine the area to be operated in the user interface according to the area to be operated field;
[0031] The operation to be executed in the user interface is determined according to the operation to be executed field and a pre-stored action mapping dictionary.
[0032] In a second aspect, another embodiment of the present application provides a human-computer interaction processing device, the device comprising:
[0033] An image acquisition module, used to acquire images of the user interface at a preset frequency;
[0034] An instruction acquisition module is used to acquire the user's operation instructions;
[0035] An image processing module, used to input the image into a pre-trained image processing model, extract content from the image, and generate structured information corresponding to the image, wherein the structured information is used to characterize the type and position of each element in the image;
[0036] A generation module, used for inputting the image, the structured information and the operation instruction into a pre-trained multimodal large model, and generating a task statement corresponding to the operation instruction, wherein the task statement at least includes an area to be operated and an operation to be performed corresponding to the operation instruction;
[0037] An execution module is used to execute the task statement on the user interface and output the execution result to the user.
[0038] In a possible implementation, before the generating module, a standardization module is further included, wherein the standardization module is used to:
[0039] The operation instructions are standardized according to the pre-trained instruction processing model to obtain standardized operation instructions.
[0040] In a possible implementation, the image processing model includes: a visual positioning model and a text recognition model; the image processing module is specifically used to:
[0041] Inputting the image into the visual positioning model to generate the type and position of each element in the image;
[0042] Inputting the image into the text recognition model to generate labels and text descriptions of each element in the image;
[0043] According to the type, position, label and text description of each element, structured information corresponding to the image is generated.
[0044] In a possible implementation, the multimodal large model includes: a preprocessing module, a splicing module and a processing module; the generation module is specifically used to:
[0045] Inputting the image into the preprocessing module for preprocessing to obtain a preprocessed image sequence;
[0046] Inputting the image sequence, the structural information and the operation instruction into the stitching module to generate a sequence to be processed;
[0047] The sequence to be processed is input into the processing module to generate a task statement corresponding to the operation instruction.
[0048] In a possible implementation, the preprocessing module includes: a segmentation module and a linear mapping layer; the generation module is specifically used to:
[0049] Inputting the image into the segmentation module, performing image segmentation processing on the image, and generating a plurality of sub-image blocks corresponding to the image;
[0050] Each of the sub-image blocks is input into the linear mapping layer for projection to generate the preprocessed image sequence.
[0051] In a possible implementation, the execution module is specifically configured to:
[0052] Determine the operation area to be operated and the operation to be performed corresponding to the task statement in the user interface;
[0053] The operation to be performed is performed in the operation area, and after the execution is completed, the execution result is output to the user.
[0054] In a possible implementation, the execution module is specifically configured to:
[0055] Extracting the task statement to obtain a region to be operated field and a field to be performed operation in the task statement;
[0056] Determine the area to be operated in the user interface according to the area to be operated field;
[0057] The operation to be executed in the user interface is determined according to the operation to be executed field and a pre-stored action mapping dictionary.
[0058] In the third aspect, another embodiment of the present application provides an electronic device, comprising: a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the storage medium through the bus, and the processor executes the machine-readable instructions to perform the steps of any method described in the first aspect above.
[0059] In a fourth aspect, another embodiment of the present application provides a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any method described in the first aspect are executed.
[0060] The beneficial effects of the present application are as follows: by acquiring images of the user interface at a preset frequency and acquiring the user's operation instructions, the image can be input into a pre-trained image processing model, content can be extracted from the image, and structured information corresponding to the image can be generated, so that the image, structured information and operation instructions can be input into a pre-trained multimodal large model, task statements corresponding to the operation instructions can be generated, and the task statements can be executed for the user interface, and the execution results can be output to the user, so that the multimodal information can be accurately integrated and reasoned and decided, and human-computer interaction under multimodality can be realized. At the same time, the human-computer interaction processing method provided in the embodiment of the present application can be executed multiple times to perform human-computer interaction processing on the user's complex operation requirements, so that the human-computer interaction processing method provided in the embodiment of the present application can also be applicable to complex industrial scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0062] Figure 1 A schematic diagram of a user interface of the human-computer interaction processing method provided in an embodiment of the present application;
[0063] Figure 2 A schematic diagram of a flow chart of a human-computer interaction processing method provided in an embodiment of the present application;
[0064] Figure 3 A structural schematic diagram of an image processing model in the human-computer interaction processing method provided in an embodiment of the present application;
[0065] Figure 4 A schematic diagram of a process for generating structured information corresponding to an image in the human-computer interaction processing method provided in an embodiment of the present application;
[0066] Figure 5 A structural schematic diagram of a multimodal large model in the human-computer interaction processing method provided in an embodiment of the present application;
[0067] Figure 6 A schematic diagram of a flow chart for generating a task statement corresponding to an operation instruction in the human-computer interaction processing method provided in an embodiment of the present application;
[0068] Figure 7 A structural schematic diagram of a preprocessing module in the human-computer interaction processing method provided in an embodiment of the present application;
[0069] Figure 8 A schematic diagram of a flow chart for obtaining a pre-processed image sequence in the human-computer interaction processing method provided in an embodiment of the present application;
[0070] Fig. 9 A schematic diagram of a flow chart for outputting execution results to a user in the human-computer interaction processing method provided in an embodiment of the present application;
[0071] Fig.10 Another flowchart diagram of outputting execution results to a user in the human-computer interaction processing method provided in an embodiment of the present application;
[0072] Fig.11 A schematic diagram of a human-computer interaction processing device provided in an embodiment of the present application;
[0073] Fig.12 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0074] To make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. It should be understood that the drawings in the present application only serve the purpose of explanation and description and are not used to limit the scope of protection of the present application. In addition, it should be understood that the schematic drawings are not drawn in real proportion. The flowchart used in this application shows the operations implemented according to some embodiments of the present application. It should be understood that the operations of the flowchart can be implemented out of sequence, and the steps without logical context can be reversed in order or implemented simultaneously. In addition, those skilled in the art can add one or more other operations to the flowchart under the guidance of the content of the present application, or remove one or more operations from the flowchart.
[0075] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.
[0076] It should be noted that the term "comprising" will be used in the embodiments of the present application to indicate the existence of the features declared thereafter, but does not exclude the addition of other features.
[0077] In the field of industrial software, traditional human-computer interaction is limited to a single linear communication channel, such as text input or mouse clicks, which not only reduces work efficiency but also easily leads to misunderstandings when handling complex business processes. In recent years, with the rapid development of deep learning and natural language processing technologies, deep learning and natural language processing technologies can be introduced into human-computer interaction in existing technologies, thereby improving the automation level and adaptability of industrial software.
[0078] However, most of the human-computer interaction solutions in existing technologies focus on single-modal interaction. For example, some intelligent robotic process automation (Robotic Process Automation) lacks the ability to integrate multi-modal information and make inference decisions, and their application in complex industrial scenarios is still immature.
[0079] Based on the above problems, an embodiment of the present application proposes a human-computer interaction processing method, which obtains images of the user interface at a preset frequency and obtains the user's operation instructions, so that the image can be input into a pre-trained image processing model, and the content of the image can be extracted to generate structured information corresponding to the image, so that the image, structured information and operation instructions can be input into a pre-trained multimodal large model to generate task statements corresponding to the operation instructions, and the task statements are executed for the user interface, and the execution results are output to the user. The multimodal information can be accurately integrated and reasoned and decided to achieve human-computer interaction under multimodality. At the same time, it can also be suitable for complex industrial scenarios.
[0080] First, the application scenarios involved in the human-computer interaction processing method provided in the embodiment of the present application are described.
[0081] It can be understood that when a user uses a certain software on a client, the human-computer interaction processing method provided by the embodiment of the present application can be pre-deployed on the client to realize intelligent human-computer interaction during the user's use of the software, thereby improving the user's work efficiency and productivity, while reducing the user's operation risks and realizing precise control.
[0082] As an example, in a specific implementation process, the human-computer interaction processing method provided in the embodiment of the present application can be deployed to the client in the form of a plug-in and executed together with the execution of the client.
[0083] For example, taking industrial software as an example, Figure 1 A user interface diagram of the human-computer interaction processing method provided in the embodiment of the present application, referring to Figure 1 As shown, the human-computer interaction processing method provided by the embodiment of the present application can be deployed in the client. When the user opens a certain industrial software, the user can enter the corresponding operation requirements in the user input box, so that the client can execute the steps of the human-computer interaction processing method provided by the embodiment of the present application after obtaining the operation requirements input by the user, and output the execution results to the user.
[0084] It is worth noting that when a user opens a certain industrial software, the user can input the operation requirements through the user input box. After obtaining the operation requirements input by the user, the human-computer interaction processing method provided in the embodiment of the present application can first process the operation requirements input by the user to generate one or more operation instructions corresponding to the operation requirements. Among them, the one or more operation instructions corresponding to the operation requirements are used to indicate one or more user interfaces to be operated and the tasks to be executed corresponding to each user interface to be operated during the human-computer interaction process.
[0085] That is to say, when the user's operation requirement is to operate in a user interface, the electronic device can obtain the operation requirement input by the user, and use the operation requirement input by the user as an operation instruction to execute the steps of the human-computer interaction processing method provided in the embodiment of the present application.
[0086] When the user's operation requirements are to perform operations on multiple user interfaces separately, the electronic device can execute the steps of the human-computer interaction processing method provided in the embodiment of the present application multiple times, that is, execute multiple operation instructions corresponding to the operation requirements in sequence, and execute the next operation instruction based on the completion of the previous operation instruction until the operation requirements are completed.
[0087] The human-computer interaction processing method provided in the embodiment of the present application is described in detail below in combination with multiple embodiments.
[0088] Figure 2 A flow chart of a human-computer interaction processing method provided in an embodiment of the present application, referring to Figure 2 As shown, the execution subject of the method can be any electronic device, such as the above-mentioned client, and the method includes:
[0089] S201: Acquire an image of a user interface at a preset frequency.
[0090] Optionally, the electronic device can take screenshots of the user interface at a preset frequency to obtain images of the user interface. When taking screenshots of the user interface, the electronic device can only take screenshots of the user interface of the software that requires human-computer interaction, or can take screenshots of the entire user interface of the electronic device.
[0091] For example, continue to refer to Figure 1 As shown, a certain interface of the industrial software may display the monitor screen of the device X monitor, multiple buttons 1-9 that can be used to control the device X and have different functions, and multiple auxiliary buttons. At this time, the image of the user interface obtained may be multiple images showing the monitor screen of the device X monitor, multiple buttons 1-9, and multiple auxiliary buttons.
[0092] S202: Obtaining the user's operation instruction.
[0093] Optionally, an operation instruction of the user is obtained, wherein the operation instruction includes an operation instruction corresponding to the image obtained in step S201.
[0094] Optionally, when the human-computer interaction processing method provided in the embodiment of the present application is executed for the first time, obtaining the user's operation instruction can be understood as obtaining the operation requirements input by the user in the user interface. The operation requirements include the operation instructions corresponding to the image obtained in step S201. The user can input the operation requirements in the form of voice or text.
[0095] Optionally, when the human-computer interaction processing method provided in the embodiment of the present application is not executed for the first time, obtaining the user's operation instructions can be understood as obtaining the task statement generated during the previous execution, wherein the task statement includes the operation instructions corresponding to the image obtained in the current execution step S201.
[0096] S203: Input the image into a pre-trained image processing model, extract content from the image, and generate structured information corresponding to the image.
[0097] It can be understood that since the image is obtained by taking a screenshot of the user interface, the image includes the interface diagram of the software that needs to be operated in the user interface. Therefore, the interface information of the software contained in the image can be extracted first to generate structured information corresponding to the image, so that the human-computer interaction processing method provided in the embodiment of the present application can better perform human-computer interaction processing in complex scenarios.
[0098] Optionally, after the image is obtained in step S201, the image may be processed by a pre-trained image processing model to extract the content in the image and generate structured information corresponding to the image.
[0099] Exemplarily, the image processing model may be a model based on computer vision and deep learning technology, such as a convolutional neural network (CNN), an object detection model (YOLO), and a semantic segmentation model (U-Net).
[0100] The structured information is used to characterize the type and position of each element in the image. Specifically, the structured information may include the position, type, and text content related to each user interface (UI) element in the image.
[0101] S204: Input the image, structured information and operation instructions into a pre-trained multimodal large model to generate a task statement corresponding to the operation instruction.
[0102] Optionally, after obtaining the structured information, the image, structured information and operation instructions can be input into a pre-trained multimodal large model, so that the multimodal large model can process the image, structured information and operation instructions, thereby generating a task statement corresponding to the operation instruction. The multimodal large model can be a multimodal large model with Qwen2 as the base.
[0103] That is to say, for the multimodal large model, images, structured information and operation instructions can be used as training data for training the multimodal large model, combined with template statements (Prompt) to input the multimodal large model, and the multimodal large model can be trained to obtain a trained multimodal large model. Among them, the multimodal large model can be only a multimodal large model with Qwen2 as the base, or it can be a multimodal large model with a preprocessing module, a splicing module and a processing module with Qwen2 as the base.
[0104] The task statement at least includes the area to be operated and the operation to be performed corresponding to the operation instruction, and the task statement may also include a task end mark. Exemplarily, the task statement may be a JSON sequence.
[0105] Exemplarily, when the human-computer interaction processing method provided in the embodiment of the present application is executed for the first time, the generated task statement can be a task statement that can be executed in the current user interface, that is, a task statement that is executed only in the current user interface, or it can be a task statement that cannot be executed in the current user interface, that is, a task statement that needs to be executed in sequence among multiple user interfaces.
[0106] S205: Execute the task statement on the user interface and output the execution result to the user.
[0107] Optionally, after obtaining the task statement, the task statement may be parsed and executed on the user interface, so that the execution result can be output to the user.
[0108] For example, continue to refer to Figure 1 As shown, assuming that the content of the user's input is "close the page of device X", the operation instruction obtained in step S202 is "close the page of device X", then in step S204, multiple monitor images showing the monitor of device X, multiple buttons 1-9 and multiple auxiliary buttons and corresponding structured information and the operation instruction with the content of "close the page of device X" can be input into the pre-trained multimodal large model to generate a task statement corresponding to the operation instruction, which includes Figure 1The interface shown includes an operation area to be operated, an operation to be performed, and a task end mark, so that after generating a task statement, the electronic device executes the operation of closing the page of device X and outputs the closed page to the user.
[0109] For example, continue to refer to Figure 1 As shown in FIG. 1 , assuming that the user input is “close the page of device X and open the page of device Y”, then the generated task statement indicates that Figure 1 The interface shown includes an operation area to be operated, an operation to be performed, "open the page of device Y" and a task end identifier, so that after the task statement is generated, the electronic device executes the operation of closing the page of device X, and outputs the closed page to the user, and then re-executes S201-S202, that is, obtains the image of the closed page, and uses "open the page of device Y" and the task end identifier as the user's operation instructions, and re-executes S203-S205 until the task statement is executed to the task end identifier.
[0110] In this embodiment, by acquiring an image of the user interface at a preset frequency and obtaining the user's operation instructions, the image can be input into a pre-trained image processing model, content can be extracted from the image, and structured information corresponding to the image can be generated, so that the image, structured information and operation instructions can be input into a pre-trained multimodal large model, task statements corresponding to the operation instructions can be generated, and the task statements can be executed for the user interface, and the execution results can be output to the user. The multimodal information can be accurately integrated and reasoned and decided to achieve human-computer interaction under multimodality. At the same time, the human-computer interaction processing method provided in the embodiment of the present application can be executed multiple times to perform human-computer interaction processing on the user's complex operation requirements, so that the human-computer interaction processing method provided in the embodiment of the present application can also be applicable to complex industrial scenarios.
[0111] In a possible implementation, before the step S204 inputs the image, structured information and operation instruction into the pre-trained multimodal large model and generates the task statement corresponding to the operation instruction, the step further includes:
[0112] The operation instructions are standardized according to the pre-trained instruction processing model to obtain standardized operation instructions.
[0113] It can be understood that in order to improve the accuracy of human-computer interaction, the operation instructions can also be standardized before being input into the multimodal large model.
[0114] Optionally, the electronic device may input the operation instruction into a pre-trained instruction processing model, perform standardization processing on the operation instruction, and thus generate a standardized operation instruction. Specifically, the standardization processing may include performing text processing on the operation instruction input by the user in text to generate a standard text, and may also include converting the operation instruction input by the user in voice into text to generate a standard text. Among them, the pre-trained instruction processing model may include: an automatic speech recognition (Automatic Speech Recognition, referred to as ASR) model and a natural language processing model.
[0115] Figure 3 A structural schematic diagram of an image processing model in the human-computer interaction processing method provided in an embodiment of the present application, Figure 4 A schematic diagram of a flow chart for generating structured information corresponding to an image in the human-computer interaction processing method provided in an embodiment of the present application.
[0116] In one possible implementation, referring to Figure 3 as well as Figure 4 As shown, the image processing model includes: a visual positioning model and a text recognition model. In the above step S203, the image is input into the pre-trained image processing model, content is extracted from the image, and structured information corresponding to the image is generated, including:
[0117] S401: Input the image into a visual positioning model to generate the type and position of each element in the image.
[0118] Optionally, the image is input into a pre-trained visual localization model to identify the type and position of each element in the image, and generate the type and position of each element in the image. The pre-trained visual localization model may be an object detection model (YOLO). The type of the element may be, for example, a button, a text box, etc. The position may be the starting coordinates of the element.
[0119] S402: Input the image into a text recognition model to generate labels and text descriptions for each element in the image.
[0120] Optionally, the image is input into a pre-trained text recognition model to recognize the label and text description of each element in the image, and generate the label and text description of each element in the image. The label may be the identifier of the element, and the text description may be the text content related to the element in the image. The pre-trained text recognition model may be an optical character recognition (OCR) model.
[0121] S403: Generate structured information corresponding to the image according to the type, position, label and text description of each element.
[0122] Optionally, the image processing model may also include a structuring module. After obtaining the type, position, label and text description of each element, the structuring module of the image processing model may perform structured processing on the type, position, label and text description of each element to generate structured information corresponding to the image.
[0123] Exemplarily, the data structure of the structured information of each element may be {label, type, text description, position}.
[0124] By inputting the image into the visual positioning model, the type and position of each element in the image are generated, and by inputting the image into the visual positioning model, the type and position of each element in the image are generated, so that the structured information corresponding to the image can be generated according to the type, position, label and text description of each element, which can improve the generation efficiency of the structured information corresponding to the image, and at the same time improve the accuracy of the obtained structured information, so that multimodal information can be accurately integrated and reasoned and decided, and human-computer interaction under multimodality can be realized.
[0125] Figure 5 A schematic diagram of a structure of a multimodal large model in the human-computer interaction processing method provided in an embodiment of the present application, Figure 6 A flowchart of generating a task statement corresponding to an operation instruction in the human-computer interaction processing method provided in an embodiment of the present application.
[0126] In one possible implementation, referring to Figure 5 as well as Figure 6 As shown, the multimodal large model includes: a preprocessing module, a splicing module and a processing module; in the above step S204, the image, structured information and operation instructions are input into the pre-trained multimodal large model to generate a task statement corresponding to the operation instruction, including:
[0127] S601: Input an image into a preprocessing module for preprocessing to obtain a preprocessed image sequence.
[0128] Optionally, the image is input into a preprocessing module for preprocessing, thereby improving the image quality, so as to improve the accuracy of the obtained task statement. The preprocessing may include image cleaning, image segmentation, image dimension conversion, etc.
[0129] S602: Input the image sequence, structured information and operation instructions into a stitching module to generate a sequence to be processed.
[0130] Optionally, the preprocessed image sequence, structured information and operation instructions are input into the splicing model so that the splicing module embeds the preprocessed image sequence, structured information and operation instructions to obtain a sequence to be processed. The splicing module may include multiple embedding layers to better splice the preprocessed image sequence, structured information and operation instructions.
[0131] Exemplarily, the dimensions of structured information and operation instructions can be processed by a splicing module to generate target structured information and target operation instructions with the same dimensions as those embedded in the processing module, and the image sequence, target structured information and target operation instructions are spliced in sequence to generate a sequence to be processed.
[0132] S603: Input the sequence to be processed into the processing module to generate a task statement corresponding to the operation instruction.
[0133] Optionally, the sequence to be processed is input into a processing module to generate a task statement corresponding to the operation instruction. The processing module may be a multi-modal large model with Qwen2 as the base.
[0134] By inputting the image into the preprocessing module for preprocessing to obtain the preprocessed image sequence, the quality of the image sequence can be improved, and the image sequence, structured information and operation instructions are input into the splicing module to generate a sequence to be processed, which can improve the readability of the sequence to be processed. After the sequence to be processed is input into the processing module, the task statement corresponding to the operation instruction can be accurately generated, and at the same time, the generation efficiency of the task statement can be improved.
[0135] Figure 7 A structural schematic diagram of a preprocessing module in the human-computer interaction processing method provided in an embodiment of the present application, Figure 8 A schematic diagram of a flow chart for obtaining a pre-processed image sequence in the human-computer interaction processing method provided in an embodiment of the present application.
[0136] In one possible implementation, referring to Figure 7 As shown, in Figure 5 Based on this, the preprocessing module includes: a segmentation module and a linear mapping layer; refer to Figure 8 As shown, the above step S601 inputs the image into the preprocessing module for preprocessing to obtain a preprocessed image sequence, including:
[0137] S801: Input an image into a segmentation module, perform image segmentation processing on the image, and generate multiple sub-image blocks corresponding to the image.
[0138] Optionally, the image may be flattened and input into a segmentation module to perform image segmentation processing (Patchify) on the image to generate a plurality of sub-image blocks corresponding to the image.
[0139] S802: Input each sub-image block into a linear mapping layer for projection to generate a preprocessed image sequence.
[0140] Optionally, each sub-image block is input into a linear mapping layer, and each sub-image block is projected into a dimension embedded in the processing module, so that the dimension of the image sequence is consistent with the dimension of the structured information and the operation instruction, so that the task statement corresponding to the operation instruction can be accurately generated, and at the same time, the generation efficiency of the task statement can be improved.
[0141] Fig. 9 A flowchart of outputting execution results to a user in the human-computer interaction processing method provided in an embodiment of the present application.
[0142] In one possible implementation, referring to Fig. 9 As shown, in the above step S205, the task statement is executed on the user interface, and the execution result is output to the user, including:
[0143] S901: Determine the operation area to be performed and the operation to be performed corresponding to the task statement in the user interface.
[0144] Optionally, after the task statement is obtained, the task statement may be read, processed and parsed to determine the area to be operated and the operation to be performed corresponding to the task statement in the user interface.
[0145] S902: Execute the operation to be executed in the operation area, and output the execution result to the user after the execution is completed.
[0146] Optionally, after determining the area to be operated and the operation to be performed corresponding to the task statement in the user interface, the corresponding software development kit (SDK) or library can be called to locate the area to be operated and execute the operation to be performed, and after the execution is completed, the execution result is output to the user.
[0147] By determining the waiting area and the waiting operation corresponding to the task statement in the user interface, executing the waiting operation in the waiting area, and outputting the execution result to the user after the execution is completed, the user can clearly see the specific location of the operation on the interface, which enhances the intuitiveness of the operation, reduces the user's cognitive burden, improves the user experience, and at the same time, improves the accuracy of the operation.
[0148] Fig.10 Another flowchart diagram of outputting execution results to a user in the human-computer interaction processing method provided in an embodiment of the present application.
[0149] In one possible implementation, referring to Fig.10As shown, the above step S901 determines the operation area and the operation to be performed corresponding to the task statement in the user interface, including:
[0150] S1001. Extract the task statement to obtain the area to be operated field and the operation to be performed field in the task statement.
[0151] Optionally, the task statement may be extracted by a regular matching algorithm to obtain a region to be operated field and an operation to be performed field in the task statement.
[0152] For example, the task statement can be {"text description": "I will use button number 6 to perform system diagnosis [49, 816, 572, 1187] to complete this task.","Position": [49, 816, 572, 1187],"Type": "click"}, and the task statement can be extracted by a regular matching algorithm to obtain the field of the area to be operated "Position": [49, 816, 572, 1187] and the field of the operation to be performed "Type": "click".
[0153] S1002: Determine the area to be operated in the user interface according to the area to be operated field.
[0154] Optionally, the center point of the area to be operated in the user interface may be calculated according to the area to be operated field, thereby determining the area to be operated in the user interface.
[0155] S1003: Determine the operation to be executed in the user interface according to the operation to be executed field and a pre-stored action mapping dictionary.
[0156] Optionally, the operation to be performed field may be matched with a pre-stored action mapping dictionary to determine the operation to be performed in the user interface.
[0157] Exemplarily, after obtaining that the operation to be performed field is "type": "click", the "click" field may be matched with a pre-stored action mapping dictionary, thereby determining that the operation to be performed in the user interface is a click.
[0158] By extracting the task statement, the field of the area to be operated and the field of the operation to be performed in the task statement are obtained, and the area to be operated in the user interface is determined according to the field of the area to be operated, and the operation to be performed in the user interface is determined according to the field of the operation to be performed and a pre-stored action mapping dictionary, thereby improving the accuracy of determining the area to be operated and the operation to be performed corresponding to the task statement in the user interface, and at the same time, it can also support complex task scenarios.
[0159] Based on the same inventive concept, a human-computer interaction processing device corresponding to the human-computer interaction processing method is also provided in the embodiment of the present application. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the above-mentioned human-computer interaction processing method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.
[0160] Fig.11 A schematic diagram of a human-computer interaction processing device provided in an embodiment of the present application, referring to Fig.11 As shown, the device includes: an image acquisition module 1101, an instruction acquisition module 1102, an image processing module 1103, a generation module 1104 and an execution module 1105;
[0161] An image acquisition module 1101 is used to acquire an image of a user interface at a preset frequency;
[0162] The instruction acquisition module 1102 is used to acquire the user's operation instruction;
[0163] The image processing module 1103 is used to input the image into a pre-trained image processing model, extract the content of the image, and generate structured information corresponding to the image. The structured information is used to characterize the type and position of each element in the image.
[0164] A generation module 1104 is used to input the image, structured information and operation instructions into a pre-trained multimodal large model to generate a task statement corresponding to the operation instruction, where the task statement at least includes an area to be operated and an operation to be performed corresponding to the operation instruction;
[0165] The execution module 1105 is used to execute the task statement on the user interface and output the execution result to the user.
[0166] In a possible implementation, before the generating module 1104, a standardization module is further included, and the standardization module is used to:
[0167] The operation instructions are standardized according to the pre-trained instruction processing model to obtain standardized operation instructions.
[0168] In a possible implementation, the image processing model includes: a visual positioning model and a text recognition model; the image processing module 1103 is specifically used to:
[0169] Input the image into the visual localization model to generate the type and position of each element in the image;
[0170] Input the image into the text recognition model to generate labels and text descriptions of each element in the image;
[0171] Generate structural information corresponding to the image based on the type, position, label and text description of each element.
[0172] In a possible implementation, the multimodal large model includes: a preprocessing module, a splicing module, and a processing module; a generating module 1104, specifically used for:
[0173] Input the image into the preprocessing module for preprocessing to obtain a preprocessed image sequence;
[0174] Input the image sequence, structural information and operation instructions into the splicing module to generate a sequence to be processed;
[0175] The sequence to be processed is input into the processing module to generate task statements corresponding to the operation instructions.
[0176] In a possible implementation, the preprocessing module includes: a segmentation module and a linear mapping layer; a generation module 1104, specifically configured to:
[0177] Input the image into the segmentation module, perform image segmentation processing on the image, and generate multiple sub-image blocks corresponding to the image;
[0178] Each sub-image block is input into the linear mapping layer for projection to generate a preprocessed image sequence.
[0179] In a possible implementation, the execution module 1105 is specifically configured to:
[0180] Determine the operation area to be performed and the operation to be performed corresponding to the task statement in the user interface;
[0181] Execute the operation to be executed in the operation area, and output the execution result to the user after the execution is completed.
[0182] In a possible implementation, the execution module 1105 is specifically configured to:
[0183] Extract the task statement to obtain the area to be operated field and the operation to be performed field in the task statement;
[0184] According to the field of the area to be operated, determine the area to be operated in the user interface;
[0185] The operation to be performed in the user interface is determined according to the operation to be performed field and a pre-stored action mapping dictionary.
[0186] For descriptions of the processing flow of each module in the device and the interaction flow between each module, reference may be made to the relevant descriptions in the above method embodiment, which will not be described in detail here.
[0187] The present application also provides an electronic device, such as Fig.12 As shown, Fig.12 The structural diagram of the electronic device provided in the embodiment of the present application includes: a processor 1201, a memory 1202, and optionally, a bus 1203. The memory 1202 stores machine-readable instructions executable by the processor 1201 (for example, Fig.11 In the device, the image acquisition module 1101, the instruction acquisition module 1102, the image processing module 1103, the generation module 1104 and the execution instructions corresponding to the execution module 1105, etc.), when the electronic device is running, the processor 1201 communicates with the memory 1202 through the bus 1203, and when the machine-readable instructions are executed by the processor 1201, the steps of the above-mentioned human-computer interaction processing method are executed.
[0188] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned human-computer interaction processing method are executed.
[0189] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, the specific working process of the system and device described above can refer to the corresponding process in the method embodiment, and will not be repeated in this application. In the several embodiments provided in this application, it should be understood that the disclosed system, device and method can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or modules, which can be electrical, mechanical or other forms.
[0190] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or part of the technical solution that contributes to the prior art or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), disk or optical disk and other media that can store program code.
[0191] The above are only specific implementation methods of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be covered by the protection scope of the present application.
Claims
1. A human-computer interaction processing method, characterized in that: include: Acquire an image of the user interface at a preset frequency; Obtain user operation instructions; Input the image into a pre-trained image processing model, extract content from the image, and generate structured information corresponding to the image, wherein the structured information is used to characterize the type and position of each element in the image, and the data structure of the structured information of each element is {label, type, text description, position}; Inputting the image, the structured information and the operation instruction into a pre-trained multimodal large model to generate a task statement corresponding to the operation instruction, wherein the task statement at least includes an area to be operated and an operation to be performed corresponding to the operation instruction; Executing the task statement on the user interface and outputting the execution result to the user; The multimodal large model includes: a preprocessing module, a splicing module and a processing module; The step of inputting the image, the structured information, and the operation instruction into a pre-trained multimodal large model to generate a task statement corresponding to the operation instruction includes: Inputting the image into the preprocessing module for preprocessing to obtain a preprocessed image sequence; Inputting the image sequence, the structural information and the operation instruction into the stitching module to generate a sequence to be processed; Inputting the sequence to be processed into the processing module to generate a task statement corresponding to the operation instruction; The step of inputting the image sequence, the structural information and the operation instruction into the stitching module to generate a sequence to be processed includes: Processing the dimensions of the structured information and the operation instructions through the splicing module to generate target structured information and target operation instructions, and splicing the image sequence, the target structured information and the target operation instructions in sequence to generate a sequence to be processed; The preprocessing module includes: a segmentation module and a linear mapping layer; The step of inputting the image into the preprocessing module for preprocessing to obtain a preprocessed image sequence includes: Inputting the image into the segmentation module, performing image segmentation processing on the image, and generating a plurality of sub-image blocks corresponding to the image; Each of the sub-image blocks is input into the linear mapping layer for projection to generate the pre-processed image sequence.
2. The human-computer interaction processing method according to claim 1, characterized in that: Before inputting the image, the structured information and the operation instruction into a pre-trained multimodal large model to generate a task statement corresponding to the operation instruction, the method further includes: The operation instructions are standardized according to the pre-trained instruction processing model to obtain standardized operation instructions.
3. The human-computer interaction processing method according to claim 1, characterized in that: The image processing model includes: a visual positioning model and a text recognition model; The step of inputting the image into a pre-trained image processing model, extracting content from the image, and generating structured information corresponding to the image includes: Inputting the image into the visual positioning model to generate the type and position of each element in the image; Inputting the image into the text recognition model to generate labels and text descriptions of each element in the image; According to the type, position, label and text description of each element, structured information corresponding to the image is generated.
4. The human-computer interaction processing method according to claim 1, characterized in that: The executing the task statement on the user interface and outputting the execution result to the user includes: Determine the operation area to be operated and the operation to be performed corresponding to the task statement in the user interface; The operation to be performed is performed in the operation area, and after the execution is completed, the execution result is output to the user.
5. The human-computer interaction processing method according to claim 4, characterized in that: The determining the area to be operated and the operation to be performed corresponding to the task statement in the user interface includes: Extracting the task statement to obtain a field of an area to be operated and a field of an operation to be performed in the task statement; Determine the area to be operated in the user interface according to the area to be operated field; The operation to be executed in the user interface is determined according to the operation to be executed field and a pre-stored action mapping dictionary.
6. A human-computer interaction processing device, characterized in that: include: An image acquisition module, used to acquire images of the user interface at a preset frequency; An instruction acquisition module is used to acquire the user's operation instructions; An image processing module, used to input the image into a pre-trained image processing model, extract content from the image, and generate structured information corresponding to the image, wherein the structured information is used to characterize the type and position of each element in the image, and the data structure of the structured information of each element is {label, type, text description, position}; A generation module, used for inputting the image, the structured information and the operation instruction into a pre-trained multimodal large model, and generating a task statement corresponding to the operation instruction, wherein the task statement at least includes an area to be operated and an operation to be performed corresponding to the operation instruction; An execution module, used for executing the task statement on the user interface and outputting the execution result to the user; The multimodal large model includes: a preprocessing module, a splicing module and a processing module; The generating module is specifically used for: Input the image into the preprocessing module for preprocessing to obtain a preprocessed image sequence; Input the image sequence, structural information and operation instructions into the splicing module to generate a sequence to be processed; Input the sequence to be processed into the processing module and generate the task statement corresponding to the operation instruction; The generating module is specifically used for: Processing the dimensions of the structured information and the operation instructions through the splicing module to generate target structured information and target operation instructions, and splicing the image sequence, the target structured information and the target operation instructions in sequence to generate a sequence to be processed; The preprocessing module includes: a segmentation module and a linear mapping layer; the generation module is specifically used for: Inputting the image into the segmentation module, performing image segmentation processing on the image, and generating a plurality of sub-image blocks corresponding to the image; Each of the sub-image blocks is input into the linear mapping layer for projection to generate the pre-processed image sequence.
7. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor executes the machine-readable instructions to perform the steps of the human-computer interaction processing method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the human-computer interaction processing method according to any one of claims 1 to 5 are executed.
Citation Information
Patent Citations
UI component analysis method and device based on visual large model
CN118312174A
Interface control method, computer program product and electronic equipment
CN118585107A