File processing method and device, equipment and storage medium

By setting prompt templates in the large language model and generating accurate prompt text, the problem of inaccurate user input prompt text manually is solved, and the ability of the large language model to quickly understand user needs and perform specific tasks is realized, improving work efficiency.

CN120104723APending Publication Date: 2025-06-06BEIJING QIHOOD TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311639576.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the prior art, users need to manually enter inaccurate prompt text, making it difficult to obtain accurate information from the large language model within a limited number of interactions, and the large language model cannot be directly used to perform specific tasks.

Method used

By setting up a prompt template, receiving user processing requests, and selecting appropriate prompt templates based on the request, generating accurate prompt text, allowing the large language model to quickly understand user needs, automatically generate code for editing files, and execute code to achieve specific tasks.

Benefits of technology

It reduces the requirements for user experience, improves the ability of large language models to understand user needs, realizes the rapid execution of specific file editing tasks, and improves the user's work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104723A_ABST
    Figure CN120104723A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a file processing method and device, equipment and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: receiving a processing request of a user for an original file, and selecting a corresponding prompt template according to the processing request; inputting the prompt template and the processing request into a large language model to obtain a prompt text output by the large language model; inputting the prompt text into the large language model to obtain a code which is output by the large language model and is used for editing the original file; executing the code based on the prompt text to obtain the target file. By adopting the method provided by the invention, the corresponding prompt template can be selected according to the processing request of the user, and the prompt text is automatically generated according to the prompt template and the processing request, so that the problem that the user needs to manually input the prompt text in related technologies is avoided, the large language model can quickly understand the requirements of the user, and the user experience is improved. The requirement for the use experience of the user is lowered, and the user can be further helped to improve the working efficiency of daily work.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a file processing method, device, equipment and storage medium. Background Art

[0002] At present, with the popularity of large language models, based on the natural language understanding capabilities of large language models including ChatGPT, users can gradually obtain the answers they need from large language models by continuously asking questions. When asking questions to large language models, due to the user's usage experience, if the user does not ask structured questions, multiple rounds of interaction with the large language model are required to get the required answers, and the operation efficiency is low, that is, the large language model only has the ability to generate dialogues and cannot be used to perform specific tasks. Summary of the invention

[0003] The present application provides a file processing method, apparatus, device and storage medium to solve the defects in the related art. By setting a prompt template, it is convenient to generate accurate prompt text, so that the user can fully utilize the intelligent capabilities of the large language model. The technical solution is as follows:

[0004] In a first aspect, an embodiment of the present application provides a file processing method, comprising:

[0005] receiving a processing request from a user for an original file, and selecting a corresponding prompt template according to the processing request;

[0006] Inputting the prompt template and the processing request into a large language model to obtain a prompt text output by the large language model;

[0007] Inputting the prompt text into the large language model to obtain a code output by the large language model for editing the original file;

[0008] The code is executed based on the prompt text, the original file is edited, and the target file is obtained.

[0009] In an optional solution of the first aspect, the receiving a processing request from a user for an original file, and selecting a corresponding prompt template according to the processing request, includes:

[0010] Receive a user's processing request for an original file, input the processing request into the large language model, parse the processing request through the large language model, obtain a function type for editing the original file, and determine a prompt template corresponding to the processing request according to the function type.

[0011] In an optional solution of the first aspect, before receiving the user's request for processing the original file, the method further includes:

[0012] According to the function type for editing the original file, a prompt template corresponding to each function type is preset.

[0013] In an optional solution of the first aspect, inputting the prompt template and the processing request into the large language model to obtain the prompt text output by the large language model includes:

[0014] Parsing the processing request by the large language model to determine the task target of the processing request, and obtaining a first prompt text including the task target output by the large language model based on the prompt template;

[0015] Inputting the first prompt text into the large language model to determine the task constraints for completing the task goal, and obtaining a second prompt text output by the large language model including the task constraints;

[0016] The first prompt text and the second prompt text are input into the large language model to obtain the prompt text output by the large language model.

[0017] In an optional solution of the first aspect, the parsing the processing request by the large language model to determine the task target of the processing request, and obtaining a first prompt text including the task target output by the large language model based on the prompt template, comprises:

[0018] Parsing the processing request through the large language model to determine the task target of the processing request and determine the cache path of the original file, obtaining the original file through the cache path, and determining the name of the large language model according to the file type of the original file;

[0019] A first prompt text including the task target, the cache path and the name of the large language model output by the large language model based on the prompt template is obtained.

[0020] In an optional solution of the first aspect, inputting the first prompt text into the large language model to determine the task constraints for completing the task goal includes:

[0021] Parsing the first prompt text by the large language model to obtain a plurality of subtask objectives after the large language model decomposes the task objective;

[0022] Determine control instructions for completing the plurality of subtask objectives through the large language model, and determine output specifications of the code through the large language model;

[0023] The task constraints include the multiple subtask objectives, the control instructions, and the output specifications.

[0024] In an optional scheme of the first aspect, the task constraint also includes a restriction condition of the control instruction, and the restriction condition is used to determine a threshold number of steps for the large language model to complete the task objective and / or a time threshold for the steps to complete the task objective.

[0025] In an optional solution of the first aspect, after executing the code based on the prompt text and before editing the original file to obtain the target file, the method further includes:

[0026] Determine whether error information is fed back after executing the code;

[0027] If the error message is not fed back, the step of editing the original file to obtain the target file is executed;

[0028] If an error message is fed back, the error message and the prompt text are input into the large language model to obtain the corrected prompt text output by the large language model, and the corrected prompt text is input into the large language model to obtain the corrected code output by the large language model, the corrected code is executed based on the corrected prompt text, and the step of determining whether an error message is fed back after executing the code is performed.

[0029] In a second aspect, an embodiment of the present application further provides a file processing device, including:

[0030] A prompt template module, used to receive a user's request for processing an original file, and select a corresponding prompt template according to the request;

[0031] A prompt text generation module, used for inputting the prompt template and the processing request into a large language model to obtain a prompt text output by the large language model;

[0032] A code generation module, used for inputting the prompt text into the large language model to obtain a code output by the large language model for editing the original file;

[0033] The file processing module is used to execute the code based on the prompt text, edit the original file, and obtain the target file.

[0034] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method provided in the first aspect of the embodiment of the present application or any one of the implementation methods of the first aspect is implemented.

[0035] In a fourth aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method provided by the first aspect of the embodiment of the present application or any one of the implementations of the first aspect.

[0036] The beneficial effects brought about by the technical solutions provided by some embodiments of the present application include at least:

[0037] A file processing method, apparatus, device and storage medium provided in an embodiment of the present application can select a corresponding prompt template according to a user's processing request, and automatically generate a prompt text based on a large language model according to the prompt template and the content of the processing request, thereby avoiding the problem in the related art that the user needs to manually input the prompt text. The prompt text manually input by the user is often not accurate enough, resulting in the inability to obtain accurate information from the large language model within a limited number of interactions. The present application allows the user to directly input natural language to describe the function requirements, and accurate prompt text can be generated according to the processing request and the corresponding prompt template, so that the large language model can quickly understand the user's needs, reducing the requirements for user experience, and further calling the large language model through the content of the prompt text to obtain the code for editing the file, and executing the code to obtain the edited file, thereby realizing the execution of specific tasks, which can further help users improve the work efficiency of daily work. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the present application or related technologies, the drawings required for use in the embodiments or related technical descriptions are briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0039] Figure 1 A schematic diagram of a system architecture of a file processing method according to an embodiment of the present application;

[0040] Figure 2 It is a flowchart of a file processing method according to an embodiment of the present application;

[0041] Figure 3 It is a flowchart of a file processing method according to an embodiment of the present application;

[0042] Figure 4 It is a flowchart of a file processing method according to an embodiment of the present application;

[0043] Figure 5 is a structural schematic diagram of a file processing device according to an embodiment of the present application;

[0044] Figure 6It is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solutions and advantages of this application clearer, the technical solutions in this application will be clearly and completely described below in conjunction with the drawings in this application. Obviously, the described embodiments are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0046] The terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device including a series of steps or modules is not limited to the listed steps or modules, but may optionally include steps or modules that are not listed, or may optionally include other steps or modules that are inherent to these processes, methods, products or devices.

[0047] It should be noted that the terms "first\second" involved in the present application are only used to distinguish similar objects, and do not represent a specific order for the objects. It is understandable that "first\second" can be interchanged with a specific order or sequence where permitted. It should be understood that the objects distinguished by "first\second" can be interchanged where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those described or illustrated herein.

[0048] Currently, users can input various questions into the large language model to obtain the required answers. The large language model (LLM) is a commonly used model in the field of natural language processing, which is used to process a variety of natural language tasks, such as text classification, question answering, dialogue, etc., to generate natural language text or understand the meaning of language text. Large language model is also a general term for deep learning models trained with large amounts of text data. For example, ChatGPT, GPT 3, GPT-4, PaLM, Galactica, LLaMA and other models are all large language models commonly used by technicians in this field.

[0049] Understandably, Big Language Model is short for Large Language Modeling (LLM), which refers to a deep learning model trained with a large amount of text data that can generate natural language text and / or understand the meaning of language text. Big Language Model can handle a variety of natural language tasks, such as text classification, question answering, dialogue, etc.

[0050] Through the natural language understanding capabilities of large language models including ChatGPT, users can set the content of their questions to obtain the information they need. However, since large language models including ChatGPT do not have the ability to call tools, they cannot directly operate the applications on the user's terminal. Only after the user continuously guides the inquiry can it be possible to get answers that meet the needs from the large language model.

[0051] Understandably, the effectiveness of using a large language model depends on the user's experience. If the user lacks experience or does not ask questions in a structured manner, the answers obtained from the large language model will be fragmented and difficult to use directly.

[0052] For example, if the user directly inputs "summary of meeting minutes", since no requirements are provided on the format, font, and content, the meeting minutes output by the large language model will automatically choose the form of the meeting minutes. If the output meeting minutes do not meet the requirements, it is necessary to repeatedly confirm the required format, font, content, etc. with the large language model.

[0053] Therefore, the current large language model is difficult to use directly in work. First, the large language model cannot directly call the tool library or has no matching tool library. Second, the large language model requires users to ask guiding questions and requires users to write appropriate prompt words in advance so that the large language model can fully understand the user's needs. However, the writing of prompt words is difficult, and repeated testing is required to determine the appropriate prompt words, which is not conducive to improving user work efficiency.

[0054] Based on this, the embodiments of the present application provide a file processing method, apparatus, device and storage medium, which can select a corresponding prompt template according to the user's processing request, and automatically generate a prompt text based on the prompt template and the content of the processing request based on the large language model, thereby avoiding the problem in the related art that the user needs to manually enter the prompt text. The prompt text manually entered by the user is often not accurate enough, resulting in the inability to obtain accurate information from the large language model within a limited number of interactions. The present application allows the user to directly enter natural language to describe the function requirements, and accurate prompt text can be generated according to the processing request and the corresponding prompt template, so that the large language model can quickly understand the user's needs, reducing the requirements for user experience, and further calling the large language model through the content of the prompt text to obtain the code for editing the file, and executing the code to obtain the edited file, thereby realizing the execution of specific tasks, which can further help users improve the work efficiency of daily work.

[0055] Please refer to the following Figure 1 , which is a schematic diagram of the architecture of a file processing system provided by an exemplary embodiment of the specification. Figure 1As shown, the file processing system may include: a terminal 110 and a server 120. Among them:

[0056] The terminal 110 may be a user terminal corresponding to one or more users, and may specifically include one or more user terminals. A user version of the software may be installed in the terminal 110, and the user may input a processing request in the terminal 110. After receiving the processing request input by the user, the terminal 110 selects a corresponding prompt template according to the processing request, inputs the processing request and the prompt template into the server 120, obtains a prompt text output by the server 120, inputs the prompt text into the server 120, obtains a code for editing a file output by the server, and the terminal 110 executes the code according to the content of the prompt text to edit the file, and after obtaining the edited target file, the target file will be fed back to the user. Among them, the terminal 110 may be, but is not limited to, a mobile phone, a tablet computer, a laptop computer, or other devices installed with the user version of the software.

[0057] Optionally, after the user inputs a processing request at the terminal 110, the processing request may be sent to the server 120 via the network, but is not limited to, so that the server 120 can subsequently output a code for editing the file based on the processing request. The server 120 may be a server that can provide a variety of file processing services, and may receive a processing request sent by the terminal 110 corresponding to the user via the network, and output a prompt text based on the processing request, and output a code for editing the file according to the prompt text. Optionally, after the server 120 outputs the prompt text and code, when receiving the processing request sent by the terminal 110 corresponding to the user, the server 120 may also send the prompt text and code to the terminal 110 corresponding to the target user for display. Among them, the server 120 may be, but is not limited to, a hardware server, a virtual server, a cloud server, etc., and a large language model is configured on the server 120, and the prompt text and code are output through the large language model.

[0058] It is understandable that the file processing method provided in the embodiment of the present application can be jointly executed by the terminal 110 and the server 120. The network can be a medium that provides a communication link between any terminal 110 and the server 120, or it can be the Internet including network equipment and transmission media, but is not limited thereto. The transmission medium can be a wired link, such as but not limited to coaxial cable, optical fiber and digital subscriber line (DSL), etc., or a wireless link, such as but not limited to wireless fidelity (WIFI), Bluetooth and mobile device network, etc.

[0059] Understandably, Figure 1The number of terminals 110 and servers 120 in the file processing system shown is only an example. In a specific implementation, the file processing system may include any number of terminals and servers.

[0060] This embodiment of the specification does not specifically limit this. For example, but not limited to, the terminal 110 can be a user terminal cluster composed of multiple user terminals, and the server 120 can be a server cluster composed of multiple servers.

[0061] The present application is described in detail below with reference to specific embodiments.

[0062] Next, combine Figure 1 , taking the AI ​​server on the terminal executing the file processing method as an example, the file processing method provided by the embodiment of the present application is introduced. Figure 2 , Figure 2 A schematic diagram of a process flow of a file processing method provided by an embodiment of the present application is shown, and the method comprises the following steps:

[0063] S201, receiving a processing request from a user for an original file, and selecting a corresponding prompt template according to the processing request.

[0064] It can be understood that before the user inputs a processing request, he or she first needs to download the necessary operating environment and the API interface of the large language model used on the terminal used. Taking AutoGPT as an example, the user needs to download the corresponding AutoGPT software from GitHub and install it on the terminal used, that is, the AI ​​server, install the runtime library including Python, and configure the API interfaces of various tools in the AI ​​server, including the API interface permissions of the required interface ChatGPT 3.5, and may also include but are not limited to the API interface permissions of Pinecone, the API interface permissions of the search engine, the voice interface permissions of Eleven Labs, and the image generation interface of HuggingFace and any one or more interface permissions of additional interfaces.

[0065] Understandably, Auto-GPT is an autonomous intelligent agent, specifically an open source project on Github, which can combine GPT-4 and GPT-3.5 technologies to create a complete project by accessing different API interfaces. Unlike large language models such as ChatGPT, ChatGPT users need to constantly ask questions to the large language model to get corresponding answers, while in AutoGPT, they only need to provide appropriate prompt text, and then AutoGPT can automatically implement the required functions based on the content of the prompt text.

[0066] Among them, AutoGPT can realize different functions through different interfaces. For example, database-related functions can be realized through the Pinecone interface, AutoGPT can conduct online queries through the search engine interface, and voice input and output and other functions can be realized through the ElevenLabs voice interface.

[0067] Specifically, after configuring the operating environment and necessary interfaces of AutoGPT, users can enter text information in the AI ​​server of the local terminal.

[0068] Specifically, the user can input text information of the processing request in natural language in the AI ​​server of the local terminal. For example, the user inputs "Please swap the values ​​in the cells of column A and column B in file 1.xlsx located on the desktop". The AI ​​server parses the content input by the user and obtains the file path as "Desktop", the file type as an Excel spreadsheet file, the file name as "1.xlsx", and the function type to be executed as data exchange. Based on the parsed file type and function type, the corresponding prompt template can be selected to convert the processing request in natural language input by the user into prompt text (prompt).

[0069] Among them, the prompt template is pre-set according to various function types and / or file types to be executed, and corresponding prompt templates are preset for different function types and / or file types, so as to facilitate the direct judgment of the most appropriate prompt template based on the text entered by the user, so as to avoid the situation where the text entered by the user cannot be accurately understood by the large language model. Through the set prompt template, the large language model can accurately understand the processing request input by the user, thereby determining the user's needs.

[0070] Specifically, after determining the prompt template, the following steps are also included:

[0071] S202, inputting the prompt template and the processing request into the large language model to obtain the prompt text output by the large language model.

[0072] Specifically, based on the prompt template determined in S201 and the processing request input by the user, the key words and / or keywords in the processing request input by the user may be filled into the corresponding positions of the prompt template.

[0073] The content to be filled in the prompt template includes but is not limited to task instructions, role positioning, the name of the file to be processed, the path of the file to be processed, output indicators, etc. The prompt template can be "You are now XXXGPT, please perform XX operation according to my description, use XXX tool, open the file XX.XXX located in path 1, specifically perform operation XX on parameter X in file XX.XXX, and save the file as XXX.XXX to path 2".

[0074] It can be understood that, depending on the type of file to be processed, ChatGPT can be positioned in different roles in the prompt text and prompt template. If an Excel file is processed, the role is defined as excelGPT, and if a Word document is processed, it is defined as wordGPT.

[0075] It is understandable that the content in the prompt template depends on the actual task requirements and does not necessarily need to include all of the above elements, and the embodiments of the present application are not limited to this.

[0076] For example, if the user inputs "Please swap the values ​​in the cells of columns A and B in file 1.xlsx located on the desktop", AutoGPT recognizes that the corresponding file type is an excel table, and the corresponding function is data exchange, and a prompt template for implementing the data exchange function in excel can be selected. It can be recognized that the input task instruction is for data exchange, so that ChatGPT can be informed of the task target to be executed. It can be recognized that the file type to be processed is an excel table, so the role type of ChatGPT can be limited to excelGPT, and the file path can be determined to be located on the desktop. The file name of the original file to be processed is 1.xlsx. The user can be prompted to enter an output indicator, or the output path of the file can be determined according to the file storage address pre-set by the user. The corresponding tool library can be selected according to the file type and the corresponding operation. For example, the excel file can be edited through the python code library. According to the various types of information identified, each type of information is filled in the corresponding position in the prompt template to generate the prompt text as follows:

[0077] You are now excelGPT, please perform the data exchange operation according to my description, use Python's openpyxl library, open the excel file 1.xlsx on the desktop, and specifically exchange the values ​​in the cells of columns A and B in file 1.xlsx, and save the file as 2.xlsx to folder F.

[0078] S203: Input the prompt text into the large language model to obtain a code output by the large language model for editing the original file.

[0079] Specifically, after obtaining the prompt text generated by ChatGPT, AutoGPT calls the interface of ChatGPT according to the content of the prompt text and inputs the prompt text into ChatGPT. ChatGPT can generate code for completing the task goal according to the content of the prompt text and feed it back to the AI ​​server of AutoGPT.

[0080] Exemplarily, ChatGPT can confirm the task goal entered by the user based on the input prompt text: swap the values ​​in the cells of columns A and B in file 1.xlsx located on the desktop. To achieve this goal, the data exchange function can be implemented through python code. ChatGPT can call the python code library, from which it can traverse the code snippets of the corresponding functions, thereby generating python code.

[0081] S204, executing the code based on the prompt text, editing the original file, and obtaining the target file.

[0082] Specifically, after generating the code, AutoGPT's AI server receives the code fed back by ChatGPT, and executes the corresponding code according to the instructions of the prompt text, thereby editing the original file input by the user to obtain the edited target file.

[0083] Specifically, the prompt text can determine the execution method of the code. For example, the corresponding control instructions can be written in the prompt text. For example, after the code is generated, the corresponding code is executed locally. The control instructions can be: commands: Creates a Python file and executes it, which means creating a Python file and executing this Python file. The code can also be saved locally and executed locally, and the edited file is stored after the code is executed. The control instructions can be: commands: Creates a Python file and executes it, then save the modified file, which means creating a Python file and executing this Python file, and storing the edited file after the code is executed.

[0084] Specifically, AutoGPT's AI server responds to the control instructions in the prompt text and automatically executes the code according to the control instructions, thereby realizing automatic editing of the file.

[0085] Specifically, there is no need to restrict the control instructions in the prompt text. After receiving the prompt text, AutoGPT will confirm the task goal by itself, and choose the way to complete the task goal according to the provided API interface. Taking the data exchange of Excel files as an example, in addition to the necessary ChatGPT interface, Python's openpyxl library is also provided. The required code can be generated according to Python specifications to realize the data exchange of Excel files. After the code is generated, the corresponding code is automatically executed.

[0086] Optionally, if the local runtime library for executing the code is missing, for example, Python's openpyxl library is required to execute Python code, if the corresponding runtime library is not installed, AutoGPT can automatically install the required runtime library.

[0087] Optionally, the processed Excel file is transferred back to the user end, and the new Excel result is displayed through the OnlyOffice online component.

[0088] In some embodiments, after receiving a processing request provided by a user, if the function type of the processing request input by the user is inconsistent with the function type corresponding to all preset prompt templates, a selection is made based on the specific function type and file type. For example, the preset prompt template does not have a setting for the function of inserting a function in Excel. For the same type of Excel file, there is a prompt template for summing two columns of data in Excel. The prompt template for summing two columns of data in Excel can be selected to construct the corresponding prompt text.

[0089] The embodiment of the present application selects a corresponding prompt template according to the user's processing request, and automatically generates a prompt text based on the prompt template and the content of the processing request based on the large language model, thereby avoiding the problem in the related art that the user needs to manually input the prompt text. The prompt text manually input by the user is often not accurate enough, resulting in the inability to obtain accurate information from the large language model within a limited number of interactions. The present application allows the user to directly input natural language to describe the function requirements, and accurate prompt text can be generated according to the processing request and the corresponding prompt template, so that the large language model can quickly understand the user's needs, reducing the requirements for user experience. The large language model is further called through the content of the prompt text to obtain the code for editing the file, and the code is executed to obtain the edited file, thereby realizing the execution of specific tasks, which can further help users improve the work efficiency of daily work.

[0090] Please refer to the following Figure 3 , Figure 3 A schematic diagram of a process flow of a file processing method provided by an embodiment of the present application is shown, and the method comprises the following steps:

[0091] S301, parsing a processing request through a large language model to determine a task target of the processing request, and obtaining a first prompt text including the task target output by the large language model based on a prompt template.

[0092] Specifically, after the processing request is input into the large language model, the large language model parses the processing request input by the user in natural language, thereby confirming the task objective corresponding to the user's processing request. For example, if the user inputs "Please swap the values ​​in the cells of columns A and B in file 1.xlsx located on the desktop", it can be determined that the task objective is to exchange data in the cells of columns A and B in the Excel file.

[0093] Specifically, after determining the task target corresponding to the user's processing request, the large language model applies the format of the prompt template, fills the task target into the prompt template, and obtains an initial prompt text, that is, the first prompt text.

[0094] Specifically, the first prompt text may also include but is not limited to information such as the name of the original file, the cache path of the original file, the role positioning of the large language model, etc. The large language model can be used to parse the processing request to determine the task target of the processing request and determine the cache path of the original file. The original file to be processed can be extracted through the name of the original file and the cache path. The role positioning of the large language model can be given by naming the large language model, and the first prompt text including the task target, cache path and name of the large language model output by the large language model based on the prompt template is obtained. For the specific steps of generating the first prompt text, please refer to the description in S202, which will not be repeated here.

[0095] S302: Input the first prompt text into the large language model to determine the task constraints for completing the task goal, and obtain a second prompt text output by the large language model including the task constraints.

[0096] Specifically, the first prompt text is used to determine the task objective corresponding to the processing request, and can further provide the constraint conditions required to complete the task objective according to the task objective of the processing request, namely, the task constraint, and include the output task constraint in the second prompt text.

[0097] Specifically, the first prompt text can be parsed by the large language model to obtain multiple subtask objectives after the large language model decomposes the task objective, wherein the multiple subtask objectives can be multiple sub-steps arranged in sequence. For example, the user inputs "Please exchange the values ​​in the cells of columns A and B in file 1.xlsx located on the desktop", and the task objective is "Exchange the data of the cells in columns A and B in the Excel file". The subtask objectives obtained after the task objective is decomposed may include:

[0098] (1) Use the openpyxl library to open the file 1.xlsx on the desktop;

[0099] (2) Swap the positions of columns A and B in the corresponding tables;

[0100] (3) Save the edited file.

[0101] The corresponding interface can be called based on each subtask target, so as to use the corresponding tool to complete the corresponding subtask. For example, in the above sub-step (1), the openpyxl library is called to execute the target of opening the Excel file, and sub-step (2) can be implemented by calling the openpyxl library to execute the corresponding Python code. Sub-step (3) can call the local resource manager to save the file to the specified location in the resource manager.

[0102] Furthermore, the control instructions for completing multiple subtask objectives can be determined through the large language model, and the output specification of the code can be determined through the large language model.

[0103] Taking AutoGPT as an example, AutoGPT needs to respond to control instructions to complete the set goals. For example, if the control instruction is set to "generate Python code file and execute Python code", then when AutoGPT responds to the control instruction, it will call the openpyxl library to generate the corresponding Python code according to the set sub-goals, save the Python code as a Python file, and execute the obtained Python file.

[0104] Optionally, the storage location of the edited file d may also be determined through a control instruction.

[0105] Specifically, the code output specification can be understood as a restriction on the code format and form, so that when the large language model generates code, it outputs the code in a specified format and form according to the given output specification. For example, if the code is output in the format of JSON schema, the output specification in the second prompt text can be expressed as:

[0106] Respond with only valid JSON conforming to the following schema:

[0107] {"$schema":"http: / / json-schema.org / draft-07 / schema#","type":"object","properties":{"thoughts":{"type":"object","properties":{ "text":{"type":"string","description":"thoughts"},"reasoning":{"type":"string"},"plan":{"type":"string","description":"-short bulleted\n-list that conveys\n-long-term plan"},"criticism":{"type":"string","description":"constructiveself-criticism"},"speak":{"type":"string","description":"thoughts summary to say to user"}},"required":["text","reasoning","plan","criticism","speak"],"additionalProperties":false},"command":{"type":"object","properties":{"name":{"type":"s tring"},"args":{"type":"object"}},"required":["name","args"],"additionalProperties":false}},"required":["thoughts","command"],"additionalProperties":false};

[0108] The above JSON schema format provides restrictions on the output code format. The large language model can confirm the user's restrictions on the code format by parsing the second prompt text, avoid grammatical errors in the code, and ensure that the code complies with the prescribed specifications, such as the JSON schema specifications given in http: / / json-schema.org / draft-07 / schema#.

[0109] Therefore, the task constraints in the second prompt text include at least a plurality of sub-task objectives, control instructions and output specifications.

[0110] Optionally, the task constraints in the second prompt text may also include restriction conditions of the control instructions, and the restriction conditions are used to determine a threshold number of steps for the large language model to complete the task objective and / or a time threshold for the steps to complete the task objective, thereby avoiding excessive decomposition of the task objective resulting in too many sub-task objectives when the large language model disassembles the task objective. Executing too many steps will waste the computing power of the large language model and reduce computing efficiency, thereby avoiding excessive iteration of the large language model and allowing the large language model to disassemble the task in the most reasonable and efficient manner.

[0111] Optionally, the task constraints may also include restrictions on the behavior of the large language model, such as prohibiting the large language model from asking questions to the user, ensuring that the large language model finds answers on its own by calling tools such as search engines, and avoiding invalid questions.

[0112] Optionally, the task constraints may also include self-evaluation of the large language model, such as allowing the large language model to complete the task goal with the least steps and lowest consumption, limiting the number of steps to complete each subtask goal, and limiting the number of iterations of the large language model.

[0113] S303: Input the first prompt text and the second prompt text into the large language model to obtain the prompt text output by the large language model.

[0114] Specifically, the first prompt text and the second prompt text can be spliced ​​through the large language model to obtain a complete prompt text. For example, the first prompt text is "You are now excelGPT, please perform data exchange operations according to my description, use Python's openpyxl library, open the excel file 1.xlsx on the desktop, and specifically exchange the values ​​in the cells of columns A and B in file 1.xlsx, and save the file as 2.xlsx to folder F", and the second prompt text includes:

[0115] Subtask objectives: (1) Use the openpyxl library to open the file 1.xlsx on the desktop; (2) Swap the positions of columns A and B in the corresponding form; (3) Save the edited file.

[0116] Restrictions: Do not ask users.

[0117] Control instructions: Generate Python files and execute Python code.

[0118] Performance evaluation: Limit the number of iterations of AutoGPT to 1.

[0119] Output Specification: Only allow feedback using JSON schema specifications.

[0120] The second prompt text may include the above-mentioned subtask objectives, constraints, control instructions, performance evaluation, and output specifications.

[0121] Specifically, the first prompt text and the second prompt text may be superimposed to obtain a complete prompt text, specifically, the first prompt text is placed above the second prompt text, that is, the complete prompt text is output.

[0122] In the embodiment of the present application, a task objective is given through a first prompt text, and a task constraint corresponding to the task objective is determined through a second prompt text. The task objective is determined according to a processing request input by a user, and the task constraint is determined by parsing the first prompt text through a large language model and calling a corresponding tool, so that the second prompt text can be more easily understood by the large language model. The task constraint given by the second prompt text avoids excessive iteration of the large language model, so that the code output by the large language model is more in line with actual needs, which is conducive to improving work efficiency.

[0123] In some embodiments, step S204 further includes:

[0124] S401, executing code based on the prompt text.

[0125] S402, determining whether error information is fed back after executing the code.

[0126] S403: If no error message is fed back, the original file is edited to obtain the target file.

[0127] S404: If an error message is fed back, the error message and the prompt text are input into the large language model to obtain the corrected prompt text output by the large language model, and the corrected prompt text is input into the large language model to obtain the corrected code output by the large language model. The corrected code is executed based on the corrected prompt text, and S402 is executed.

[0128] Specifically, the code output by the large language model may contain errors, so it is necessary to judge the execution status of the code after the code is executed. If there is no error feedback on the user's terminal, it means that the code runs successfully, and the user's terminal can receive the edited file obtained after the code runs.

[0129] Specifically, if an error message is fed back, it indicates that there is a bug in the code operation, and the code needs to be modified based on the specific error message fed back. The error message, prompt text, and the code where the error message appears can be input into the large language model together, and the large language model can confirm the corresponding error cause and error location by itself, and re-output the corrected code through the large language model. Specifically, a new prompt text can be generated, and the prompt text can be used to give the task goal of finding the error cause and error location of the code, so that the large language model can correct the code by itself.

[0130] Specifically, after the code is corrected, it is necessary to execute the code again and determine whether an error is reported through step S402. If the code is executed successfully without error information, step S403 is executed. If the code feedbacks an error message, step S404 is continued.

[0131] Optionally, the iterative process of code correction can be limited to limit the number of corrections for the same code, for example, limiting the number of corrections to 3 times. If an error message still appears after the same code has been corrected 3 times, the iteration is stopped, and the large language model is commanded to regenerate new code and discard the original erroneous code.

[0132] The embodiment of the present application further realizes the automation of code generation by judging whether there is error information after the code is executed. If there is error information, the large language model is used to continue to correct the code, which is beneficial to improving the accuracy of the generated code. By limiting the number of iterations of the large language model, efficiency problems caused by the code still reporting errors after being corrected by the large language model for multiple times are avoided, thereby avoiding the waste of computing power of the large language model, so that the large language model can change the way of generating code in time, avoid the occurrence of stubborn errors, and thus improve the user experience.

[0133] The following are device embodiments of the present application, which can be used to execute the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0134] See next Figure 5 , is a schematic diagram of the structure of a file processing device provided by an exemplary embodiment of the present application. The device can be implemented as all or part of a terminal through software, hardware, or a combination of both, and can also be integrated on a server as an independent module. The file processing device in the embodiment of the present application can be applied to a terminal or a cloud. The device 500 includes a prompt template module 510, a prompt text generation module 520, a code generation module 530, and a file processing module 540, wherein:

[0135] The prompt template module 510 is used to receive a user's request for processing an original file and select a corresponding prompt template according to the request;

[0136] A prompt text generating module 520, configured to input the prompt template and the processing request into a large language model to obtain a prompt text output by the large language model;

[0137] A code generation module 530, used for inputting the prompt text into the large language model to obtain a code output by the large language model for editing the original file;

[0138] The file processing module 540 is used to execute the code based on the prompt text, edit the original file, and obtain the target file.

[0139] In some embodiments, when the prompt template module 510 receives a user's request for processing an original file and selects a corresponding prompt template according to the request, it includes:

[0140] The prompt template module 510 receives a user's processing request for an original file, inputs the processing request into the large language model, parses the processing request through the large language model, obtains a function type for editing the original file, and determines a prompt template corresponding to the processing request according to the function type.

[0141] In some embodiments, before the prompt template module 510 receives the user's request for processing the original file, the prompt template module 510 further includes:

[0142] The prompt template module 510 presets prompt templates corresponding to each function type according to the function type to be edited for the original file.

[0143] In some embodiments, the prompt text generation module 520 inputs the prompt template and the processing request into a large language model to obtain a prompt text output by the large language model, including:

[0144] The prompt text generating module 520 parses the processing request through the large language model to determine the task target of the processing request, and obtains a first prompt text including the task target output by the large language model based on the prompt template;

[0145] The prompt text generation module 520 inputs the first prompt text into the large language model to determine the task constraints for completing the task goal, and obtains a second prompt text output by the large language model including the task constraints;

[0146] The prompt text generation module 520 inputs the first prompt text and the second prompt text into the large language model to obtain the prompt text output by the large language model.

[0147] In some embodiments, the prompt text generation module 520 is further used to parse the processing request through the large language model to determine the task target of the processing request and determine the cache path of the original file, obtain the original file through the cache path, and determine the name of the large language model according to the file type of the original file;

[0148] A first prompt text including the task target, the cache path and the name of the large language model output by the large language model based on the prompt template is obtained.

[0149] In some embodiments, the prompt text generation module 520 is further used to parse the first prompt text through the large language model to obtain a plurality of sub-task objectives after the large language model decomposes the task objective;

[0150] Determine control instructions for completing the plurality of subtask objectives through the large language model, and determine output specifications of the code through the large language model;

[0151] The task constraints include the multiple subtask objectives, the control instructions, and the output specifications.

[0152] In some embodiments, the task constraints also include restriction conditions of the control instructions, and the restriction conditions are used to determine a threshold number of steps for the large language model to complete the task goal and / or a time threshold for the steps to complete the task goal.

[0153] In some embodiments, after the code generation module 530 executes the code based on the prompt text, before editing the original file and obtaining the target file, the code generation module 530 further includes:

[0154] The file processing module 540 determines whether error information is fed back after executing the code;

[0155] If the file processing module 540 does not feedback the error information, the file processing module 540 executes the editing of the original file and obtains the target file;

[0156] If the file processing module 540 feeds back an error message, the prompt text generation module 520 inputs the error message and the prompt text into the large language model to obtain the corrected prompt text output by the large language model, the code generation module 530 inputs the corrected prompt text into the large language model to obtain the corrected code output by the large language model, the file processing module 540 executes the corrected code based on the corrected prompt text, and the file processing module 540 continues to determine whether an error message is fed back after executing the code.

[0157] It should be noted that, when the device 500 provided in the above embodiment executes the file processing method, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the file processing method embodiment belong to the same concept, and the implementation process thereof is detailed in the method embodiment, which will not be repeated here.

[0158] An embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above-mentioned method embodiments when executing the program.

[0159] See also Figure 6 , which is a structural block diagram of an electronic device provided in an embodiment of the present application.

[0160] like Figure 6 As shown, the electronic device 600 includes: a processor 601 and a memory 602 .

[0161] In the embodiment of the present application, the processor 601 is the control center of the computer system, which can be a processor of a physical machine or a processor of a virtual machine. The processor 601 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 601 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array).

[0162] The processor 601 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also called a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state.

[0163] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In some embodiments of the present application, the non-transitory computer-readable storage medium in the memory 602 is used to store at least one instruction, which is used to be executed by the processor 601 to implement the method in the embodiment of the present application.

[0164] In some embodiments, the electronic device 600 further includes: a peripheral device interface 603 and at least one peripheral device. The processor 601, the memory 602 and the peripheral device interface 603 can be connected via a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 603 via a bus, a signal line or a circuit board. Specifically, the peripheral devices include: a display screen 604, a camera 605 and an audio circuit 606. The peripheral device interface 603 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 601 and the memory 602.

[0165] In some embodiments of the present application, the processor 601, the memory 602, and the peripheral device interface 603 are integrated on the same chip or circuit board; in some other embodiments of the present application, any one or two of the processor 601, the memory 602, and the peripheral device interface 603 can be implemented on a separate chip or circuit board. This embodiment of the present application does not specifically limit this.

[0166] The display screen 604 is used to display a UI. The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 604 is a touch display screen, the display screen 604 also has the ability to collect touch signals on the surface or above the surface of the display screen 604. The touch signal may be input as a control signal to the processor 601 for processing. At this time, the display screen 604 may also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards.

[0167] In some embodiments of the present application, the display screen 604 may be one, which is arranged on the front panel of the electronic device 600; in other embodiments of the present application, the display screen 604 may be at least two, which are arranged on different surfaces of the electronic device 600 or are folded; in still other embodiments of the present application, the display screen 604 may be a flexible display screen, which is arranged on the curved surface or folded surface of the electronic device 600. Even more, the display screen 604 may be arranged in a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 604 may be made of materials such as LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode) and the like.

[0168] Camera 605 is used to capture images or videos. Optionally, camera 605 includes a front camera and a rear camera. Usually, the front camera is arranged on the front panel of the electronic device, and the rear camera is arranged on the back of the electronic device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize the panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments of the present application, camera 605 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. Dual-color temperature flash refers to a combination of warm light flash and cold light flash, which can be used for light compensation at different color temperatures.

[0169] The audio circuit 606 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 601 for processing. For the purpose of stereo sound collection or reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 600. The microphone may also be an array microphone or an omnidirectional collection microphone.

[0170] The power supply 607 is used to power various components in the electronic device 600. The power supply 607 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 607 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.

[0171] The electronic device structure block diagram shown in the embodiment of the present application does not constitute a limitation on the electronic device 600. The electronic device 600 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0172] The embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method of any of the above embodiments are implemented. The computer-readable storage medium may include, but is not limited to, any type of disk, including a floppy disk, an optical disk, a DVD, a CD-ROM, a micro drive, and a magneto-optical disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, a magnetic card or an optical card, a nanosystem (including a molecular memory IC), or any type of medium or device suitable for storing instructions and / or data.

[0173] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, by hardware. Based on this understanding, the above technical solution can be essentially or in other words, the part that contributes to the relevant technology can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiment.

[0174] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A file processing method, It is characterized in that include: receiving a processing request from a user for an original file, and selecting a corresponding prompt template according to the processing request; Inputting the prompt template and the processing request into a large language model to obtain a prompt text output by the large language model; Inputting the prompt text into the large language model to obtain a code output by the large language model for editing the original file; The code is executed based on the prompt text, the original file is edited, and the target file is obtained.

2. The method according to claim 1, It is characterized in that The receiving a user's request for processing the original file and selecting a corresponding prompt template according to the request comprises: Receive a user's processing request for an original file, input the processing request into the large language model, parse the processing request through the large language model, obtain a function type for editing the original file, and determine a prompt template corresponding to the processing request according to the function type.

3. The method according to claim 1, It is characterized in that Before receiving the processing request of the user for the original file, the method further includes: According to the function type for editing the original file, a prompt template corresponding to each function type is preset.

4. The method according to claim 1, It is characterized in that The step of inputting the prompt template and the processing request into a large language model to obtain a prompt text output by the large language model includes: Parsing the processing request by the large language model to determine the task target of the processing request, and obtaining a first prompt text including the task target output by the large language model based on the prompt template; Inputting the first prompt text into the large language model to determine the task constraints for completing the task goal, and obtaining a second prompt text output by the large language model including the task constraints; The first prompt text and the second prompt text are input into the large language model to obtain the prompt text output by the large language model.

5. The method according to claim 4, It is characterized in that The step of parsing the processing request by the large language model to determine the task target of the processing request, and obtaining a first prompt text including the task target output by the large language model based on the prompt template, comprises: Parsing the processing request through the large language model to determine the task target of the processing request and determine the cache path of the original file, obtaining the original file through the cache path, and determining the name of the large language model according to the file type of the original file; A first prompt text including the task target, the cache path and the name of the large language model output by the large language model based on the prompt template is obtained.

6. The method according to claim 4, It is characterized in that The step of inputting the first prompt text into the large language model to determine the task constraints for completing the task goal includes: Parsing the first prompt text by the large language model to obtain a plurality of subtask objectives after the large language model decomposes the task objective; Determine control instructions for completing the plurality of subtask objectives through the large language model, and determine output specifications of the code through the large language model; The task constraints include the multiple subtask objectives, the control instructions, and the output specifications.

7. The method according to claim 6, It is characterized in that The task constraint also includes a restriction condition of the control instruction, and the restriction condition is used to determine a threshold value of the number of steps for the large language model to complete the task goal and / or a time threshold for completing the steps of the task goal.

8. A file processing device, It is characterized in that include: A prompt template module, used to receive a user's request for processing an original file, and select a corresponding prompt template according to the request; A prompt text generation module, used for inputting the prompt template and the processing request into a large language model to obtain a prompt text output by the large language model; A code generation module, used for inputting the prompt text into the large language model to obtain a code output by the large language model for editing the original file; The file processing module is used to execute the code based on the prompt text, edit the original file, and obtain the target file.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, It is characterized in that When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, It is characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Method for enhancing json generation capability of large model and plug-in tool

    CN121436150A