A document-based intelligent dialogue method, device, and storage medium
By creating file assistants and conversation threads in ChatGPT, the problem of slow uploading and answer acquisition of users when using ChatGPT is solved, and the effect of quickly obtaining document information is achieved and time cost is reduced.
Patent Information
- Application Number
- CN202410918592.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-05
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-07-05
AI Technical Summary
When users use ChatGPT to analyze the target document, the speed of uploading the target document and obtaining answers is slow, resulting in higher time costs.
By obtaining the target document and sending it to a large language model, receiving the file ID, creating assistants and conversation threads associated with the file ID with the endpoint, creating execution commands containing user questions, and sending instructions to run the conversation thread to generate the answer text.
By creating file assistants and conversation threads, users can quickly obtain document information in the corresponding threads, significantly reducing time costs and improving information acquisition efficiency.
Smart Images

Figure CN118760754B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a document-based intelligent dialogue method, device, and storage medium. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, the document-assisted reading function based on artificial intelligence technology has played an important role in people's rapid access to information. For example, the file reading assistant based on ChatGPT provided by OpenAI can perform intelligent analysis on the content of documents, thereby helping users quickly extract, locate, and summarize information from the documents.
[0003] However, when users use ChatGPT to perform content analysis on the target document, the speed of uploading the target document and obtaining the answers generated by ChatGPT is often slow, resulting in a high time cost for users to use this function. Summary of the Invention
[0004] To solve the above technical problems, this application provides a document-based intelligent dialogue method, device, and storage medium.
[0005] An embodiment of this application provides a document-based intelligent dialogue method, including:
[0006] Obtain the target document and send it to the large language model, and receive the returned file ID;
[0007] Use the endpoint to create an assistant associated with the file ID and obtain the corresponding assistant ID;
[0008] Send a setting instruction containing the file ID to the large language model to generate a dialogue thread based on the target document and the corresponding thread ID;
[0009] When receiving the target question sent by the user, create an execution command containing the target question in the dialogue thread;
[0010] Send an instruction to run the dialogue thread to the large language model, so that the large language model runs the execution command and generates a response text.
[0011] In some embodiments, the step of when receiving the target question sent by the user, creating an execution command containing the target question in the dialogue thread specifically includes:
[0012] Obtain the target question and the corresponding file ID;
[0013] Match the corresponding thread ID and assistant ID according to the corresponding file ID, and create an execution command containing the target question in the dialogue thread corresponding to the thread ID.
[0014] In some embodiments, when sending an instruction to run the conversation thread to the large language model, the target question is synchronously transmitted by using the assistant corresponding to the assistant ID.
[0015] In some embodiments, after sending an instruction to run the conversation thread to the large language model, it further includes: sending an instruction to request to obtain an answer to the large language model; receiving the response text generated by the large language model.
[0016] In some embodiments, after obtaining the target document, it further includes: uploading the target document to a cloud storage platform for storage.
[0017] In some embodiments, the obtaining of the target document includes: receiving the target document uploaded by the user, or; receiving a URL uploaded by the user, where the URL contains the target document.
[0018] In some embodiments, the target question sent by the user includes an associated question composed of a number of sub-questions.
[0019] In some embodiments, when receiving the associated question sent by the user, the response text generated by the large language model includes response contents corresponding to answering each of the sub-questions.
[0020] The embodiment of the present application further provides a document-based intelligent dialogue device, including:
[0021] A creation module, configured to obtain a target document and send it to a large language model, and receive the returned file ID; use an endpoint to create an assistant associated with the file ID, and obtain the corresponding assistant ID; send a setting instruction including the file ID to the large language model to generate a conversation thread based on the target document and the corresponding thread ID;
[0022] A dialogue module, configured to, when receiving a target question sent by the user, create an execution command including the target question in the conversation thread; send an instruction to run the conversation thread to the large language model, so that the large language model runs the execution command and generates a response text.
[0023] In addition, the embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, any document-based intelligent dialogue method provided by the embodiment of the present application is implemented.
[0024] Compared with the prior art, the beneficial effects of the embodiment of the present application are as follows:
[0025] The document-based intelligent dialogue method provided by this application includes obtaining a target document and sending it to a large language model, and receiving the returned file ID; creating an assistant associated with the file ID using an endpoint, and obtaining the corresponding assistant ID; sending a setting instruction containing the file ID to the large language model to generate a dialogue thread based on the target document and the corresponding thread ID; when receiving a target question sent by a user, creating an execution command containing the target question in the dialogue thread; sending an instruction to run the dialogue thread to the large language model, so that the large language model runs the execution command and generates a response text. By creating a file assistant and a dialogue thread, the above method enables users to upload a target document and then conduct accurate and fast intelligent Q&A with the large language model in the corresponding dialogue thread, effectively reducing the user's time cost. Description of the Drawings
[0026] To more clearly illustrate the technical solutions of this application, the drawings required for the implementation will be briefly introduced below. Obviously, the drawings in the following description are only some implementations of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 It is a schematic flowchart of the document-based intelligent dialogue method provided by an embodiment of this application;
[0028] Figure 2 It is a schematic structural diagram of the document-based intelligent dialogue device provided by an embodiment of this application. Detailed Embodiments
[0029] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts fall within the scope of protection of this application.
[0030] It should be understood that the step numbers used in the text are only for convenient description and do not limit the execution order of the steps.
[0031] It should be understood that the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification of this application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.
[0032] The terms "comprising" and "including" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0033] The term "and / or" refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0034] Since the advent of ChatGPT, the function of using artificial intelligence technology to improve the efficiency of document reading has attracted much attention from users. Generally, ChatGPT extracts and summarizes the content of the documents uploaded by users based on large language models. Users can deeply explore the document structure and content through conversational learning with it.
[0035] However, in the process of using ChatGPT, the speed of uploading the target document and obtaining the answers generated by ChatGPT is often slow, which requires a high time cost, resulting in a low efficiency for users to obtain information from the document content.
[0036] Based on the above technical problems, this application provides an intelligent conversation method based on documents.
[0037] Please refer to Figure 1 , the intelligent conversation method includes the following steps:
[0038] S1: Obtain the target document and send it to the large language model, and receive the returned file ID.
[0039] S2: Use the endpoint to create an assistant associated with the file ID, and obtain the corresponding assistant ID.
[0040] S3: Send the setting instruction containing the file ID to the large language model to generate a conversation thread based on the target document and the corresponding thread ID.
[0041] S4: When receiving the target question sent by the user, create an execution command containing the target question in the conversation thread.
[0042] S5: Send an instruction to run the conversation thread to the large language model so that the large language model runs the execution command and generates a response text.
[0043] The large language model includes ChatGPT launched by OpenAI. It should be noted that the relevant embodiments of this application will be described by taking ChatGPT as an example, but it does not constitute a limitation on the large language model.
[0044] In this embodiment, the method for obtaining the target document in step S1 includes receiving the target document uploaded by the user, or receiving the URL uploaded by the user, where the URL contains the target document; among them, when the user uploads the target document through the URL, it is necessary to enter the corresponding page to download the target document.
[0045] It should be noted that after obtaining the target document, the target document can be uploaded to the cloud storage platform for storage.
[0046] Specifically, the obtained document can be uploaded to OSS for data storage. At this time, the user can obtain the uploaded document and its corresponding session records by logging in to different devices with an account, without repeatedly obtaining relevant data when updating the device.
[0047] In this embodiment, after uploading the document to OSS, a corresponding link will be generated. This link can be sent to the front end for re-downloading the document. At the same time, this link will also be sent to ChatGPT.
[0048] Exemplarily, the Files / Create endpoint can be used to upload the document. The following are the relevant steps for uploading the document using Postman:
[0049] (1) Set the authorization header to Bearer <MY GPT-4API KEY>;
[0050] (2) Set the request type to POST;
[0051] (3) Set the request URL to https: / / api.openai.com / v1 / files;
[0052] (4) Set the body to the type form-data;
[0053] (5) Create a file key and set the purpose attribute to assistants;
[0054] After uploading the above code and the target document to ChatGPT together, the file_ID (i.e., the file ID) returned by ChatGPT will be obtained, and this file ID will be recorded.
[0055] Step S2 of this embodiment involves assistant creation. Specifically, the createAssistant endpoint can be used to create an assistant: send the specific settings of the assistant to the createAssistant endpoint, and the settings include name, description, model used, type, and file ID to be read, etc. After creation, obtain the assistant ID returned by the endpoint and record this assistant ID.
[0056] Step S3 of this embodiment involves creating a conversation thread and assigning an assistant. Specifically, the file ID and the user instruction recorded in Step S1 are concatenated to obtain a setting instruction, and the setting instruction is sent to ChatGPT to generate a conversation thread based on the content of the target document.
[0057] Specifically, the setting instruction includes a creation instruction, a user instruction, and a file ID, where the user instruction includes the question sent by the user concatenated with the corresponding function instruction.
[0058] Exemplarily, when the target document uploaded by the user is received, the following code is executed:
[0059] const thread = await openai.beta.threads.create(); / / Creation instruction
[0060] {"messages":
[0061] {"role":"user",
[0062] "content":"Give me a summary based on the document content.?", / / User instruction
[0063] "file_ids":["file-ethprices"] / / File ID
[0064] }]}
[0065] ChatGPT will return the corresponding thread ID based on the above code. At this time, all subsequent conversations between ChatGPT and the user will be carried out in the conversation thread corresponding to this thread ID, ensuring that for any relevant questions subsequently raised by the user, ChatGPT will generate accurate response content based on this target document.
[0066] In this embodiment, Steps S4 and S5 involve: when the target question sent by the user is received, an execution command containing the target question is created in the corresponding conversation thread, and an instruction to run the conversation thread is sent to ChatGPT, so that ChatGPT runs the execution command and generates a response text.
[0067] Among them, the target question includes the question edited by the user or the instruction selected by the user.
[0068] Exemplarily, the instruction to run the conversation thread can be:
[0069]
[0070] Specifically, after sending the instruction to run the conversation thread to ChatGPT, it is necessary to send another instruction to request an answer to receive the response text generated by ChatGPT and send the response text to the front end.
[0071] It can be understood that when the response file is received, the response text can be stored in the database.
[0072] In the above embodiments of the present application, by creating an assistant and a conversation thread, the target document uploaded by the user can be associated with the corresponding conversation thread. When the user conducts an intelligent conversation with the large language model based on this document, a response can be quickly obtained through the corresponding conversation thread, thereby significantly improving the efficiency of obtaining document information and greatly reducing the time cost of the user.
[0073] In another embodiment, when the user issues a question through the front end, at this time, the question and its corresponding file ID are received, and the corresponding thread ID and assistant ID are matched based on the file ID to create an execution command containing the question in the conversation thread corresponding to the thread ID.
[0074] Exemplarily, the execution command may be:
[0075] from openai import OpenAI
[0076] client = OpenAI()
[0077] thread_message = client.beta.threads.messages.create(
[0078] "thread_abc123", / / Pass the thread ID
[0079] role = "user",
[0080] content = "How does AI work?Explain it in simple terms.", / / User question )
[0082] At this time, when sending the instruction to run the conversation thread to ChatGPT, the assistant corresponding to the assistant ID will also synchronously pass the target question.
[0083] It should be noted that the URL must contain the corresponding thread ID.
[0084] In yet another embodiment, the target question sent by the user may include a related question composed of several sub-questions. When the related question sent by the user is received, the response text generated by ChatGPT will also include the response content corresponding to each sub-question.
[0085] Exemplarily, after receiving the target question sent by the user and concatenating it into the following information:
[0086] List three related questions based on what I sent. Specific requirements are as follows.
[0087] The language of the question must match what I sent.
[0088] What I sent is [questions]
[0089] The generated questions must be highly relevant. There is no need to list keywords and summary content.
[0090] Limit each relevant question to 20 words or less. Three questions must be generated.
[0091] Send the instruction to run the conversation thread containing the above information to ChatGPT so that ChatGPT generates response texts for answering the above three questions. At this time, store these three questions in the database and synchronously send them to the front end, and the user can see a set of complete answers through the front end.
[0092] Please refer to Figure 2 , another embodiment of the present application also provides an intelligent conversation device based on a document, including a creation module 101 and a conversation module 102.
[0093] The creation module 101 is used to obtain the target document and send it to the large language model, and receive the returned file ID; use the endpoint to create an assistant associated with the file ID and obtain the corresponding assistant ID; send the setting instruction containing the file ID to the large language model to generate a conversation thread based on the target document and the corresponding thread ID.
[0094] The dialogue module 102 is used to create an execution command containing the target question in the dialogue thread when receiving the target question sent by the user; send an instruction to run the dialogue thread to the large language model, so that the large language model runs the execution command and generates a response text.
[0095] Regarding the information interaction and execution process among the modules in the above device, since they are based on the same concept as the method embodiments of the present application, the specific content can be referred to the description in the method embodiments of the present application, and will not be elaborated here.
[0096] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for intelligent dialogue based on documents as described above is implemented.
[0097] Those of ordinary skill in the art can understand that all or part of the processes in the above method embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above method embodiments. Among them, the storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM), or a random access memory (RAM), etc.
[0098] The above is the preferred implementation manner of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present application.
Claims
1. A document-based intelligent dialogue method, characterized in that: include: Get the target document and send it to the large language model, and receive the document ID in return; Use the endpoint to create an assistant associated with the file ID and obtain the corresponding assistant ID; Sending a setting instruction including the document ID to the large language model to generate a conversation thread based on the target document and a corresponding thread ID; When receiving a target question sent by a user, creating an execution command containing the target question in the conversation thread, wherein the target question sent by the user includes an associated question consisting of a plurality of sub-questions; sending an instruction to run the conversation thread to the large language model, so that the large language model runs the execution command and generates a response text, wherein when receiving the associated question sent by the user, the response text generated by the large language model includes the response content corresponding to each of the sub-questions; When receiving the target question sent by the user, creating an execution command containing the target question in the conversation thread specifically includes: Obtain the target question and the corresponding file ID; Match the corresponding thread ID and assistant ID according to the corresponding file ID, and create an execution command containing the target question in the dialogue thread corresponding to the thread ID, wherein when the instruction to run the dialogue thread is sent to the large language model, the assistant corresponding to the assistant ID is used to synchronously transmit the target question.
2. The document-based intelligent dialogue method according to claim 1, characterized in that: After sending the instruction to run the conversation thread to the large language model, the method further includes: Sending a request for an answer to the large language model; The response text generated by the large language model is received.
3. The document-based intelligent dialogue method according to claim 1, characterized in that: After obtaining the target document, the method further includes: The target document is uploaded to a cloud storage platform for storage.
4. The document-based intelligent dialogue method according to claim 1, characterized in that: The obtaining of the target document comprises: receiving the target document uploaded by the user, or; A URL uploaded by a user is received, wherein the URL includes the target document.
5. A document-based intelligent dialogue device, characterized in that: include: Create a module to get the target document and send it to the large language model, and receive the document ID in return; Using the endpoint to create an assistant associated with the document ID and obtain the corresponding assistant ID; sending a setting instruction containing the document ID to the large language model to generate a conversation thread based on the target document and a corresponding thread ID; The dialogue module is configured to, when receiving a target question sent by a user, create an execution command containing the target question in the dialogue thread, wherein the target question sent by the user includes an associated question consisting of a plurality of sub-questions; send an instruction to run the dialogue thread to the large language model, so that the large language model runs the execution command and generates a response text, wherein when receiving the associated question sent by the user, the response text generated by the large language model includes the response content corresponding to each of the sub-questions; Wherein, the dialogue module is also used to obtain the target question and the corresponding file ID; Match the corresponding thread ID and assistant ID according to the corresponding file ID, and create an execution command containing the target question in the dialogue thread corresponding to the thread ID, wherein when the instruction to run the dialogue thread is sent to the large language model, the assistant corresponding to the assistant ID is used to synchronously transmit the target question.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the document-based intelligent dialogue method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Automatic conversation techniques
CN102067167A
Railway industry intelligent question and answer assistant system
CN117633179A