Method and system for generating queries for querying a neural network language model
The method automates query generation for neural network language models by integrating search results from contextually relevant documents, enhancing response relevance through sequential user and system query formation.
Patent Information
- Application Number
- PCT/RU2024/000164
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-24
- Filing Date
- 2024-05-17
- Publication Date
- 2025-10-30
AI Technical Summary
Existing solutions lack automation in generating queries for accessing neural network language models using the results of information search from the context of provided or found documents, leading to reduced relevance of responses.
A method and system for generating queries that involve receiving a user request, processing it with a search engine, forming a system part of the request, sequentially adding user parts based on search results, and sending these to a neural network language model to generate responses, which are then combined into a final query, enhancing relevance.
This approach automates query generation and significantly increases the relevance of responses from the neural network language model by leveraging information from contextually relevant documents.
Smart Images

Figure RU2024000164_30102025_PF_FP_ABST
Abstract
Description
[0001] METHOD AND SYSTEM FOR GENERATING QUERYS TO CONTACT A NEURAL NETWORK LANGUAGE MODEL.
[0002] AREA OF TECHNOLOGY
[0003] [1] This technical solution generally relates to methods of processing data with the preparation of queries adapted to the needs of the user, namely to a method and system for generating queries for accessing a neural network language model.
[0004] LEVEL OF TECHNOLOGY
[0005] [2] A language model is a probability distribution over sequences of words. Language models generate probabilities by training on a corpus of texts in one or more languages. Given that languages can be used to express a vast number of valid sentences (the so-called digital infinity), language modeling faces the challenge of assigning non-zero probabilities to linguistically valid sequences that may never appear in the training data. To overcome this problem, several modeling approaches have been developed, such as the use of Markov chains or the use of neural architectures such as recurrent neural networks or transformers.
[0006] [3] Language models are useful for a wide range of computational linguistics problems, from initial applications in speech recognition to avoid generating meaningless (i.e., unlikely) word sequences, to more widespread use in machine translation (e.g., evaluating candidate translations), natural language generation (generating text that is more human-like), part-of-speech tagging, syntactic analysis, optical character recognition, handwriting recognition, grammar inference, information retrieval, and other applications.
[0007] [4] One application of language models is chatbots. For example, ChatGPT is a chatbot with generative artificial intelligence developed by OpenAI that can operate in a conversational mode and supports queries in natural languages. The system can answer questions and generate texts in various languages, including Russian, covering various subject areas.
[0008] [5] When working with such services, it's essential to send accurate requests so that the chatbot can accurately evaluate all input and provide the most relevant response. One solution to this problem is to enrich the request by attaching relevant documents related to the topic of the request and automating this process.
[0009] [6] Known solutions include DocuChat, a chatbot that learns from client data and documents and provides human-readable answers to user questions based on information from these documents. Also known is ChatDOC, a chatbot that generates brief descriptions of documents in uploaded files, online files, and websites. Documents can be accessed with specific queries (the service offers suggestions).
[0010] [7] The disadvantages of the solutions known from the prior art include the lack of automation of query generation using the results of information search from the context of the provided or found documents. DISCLOSURE OF THE INVENTION
[0011] [8] This technical solution is aimed at eliminating the shortcomings inherent in existing solutions known from the prior art.
[0012] [9] The technical problem solved in this technical solution is the lack of automation of query generation for accessing the neural network language model, using the results of information search from the context of the provided or found documents.
[0013]
[0010] The main technical result that emerges from solving the above-mentioned problem is the provision of automation of the generation of queries for accessing the neural network language model, using the results of searching for information from the context of the documents provided or found.
[0014]
[0011] An additional technical result that appears when solving the above-mentioned problem is an increase in the relevance of the responses of the neural network language model when sending generated requests.
[0015]
[0012] The specified technical results are achieved by implementing a method for generating queries for accessing a neural network language model, implemented using a processor and a data storage device, including the following steps:
[0016] • receive a user request to access the neural network language model and send it to the search engine for processing;
[0017] • the relevant objects found as a result of the search engine processing and the initial user request are used to generate requests to access the neural network language model, whereby: o the system part of the request is formed and sent to the neural network language model, indicating the conditions for fulfilling the initial user request; z then, based on the relevant objects found as a result of the search engine processing, the user parts of the request are sequentially formed and sent to the neural network language model, whereby the neural network language model sequentially forms a response to each user part of the request, which is taken into account with each subsequent formation and sending of the user part of the request;o generate and send to the neural network language model the final iteration of the query to access the neural network language model, which contains the system part of the query, all previously generated user parts of the query and the responses to them of the neural network language model.
[0018]
[0013] In one particular example of the implementation of the method, the responses from the neural network language model are subject to post-processing.
[0019]
[0014] In another particular example of the implementation of the method, the search system used is an external system with access to it through a hardware-software interface.
[0020]
[0015] Furthermore, the claimed technical result is achieved through the operation of a query generation system for accessing a neural network language model, comprising: at least one data processing device; at least one data storage device; at least one program, where one or more programs are stored on one or more data storage devices and executed on one or more data processing devices, wherein one or more programs ensure the execution of the following steps: • receive a user request for accessing a neural network language model and send it for processing to a search engine;
[0021] • the relevant objects found as a result of the search engine processing and the initial user request are used to generate requests to access the neural network language model, whereby: o the system part of the request is formed and sent to the neural network language model, indicating the conditions for fulfilling the initial user request; o then, based on the relevant objects found as a result of the search engine processing, the user parts of the request are sequentially formed and sent to the neural network language model, whereby the neural network language model sequentially forms a response to each user part of the request, which is taken into account with each subsequent formation and sending of the user part of the request;o generate and send to the neural network language model the final iteration of the query to access the neural network language model, which contains the system part of the query, all previously generated user parts of the query and the responses to them of the neural network language model.
[0022]
[0016] In one particular example of the system implementation, responses from the neural network language model are subject to post-processing.
[0023]
[0017] In another particular example of the implementation of the system, the search system used is an external system with access to it through a hardware-software interface. BRIEF DESCRIPTION OF THE DRAWINGS
[0024]
[0018] The features and advantages of the present technical solution will become apparent from the following detailed description and the accompanying drawings, in which:
[0019] Fig. 1 illustrates a block diagram of the implementation of the claimed method.
[0025]
[0020] Fig. 2 illustrates an example of information found during a search.
[0026]
[0021] Fig. 3 illustrates an example of a response from a virtual assistant.
[0027]
[0022] Fig. 4 illustrates a high-level architectural diagram of the operation of the described particular embodiment of the claimed technical solution.
[0023] Fig. 5 illustrates a system for implementing the claimed method.
[0028] IMPLEMENTATION OF THE INVENTION
[0029]
[0024] Below, the terms and concepts necessary for the implementation of this technical solution will be described.
[0030]
[0025] A query (or prompt) is a question or statement presented to a neural network language model to generate a response. To create such a query, the user must define the topic, category, and structure of the question. The relevance of the output information depends on the correct formulation of the query, i.e., a valid query.
[0031]
[0026] A virtual assistant is a chatbot that can interact with a user on any chat platform (telegram, jivo, edna, etc.), and can also include a search model and a query generator for a neural network language model.
[0032]
[0027] The claimed technical solution may be implemented, for example, by a system, a machine-readable medium, a server, etc. In this technical solution, the term “system” means, among other things, a computer system, a computer (electronic computer), a CNC (computer numerical control), a PLC (programmable logic controller), computerized control systems, and any other devices capable of performing a given, clearly defined sequence of operations (actions, instructions).
[0033]
[0028] A command processing unit is an electronic unit or integrated circuit (microprocessor) that executes machine instructions (programs).
[0034]
[0029] The command processing unit reads and executes machine instructions (programs) from one or more data storage devices, such as random access memory (RAM) and / or read-only memory (ROM). ROM may include, but is not limited to, hard disk drives (HDD), flash memory, solid-state drives (SSD), optical storage media (CD, DVD, BD, MD, etc.), etc.
[0035]
[0030] A program is a sequence of instructions intended for execution by a computer control device or a command processing device.
[0036]
[0031] The term "instructions" as used in this application may generally refer to software instructions or software commands that are written in a given programming language to perform a specific function, such as, for example, receiving and processing data, generating a user profile, receiving and transmitting signals, analyzing received data, identifying a user, etc. The instructions may be implemented in a variety of ways, including, for example, object-oriented methods. For example, the instructions may be implemented using the C++ programming language, Java, Python, various libraries (e.g., Microsoft Foundation Classes), etc. The instructions that perform the processes described in this solution may be transmitted via both wired and wireless data transmission channels, for example, Wi-Fi, Bluetooth, USB, WLAN, LAN, etc.
[0037]
[0032] The presented method for generating queries for accessing a neural network language model (Fig. 1 shows a diagram of the method) solves the problem of ensuring the automation of generating queries for accessing a neural network language model, using the results of searching for information from the context of provided or found documents, and increasing the relevance of the responses of the neural network language model when sending generated queries by sequentially performing the following steps:
[0038] • receive a user request for accessing the neural network language model and send it to the search engine for processing; • the relevant objects found as a result of search engine processing and the initial user request are used to generate requests for accessing the neural network language model, whereby: o the system part of the request is formed and sent to the neural network language model, indicating the conditions for fulfilling the initial user request; o then, based on the relevant objects found as a result of search engine processing, the user parts of the request are sequentially formed and sent to the neural network language model, whereby the neural network language model sequentially forms a response to each user part of the request, which is taken into account with each subsequent formation and sending of the user part of the request;o generate and send to the neural network language model the final iteration of the query to access the neural network language model, which contains the system part of the query, all previously generated user parts of the query and the responses to them from the neural network language model.
[0039]
[0033] Query generation for the neural network language model is initiated after the user asks a question, and a separate search function finds several documents containing the desired information. To initiate query generation for the neural network language model, the client's question (in text form), as well as documents or document fragments, are passed as input.
[0040]
[0034] In a particular example of the implementation of the claimed technical solution, the search system is a separate software module consisting of: • an interface for uploading documents and a database for storing documents, with the help of which the client can upload his documents (txt, pdf, doc, etc.) or indicate the site on which the documents are stored for their further parsing;
[0041] • a module for converting documents into vector format, with the help of which documents are converted into a numerical format, which will subsequently be required to search the content of documents;
[0042] • a module for converting a question into a vector format, with the help of which the question is converted into a numerical format, which will subsequently be required to search for an answer based on the content of documents;
[0043] • a module for searching documents by the topic and meaning of the question, with the help of which the search for the question vector is carried out by the document vectors;
[0044] • a search engine orchestrator, which initiates the process and calls all systems in turn, including the query generator.
[0045]
[0035] After the client uploads documents to a dedicated interface, they are converted to vector format. When a question is received, a search is performed, specifically, the question vector is compared with the vectors of all uploaded documents. During the search, the most relevant documents, based on the topic and meaning of the question, are selected from among those previously uploaded.
[0046]
[0036] In another particular example of the implementation of the claimed technical solution, the search system used is an external system with access to it through a hardware-software interface.
[0047]
[0037] The incoming question and the retrieved documents in their original text form are sent to the query generator for the neural network language model.
[0038] When launched, the generator creates a sequence of queries consisting of system and user queries.
[0048] Example of a system request:
[0049] "Answer like an airline customer service representative."
[0050] Example of a user query:
[0051] "How to transport a cat?"
[0052]
[0039] First of all, the generator forms and sends a system request to the neural network language model, which indicates that the neural network language model plays a certain contextual role in communication with the client, and also indicates the task that must be performed - to answer the question using documents and other additional response conditions.
[0053]
[0040] After the system request, user requests are sent. Each user request contains a question and the next document fragment. A user request is sent for each document fragment that was previously found.
[0054]
[0041] Starting from the second user query, the neural network language model must take into account the response to the previous user query, thereby achieving a summation of all information from previous documents.
[0055]
[0042] The last user query contains information on the responses to all previous user queries. The last response from the neural network language model takes into account information on all document fragments, as well as summarized information from the neural network language model on the client's question. Thus, the last response from the neural network language model is the final answer to the user's question. The final response from the neural network language model can be sent for post-processing, after which it can be sent to the end user or directly to the end user.
[0043] Post-processing of responses from the neural network language model can include: removal of extra spaces, empty answers, unnecessary characters or system messages, errors, etc., meaning checking, censoring.
[0056]
[0044] As an example of the operation of a particular embodiment of the claimed technical solution, a variant of working with the Gigachat neural network language model is given (to implement the functionality, GigaChain was used - a Python library that allows for simplifying and automating work with the Gigachat neural network language model and other large language models (LLM). GigaChain is a version of the LangChain library that is adapted for working with the Russian language).
[0057] In the example described:
[0058] A virtual assistant is a chatbot that can interact with users on any chat platform (Telegram, Jivo, Edna, etc.), and also includes a search model and a Gigachat query generator.
[0059] The client is a legal entity that uses a virtual assistant for commercial purposes (for example, an airline).
[0060] The virtual assistant is trained using the Client's documents, for example, 1,000 of the Client's regulatory documents on how cargo transportation is carried out are loaded.
[0061] User - an individual who has contacted the Client's Virtual Assistant with a question, the answer to which is contained in the Client's documents.
[0062]
[0045] The work process is as follows:
[0063] Prerequisite: The Client (legal entity) has uploaded a database of their documents (for example, 1000 documents in .pdf format) to our Virtual Assistant.
[0064] Process initiation: The user asks the Virtual Assistant a question. The process executed:
[0065] 1. The search engine searches for documents among the uploaded ones that clearly mention the essence of the question or contain a direct answer to the question.
[0066] After the client uploads documents to a dedicated interface, they are converted into vector format. When a question is received, a search is performed, or more precisely, the question vector is compared with the vectors of all uploaded documents. During the search, up to eight documents from the previously uploaded documents are selected that are most relevant to the topic and meaning of the question.
[0067] The received question and the found documents in their original text form are sent to the query generator.
[0068] 2. The query generator sends the User's question and the documents found in the previous step (8 out of 1000) to Gigachat:
[0069] 1) The query generator is launched after the user asks a question, and a separate search function will find multiple documents containing the desired information. To launch the query generator, the client's question (in text form) and documents or document fragments (also in text form) are passed as input.
[0070] 2) When launched, the generator creates a sequence of queries consisting of system and user queries
[0071] 3) First, the generator creates and sends a system request to Gigachat, which specifies that Gigachat will act as the airline's virtual assistant in communicating with the client, and also specifies the task that must be completed—answering the question using documents, and other additional response conditions.
[0072] 4) After the system request, user requests are sent. Each user request contains a question and a subsequent document fragment. A user request is sent for each document fragment previously found (in step 1). Currently, the maximum number of document fragments is 8, so 8 user requests are sent to Gigachat.
[0073] 5) Starting from the second user request, Gigachat must take into account the response to the previous user request, thereby achieving a summation of all information from previous documents.
[0074] 6) The most recent user request contains information on responses to all previous user requests. The most recent response from Gigachat includes information on all document fragments, as well as Gigachat's summary of the client's question.
[0075] 7) Thus, the last answer from Gigachat is the final answer to the user's question.
[0076] 8) The final response from Gigachat is sent for post-processing, and then to the end user.
[0077] 3. The final response from Gigachat is sent to the end user in the same channel in which the user asked their question.
[0078]
[0046] Example:
[0079] A user asked the airline's Virtual Assistant, "How much does it cost to transport a cat?", in a chat on the airline's website.
[0080] The virtual assistant found several documents, one of which contains information (shown in Fig. 2).
[0081] The virtual assistant used the request generator to send documents and a question to Gigachat (following the process described above).
[0082] Ultimately, the user received a response via the airline's website chat (shown in Fig. 3). Fig. 4 shows a high-level architectural diagram of the operation of the described specific implementation of the claimed technical solution.
[0083]
[0047] In general (see Fig. 5), the system for generating queries for accessing the neural network language model (500) contains one or more processors (501), memory means such as RAM (502) and ROM (503), and input / output interfaces (504), connected by a common information exchange bus.
[0084]
[0048] The processor (501) (or several processors, a multi-core processor, etc.) can be selected from a range of devices that are widely used at present, for example, from manufacturers such as: Intel™, AMD™, Apple™, Samsung Exynos™, MediaTEK™, Qualcomm Snapdragon™, etc. Under the processor or one of the processors used in the system (500), it is also necessary to take into account a graphic processor, for example, an NVIDIA GPU with a software model compatible with CUDA, or Graphcore, the type of which is also suitable for the full or partial implementation of the method, and can also be used for training and applying machine learning models in various information systems.
[0085]
[0049] RAM (502) is a random access memory and is intended for storing machine-readable instructions executable by the processor (501) for performing the necessary operations for logical data processing. RAM (502), as a rule, contains executable instructions of the operating system and the corresponding software components (applications, software modules, etc.). In this case, the available memory capacity of a graphics card or graphics processor may serve as RAM (502).
[0086]
[0050] ROM (503) represents one or more permanent data storage devices, such as a hard disk drive (HDD), a solid state drive (SSD), flash memory (EEPROM, NAND, etc.), optical storage media (CD-R / RW, DVD-R / RW, BlueRay Disc, MD), etc.
[0087]
[0051] To organize the operation of the components of the device (500) and to organize the operation of external connected devices, various types of I / O interfaces (504) are used. The selection of the corresponding interfaces depends on the specific design of the computing device, which may be, but are not limited to: PCI, AGP, PS / 2, IrDa, FireWire, LPT, COM, SATA, IDE, Lightning, USB (2.0, 3.0, 3.1, micro, mini, type C), TRS / Audio jack (2.5, 3.5, 6.35), HDMI, DVI, VGA, Display Port, RJ45, RS232, etc.
[0088]
[0052] To ensure user interaction with the device (500), various I / O information means (505) are used, for example, a keyboard, a display (monitor), a touch display, a touchpad, a joystick, a mouse, a light pen, a stylus, a touch panel, a trackball, speakers, a microphone, augmented reality means, optical sensors, a tablet, light indicators, a projector, a camera, biometric identification means (a retinal scanner, a fingerprint scanner, a voice recognition module), etc.
[0089]
[0053] The network interaction means (506) ensures the transmission of data via an internal or external computer network, for example, an Intranet, the Internet, a LAN, etc. One or more means (506) may be, but are not limited to: an Ethernet card, a GSM modem, a GPRS modem, an LTE modem, a 5G modem, a satellite communication module, an NFC module, a Bluetooth and / or BLE module, a Wi-Fi module, etc.
[0090]
[0054] The specific selection of device elements (500) for implementing various hardware and software architectural solutions may vary while maintaining the required functionality. In particular, such an implementation may be accomplished using electronic components used to create digital integrated circuits. This includes, but is not limited to, microcircuits whose operating logic is determined during manufacture, or programmable logic integrated circuits (FPGAs), whose operating logic is specified through programming. Programmers and debugging environments are used for programming, allowing the desired structure of the digital device to be specified in the form of a circuit diagram or a program in specialized hardware description languages: Verilog, VHDL, AHDL, etc.Alternatives to FPGAs include programmable logic controllers (PLCs), basic matrix chips (BMCs), which require a factory production process for programming, and ASICs (specialized custom large-scale integrated circuits), which are significantly more expensive for small-scale and single-unit production. Therefore, implementation can be achieved using standard tools based on classical principles of computing.
[0091]
[0055] The submitted application materials disclose preferred examples of the implementation of the technical solution and should not be interpreted as limiting other, particular examples of its implementation that do not go beyond the scope of the requested legal protection, which are obvious to specialists in the relevant field of technology.
Claims
FORMULA 1. A method for generating queries to access a neural network language model, implemented using a processor and a data storage device, comprising the following steps: • receive a user request to access the neural network language model and send it to the search engine for processing; • the relevant objects found as a result of the search engine processing and the initial user request are used to generate requests to access the neural network language model, whereby: o the system part of the request is formed and sent to the neural network language model, indicating the conditions for fulfilling the initial user request; o then, based on the relevant objects found as a result of the search engine processing, the user parts of the request are sequentially formed and sent to the neural network language model, whereby the neural network language model sequentially forms a response to each user part of the request, which is taken into account with each subsequent formation and sending of the user part of the request;o generate and send to the neural network language model the final iteration of the query to access the neural network language model, which contains the system part of the query, all previously generated user parts of the query and the responses to them of the neural network language model.
2. A method for generating queries for accessing a neural network language model according to paragraph 1, characterized in that the responses from the neural network language model are subject to post-processing.
3. A method for generating queries for accessing a neural network language model according to paragraph 1, characterized in that the search system used is an external system with access to it through a hardware-software interface.
4. A query generation system for accessing a neural network language model, comprising: at least one data processing device; at least one data storage device; at least one program, where one or more programs are stored on one or more data storage devices and executed on one or more data processing devices, wherein the one or more programs ensure the execution of the following steps: • receive a user request to access the neural network language model and send it to the search engine for processing; • the relevant objects found as a result of the search engine processing and the user's initial query are used to generate queries to access the neural network language model, whereby: o the system part of the query is generated and sent to the neural network language model, indicating the conditions for executing the user's initial query; o then, based on the relevant objects found as a result of the search engine processing, the queries are sequentially generated and sent to the neural network language model The model generates user parts of the query, whereby the neural network language model sequentially generates a response for each user part of the query, which is taken into account with each subsequent generation and sending of the user part of the query; and generates and sends to the neural network language model the final iteration of the query to access the neural network language model, which contains the system part of the query, all previously generated user parts of the query, and the responses of the neural network language model to them.
5. A query generation system for accessing a neural network language model according to paragraph 4, characterized in that the responses from the neural network language model are subject to post-processing.
6. A system for generating queries for accessing the neural network language model according to paragraph 4, characterized in that the search system used is an external system with access to it through a hardware-software interface.
Citation Information
Patent Citations
Method of creating model for analysing dialogues based on artificial intelligence for processing user requests and system using such model
EA038264B1
Interactive conversation assistance using semantic search and generative AI
US11960514B1
Systems and methods for real-time search based generative artificial intelligence
US20240020538A1
Conversational document question answering
US20240126795A1