Token optimization with minimized payload for large language models
By using a minimized intermediate file format to process the input and output of LLM, the problem of limiting the number of tokens during text generation is solved, and more efficient and accurate text generation is achieved.
Patent Information
- Application Number
- CN202311737802.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-10
- Filing Date
- 2023-12-15
- Publication Date
- 2025-05-13
AI Technical Summary
Large language models limit the number of input and output tokens when text is generated, and limiting the number of tokens is technically challenging without sacrificing accuracy.
The intermediate file format with minimized attributes, such as JSON, is used to process the input and output of LLM, by converting natural language text into a minimized intermediate file format, reducing the number of tokens, and parsing the generated text using a dedicated parser.
By reducing the number of tokens, the cost and generation time of LLM is reduced, while improving the speed and accuracy of text generation, avoiding performance problems caused by the limit on the number of tokens.
Smart Images

Figure CN119990061A_ABST
Abstract
Description
Technical Field
[0001] This document generally relates to computer systems and more particularly to the use of large language models. Background Art
[0002] Large Language Models (LLMs) refer to artificial intelligence (AI) systems that have been trained on extensive datasets to understand and generate human language. These models are designed to process and understand natural language in a way that allows them to answer questions, engage in conversations, generate text, and perform a variety of language-related tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] The present disclosure is illustrated by way of example, and not by way of limitation, in the figures of the accompanying drawings in which like references indicate similar elements.
[0004] Figure 1 is a block diagram illustrating a system for automatically generating text from natural language text according to an example embodiment.
[0005] Figure 2 is a flow chart illustrating a method for automatically generating text from natural language text according to an example embodiment.
[0006] Figure 3 is a block diagram illustrating the architecture of software that may be installed on any one or more of the above-described devices.
[0007] Figure 4 A diagrammatic representation of a machine in the form of a computer system according to an example embodiment is shown within which sets of instructions may be executed to cause the machine to perform any one or more of the methodologies discussed herein. DETAILED DESCRIPTION
[0008] The following description discusses illustrative systems, methods, techniques, instruction sequences, and computing machine program products. In the following description, for the purpose of explanation, many specific details are set forth in order to provide an understanding of various example embodiments of the present subject matter. However, it will be apparent to those skilled in the art that various example embodiments of the present subject matter may be practiced without these specific details.
[0009] When using an LLM to generate text, the amount of input tokens fed to the LLM and the amount of output tokens generated by the LLM can affect both cost and performance. On the input side, tokens are often used as input to the LLM to help provide context for the LLM to generate text. If the user provides more guidance, such as by specifying the expected data structure with examples or specifications, the results are more robust.
[0010] LLMs have an absolute limit on the number of input tokens they will accept (e.g., no more than 500 tokens). As a result, it is beneficial to limit the number of input tokens passed to the LLM when requesting text generation. However, it can be technically challenging to accomplish this without sacrificing the accuracy provided by the additional context of those additional tokens.
[0011] Similar issues arise with the tokens output by the LLM. The more tokens the LLM produces, the longer the generation process takes, so there is usually an upper limit or toll (in speed or cost), or both, on large numbers of output tokens, just as there is with input tokens. As a result, when performing text generation, it can also be beneficial to limit the number of output tokens generated by the LLM. This can also be technically challenging to accomplish without sacrificing the accuracy that occurs with the generation of the extra context for those additional tokens.
[0012] In an example embodiment, a solution is provided by using an intermediate file format for input prompts to the LLM or output from the LLM (or both). The intermediate file format is a file format with a minimizeable attribute, which means that the tokens contained in the file in the intermediate file format can be stripped or otherwise removed without changing the semantic meaning of the file. An example of such an intermediate file format is Javascript Object Notation (JSON), but there are other formats, such as Extensible Markup Language (XML) and Yet Another Markup Language (YAML). Input prompts can be created in or converted to the intermediate file format and then minimized before being sent to the LLM for text generation. In addition, the system message included in the input prompt can guide the LLM to generate text in the intermediate file format in a minimized form. A dedicated parser can then be included to parse the minimized output generated by the LLM.
[0013] LLMs for generating information are often referred to as Generative Artificial Intelligence (GAI) models. GAI models can be implemented as Generative Pre-trained Transformer (GPT) models or bidirectional encoders. The GPT model is a machine learning model that uses a transformer architecture, which is a deep neural network that excels at processing sequence data such as natural language.
[0014] A bidirectional encoder is a neural network architecture in which the input sequence is processed in two directions: forward and backward. The forward direction starts at the beginning of the sequence and processes one input token at a time, while the backward direction starts at the end of the sequence and processes the input in reverse order.
[0015] By processing the input sequence in two directions, the bidirectional encoder can capture more contextual information and dependencies between words, leading to better performance.
[0016] The bidirectional encoder can be implemented as a Bidirectional Long Short-Term Memory (BiLSTM) or a Bidirectional Encoder Representations from Transformers (BERT) model.
[0017] Each direction has its own hidden state, and the final output is the combination of the two hidden states.
[0018] Long Short-Term Memory (LSTM) is a recurrent neural network (RNN) designed to overcome the vanishing gradient problem in traditional RNNs that can make it difficult to learn long-term dependencies in sequence data.
[0019] LSTM includes a cell state, which acts as a memory to store information over time. The cell state is controlled by three gates: input gate, forget gate, and output gate. The input gate determines how much new information is added to the cell state, while the forget gate determines how much old information is discarded. The output gate determines how much of the cell state is used to calculate the output. Each gate is controlled by a sigmoid activation function, which outputs a value between 0 and 1 that determines the amount of information that passes through the gate.
[0020] In BiLSTM, there are separate LSTMs for the forward direction and the backward direction. At each timestep, the forward and backward LSTM units receive the current input token and the hidden state from the previous timestep. The forward LSTM processes the input tokens from left to right, while the backward LSTM processes them from right to left.
[0021] The output of each LSTM cell at each time step is a combination of the input token and the previous hidden state, which allows the model to capture both short-term and long-term dependencies between input tokens.
[0022] BERT applies bidirectional training of models called transformers to language modeling. This is in contrast to prior art solutions, which look at text sequences from left to right or a combination of left to right and right to left. Bidirectionally trained language models have a deeper sense of language context and speech flow than unidirectional language models.
[0023] More specifically, the Transformer encoder reads the entire sequence of information at once and is therefore considered bidirectional (although one could argue that it is actually non-directional). This property allows the model to learn the context of a piece of information based on all of its surroundings.
[0024] In other example embodiments, a Generative Adversarial Network (GAN) embodiment may be used. A GAN is a supervised machine learning model with two sub-models: a generator model, which is trained to generate new examples, and a discriminator model, which attempts to classify examples as real or generated. The two models are trained together in an adversarial manner (using a zero-sum game according to game theory) until the discriminator model is fooled about half of the time, meaning that the generator model is generating reasonable examples.
[0025] The generator model takes as input a random vector of fixed length and generates samples in the domain in question. The vectors are randomly drawn according to a Gaussian distribution, and the vectors are used to seed the generation process. After training, the points in this multidimensional vector space will correspond to points in the problem domain, forming a compressed representation of the data distribution. This vector space is called a latent space, or a vector space composed of latent variables. Latent variables or hidden variables are those variables that are important to the domain but cannot be directly observed.
[0026] The discriminator model takes as input examples from the domain (real or generated) and predicts a binary class label of real or fake (generated).
[0027] Generative modeling is an unsupervised learning problem, although a clever property of the GAN architecture is that the training of the generative model is framed as a supervised learning problem.
[0028] The two models, the generator and the discriminator, are trained together. The generator generates batches of samples and these samples along with real examples from the domain are fed to the discriminator and classified as real or fake.
[0029] The discriminator is then updated to better discriminate between real and fake samples in the next round, and importantly, the generator is updated based on how well the generated samples fool the discriminator.
[0030] In another example embodiment, the GAI model is a variational autoencoder (VAE) model. A VAE includes an encoder network that compresses input data into a low-dimensional representation called a latent code, and a decoder network that generates new data from the latent code. In either case, the GAI model includes a generative classifier, which can be implemented as, for example, a naive Bayes classifier.
[0031] The present solution works with any type of GAI model, although an implementation will be described specifically for use with a GPT model.
[0032] Figure 1 1 is a block diagram illustrating a system 100 for automatically generating text from natural language text according to an example embodiment. Here, application 102 provides a mechanism that allows users 104A and 104B (referred to herein as users 104) to request automatic creation of text by providing natural language text requests to application 102. In some example embodiments, the text requested to be created may be computer code, but is not necessarily limited thereto.
[0033] Along with the natural language text request, the application 102 may also receive context information related to the natural language text request. The context information may be provided as part of the natural language text request, as separate information from the user 104 (in the same communication as the natural language text request or in a different communication from the natural language text request), or accessed or otherwise obtained by the application 102 itself.
[0034] Application 102 includes prompt creator 106. Prompt creator 106 is used to obtain natural language text request and optionally obtain context information (if provided separately from natural language text request) and create LLM prompt according to them. This can include attaching system message to natural language text request and optionally attaching to context information. In an exemplary embodiment, whether it comes from natural language text request, context data or system message (or any combination thereof), some part of the prompt can be created or converted to intermediate file format in minimized form in intermediate file format. Thus, for example, the system message can include a part written in intermediate file format in minimized form (e.g., spaces, slashes and carriage returns are removed). Alternatively, context information can be converted into minimized form of intermediate file format. Alternatively, natural language text request can be converted into intermediate file format in some non-minimized form (e.g., spaces and carriage returns are present), and then the file can be converted into minimized form.
[0035] Regardless of how it is performed, some portion of the prompt is written in a minimized form in an intermediate file format. Thus, when the prompt is then passed to LLM 108 for processing, the input tokens to the LLM have already been minimized, which increases the speed of text generation by the LLM and may also reduce the cost of text generation in addition to increasing the likelihood that the number of tokens remains under any hard limits imposed by the LLM.
[0036] In addition, the system 100 is designed to minimize the output tokens generated by the LLM 108. To accomplish this, system messages may include instructions that direct the LLM 108 to return its output in an intermediate file format in a minimized form. In addition, in order to be able to understand this returned output, a dedicated parser 110 is provided that understands how to parse a file in the minimized form of the intermediate file format. The parser 110 then parses the returned text from the LLM based on this understanding, thereby producing a parsed output.
[0037] The following is an example of a system message that may be used for this embodiment:
[0038]
[0039]
[0040] You are an expert in domain modeling. Based on the following prompts, create JSON.
[0041] The JSON must conform to the following JSON schema definition:
[0042] \`\`\`json_schema
[0043] ${JSON.stringify(schema)}
[0044] \`\`\`
[0045] Don't output any text, just minified JSON, format:
[0046] \`\`\`json
[0047] {{generated JSON}}
[0048] \`\`\`
[0049] It cannot be formatted, everything is on one line, minimize token count`.trim()
[0050]
[0051] The following is an example of a file in the intermediate file format, but not in the minimized form. Note the presence of spaces, slashes, and carriage returns:
[0052]
[0053]
[0054]
[0055]
[0056] Here is an example of the above file after it has been transformed into a minimized form:
[0057] [{"id":"0001","type":"donut","name":"Cake","ppu":0.55,"batters":{"batter":[{"id":"10 01","type":"Regular"},{"id":"1002","type":"Chocolate"},{"id":"1003","type":"Blueberr y"},{"id":"1004","type":"Devil′sFood"}]},"topping":[{"id":"5001","type":"None"},{"id ":"5002","type":"Glazed"},{"id":"5005","type":"Sugar"},{"id":"5007","type":"Powdered
[0058] Sugar″},{″id″:″5006″,″type″:″Chocolate withSprinkles"},{"id":"5003","type":"Chocolate"},{"id":"5004","type":"Maple"}]},{"id":"0002","type":"donut","name":"Raised","ppu":0.55,"batters":{"batter":[ {″id″:″1001″,″type″:″Regular″}]},″topping″:[{″id″:″5001″,″type″:″None”},{″id”: "5002","type":"Glazed"},{"id":"5005","type":"Sugar"},{"id":"5003","type":"Choc olate"},{"id":"5004","type":"Maple"}]},{"id":"0003","type":"donut","name":"OldFashioned","ppu":0.55,"batters":{"batter":[{"id":"1001","type":"Regular"},{"id ″:″1002″,″type″:″Chocolate″}]},″topping″:[{″id″:″5001″,″type″:″None″},{″id″:”5 002","type":"Glazed"},{"id":"5003","type":"Chocolate"},{"id":"5004","type":"Ma ple"}]}]
[0059] It should be noted that the application 102 itself can be located on the user's device, or can be centrally located on a server with which the user interacts. For server-based embodiments, the application 102 can be run in the cloud, for example, on an application server, so that many different users can access it simultaneously.
[0060] Figure 2200 is a flowchart illustrating a method 200 for automatically generating text according to natural language text according to an example embodiment. At operation 202, a natural language text of a request for a text to be generated is received. The natural language text can be a text directly input by a user, wherein the user specifies what text the user wants to be generated using a speakable language (such as English). At operation 204, contextual information about the natural language text is accessed. The contextual information can be provided as part of a natural language text request, as separate information from the user (in the same communication as the natural language text request or in a communication different from the natural language text request), or accessed or otherwise obtained by the application itself. Contextual information can be any information that adds additional context to the natural language text, such as information obtained from a user profile, article, video, audio recording, usage data, etc.
[0061] At operation 206, a prompt is generated using the natural language text, the context information, and a system message. The system message includes instructions for generating the generated text in a minimized form of an intermediate file format, which is a format that enables portions of a file stored in the intermediate file format to be removed without changing the semantic meaning of the file.
[0062] At operation 208, the hint is passed to a large language model (LLM). At operation 210, the generated text in the intermediate file format in a minimized form is received from the LLM. At operation 212, the generated text is parsed using a parser configured to parse files in the intermediate file format, thereby producing parsed generated text.
[0063] Parsing a file involves interpreting the file's data following the rules of the file's format. In the context of software, parsing means breaking a block of data into smaller parts and interpreting it according to certain rules.
[0064] For example, when parsing an XML file, the parser identifies tags that mark the beginning and end of an element, the content of an element, and other structures. The parsing of the file is not affected by adding characters that have no semantic meaning to the file or removing characters that have no semantic meaning from the file.
[0065] In some example embodiments, the generated text is a compilable computer code.
[0066] In view of the above disclosure, various examples are set forth below. It should be noted that one or more features of the examples taken alone or in combination should be considered within the disclosure of the present application.
[0067] Example 1. A system comprising:
[0068] at least one hardware processor; and
[0069] A computer readable medium storing instructions that, when executed by at least one hardware processor, cause the at least one hardware processor to perform operations comprising:
[0070] receiving natural language text describing a request to generate text;
[0071] Access contextual information about natural language text;
[0072] generating a prompt using the natural language text, the context information, and a system message, the system message including instructions for producing the generated text in a minimized form of an intermediate file format, the intermediate file format being a format that enables portions of a file stored in the intermediate file format to be removed without changing the semantic meaning of the file;
[0073] Pass the hint to the Large Language Model (LLM);
[0074] receiving the generated text in a minimized form of an intermediate file format from the LLM; and
[0075] The generated text is parsed using a parser configured to parse a file in the intermediate file format, thereby producing parsed generated text.
[0076] Example 2. The system of Example 1, wherein the generated text is a compilable computer code.
[0077] Example 3. A system according to example 1 or example 2, wherein the prompt includes at least some text in the intermediate file format in a minimized form.
[0078] Example 4. A system according to Example 3, wherein the LLM has a hard limit on the number of input tokens it will process for a single prompt, and due to the presence of at least some text in the intermediate file format in a minimized form, the prompt includes fewer tokens than the hard limit.
[0079] Example 5. The system of any one of Examples 1-4, wherein the minimized form includes none.
[0080] Example 6. A system according to any of Examples 1-5, wherein the minimized form includes no carriage returns.
[0081] Example 7. A system according to any one of Examples 1-6, wherein the intermediate file format is JavaScript Object Notation (JSON).
[0082] Example 8. A method comprising:
[0083] receiving natural language text describing a request to generate text;
[0084] Access contextual information about natural language text;
[0085] generating a prompt using the natural language text, the context information, and a system message, the system message including instructions for producing the generated text in a minimized form of an intermediate file format, the intermediate file format being a format that enables portions of a file stored in the intermediate file format to be removed without changing the semantic meaning of the file;
[0086] Pass the hint to the Large Language Model (LLM);
[0087] receiving the generated text in a minimized form of an intermediate file format from the LLM; and
[0088] The generated text is parsed using a parser configured to parse a file in the intermediate file format, thereby producing parsed generated text.
[0089] Example 9. The method of Example 8, wherein the generated text is a compilable computer code.
[0090] Example 10. The method of example 8 or 9, wherein the hint includes at least some text in the intermediate file format in a minimized form.
[0091] Example 11. A method according to Example 10, wherein the LLM has a hard limit on the number of input tokens it will process for a single prompt, and due to the presence of at least some text in the intermediate file format in a minimized form, the prompt includes fewer tokens than the hard limit.
[0092] Example 12. The method of any one of Examples 8-11, wherein the minimized form includes no spaces.
[0093] Example 13. A method according to any one of Examples 8-12, wherein the minimized form includes no carriage returns.
[0094] Example 14. The method of any one of Examples 8-13, wherein the intermediate file format is JavaScript Object Notation (JSON).
[0095] Example 15. A non-transitory machine-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:
[0096] receiving natural language text describing a request to generate text;
[0097] Access contextual information about natural language text;
[0098] generating a prompt using the natural language text, the context information, and a system message, the system message including instructions for producing the generated text in a minimized form of an intermediate file format, the intermediate file format being a format that enables portions of a file stored in the intermediate file format to be removed without changing the semantic meaning of the file;
[0099] Pass the hint to the Large Language Model (LLM);
[0100] receiving the generated text in a minimized form of an intermediate file format from the LLM; and
[0101] The generated text is parsed using a parser configured to parse a file in the intermediate file format, thereby producing parsed generated text.
[0102] Example 16. The non-transitory machine-readable medium of Example 15, wherein the generated text is a compilable computer code.
[0103] Example 17. The non-transitory machine-readable medium of example 15 or 16, wherein the hint comprises at least some text of the intermediate file format in a minimized form.
[0104] Example 18. A non-transitory machine-readable medium according to Example 17, wherein the LLM has a hard limit on the number of input tokens it will process for a single prompt, and due to the presence of at least some text in the minimized form of the intermediate file format, the prompt includes fewer tokens than the hard limit.
[0105] Example 19. The non-transitory machine-readable medium of any one of Examples 15-18, wherein the minimized form includes no spaces.
[0106] Example 20. The non-transitory machine-readable medium of any one of Examples 15-19, wherein the minimized form includes no carriage returns.
[0107] Figure 3 is a block diagram 300 illustrating a software architecture 302 that may be installed on any one or more of the above-described devices. Figure 3 This is merely a non-limiting example of a software architecture, and it should be understood that many other architectures may be implemented to facilitate the functionality described herein. In various embodiments, the software architecture 302 is comprised of, for example, Figure 44, and an embodiment of the present invention is a hardware implementation of a machine 400, which includes a processor 410, a memory 430, and an input / output (I / O) component 450. In this example architecture, the software architecture 302 can be conceptualized as a stack of layers, where each layer can provide specific functionality. For example, the software architecture 302 includes layers such as an operating system 304, a library 306, a framework 308, and an application 310. In operation, consistent with some embodiments, the application 310 calls an API call 312 through the software stack and receives a message 314 in response to the API call 312.
[0108] In various embodiments, operating system 304 manages hardware resources and provides common services. Operating system 304 includes, for example, kernel 320, services 322, and drivers 324. Consistent with some embodiments, kernel 320 acts as an abstraction layer between hardware and other software layers. For example, kernel 320 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security settings. Services 322 can provide other common services to other software layers. According to some embodiments, drivers 324 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 324 may include display drivers, camera drivers, Bluetooth drivers, etc. or Bluetooth Low-power drivers, Flash memory drivers, Serial communication drivers (e.g., Universal Serial Bus (USB) drivers), Wi-Fi Drivers, audio drivers, power management drivers, etc.
[0109] In some embodiments, library 306 provides a low-level common infrastructure utilized by application 310. Library 306 may include system library 330 (eg, C standard library), which may provide functions such as memory allocation functions, string manipulation functions, math functions, and the like. In addition, the library 306 may include an API library 332, such as a media library (e.g., a library that supports the presentation and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG) or Portable Network Graphic (PNG)), a graphics library (e.g., an OpenGL framework for rendering in two dimensions (2D) and three dimensions (3D) in a graphics context on a display). A database library (e.g., SQLite that provides various relational database functions), a web library (e.g., WebKit that provides web browsing functions), etc. The library 306 may also include a variety of other libraries 334 to provide many other APIs to the application 310.
[0110] According to some embodiments, the framework 308 provides a high-level, general-purpose infrastructure that can be utilized by applications 310. For example, the framework 308 provides various GUI functions, advanced resource management, advanced location services, etc. The framework 308 can provide a wide range of other APIs that can be utilized by applications 310, some of which can be specific to a particular operating system 304 or platform.
[0111] In an example embodiment, applications 310 include a home application 350, a contacts application 352, a browser application 354, a book reader application 356, a location application 358, a media application 360, a messaging application 362, a game application 364, and a variety of other applications, such as third-party applications 366. According to some embodiments, applications 310 are programs that perform functions defined in the program. Various programming languages can be used to create one or more applications 310 structured in various ways, such as an object-oriented programming language (e.g., Objective-C, Java, or C++) or a procedural programming language (e.g., C or assembly language). In a specific example, third-party applications 366 (e.g., used by an entity other than the vendor of a particular platform using ANDROID TM or IOS TM Software Development Kit (SDK) can be used to develop applications on mobile operating systems such as IOS TM ANDROID TM , Mobile software running on the iPod Phone or another mobile operating system). In this example, third-party applications 366 can call API calls 312 provided by the operating system 304 to facilitate the functions described herein.
[0112] Figure 4 A diagrammatic representation of a machine 400 in the form of a computer system according to an example embodiment is shown, wherein a set of instructions may be executed to cause the machine 400 to perform any one or more of the methodologies discussed herein. In particular, Figure 4 A diagrammatic representation of a machine 400 in the example form of a computer system is shown in which instructions 416 (e.g., software, programs, applications, applet, app, or other executable code) may be executed to cause the machine 400 to perform any one or more of the methodologies discussed herein. For example, the instructions 416 may cause the machine 400 to perform Figure 2 Additionally or alternatively, instruction 416 may implement Figure 1-Figure 2Etc. Instructions 416 transform a general unprogrammed machine 400 into a specific machine 400 that is programmed to implement the described and illustrated functions in the described manner. In alternative embodiments, the machine 400 operates as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 400 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer (or distributed) network environment. The machine 400 may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (Set-Top Box STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smart phone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, web devices, network routers, network switches, bridges, or any machine capable of sequentially or otherwise executing instructions 416, which specifies the actions to be taken by the machine 400. Further, while a single machine 400 is illustrated, the term "machine" shall also be taken to include any collection of machines 400 that individually or jointly execute instructions 416 to perform any one or more of the methodologies discussed herein.
[0113] Machine 400 may include processor 410, memory 430, and I / O components 450, which may be configured to communicate with each other, such as via bus 402. In an example embodiment, processor 410 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processor 412 and processor 414 that may execute instructions 416. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that may execute instructions 416 simultaneously. Although Figure 4 Multiple processors 410 are shown, but the machine 400 may include a single processor 412 with a single core, a single processor 412 with multiple cores (e.g., a multi-core processor 412), multiple processors 412, 414 with a single core, multiple processors 412, 414 with multiple cores, or any combination thereof.
[0114] The memory 430 may include a main memory 432, a static memory 434, and a storage unit 436, each of which may be accessed by the processor 410, such as via the bus 402. The main memory 432, the static memory 434, and the storage unit 436 store instructions 416 that embody any one or more of the methodologies or functionality described herein. The instructions 416 may also reside, in whole or in part, within the main memory 432, within the static memory 434, within the storage unit 436, within at least one of the processors 410 (e.g., within a cache memory of a processor), or any suitable combination thereof during execution by the machine 400.
[0115] I / O components 450 may include a wide variety of components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurements, etc. The specific I / O components 450 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone will likely include a touch input device or other such input mechanism, while a headless server machine will likely not include such a touch input device. It should be understood that I / O components 450 may include Figure 4 4 and 5. Many other components not shown in the drawings. The I / O components 450 are grouped according to function only to simplify the following discussion, and the grouping is by no means limiting. In various example embodiments, the I / O components 450 may include output components 452 and input components 454. The output components 452 may include visual components (e.g., displays such as plasma display panels (PDPs), light emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), acoustic components (e.g., speakers), tactile components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. The input components 454 may include alphanumeric input components (e.g., keyboards, touch screens configured to receive alphanumeric input, photo-optical keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touch pads, trackballs, joysticks, motion sensors, or other pointing instruments), tactile input components (e.g., physical buttons, touch screens or other tactile input components that provide the location and / or force of touch or touch gestures), audio input components (e.g., microphones), etc.
[0116] In other example embodiments, the I / O component 450 may include a biometric component 456, a motion component 458, an environmental component 460, or a positioning component 462, as well as various other components. For example, the biometric component 456 may include a component for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition), etc. The motion component 458 may include an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. The environment component 460 may include, for example, an illumination sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas detection sensor that detects concentrations of hazardous gases for safety or measures pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. The positioning component 462 may include a position sensor component (e.g., a Global Positioning System (GPS) receiver component), an altitude sensor component (e.g., an altimeter or a barometer that detects air pressure from which altitude can be derived), an orientation sensor component (e.g., a magnetometer), and the like.
[0117] A variety of technologies may be used to implement communications. I / O components 450 may include a communication component 464 operable to couple machine 400 to network 480 or device 470 via coupling 482 and coupling 472, respectively. For example, communication component 464 may include a network interface component or other suitable device that interfaces with network 480. In other examples, communication component 464 may include a wired communication component, a wireless communication component, a cellular communication component, a Near Field Communication (NFC) component, a Bluetooth® communication component, or a wireless communication component. Components (e.g. Bluetooth Low power consumption), Wi-Fi Components and other communication components that provide communication via other modalities. Device 470 can be another machine or any of a variety of peripheral devices (eg, coupled via USB).
[0118] In addition, the communication component 464 can detect an identifier or include a component operable to detect an identifier. For example, the communication component 464 can include a Radio-Frequency Identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as QR codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying an audio signal of a tag). In addition, various information can be derived via the communication component 464, such as location via Internet Protocol (IP) geolocation, location via signal triangulation, Wi-Fi via detection of NFC beacon signals that can indicate a specific location, and location via the NFC beacon signal. Location, etc.
[0119] Various memories (e.g., 430, 432, 434 and / or memory in processor 410) and / or storage unit 436 may store one or more sets and data structures (e.g., software) of instructions 416 embodying or utilized by any one or more of the methods or functions described herein. When these instructions (e.g., instructions 416) are executed by processor 410, various operations are caused to implement the disclosed embodiments.
[0120] As used herein, the terms "machine storage medium", "device storage medium" and "computer storage medium" mean the same thing and are used interchangeably. These terms refer to a single or multiple storage devices and / or media (e.g., centralized or distributed databases and / or associated caches and servers) that store executable instructions and / or data. Therefore, these terms should be considered to include, but are not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media and / or device storage media include non-volatile memory, such as semiconductor memory devices, such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), field programmable gate array (FPGA) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium", "computer storage medium" and "device storage medium" specifically exclude carrier waves, modulated data signals and other such media, at least some of which are covered by the term "signal media" discussed below.
[0121] In various example embodiments, one or more portions of network 480 may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (Metropolitan-Area Network MAN), the Internet, a portion of the Internet, a portion of a Public Switched Telephone Network (PSTN), a Plain Old Telephone Service (POTS) network, a cellular telephone network, a wireless network, a Wi-Fi For example, network 480 or a portion of network 480 may include a wireless or cellular network, and coupling 482 may be a code division multiple access (CDMA) connection, a Global System for Mobile communications (GSM) connection, or other type of cellular or wireless coupling. In this example, coupling 482 can implement any of various types of data transmission technologies, such as Single Carrier Radio Transmission Technology (1xRTT), Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data rates for GSM Evolution (EDGE) technology, the Third Generation Partnership Project (3GPP) including 3G, the Fourth Generation Wireless (4G) network, the Universal Mobile Telecommunications System (UMTS), High-Speed Packet Access (HSPA), Worldwide Interoperability for Microwave Access (WiMAX), Long Term Evolution (LTE) standards, other standards defined by various standards setting organizations, other remote protocols, or other data transmission technologies.
[0122] Instructions 416 may be sent or received over network 480 using a transmission medium via a network interface device (e.g., a network interface component included in communication component 464) and utilizing any of a number of well-known transfer protocols (e.g., HTTP). Similarly, instructions 416 may be sent or received using a transmission medium via coupling 472 (e.g., a peer-to-peer coupling) to device 470. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure. The terms "transmission medium" and "signal medium" should be deemed to include any intangible medium capable of storing, encoding, or carrying instructions 416 for execution by machine 400, and include digital or analog communication signals or other intangible media to facilitate the communication of such software. Thus, the terms "transmission medium" and "signal medium" should be deemed to include any form of modulated data signals, carrier waves, and the like. The term "modulated data signal" means a signal whose characteristics are set or changed in a manner that encodes information in the signal.
[0123] The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" mean the same thing and are used interchangeably in this disclosure. These terms are defined to include both machine storage media and transmission media. Thus, the term includes both storage devices / media and carrier / modulated data signals.
Claims
1. A system comprising: at least one hardware processor; as well as A computer readable medium storing instructions that, when executed by the at least one hardware processor, cause the at least one hardware processor to perform operations comprising: receiving natural language text describing a request to generate text; accessing contextual information about the natural language text; generating a prompt using the natural language text, the context information, and a system message, the system message including instructions for producing the generated text in a minimized form of an intermediate file format, the intermediate file format being a format that enables portions of a file stored in the intermediate file format to be removed without changing the semantic meaning of the file; Passing the hint to a Large Language Model (LLM); receiving, from the LLM, the generated text in the intermediate file format in a minimized form; and The generated text is parsed using a parser configured to parse a file in the intermediate file format, thereby producing parsed generated text.
2. The system according to claim 1, wherein: The generated text is a compilable computer code.
3. The system according to claim 1, wherein: The hint includes at least some text of the intermediate file format in minimized form.
4. The system according to claim 3, wherein: The LLM has a hard limit on the number of input tokens it will process for a single prompt, and due to the presence of the at least some text in the intermediate file format in minimized form, the prompt includes fewer tokens than the hard limit.
5. The system according to claim 1, wherein: The minimized form includes no spaces.
6. The system according to claim 1, wherein: The minimized form includes no carriage returns.
7. The system according to claim 1, wherein: The intermediate file format is JavaScript Object Notation (JSON).
8. A method comprising: receiving natural language text describing a request to generate text; accessing contextual information about the natural language text; generating a prompt using the natural language text, the context information, and a system message, the system message including instructions for producing the generated text in a minimized form of an intermediate file format, the intermediate file format being a format that enables portions of a file stored in the intermediate file format to be removed without changing the semantic meaning of the file; Passing the hint to a Large Language Model (LLM); receiving, from said LLM, generated text in said intermediate file format in a minimized form; as well as The generated text is parsed using a parser configured to parse a file in the intermediate file format, thereby producing parsed generated text.
9. The method according to claim 8, wherein: The generated text is a compilable computer code.
10. The method according to claim 8, wherein: The hint includes at least some text of the intermediate file format in minimized form.
11. The method according to claim 10, wherein: The LLM has a hard limit on the number of input tokens it will process for a single prompt, and due to the presence of the at least some text in the intermediate file format in minimized form, the prompt includes fewer tokens than the hard limit.
12. The method according to claim 8, wherein: The minimized form includes no spaces.
13. The method according to claim 8, wherein: The minimized form includes no carriage returns.
14. The method according to claim 8, wherein: The intermediate file format is JavaScript Object Notation (JSON).
15. A non-transitory machine-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising: receiving natural language text describing a request to generate text; accessing contextual information about the natural language text; generating a prompt using the natural language text, the context information, and a system message, the system message including instructions for producing the generated text in a minimized form of an intermediate file format, the intermediate file format being a format that enables portions of a file stored in the intermediate file format to be removed without changing the semantic meaning of the file; Passing the hint to a Large Language Model (LLM); receiving, from said LLM, generated text in said intermediate file format in a minimized form; as well as The generated text is parsed using a parser configured to parse a file in the intermediate file format, thereby producing parsed generated text.
16. The non-transitory machine-readable medium of claim 15, wherein: The generated text is a compilable computer code.
17. The non-transitory machine-readable medium of claim 15, wherein: The hint includes at least some text of the intermediate file format in minimized form.
18. The non-transitory machine-readable medium of claim 17, wherein: The LLM has a hard limit on the number of input tokens it will process for a single prompt, and due to the presence of the at least some text in the intermediate file format in minimized form, the prompt includes fewer tokens than the hard limit.
19. The non-transitory machine-readable medium of claim 15, wherein: The minimized form includes no spaces.
20. The non-transitory machine-readable medium of claim 15, wherein: The minimized form includes no carriage returns.