Prompt input tuning with any unstructured data
The prompt tuning system with a fusion GNN enhances LLMs' ability to process diverse unstructured data, addressing context length constraints and improving output accuracy by generating a summarized embedding from multiple data types.
Patent Information
- Application Number
- PCT/CN2024/102847
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-01
- Publication Date
- 2026-01-08
AI Technical Summary
Large language models (LLMs) face limitations in efficiently processing and understanding unstructured data types such as images, tables, and diagrams, and are constrained by context length, leading to inefficiencies in communication and output accuracy.
A prompt tuning system utilizing a fusion graph neural network (GNN) generates a fusion prompt input embedding that combines embeddings from various unstructured data types, including text, images, and transaction history, to provide context for LLMs, allowing them to generate more accurate and relevant outputs.
The system enables LLMs to effectively process and understand a broader range of unstructured data, overcoming context length limitations and improving output accuracy by aggregating diverse data types into a summarized embedding.
Smart Images

Figure CN2024102847_08012026_PF_FP_ABST
Abstract
Description
PROMPT INPUT TUNING WITH ANY UNSTRUCTURED DATATECHNICAL FIELD
[0001] The subject matter disclosed herein generally relates to use of large language models. Specifically, the present disclosure addresses systems and methods that tune prompt inputs using any unstructured data.BACKGROUND
[0002] Recently, large language models (LLMs) have become the de-facto standard for generative artificial intelligence (AI) solutions. Thus, LLMs or generic, generative AI are being used more in applications. LLMs can follow text and understands sentiment behind it. LLMs can do some reasoning given some evidence, in context learning, or known knowledge. With classic LLM, instructions or language is treated as a kind of prompt. Given the prompt, the LLM provides an output.BRIEF DESCRIPTION OF THE DRAWINGS
[0003] FIG. 1 is a diagram illustrating an example network environment suitable for tuning a prompt input using unstructured data, according to example implementations.
[0004] FIG. 2 is a diagram illustrating components of the prompt tuning system, according to example implementations.
[0005] FIG. 3 is a diagram illustrating components and a data flow within and between a prompt tuning system and a large language model (LLM) , according to example implementations.
[0006] FIG. 4 is a diagram illustrating components and a data flow within and between the prompt tuning system and the LLM with fine-tuning, according to example implementations.
[0007] FIG. 5 is a diagram illustrating a data flow for an example user case that uses sequences of events as additional context, according to example implementations.
[0008] FIG. 6 is a flowchart illustrating a method for tuning a prompt input using unstructured data, according to example implementations.
[0009] FIG. 7 is a block diagram illustrating components of a machine, according to some examples, able to read instructions from a machine-storage medium and perform any one or more of the methodologies discussed herein.DETAILED DESCRIPTION
[0010] The description that follows describes systems, methods, techniques, instruction sequences, and computing machine program products that illustrate examples of the present subject matter. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide an understanding of various examples of the present subject matter. It will be evident, however, to those skilled in the art, that examples of the present subject matter may be practiced without some or other of these specific details. Examples merely typify possible variations. Unless explicitly stated otherwise, structures (e.g., structural components) are optional and may be combined or subdivided, and operations (e.g., in a procedure, algorithm, or other function) may vary in sequence or be combined or subdivided.
[0011] Systems and methods that tune prompt inputs using unstructured data is discussed herein. One use case is in fraud detection. In fraud detection, clues are represented in many different ways such as text, tables, charts, diagrams, and images. LLMs have been proven useful to solve problems with chains of thoughts. However, communicating through natural language and LLMs is sometimes not an efficient form. Additionally, LLMs are also limited by context length. In order to make efficient use of LLMs, it is necessary to have a more efficient content organization. Compressing pure text into embeddings is a natural approach. In many cases, the information summary is not just text-based but multimodal, including images, tables, charts, diagrams, and other types of data. Example implementations provide a universal, unstructured data organization method that can efficiently communicate with LLMs by tuning prompt inputs.
[0012] In particular, example implementations provide a prompt tuning system that includes a fusion graph neural network (GNN) . The fusion GNN comprises a transformer encoder, which, based on given data, associations, and entity information, trains the encoder to generate a prefix (referred to herein as a “fusion prompt input embedding) that can be understood by LLMs. This prefix enables LLMs to comprehend and complete corresponding responses effectively. A graph representation based on multiple raw data inputs helps to eliminate the constraints on context length of text but also extends the variety of the input data types.
[0013] As a result, example implementations provide a technical solution to the technical problem of providing additional context and information to an LLM in order to obtain a more accurate and / or relevant output. In particular, the technical solution access multiple types of unstructured data and generates respective embeddings for each of the multiple types of unstructured data. The respective embeddings are inputted to a fusion graph neural network (GNN) , which generates a fusion prompt input embedding by combining information from the respective embedding for each of the multiple types of unstructured data into a single embedding that provides context for a pure text input (e.g., a request for an output) to the LLM. The fusion prompt input embedding and a token embedding representing the pure text input can then be provided to the LLM and a more accurate / relevant output can be received in response.
[0014] FIG. 1 is a diagram illustrating an example network environment 100 suitable for tuning a prompt input using a plurality of different unstructured data, according to example implementations. A network system 102 provides server-side functionality via a communication network 104 (e.g., the Internet, wireless network, cellular network, or a Wide Area Network (WAN) ) to a client device 106. The network system 102 is configured to trigger various operations at a large language model (LLM) 108, as will be discussed in more detail below.
[0015] In various cases, the client device 106 is a device associated with a user of the network system 102 that wants to obtain an output from the LLM 108 using functionalities of the network system 102. In other cases, the client device 106 is a device associated with a user that uses the network system 102 to conduct a transaction. The client device 106 may comprise, but is not limited to, a smartphone, a tablet, a laptop, multi-processor systems, microprocessor-based or programmable consumer electronics, a desktop computer, a server, or any other communication device that can access the network system 102. The client device 106 can include an application that exchanges data, via the network 104, with the network system 102. For example, the application can be browser application or a local version of an application associated with the network system 102 that can provide data to and access data from one or more components at the network system 102.
[0016] In example implementations, the client device 106 interfaces with the network system 102 via a connection with the network 104. Depending on the form of the client device 106, any of a variety of types of connections and networks 104 may be used. For example, the connection may be Code Division Multiple Access (CDMA) connection, a Global System for Mobile communications (GSM) connection, or another type of cellular connection. Such a connection may implement any of a variety of types of data transfer technology, such as Single Carrier Radio Transmission Technology (1xRTT) , Evolution-Data Optimized (EVDO) technology, General Packet Radio Service (GPRS) technology, Enhanced Data rates for GSM Evolution (EDGE) technology, or other data transfer technology (e.g., fourth generation wireless, 4G networks, 5G networks) . When such technology is employed, the network 104 includes a cellular network that has a plurality of cell sites of overlapping geographic coverage, interconnected by cellular telephone exchanges. These cellular telephone exchanges are coupled to a network backbone (e.g., the public switched telephone network (PSTN) , a packet-switched data network, or other types of networks.
[0017] In another example, the connection to the network 104 is a Wireless Fidelity (e.g., Wi-Fi, IEEE 802.11x type) connection, a Worldwide Interoperability for Microwave Access (WiMAX) connection, or another type of wireless data connection. In such an example, the network 104 includes one or more wireless access points coupled to a local area network (LAN) , a wide area network (WAN) , the Internet, or another packet-switched data network. In yet another example, the connection to the network 104 is a wired connection (e.g., an Ethernet link) and the network 104 is a LAN, a WAN, the Internet, or another packet-switched data network. Accordingly, a variety of different configurations are expressly contemplated.
[0018] Additionally, the client device 106 comprises a display component (not shown) to display information (e.g., in the form of user interfaces) as will be discussed in more detail below. The client device 106 can be operated by a human user and / or a machine user.
[0019] The LLM 108 is a trained model that is configured to generate text and perform natural language processing tasks. Generally, the LLM 108 learns relationships from a large data set during a training process. The LLM 108 can then be used to generate text by taking an input and repeatedly predicting a next token or word, for example. For instance, the LLM 108 can generate a probability for the next tokens and select a proper one (e.g., highest probability) for output. In some implementations, the LLM 108 is a closed-source LLM. While the LLM 108 is shown outside of the network system 102 and communicatively coupled via the network 104, the LLM 108 can, in some implementations, be a part of the network system 102 (e.g., be located within the network system 102) .
[0020] Turning specifically to the network system 102, an application programing interface (API) server 110 and a web server 112 are coupled to and provide programmatic and web interfaces respectively to one or more networking servers 114. The networking servers 114 host various systems including a prompt tuning system 116 and a transaction system 118, each comprising a plurality of components and each of which can be embodied as a combination of hardware, software, and / or firmware.
[0021] The prompt tuning system 116 is configured to tune a prompt input by generating a fusion prompt input embedding from different types of unstructured data. The different types of unstructured data can include any combination of, for example, text, images, charts, notes, tables, and / or diagrams. The fusion prompt input embedding can provide context to, and is provided with, a token embedding representing a text input. The prompt tuning system 116 will be discussed in more detail in connection with FIG. 2 – FIG. 5 below.
[0022] The transaction system 118 is configured to perform and / or monitor transactions associated with the network system 102. For example, the transactions can comprise authentication services performed using the network system 102, commerce transactions completed via the network system 102, banking transactions completed via the network system 102, or any other types of operations that can occur using the network system 102 or monitored by the network system 102. In the case of commerce transactions, the transaction system 118 can provide e-commerce functionalities. Similarly for a banking transaction, the transaction system 118 can provide banking functionalities. The transaction system 118 monitors sequences of events (interactions) performed by users and their devices during these transactions and stores indications of these sequences of events for use by the prompt tuning system 116 as will be discussed further below.
[0023] The networking servers 114 can be, in turn, coupled to one or more database servers 120 that facilitate access to one or more storage repositories or data storage 122. The data storage 122 is a storage device storing, for example, user accounts including user profiles of users of the network system 102, the sequences of events detected by the transaction system 118, and / or insight data that provides information from previous transactions.
[0024] Any of the systems, data storage, servers, or devices (collectively referred to as “components” ) shown in, or associated with, FIG. 1 may be, include, or otherwise be implemented in a special-purpose (e.g., specialized or otherwise non-generic) computer that can be modified (e.g., configured or programmed by software, such as one or more software components of an application, operating system, firmware, middleware, or other program) to perform one or more of the functions described herein for that system or machine. For example, a special-purpose computer system able to implement any one or more of the methodologies described herein is discussed below with respect to FIG. 7, and such a special-purpose computer is a means for performing any one or more of the methodologies discussed herein. Within the technical field of such special-purpose computers, a special-purpose computer that has been modified by the structures discussed herein to perform the functions discussed herein is technically improved compared to other special-purpose computers that lack the structures discussed herein or are otherwise unable to perform the functions discussed herein. Accordingly, a special-purpose machine configured according to the systems and methods discussed herein provides an improvement to the technology of similar special-purpose machines.
[0025] Moreover, any two or more of the components illustrated in FIG. 1 may be combined, and the functions described herein for any single component may be subdivided among multiple components. Functionalities of one component may, in alternative examples, be embodied in a different component. Additionally, any number of client devices 106 and data storage 120 may be embodied within the network environment 100. While only a single network system 102 is shown, alternatively, more than one network system 102 can be included (e.g., localized to a particular region) .
[0026] FIG. 2 is a diagram illustrating components of the prompt tuning system 116, according to example implementations. In example implementations, the prompt tuning system 116 comprises a server that triggers (e.g., provides instructions or prompts) the LLM 108 to generate an output based on context information and a text input comprising an instruction. To enable these operations, the prompt tuning system 116 comprises a user interface component 202 and an embedding component 204 configured in communication with one another (e.g., via a bus, shared memory, or a switch) . The prompt tuning system 116 may also comprise a transaction component 206 and other components (not shown) that are not germane to example implementations.
[0027] The user interface component 202 is configured to generate and manage user interfaces that are displayed on the client device 106. The user interface component 202 can receive inputs via the user interface from the client device 106. For example, the user interface can receive a pure text input that provides instructions for an output of the LLM 108. The user interface can also receive indications of unstructured data to be used in generating or fine tuning the prompt fusion input embedding. The user interface component 202 also generates and / or updates user interfaces to display the output generated by the LLM 108.
[0028] The embedding component 204 is configured to manage generation of prompts (e.g., instructions, questions, information, or coding that communicates to the LLM 108 what response is needed; context information) and exchanges data with the LLM 108 in order to trigger the LLM 108 to generate the output. Accordingly, the embedding component 204 includes one or more vocabulary tokenizers 208, the large language model (LLM) 108, one or more visual encoders 212, and a fusion graphic neural network (GNN) 214. While the implementation of FIG. 2 includes vocabular tokenizers 208 to tokenize text and visual encoders 212 to encode images, alternative implementations can comprise other types of embedding generating components that converts corresponding unstructured data into a corresponding embedding. For example, a sound embedding generator can be included to generate a sound embedding. While the implementation of FIG. 2 includes the LLM 108, alternative implementations can have the LLM 108 external to the network system 102 but be accessible (e.g., via the network 104) by the network system 102 (e.g., as shown in FIG. 1) .
[0029] The vocabulary tokenizers 208 are configured to tokenize text into tokens. For example, if the text includes “Microsoft Surface tablet, ” tokens are generated for “Microsoft, ” “Surface, ” and “tablet. ” Each token comprises a token identifier (ID) generated by an encoding block that converts a string value into a long integer. The token IDs (e.g., list of long integers) for the tokens are then provided to an embedding layer of the LLM 108 that converts each long integer into a vector of numbers. In some cases, the vector of numbers comprises a matrix. The embedding layer can capture semantic and syntactic meaning of the input (e.g., token IDs) , so as to provide context. The result is a text embedding for each piece of text. It is noted that the prompt tuning system 116 can comprise any number of vocabulary tokenizers 208.
[0030] The visual encoders 212 are configured to transform the image raw pixel data into image or visual embeddings. In some implementations, the visual encoder 212 comprises a convolutional neural network or a machine-learning model. The visual embeddings are numeric representations of the images that encode semantics of contents of the images into lower-dimensional vector representations. The prompt tuning system 116 can comprise any number of visual encoders 212.
[0031] The fusion GNN 214 is configured to aggregate all the different aspects of information from the unstructured data and generate a fusion prompt input embedding. In particular, the fusion GNN 214 can take different types of data and combine them to derive an understanding or context that the LLM 108 can understand. For example, if there are multiple vectors or embeddings generated by the language models 210 and the visual encoders 212, the fusion GNN 214 aggregates these embeddings into a single embedding (the fusion prompt input embedding) . More specifically, the fusion GNN 214 compresses the information from each of the multiple embeddings and generates a summarized embedding that provides context to the LLM 108.
[0032] The transaction component 206 is configured to access and provide transaction history to the prompt tuning system 116. In example implementations, the transaction component 206 accesses the data storage 122 to retrieve transaction history stored by the transaction system 118. The transaction history can include sequences of events performed during transactions conducted or monitored by the transaction system 118. The transaction history is used to provide further context that can fine-tune the fusion prompt input embedding as will be discussed in more detail below.
[0033] FIG. 3 is a diagram illustrating components and a data flow within and between the prompt tuning system 116 and the large language model (LLM) 108, according to example implementations. In example implementations, the prompt tuning system 116 comprises a server that generates a fusion prompt input embedding that can be combined with a token embedding representing a text input as inputs to the LLM 108. The fusion prompt input embedding, which is generated from a plurality of unstructured data, can provide context for the text input.
[0034] Example implementations leverage the idea of P-tuning (or prompt tuning) in which a prompt prefix embedding is generated and added to an input prompt before the input prompt is provided to a LLM. With P-tuning, the prompt can be text (e.g., word, phrase, sentence) that guides the LLM towards generating a specific type of output. Thus, P-tuning is confined to text. Example implementations build on this concept by generating a fusion prompt input embedding from a combination of text, images, notes, diagrams, tables, and / or any other type of unstructured data. Because the LLM 108 only sees the embeddings, the type of token from which the embeddings are generated does not matter. As such, any type of unstructured data can be used. While the example implementation of FIG. 3 is described below as generating the fusion prompt input embedding from a combination of text and images, any combination of unstructured data can be used. For example, the fusion prompt input embedding can be based on images, sounds, and data from spreadsheets. It is noted that text can be text from charts, tables, notes, lists, or other types of textural information.
[0035] Referring to FIG. 3, text 302 is accessed from a source (e.g., the data storage 120) or provided directly by a user (e.g., as a file) . For example, the user can provide an indication of the text 302 that should be used and the embedding component 204 can retrieve the text 302 from data storage. The text 302 is a raw input that comprises any textural information. For example, the text 302 can include a title, a sentence, a description, entries in a table, lists of information, a spreadsheet, and so forth. Each piece of the text 302 is provided to the vocabulary tokenizer 208 that tokenizes the text 302 into tokens. Each token comprises a token identifier (ID) generated by an encoding block that converts a string value into a long integer. The token IDs (e.g., list of long integers) for the tokens are then provided to an embedding layer of the LLM 108 that converts each long integer into a vector of numbers. The embedding layer can capture semantic and syntactic meaning of the input (e.g., token IDs) , so as to provide context. The result is a text embedding 304 for each piece of the text 302. These text embeddings 304 are provided to the fusion GNN 214.
[0036] Similarly, images 306 are accessed from a source (e.g., the data storage 120) or provided directly by a user. The images 306 are raw inputs that can comprise any visual information that is not in textual format. For example, the images 306 can include photographs, pie charts, diagrams, drawings, paintings, and so forth. Each image 306 is provided to the visual encoder 212 which transforms raw pixel data of the image 306 into a visual embeddings 308. In some implementations, the visual encoder 212 comprises a convolutional neural network or a machine-learning model. The visual embeddings 308 are numeric representations of the images 306 that encode semantics of contents of the images 306 into lower-dimensional vector representations. These visual embeddings 308 are also provided to the fusion GNN 214.
[0037] The fusion GNN 214 is configured to generate a fusion prompt input embedding 310 by combining and summarizing information from the respective text embeddings 304 and visual embeddings 308 into a single embedding, by, for example, the equation Hfusion=Agg (Wtrans·Htrans, Wtext·Htext, Wvision·Hvision, …) . Hfusion is the embedding generated by the fusion GNN 214, W is a weight applied to each respective type of embedding, and H is the respective embeddings (e.g., transaction embedding, text embedding, visual embedding) . The fusion GNN 214 uses relationships of a graph representing connections between the different embeddings to generate the fusion prompt input embedding. For example, if the embeddings 304, 308 are for an item title and picture pair, the embeddings 304, 308 can be aggregated through relationships of a graph representing connections between the embeddings 304, 308 of the item title and picture. The fusion prompt input embedding 310 can provide context for a text input 312 to the LLM 108.
[0038] The text input 312 can also be received from the client device 106 (e.g., via the user interface component 212) . The text input 312 is purely text that comprises instructions for an output from the LLM 108. The text input 312 can ask a question or provide a command that can be based on the text 302 and images 306. For example, in the context of fraud detection, the text input 312 can ask, given the text 302 (e.g., an item listing, communications between the parties) and image (s) 306 (e.g., of the item in the item listing) , do you think this is an authentic item? In an AI generative example, the text input 312 can command, given the text 302 (e.g., name of an item) and image (s) 306 (e.g., of the item) , generate an item listing that describes the item.
[0039] Similar to the text 302, the text input 312 is provided to the vocabulary tokenizer 208 that tokenizes the text input 312 into tokens that each comprises a token ID 314. The token IDs 314 (e.g., list of long integers) for the tokens are then provided to an embedding layer 316 of the LLM 108 that creates token embeddings 318 from the input token IDs 314. The embedding layer 316 can capture semantic and syntactic meaning of the input, so that the LLM 108 can understand context. These token embeddings 318 can then be combined with the fusion prompt input embedding 310 and provided to the LLM 108.
[0040] In example implementations, the LLM 108 is a transformer model that comprises an encoder / decoder 320. Using the token embeddings 318 and fusion prompt input embedding 310, the transformer model can pre-process data as numerical representations through the encoder and understand the context (e.g., of words and phrases with similar meanings) as well as other relationships (e.g., between words, between images, or combination of words and images) . The LLM 108 can then apply this knowledge through the decoder to produce a unique output. The output of the LLM 108 is returned to the prompt tuning system 116 and can then be displayed (e.g., on a display of the client device 106) via the user interface component 202.
[0041] While the language model 210 is shown as a different component than the LLM 108, alternative implementations may have the language model 210 be the LLM 108. Additionally, while the implementation of FIG. 3 generated the fusion prompt input embedding 310 based on text 302 and images 306, alternative implementations can generate the fusion prompt input embedding 310 based on other combinations of unstructured data. For example, the fusion prompt input embedding 310 can be based on text 302, images 306, and sound.
[0042] FIG. 4 is a diagram illustrating components and a data flow within and between the prompt tuning system 116 and the LLM 108 with fine-tuning, according to example implementations. The implementation shown in FIG. 4 provides additional components to the implementation of FIG. 3 that perform the fine-tuning of the inputs to the LLM 108. In particular, the implementation of FIG. 4 helps the semantic context in the LLM 108 understand the fusion prompt input embedding 310.
[0043] A data storage 402 comprises insight data providing insight into previous transactions that have occurred in association with the network system 102. For example, the transaction can comprise authentication services performed using the network system 102, commerce transactions completed via the network system 102, banking transactions completed via the network system 102, or any other types of operations that can occur using the network system 102. In some implementations, the data storage 402 can be the data storage 120.
[0044] When, for example, a case associated with the text 302 and the images 306 is presented, the insight data comprises notes that provide insight into the case. The insight data can comprise human or machine generated notes based on previous transactions performed using the transaction system 118 that can be relevant to the present case as context. For example, if the prompt tuning system 116 and LLM 108 are used to detect fraudulent transactions, the insight data can include notes based on previous transactions that indicate, for example, activities, devices, or locations that are detected to be prevalent for fraud. For instance, the insight data can indicate that a particular Internet Protocol (IP) address is used in many chargeback or fraudulent cases or by many fraudulent users. The text 302 and the images 306 can be associated with a new case or transaction that appears suspicious (e.g., can be a case of buyer fraud or use of stolen financial information) for which the LLM 108 is asked to determine whether it is likely a fraudulent transaction or not. This insight data can be used to fine tune the fusion prompt input embedding 310 such that the LLM 108 will have some insight based on previous history.
[0045] In example implementations, the insight data is accessed from the data storage 402 and applied to an embedding layer of the language model 210. While not shown, the insight data may be tokenized by a vocabulary tokenizer 208 prior to being applied to the LLM 108. Text embeddings 404 of the insight data is then outputted from the LLM 108. In example implementations, the text embeddings 404 are then used to pretrain or fine tune the fusion GNN 214 in generating the fusion prompt input embedding 310. In example implementations, the fusion GNN 214 is fine-tuned to minimize a gap between the text / visual embeddings 304, 308 and the fusion prompt input embedding so that they are understandable to the LLM 108 with similar context.
[0046] Based on current context information from the fusion prompt input embedding 310 along with the instruction or context token embedding 318, the encoder / decoder 320 of the LLM 108 can provide an output. For example, the output can indicate a decision or answer to the text input 312 (e.g., “Is this transaction a fraudulent buyer case? ” ) that is represented by the token embedding 318. The output can also include rationale for the decision or answer.
[0047] Referring now to FIG. 5, a data flow for an example use case that uses sequences of events (e.g., user interactions) as additional context is illustrated, according to example implementations. The sequences of events can represent user or device behaviors that can provide the additional context. The example use case is for a fraud investigation that includes use of detected event sequences to determine whether the transaction is fraudulent. Because example implementations allow different types of data to be introduced and organized to provide context for the LLM 108, monitored sequences of events associated with a user or their device can be used to provide the additional context. A monitoring component of the transaction system 118 monitors events performed on a website associated with the network system 102. For instance, the monitoring component can detect, pages viewed (or items on the pages viewed) , items searched, mouse movements, click actions, and / or any other actions performed (e.g., complete a transaction, initiate a claim) on the website and store this information for use by the prompt tuning system 116.
[0048] For the example use case, assume an entity wants to check a probability behind a transaction being fraudulent or a buyer being fraudulent. The monitoring component monitors a flow of how the user (e.g., buyer) behaved while on the website and how the user arrived at a transaction checkpoint. As shown in FIG. 5, the user viewed an item (view event 502) , selected “like it” (like it event 504) , viewed the item again (view item event 506) , performed a search for something else (search event 508) , and then arrived at a checkpoint (user node 510) .
[0049] The client device of the user can also be monitored. For instance, the client device checked out five minutes ago (checkout event 512) , created a claim ticket (claim event 514) , checked out again (checkout event 516) , and then performed a search (search event 518) . The device can be detected to be used by the user and thus, arrived at both the user node 510 and device node 520.
[0050] Example implementations can use message passing. Each event or interaction can be represented as a vector comprising hundreds of numbers and each event can be converted into an embedding. Embedding messages are passed through a pathway created by the sequences of events. For example, a message with a view item embedding is passed to the like it event 504. The message is aggregated with a like it embedding and passed to the view item event 506 where the message is aggregated with a view item embedding. The message is then passed to the search event 508 and aggregated with a search embedding and then sent to the user node 510. Similarly, an embedding message for the checkout event 512 is passed to the claim event 514 and the message is aggregated with a claim embedding. The message is then passed to the checkout event 516 and aggregated with a checkout embedding. The message is then sent to the search event 518 and aggregated with a search embedding. The message is then provided to both the user node 510 and the device node 520.
[0051] Just like how information is sent through different network devices, the embedding messages are sent through graph pathways and come to a final node (e.g., the fusion GNN 214) . The user node 510 can comprise a list of user activities / events over a period of time. Similarly, the device node 520 can comprise a list of activities / events performed by the device over a period of time. Since all the data are embeddings, the transaction node 522 can aggregate all the information or message embeddings sent from the upstream nodes (e.g., the user node 510 and device node 520) . When the embeddings arrive at the transaction node 522, the transaction node 522 can check the user history embedding (from the user node 510) and the device history embedding (from the device node 512) . The user history embedding and the device history embedding are then aggregated into a transaction embedding at the transaction node 522. The transaction embedding is then passed to the fusion GNN 214, which consumes the different data sources (e.g., embeddings) represented in graph structure.
[0052] The prompt tuning system 116 can extract relationship information from logins and construct a graph to see both the user and the device are relevant to the transaction. The item title (text) and images that the user is checking out (e.g., a listing or publication that includes the item time and images) can also be linked to the transaction. In the present use case, the text can also include, for example, a timeline of communications between the buyer and seller, a timeline of money movement for the transactions, and / or a timeline of other events that occur for the transaction. The images can be obtained from the listing or publication involved in the transaction. All of the lines in the use case flow represent some relationship information which a graph neural network (e.g., the fusion GNN 214) can aggregate.
[0053] The fusion GNN 214 is an aggregation node that can aggregate all the different aspects of the information. For example, if there are three vectors or embeddings (e.g., transaction embedding, text embedding 304, visual embedding 308) , the fusion GNN 214 aggregates these embeddings into a single embedding (the fusion prompt input embedding 310) . That is, the fusion GNN 214 compresses the information from the vectors / embeddings and generates a summarized vector / embedding. In the fraud use case, the fusion prompt input embedding 214 summarizes how the user behaved, how the device behaved, how the item that the user transacts appears (e.g., images) , and what are the words associated with the item (e.g., text) . Thus, the fusion prompt input embedding 310 can be a description of the use case. It is not easy to represent in plain text, but by using graph connections, clues / context with connections can be aggregated and provided to the LLM 108.
[0054] The fusion prompt input embedding 310 is combined with the token embeddings 318 when provided to the LLM 108. Thus, the text input may ask “does the transaction look like a buyer transaction fraud or is it an abusive buyer case? ” This text input is encoded into the token embedding 318 and presented with all the context information in the fusion prompt input embedding 310.
[0055] Based on the fusion prompt input embedding 310 and the token embedding 318, the encoder / decoder 320 of the LLM 108 can provide an output. For example, the output can indicate a decision or answer to the text input 312 (e.g., “Does the transaction look like a buyer transaction fraud or is it an abusive buyer case? ” ) that is represented by the token embedding 318. The output can also include rationale for the decision or answer. For example, the rationale can indicate facts derived from the context (e.g., fusion prompt input embedding 310) that justify the decision or answer.
[0056] FIG. 6 is a flowchart illustrating operations of a method 600 for tuning a prompt input using unstructured data, according to example implementations. Operations in the method 600 may be performed by the prompt tuning system 116, using components described above in part with respect to FIG. 2. Accordingly, the method 600 is described by way of example with reference to the prompt tuning system 116. However, it shall be appreciated that at least some of the operations of the method 600 may be deployed on various other hardware configurations or be performed by similar components residing elsewhere in the network environment 100. Therefore, the method 600 is not intended to be limited to the prompt tuning system 116.
[0057] In operation 602, a text input and indication of unstructured data is received by the prompt tuning system 116. In example implementations, the user interface component 202 receives this information via a user interface presented to a user (e.g., at the client device 106) . The operations of the method 600 are triggered by the user wanting to perform some analysis using the LLM 108. In some implementations, the user provides information regarding a case (e.g., a transaction) that provides context for a text input. As such, the user triggers the prompt tuning system 116 by providing the information and the text input. The text input comprises pure text that provides instructions to the LLM 108 to perform some analysis and return an output (e.g., decision and rationale) . The indication of the unstructured data comprises an indication of data that can provide context for the LLM 108 in performing the analysis based in the instructions. For example, the indication can point to a case or transaction data stored in a data storage that can provide some context.
[0058] In operation 604, the unstructured data is accessed by the prompt tuning system 116. For example, the embedding component 204 can access textural data (e.g., publications, listings, tables, spreadsheets) and images (e.g., photographs, charts, drawings) from a data storage based on the indication received in operation 602. In implementations that use transaction history and insight, the transaction component 206 accesses the insight data from a data storage.
[0059] In operation 606, embeddings for each type of unstructured data are generated. In some cases, the unstructured data comprises text. Text can comprise any textural information. For example, the text can include a title, a sentence, a description, entries in a table or list, and so forth. The text is provided to the vocabulary tokenizer 208 that tokenizes the text into tokens that each comprises a token identifier (ID) . The token IDs are then provided to an embedding layer of the LLM 108 that converts the list of token IDS into a vector of numbers. The embedding layer can capture semantic and syntactic meaning of the input (e.g., token IDs) , so as to provide context. The result is a text embedding for each piece of text.
[0060] In some cases, the unstructured data comprises images. The images can comprise any visual information that is not in textual format. For example, the images can include photographs, charts, diagrams, drawings, paintings, and so forth. The images are provided to the visual encoder 212 which transforms raw pixel data into visual embeddings. The visual embeddings are numeric representations of the images that encode semantics of contents of the images into lower-dimensional vector representations.
[0061] Furthermore, the unstructured data can comprise insight data. The insight data provides insight into previous transactions that have occurred in association with the network system 102. The insight data can comprise notes that provide insight with respect to the unstructured data (e.g., text and images) for which embeddings are concurrently being generated. For example, the insight data can comprise human or machine generated notes based on previous transactions performed using the transaction system 118 that can be relevant to the unstructured data as context. In example implementations, the insight data is applied to an embedding layer of the LLM 108. This can occur after the insight data has been tokenized (e.g., for insight data that is pure text) to generate token IDs. Text embeddings of the insight data is then outputted from the LLM 108.
[0062] In operation 608, all of the generated embeddings are provided to the fusion GNN 214. The fusion GNN 214 then generates a fusion prompt input embedding in operation 610 by combining the information from the respective embeddings generated in operation 606 into a single embedding – the fusion prompt input embedding. More specifically, the fusion GNN 214 compresses the information from each of the multiple embeddings and generates a summarized embedding that provides context to the LLM 108. For example, if a text embedding is for an item title and a visual embedding is for a picture, the text and visual embeddings can be aggregated through relationships of a graph representing connections between the text embedding and visual embedding.
[0063] In operation 612, token embeddings for the text input received in operation 602 are generated. Similar to text, the text input is provided to the vocabulary tokenizer 208 that tokenizes the text input into tokens that each comprises a token ID. The token IDs (e.g., list of long integers) for the tokens are then provided to an embedding layer of a language model. In some cases, the language model is the LLM 108. The embedding layer creates token embeddings from the input token IDs. The embedding layer can capture semantic and syntactic meaning of the input, so that the LLM 108 can understand context. It is noted that operation 612 can occur prior to or concurrently with operations 604-610.
[0064] In operation 614, the fusion prompt input embedding and the token embeddings are provided to the LLM 108. Based on current context information from the fusion prompt input embedding and the token embedding, the encoder / decoder 320 of the LLM 108 can provide an output. For example, the output can indicate a decision or answer to the text input that is represented by the token embeddings 318. The output can also include rationale for the decision or answer.
[0065] In operation 616, the prompt tuning system 116 receives the output and causes display of the output to the user. Thus, the user interface component 202 can generate / update (or provide instructions for the generation / updating of) a user interface that displays the output.
[0066] FIG. 7 illustrates components of a machine 700, according to some example implementations, that is able to read instructions from a machine-storage medium (e.g., a machine-storage device, a non-transitory machine-storage medium, a computer-storage medium, or any suitable combination thereof) and perform any one or more of the methodologies discussed herein. Specifically, FIG. 7 shows a diagrammatic representation of the machine 700 in the example form of a computer device (e.g., a computer) and within which instructions 724 (e.g., software, a program, an application, an applet, an app, or other executable code) for causing the machine 700 to perform any one or more of the methodologies discussed herein may be executed, in whole or in part.
[0067] For example, the instructions 724 may cause the machine 700 to execute the flow diagram of FIG. 6. In one implementation, the instructions 724 can transform the machine 700 into a particular machine (e.g., specially configured machine) programmed to carry out the described and illustrated functions in the manner described.
[0068] In alternative implementations, the machine 700 operates as a standalone device or may be connected (e.g., networked) to other machines. In a networked deployment, the machine 700 may operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 700 may be a server computer, a client computer, a personal computer (PC) , a tablet computer, a laptop computer, a netbook, a set-top box (STB) , a personal digital assistant (PDA) , a cellular telephone, a smartphone, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing the instructions 724 (sequentially or otherwise) that specify actions to be taken by that machine. Further, while only a single machine is illustrated, the term “machine” shall also be taken to include a collection of machines that individually or jointly execute the instructions 724 to perform any one or more of the methodologies discussed herein.
[0069] The machine 700 includes a processor 702 (e.g., a central processing unit (CPU) , a graphics processing unit (GPU) , a digital signal processor (DSP) , an application specific integrated circuit (ASIC) , a radio-frequency integrated circuit (RFIC) , or any suitable combination thereof) , a main memory 704, and a static memory 706, which are configured to communicate with each other via a bus 708. The processor 702 may contain microcircuits that are configurable, temporarily or permanently, by some or all of the instructions 724 such that the processor 702 is configurable to perform any one or more of the methodologies described herein, in whole or in part. For example, a set of one or more microcircuits of the processor 702 may be configurable to execute one or more components described herein.
[0070] The machine 700 may further include a graphics display 710 (e.g., a plasma display panel (PDP) , a light emitting diode (LED) display, a liquid crystal display (LCD) , a projector, or a cathode ray tube (CRT) , or any other display capable of displaying graphics or video) . The machine 700 may also include an input device 712 (e.g., a keyboard) , a cursor control device 714 (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instrument) , a storage unit 716, a signal generation device 718 (e.g., a sound card, an amplifier, a speaker, a headphone jack, or any suitable combination thereof) , and a network interface device 720.
[0071] The storage unit 716 includes a machine-storage medium 722 (e.g., a tangible machine-storage medium) on which is stored the instructions 724 (e.g., software) embodying any one or more of the methodologies or functions described herein. The instructions 724 may also reside, completely or at least partially, within the main memory 704, within the processor 702 (e.g., within the processor’s cache memory) , or both, before or during execution thereof by the machine 700. Accordingly, the main memory 704 and the processor 702 may be considered as machine-storage media (e.g., tangible and non-transitory machine-storage media) . The instructions 724 may be transmitted or received over a network 726 via the network interface device 720.
[0072] In some example implementations, the machine 700 may be a portable computing device and have one or more additional input components (e.g., sensors or gauges) . Examples of such input components include an image input component (e.g., one or more cameras) , an audio input component (e.g., a microphone) , a direction input component (e.g., a compass) , a location input component (e.g., a global positioning system (GPS) receiver) , an orientation component (e.g., a gyroscope) , a motion detection component (e.g., one or more accelerometers) , an altitude detection component (e.g., an altimeter) , and a gas detection component (e.g., a gas sensor) . Inputs harvested by any one or more of these input components may be accessible and available for use by any of the components described herein.
[0073] EXECUTABLE INSTRUCTIONS AND MACHINE-STORAGE MEDIUM
[0074] The various memories (e.g., 704, 706, and / or memory of the processor (s) 702) and / or storage unit 716 may store one or more sets of instructions and data structures (e.g., software) 724 embodying or utilized by any one or more of the methodologies or functions described herein. These instructions, when executed by processor (s) 702 cause various operations to implement the disclosed implementations.
[0075] As used herein, the terms “machine-storage medium, ” “device-storage medium, ” “computer-storage medium” (referred to collectively as “machine-storage medium 722” ) mean the same thing and may be used interchangeably in this disclosure. The terms refer to a single or multiple storage devices and / or media (e.g., a centralized or distributed database, and / or associated caches and servers) that store executable instructions and / or data, as well as cloud-based storage systems or storage networks that include multiple storage apparatus or devices. The terms shall accordingly be taken to include, but not be limited to, solid-state memories, and optical and magnetic media, including memory internal or external to processors. Specific examples of machine-storage media, computer-storage media, and / or device-storage media 722 include non-volatile memory, including by way of example semiconductor memory devices, for example, erasable programmable read-only memory (EPROM) , electrically erasable programmable read-only memory (EEPROM) , FPGA, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms machine-storage medium or media, computer-storage medium or media, and device-storage medium or media 722 specifically exclude carrier waves, modulated data signals, and other such media, at least some of which are covered under the term “signal medium” discussed below. In this context, the machine-storage medium is non-transitory.
[0076] SIGNAL MEDIUM
[0077] The term “signal medium” or “transmission medium” shall be taken to include any form of modulated data signal, carrier wave, and so forth. The term “modulated data signal” means a signal that has one or more of its characteristics set or changed in such a matter as to encode information in the signal.
[0078] COMPUTER READABLE MEDIUM
[0079] The terms “machine-readable medium, ” “computer-readable medium” and “device-readable medium” mean the same thing and may be used interchangeably in this disclosure. The terms are defined to include both machine-storage media and signal media. Thus, the terms include both storage devices / media and carrier waves / modulated data signals.
[0080] The instructions 724 may further be transmitted or received over a communications network 726 using a transmission medium via the network interface device 720 and utilizing any one of a number of well-known transfer protocols (e.g., HTTP) . Examples of communication networks 726 include a local area network (LAN) , a wide area network (WAN) , the Internet, mobile telephone networks, plain old telephone service (POTS) networks, and wireless data networks (e.g., Wi-Fi, LTE, and WiMAX networks) . The term “transmission medium” shall be taken to include any intangible medium that is capable of storing, encoding, or carrying instructions 724 for execution by the machine 700, and includes digital or analog communications signals or other intangible medium to facilitate communication of such software.
[0081] Throughout this specification, plural instances may implement components, operations, or structures described as a single instance. Although individual operations of one or more methods are illustrated and described as separate operations, one or more of the individual operations may be performed concurrently, and nothing requires that the operations be performed in the order illustrated. Structures and functionality presented as separate components in example configurations may be implemented as a combined structure or component. Similarly, structures and functionality presented as a single component may be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter herein.
[0082] "Component" refers, for example, to a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, APIs, or other technologies that provide for the partitioning or modularization of particular processing or control functions. Components may be combined via their interfaces with other components to carry out a machine process. A component may be a packaged functional hardware unit designed for use with other components and a part of a program that usually performs a particular function of related functions. Components may constitute either software components (e.g., code embodied on a machine-readable medium) or hardware components.
[0083] A “hardware component” is a tangible unit capable of performing certain operations and may be configured or arranged in a certain physical manner. In various example implementations, one or more computer systems (e.g., a standalone computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or application portion) as a hardware component that operates to perform certain operations as described herein.
[0084] In some implementations, a hardware component may be implemented mechanically, electronically, or any suitable combination thereof. For example, a hardware component may include dedicated circuitry or logic that is permanently configured to perform certain operations. For example, a hardware component may be a special-purpose processor, such as a field programmable gate array (FPGA) or an ASIC. A hardware component may also include programmable logic or circuitry that is temporarily configured by software to perform certain operations. For example, a hardware component may include software encompassed within a general-purpose processor or other programmable processor. Once configured by such software, hardware components become specific machines (or specific components of a machine) uniquely tailored to perform the configured functions and are no longer general-purpose processors. It will be appreciated that the decision to implement a hardware component mechanically, in dedicated and permanently configured circuitry, or in temporarily configured circuitry (e.g., configured by software) , may be driven by cost and time considerations.
[0085] Accordingly, the term “hardware component” should be understood to encompass a tangible entity, be that an entity that is physically constructed, permanently configured (e.g., hardwired) , or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. Considering examples in which hardware components are temporarily configured (e.g., programmed) , each of the hardware components need not be configured or instantiated at any one instance in time. For example, where the hardware component comprises a general-purpose processor configured by software to become a special-purpose processor, the general-purpose processor may be configured as respectively different special-purpose processors (e.g., comprising different hardware components) at different times. Software may accordingly configure a processor, for example, to constitute a particular hardware component at one instance of time and to constitute a different hardware component at a different instance of time.
[0086] Hardware components can provide information to, and receive information from, other hardware components. Accordingly, the described hardware components may be regarded as being communicatively coupled. Where multiple hardware components exist contemporaneously, communications may be achieved through signal transmission (e.g., over appropriate circuits and buses) between or among two or more of the hardware components. In examples in which multiple hardware components are configured or instantiated at different times, communications between such hardware components may be achieved, for example, through the storage and retrieval of information in memory structures to which the multiple hardware components have access. For example, one hardware component may perform an operation and store the output of that operation in a memory device to which it is communicatively coupled. A further hardware component may then, at a later time, access the memory device to retrieve and process the stored output. Hardware components may also initiate communications with input or output devices, and can operate on a resource (e.g., a collection of information) .
[0087] The various operations of example methods described herein may be performed, at least partially, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented components that operate to perform one or more operations or functions described herein. As used herein, “processor-implemented component” refers to a hardware component implemented using one or more processors.
[0088] Similarly, the methods described herein may be at least partially processor-implemented, a processor being an example of hardware. For example, at least some of the operations of a method may be performed by one or more processors or processor-implemented components. Moreover, the one or more processors may also operate to support performance of the relevant operations in a “cloud computing” environment or as a “software as a service” (SaaS) . For example, at least some of the operations may be performed by a group of computers (as examples of machines including processors) , with these operations being accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., an application program interface (API)) .
[0089] The performance of certain of the operations may be distributed among the one or more processors, not only residing within a single machine, but deployed across a number of machines. In some example implementations, the one or more processors or processor-implemented components may be located in a single geographic location (e.g., within a home environment, an office environment, or a server farm) . In other example implementations, the one or more processors or processor-implemented components may be distributed across a number of geographic locations.
[0090] EXAMPLES
[0091] Example 1 is a method for tuning prompt inputs for a large language model using any unstructured data. The method comprises accessing multiple types of unstructured data for use in generating a fusion prompt input embedding to the LLM;
[0092] generating a respective embedding for each of the multiple types of unstructured data; inputting the respective embedding for each of the multiple types of unstructured data to a fusion graph neural network (GNN) ; generating, by the fusion GNN, the fusion prompt input embedding by combining information from the respective embedding for each of the multiple types of unstructured data into a single embedding that is the fusion prompt input embedding, the fusion prompt input embedding providing context for a pure text input to the LLM; providing the fusion prompt input embedding and a token embedding representing the pure text input to the LLM, the pure text input comprising instructions for an output from the LLM; and causing display of the output of the LLM.
[0093] In example 2, the subject matter of example 1 can optionally include wherein the multiple types of unstructured data includes an image; the respective embedding is a visual embedding; and the method further comprises generating the visual embedding by applying the image to a visual encoder.
[0094] In example 3, the subject matter of any of examples 1-2 can optionally include wherein the multiple types of unstructured data includes text; the respective embedding is a text embedding of the text; and the method further comprises generating, by a vocabulary tokenizer, vocabulary tokens from the text; and generating the text embedding of the text by applying the vocabulary tokens to the LLM.
[0095] In example 4, the subject matter of any of examples 1-3 can optionally include wherein the multiple types of unstructured data includes a sequence of events associated with a user or device of the user; and the respective embedding is a transaction embedding.
[0096] In example 5, the subject matter of any of examples 1-4 can optionally include generating the transaction embedding based on appending an embedding representing each event associated with the user or the device in the sequence of events to an embedding message sent through the sequence of events.
[0097] In example 6, the subject matter of any of examples 1-5 can optionally include accessing insight data providing insight into previous transactions; generating a text embedding of the insight data; and fine-tuning the fusion prompt input embedding based on the text embedding of the insight data.
[0098] In example 7, the subject matter of any of examples 1-6 can optionally include wherein the multiple types of unstructured data includes text associated with an item and one or more images of the item; the respective embedding for each of the multiple types of unstructured data comprises one or more text embeddings based on the text and one or more image embeddings based on the one or more images; and the fusion prompt input embedding combines information from the one or more text embedding and the one or more image embeddings.
[0099] In example 8, the subject matter of any of examples 1-7 can optionally include wherein the item is associated with a transaction that is monitored by a network system associated with the fusion GNN.
[0100] In example 9, the subject matter of any of examples 1-8 can optionally include generating a transaction embedding based on a sequence of events associated with a user or device of the user associated with the transaction, wherein the fusion prompt input embedding further combines information from the transaction embedding with the information from the one or more text embeddings and the one or more image embeddings.
[0101] In example 10, the subject matter of any of examples 1-9 can optionally include wherein the multiple types of unstructured data includes one or more charts, tables, sounds, graphs, or diagrams; the respective embedding is a corresponding embedding of the one or more charts, tables, sounds, graphs, or diagrams; and the method further comprises generating the corresponding embedding.
[0102] Example 11 is a system for tuning prompt inputs for a large language model. The system comprises one or more processors and a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising accessing multiple types of unstructured data for use in generating a fusion prompt input embedding to the LLM; generating a respective embedding for each of the multiple types of unstructured data; inputting the respective embedding for each of the multiple types of unstructured data to a fusion graph neural network (GNN) ; generating, by the fusion GNN, the fusion prompt input embedding by combining information from the respective embedding for each of the multiple types of unstructured data into a single embedding that is the fusion prompt input embedding, the fusion prompt input embedding providing context for a pure text input to the LLM; providing the fusion prompt input embedding and a token embedding representing the pure text input to the LLM, the pure text input comprising instructions for an output from the LLM; and causing display of the output of the LLM.
[0103] In example 12, the subject matter of example 11 can optionally include wherein the multiple types of unstructured data includes an image; the respective embedding is a visual embedding; and the operations further comprise generating the visual embedding by applying the image to a visual encoder.
[0104] In example 13, the subject matter of any of examples 11-12 can optionally include wherein the multiple types of unstructured data includes text; the respective embedding is a text embedding of the text; and the operations further comprise generating, by a vocabulary tokenizer, vocabulary tokens from the text; and generating the text embedding of the text by applying the vocabulary tokens to the LLM.
[0105] In example 14, the subject matter of any of examples 11-13 can optionally include wherein the multiple types of unstructured data includes a sequence of events associated with a user or device of the user; and the respective embedding is a transaction embedding.
[0106] In example 15, the subject matter of any of examples 11-14 can optionally include wherein the operations further comprise generating the transaction embedding based on appending an embedding representing each event associated with the user or the device in the sequence of events to an embedding message sent through the sequence of events.
[0107] In example 16, the subject matter of any of examples 11-15 can optionally include wherein the operations further comprise accessing insight data providing insight into previous transactions; generating a text embedding of the insight data; and fine-tuning the fusion prompt input embedding based on the text embedding of the insight data.
[0108] In example 17, the subject matter of any of examples 11-16 can optionally include wherein the multiple types of unstructured data includes text associated with an item and one or more images of the item; the respective embedding for each of the multiple types of unstructured data comprises one or more text embeddings based on the text and one or more image embeddings based on the one or more images; and the fusion prompt input embedding combines information from the one or more text embedding and the one or more image embeddings.
[0109] In example 18, the subject matter of any of examples 11-17 can optionally include wherein the item is associated with a transaction that is monitored by a network system associated with the fusion GNN; and the operations further comprise generating a transaction embedding based on a sequence of events associated with a user or device of the user associated with the transaction, wherein the fusion prompt input embedding further combines information from the transaction embedding with the information from the one or more text embeddings and the one or more image embeddings.
[0110] In example 19, the subject matter of any of examples 11-18 can optionally include wherein the multiple types of unstructured data includes one or more charts, tables, sounds, graphs, or diagrams; the respective embedding is a corresponding embedding of the one or more charts, tables, sounds, graphs, or diagrams; and the operations further comprise generating the corresponding embedding.
[0111] Example 20 is a computer-storage medium comprising instructions which, when executed by one or more processors of a machine, cause the machine to perform operations for tuning prompt inputs for a large language model using any unstructured data. The operations comprise accessing multiple types of unstructured data for use in generating a fusion prompt input embedding to the LLM; generating a respective embedding for each of the multiple types of unstructured data; inputting the respective embedding for each of the multiple types of unstructured data to a fusion graph neural network (GNN) ; generating, by the fusion GNN, the fusion prompt input embedding by combining information from the respective embedding for each of the multiple types of unstructured data into a single embedding that is the fusion prompt input embedding, the fusion prompt input embedding providing context for a pure text input to the LLM; providing the fusion prompt input embedding and a token embedding representing the pure text input to the LLM, the pure text input comprising instructions for an output from the LLM; and causing display of the output of the LLM.
[0112] Some portions of this specification may be presented in terms of algorithms or symbolic representations of operations on data stored as bits or binary digital signals within a machine memory (e.g., a computer memory) . These algorithms or symbolic representations are examples of techniques used by those of ordinary skill in the data processing arts to convey the substance of their work to others skilled in the art. As used herein, an “algorithm” is a self-consistent sequence of operations or similar processing leading to a desired result. In this context, algorithms and operations involve physical manipulation of physical quantities. Typically, but not necessarily, such quantities may take the form of electrical, magnetic, or optical signals capable of being stored, accessed, transferred, combined, compared, or otherwise manipulated by a machine. It is convenient at times, principally for reasons of common usage, to refer to such signals using words such as “data, ” “content, ” “bits, ” “values, ” “elements, ” “symbols, ” “characters, ” “terms, ” “numbers, ” “numerals, ” or the like. These words, however, are merely convenient labels and are to be associated with appropriate physical quantities.
[0113] Unless specifically stated otherwise, discussions herein using words such as “processing, ” “computing, ” “calculating, ” “determining, ” “presenting, ” “displaying, ” or the like may refer to actions or processes of a machine (e.g., a computer) that manipulates or transforms data represented as physical (e.g., electronic, magnetic, or optical) quantities within one or more memories (e.g., volatile memory, non-volatile memory, or any suitable combination thereof) , registers, or other machine components that receive, store, transmit, or display information. Furthermore, unless specifically stated otherwise, the terms “a” or “an” are herein used, as is common in patent documents, to include one or more than one instance. Finally, as used herein, the conjunction “or” refers to a non-exclusive “or, ” unless specifically stated otherwise.
[0114] Although an overview of the present subject matter has been described with reference to specific examples, various modifications and changes may be made to these examples without departing from the broader scope of examples of the present invention. For instance, various examples or features thereof may be mixed and matched or made optional by a person of ordinary skill in the art. Such examples of the present subject matter may be referred to herein, individually or collectively, by the term “invention” merely for convenience and without intending to voluntarily limit the scope of this application to any single invention or present concept if more than one is, in fact, disclosed.
[0115] The examples illustrated herein are believed to be described in sufficient detail to enable those skilled in the art to practice the teachings disclosed. Other examples may be used and derived therefrom, such that structural and logical substitutions and changes may be made without departing from the scope of this disclosure. The Detailed Description, therefore, is not to be taken in a limiting sense, and the scope of various examples is defined only by the appended claims, along with the full range of equivalents to which such claims are entitled.
[0116] Moreover, plural instances may be provided for resources, operations, or structures described herein as a single instance. Additionally, boundaries between various resources, operations, modules, engines, and data stores are somewhat arbitrary, and particular operations are illustrated in a context of specific illustrative configurations. Other allocations of functionality are envisioned and may fall within a scope of various examples of the present invention. In general, structures and functionality presented as separate resources in the example configurations may be implemented as a combined structure or resource. Similarly, structures and functionality presented as a single resource may be implemented as separate resources. These and other variations, modifications, additions, and improvements fall within a scope of examples of the present invention as represented by the appended claims. The specification and drawings are, accordingly, to be regarded in an illustrative rather than a restrictive sense.
Claims
1.A method for tuning prompt inputs for a large language model (LLM) , the method comprising:accessing multiple types of unstructured data for use in generating a fusion prompt input embedding to the LLM;generating a respective embedding for each of the multiple types of unstructured data;inputting the respective embedding for each of the multiple types of unstructured data to a fusion graph neural network (GNN) ;generating, by the fusion GNN, the fusion prompt input embedding by combining information from the respective embedding for each of the multiple types of unstructured data into a single embedding that is the fusion prompt input embedding, the fusion prompt input embedding providing context for a pure text input to the LLM;providing the fusion prompt input embedding and a token embedding representing the pure text input to the LLM, the pure text input comprising instructions for an output from the LLM; andcausing display of the output of the LLM.2.The method of claim 1, wherein:the multiple types of unstructured data includes an image;the respective embedding is a visual embedding; andthe method further comprises generating the visual embedding by applying the image to a visual encoder.3.The method of claim 1, wherein:the multiple types of unstructured data includes text;the respective embedding is a text embedding of the text; andthe method further comprises:generating, by a vocabulary tokenizer, vocabulary tokens from the text; andgenerating the text embedding of the text by applying the vocabulary tokens to the LLM.4.The method of claim 1, wherein:the multiple types of unstructured data includes a sequence of events associated with a user or device of the user; andthe respective embedding is a transaction embedding.5.The method of claim 4, further comprising:generating the transaction embedding based on appending an embedding representing each event associated with the user or the device in the sequence of events to an embedding message sent through the sequence of events.6.The method of claim 1, further comprising:accessing insight data providing insight into previous transactions;generating a text embedding of the insight data; andfine-tuning the fusion prompt input embedding based on the text embedding of the insight data.7.The method of claim 1, wherein:the multiple types of unstructured data includes text associated with an item and one or more images of the item;the respective embedding for each of the multiple types of unstructured data comprises one or more text embeddings based on the text and one or more image embeddings based on the one or more images; andthe fusion prompt input embedding combines information from the one or more text embedding and the one or more image embeddings.8.The method of claim 7, wherein the item is associated with a transaction that is monitored by a network system associated with the fusion GNN.9.The method of claim 8, further comprising:generating a transaction embedding based on a sequence of events associated with a user or device of the user associated with the transaction, wherein the fusion prompt input embedding further combines information from the transaction embedding with the information from the one or more text embeddings and the one or more image embeddings.10.The method of claim 1, wherein:the multiple types of unstructured data includes one or more charts, tables, sounds, graphs, or diagrams;the respective embedding is a corresponding embedding of the one or more charts, tables, sounds, graphs, or diagrams; andthe method further comprises generating the corresponding embedding.11.A system comprising:one or more processors; anda memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:accessing multiple types of unstructured data for use in generating a fusion prompt input embedding to the LLM;generating a respective embedding for each of the multiple types of unstructured data;inputting the respective embedding for each of the multiple types of unstructured data to a fusion graph neural network (GNN) ;generating, by the fusion GNN, the fusion prompt input embedding by combining information from the respective embedding for each of the multiple types of unstructured data into a single embedding that is the fusion prompt input embedding, the fusion prompt input embedding providing context for a pure text input to the LLM;providing the fusion prompt input embedding and a token embedding representing the pure text input to the LLM, the pure text input comprising instructions for an output from the LLM; andcausing display of the output of the LLM.12.The system of claim 11, wherein:the multiple types of unstructured data includes an image;the respective embedding is a visual embedding; andthe operations further comprise generating the visual embedding by applying the image to a visual encoder.13.The system of claim 11, wherein:the multiple types of unstructured data includes text;the respective embedding is a text embedding of the text; andthe operations further comprise:generating, by a vocabulary tokenizer, vocabulary tokens from the text; andgenerating the text embedding of the text by applying the vocabulary tokens to the LLM.14.The system of claim 11, wherein:the multiple types of unstructured data includes a sequence of events associated with a user or device of the user; andthe respective embedding is a transaction embedding.15.The system of claim 14, wherein the operations further comprise:generating the transaction embedding based on appending an embedding representing each event associated with the user or the device in the sequence of events to an embedding message sent through the sequence of events.16.The system of claim 11, wherein the operations further comprise:accessing insight data providing insight into previous transactions;generating a text embedding of the insight data; andfine-tuning the fusion prompt input embedding based on the text embedding of the insight data.17.The system of claim 11, wherein:the multiple types of unstructured data includes text associated with an item and one or more images of the item;the respective embedding for each of the multiple types of unstructured data comprises one or more text embeddings based on the text and one or more image embeddings based on the one or more images; andthe fusion prompt input embedding combines information from the one or more text embedding and the one or more image embeddings.18.The system of claim 17, wherein:the item is associated with a transaction that is monitored by a network system associated with the fusion GNN; andthe operations further comprise generating a transaction embedding based on a sequence of events associated with a user or device of the user associated with the transaction, wherein the fusion prompt input embedding further combines information from the transaction embedding with the information from the one or more text embeddings and the one or more image embeddings.19.The system of claim 11, wherein:the multiple types of unstructured data includes one or more charts, tables, sounds, graphs, or diagrams;the respective embedding is a corresponding embedding of the one or more charts, tables, sounds, graphs, or diagrams; andthe operations further comprise generating the corresponding embedding.20.A machine-storage medium comprising instructions which, when executed by one or more processors of a machine, cause the machine to perform operations comprising:accessing multiple types of unstructured data for use in generating a fusion prompt input embedding to the LLM;generating a respective embedding for each of the multiple types of unstructured data;inputting the respective embedding for each of the multiple types of unstructured data to a fusion graph neural network (GNN) ;generating, by the fusion GNN, the fusion prompt input embedding by combining information from the respective embedding for each of the multiple types of unstructured data into a single embedding that is the fusion prompt input embedding, the fusion prompt input embedding providing context for a pure text input to the LLM;providing the fusion prompt input embedding and a token embedding representing the pure text input to the LLM, the pure text input comprising instructions for an output from the LLM; andcausing display of the output of the LLM.
Citation Information
Patent Citations
Risk prediction method and device, equipment and storage medium
CN117764373A
Detecting fraudulent transactions
US20220101192A1