System and method for efficient language model editing using contextual hint generator

By using the context prompt generator (CPG) to generate prompt embeddings in large language model (LLM) editing and editing in combination with word embedding, the problems of reduction in accuracy and overhead in the existing technology are solved, and efficient and accurate language model editing is achieved.

CN119948489APending Publication Date: 2025-05-06SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380068845.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-12
Filing Date
2023-09-26
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When editing large language models (LLM), the existing technology has problems such as degradation in accuracy, irrelevant input and output errors, and large training and inference overhead, which is difficult to meet the needs of commercial deployment.

Method used

The context prompt generator (CPG) is used to generate prompt embeddings, combined with word embedding input into the LLM, and the updated output is generated through the LLM, thereby achieving efficient language model editing.

Benefits of technology

It realizes efficient and controllable language model editing, improves editing accuracy, reduces costs, and is suitable for large-scale LLM editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948489A_ABST
    Figure CN119948489A_ABST
Patent Text Reader

Abstract

A method includes receiving an input of a large language model (LLM) from a user. The method also includes generating one or more lexical embedding based on the input. The method also includes generating, using a contextual cue generator (CPG), one or more cue inserts based on the input, the one or more cue inserts representing new or updated information that is not included in the existing knowledge of the LLM. The method also includes providing one or more lexical embedding and one or more hint embedding to the LLM. Further, the method includes outputting, using the LLM, a prediction based on the one or more lexical embedding and the one or more hint embedding, where the prediction reflects new or updated information represented by the one or more hint embedding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to natural language processing and more particularly to systems and methods for efficient language model editing using a contextual hint generator. Background Art

[0002] A large language model (LLM) is a neural network model that takes a piece of natural language as input and generates one or more words as output. LLMs can contain many parameters (such as billions of parameters) and can be trained with Internet-scale text data. The scale of the model and data can improve the capabilities of LLMs and improve the quality of the output, making LLMs a viable option for commercial use. The industry has widely discussed and explored various use cases for LLM applications, such as chatbots for customer service, intelligent assistants for software tools, question-answering for knowledge query, interfaces for interactive communication between people and devices, etc. In general, there is a high demand for commercial deployment of LLMs. Summary of the invention

[0003] Technical Solution

[0004] The present disclosure provides a system and method for efficient language model editing using a contextual hint generator.

[0005] In a first embodiment, a method includes receiving input to a large language model (LLM) from a user. The method also includes generating one or more token embeddings based on the input. The method also includes using a contextual prompt generator (CPG) to generate one or more prompt embeddings based on the input, wherein the one or more prompt embeddings represent new or updated information that is not included in the existing knowledge of the LLM. The method also includes providing the one or more token embeddings and the one or more prompt embeddings to the LLM. In addition, the method includes using the LLM to output a prediction based on the one or more token embeddings and the one or more prompt embeddings, wherein the prediction reflects the new or updated information represented by the one or more prompt embeddings.

[0006] In a second embodiment, an electronic device includes at least one processing device configured to receive input of an LLM from a user. The at least one processing device is also configured to generate one or more word-meta embeddings based on the input. The at least one processing device is also configured to generate one or more prompt embeddings based on the input using a CPG, wherein the one or more prompt embeddings represent new or updated information that is not included in the existing knowledge of the LLM. The at least one processing device is also configured to provide the one or more word-meta embeddings and the one or more prompt embeddings to the LLM. In addition, the at least one processing device is configured to output a prediction based on the one or more word-meta embeddings and the one or more prompt embeddings using the LLM, wherein the prediction reflects the new or updated information represented by the one or more prompt embeddings.

[0007] In a third embodiment, a non-transitory machine-readable medium includes instructions that, when executed, cause at least one processor of an electronic device to receive input to an LLM from a user. The non-transitory machine-readable medium also includes instructions that, when executed, cause at least one processor to generate one or more word element embeddings based on the input. The non-transitory machine-readable medium also includes instructions that, when executed, cause at least one processor to use a CPG to generate one or more prompt embeddings based on the input, wherein the one or more prompt embeddings represent new or updated information that is not included in the existing knowledge of the LLM. The non-transitory machine-readable medium also includes instructions that, when executed, cause at least one processor to provide the one or more word element embeddings and the one or more prompt embeddings to the LLM. In addition, the non-transitory machine-readable medium includes instructions that, when executed, cause at least one processor to use the LLM to output a prediction based on the one or more word element embeddings and the one or more prompt embeddings, wherein the prediction reflects the new or updated information represented by the one or more prompt embeddings.

[0008] Other technical features may be apparent to those skilled in the art from the following drawings, descriptions and claims.

[0009] Before proceeding to the following detailed description, it may be advantageous to set forth definitions of certain words and phrases used throughout this patent document. The terms "send," "receive," and "communicate," and their derivatives, cover both direct and indirect communications. The terms "include," and "comprise," and their derivatives, mean including, but not limited to. The term "or" is inclusive, meaning and / or. The phrase "associated with," and its derivatives, means including, included within, interconnected with, containing, contained within, connected to or connected with, coupled to or coupled with, communicable with, collaborate with, interlace, juxtapose, be close to, bound to or bound with, have, have the nature of, have to or with, etc., a relationship, etc.

[0010] In addition, the various functions described below may be implemented or supported by one or more computer programs, each of which is formed by a computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, processes, functions, objects, classes, instances, related data, or a portion thereof, suitable for implementation in a suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as a read-only memory (ROM), a random access memory (RAM), a hard drive, a compact disk (CD), a digital video disk (DVD), or any other type of memory. "Non-transitory" computer-readable media excludes wired, wireless, optical, or other communication links that transmit temporary electrical or other signals. Non-transitory computer-readable media include media in which data can be permanently stored and media in which data can be stored and later rewritten, such as rewritable optical disks or erasable memory devices.

[0011] As used herein, terms and phrases such as "having", "may have", "include", or "may include" a feature (such as a number, function, operation, or component such as a part) indicate the presence of the feature without excluding the presence of other features. In addition, as used herein, the phrases "A or B", "at least one of A and / or B", or "one or more of A and / or B" may include all possible combinations of A and B. For example, "A or B", "at least one of A and B", and "at least one of A or B" may indicate all of the following: (1) including at least one A, (2) including at least one B, or (3) including at least one A and at least one B. In addition, as used herein, the terms "first" and "second" may modify various components regardless of importance and do not limit the components. These terms are only used to distinguish one component from another. For example, a first user device and a second user device may indicate user devices that are different from each other, regardless of the order or importance of the devices. Without departing from the scope of the present disclosure, a first component may be represented as a second component, and vice versa.

[0012] It will be understood that when an element (such as a first element) is referred to as being (operably or communicatively) "coupled" / "connected" / "coupled to" or "connected to" another element (such as a second element), it may be coupled to or connected to the other element directly or via a third element. Conversely, it will be understood that when an element (such as a first element) is referred to as being "directly coupled" or "directly connected" or "directly coupled" or "directly connected" to another element (such as the second element), no other elements (such as a third element) are interposed between the element and the other element.

[0013] As used herein, the phrase "configured (or set) to" may be used interchangeably with the phrases "suitable for", "capable of", "designed to", "adapted to", "manufactured to", or "capable of", depending on the circumstances. The phrase "configured (or set) to" does not essentially mean "specially designed in hardware". On the contrary, the phrase "configured to" may indicate that a device may perform an operation together with another device or part. For example, the phrase "a processor configured (or set) to perform A, B, and C" may indicate a general-purpose processor (such as a CPU or an application processor) that can perform operations by running one or more software programs stored in a memory device, or a dedicated processor (such as an embedded processor) for performing operations.

[0014] The terms and phrases used herein are only used to describe some embodiments of the present disclosure, rather than to limit the scope of other embodiments of the present disclosure. It should be understood that, unless the context clearly states otherwise, the singular forms "one", "an" and "the" include plural references. All terms and phrases used herein (including technical and scientific terms and phrases) have the same meanings as those generally understood by those of ordinary skill in the art to which the embodiments of the present disclosure belong. It will be further understood that terms and phrases (such as those defined in commonly used dictionaries) should be interpreted as having meanings consistent with their meanings in the context of the relevant art, and will not be interpreted in an idealized or overly formal sense, unless explicitly defined herein. In some cases, the terms and phrases defined herein may be interpreted as excluding embodiments of the present disclosure.

[0015] Examples of "electronic devices" according to embodiments of the present disclosure may include at least one of a smart phone, a tablet personal computer (PC), a mobile phone, a video phone, an e-book reader, a desktop PC, a laptop computer, a netbook computer, a workstation, a personal digital assistant (PDA), a portable multimedia player (PMP), an MP3 player, a mobile medical device, a camera, or a wearable device (such as smart glasses, a head mounted device (HMD), electronic clothing, an electronic bracelet, an electronic necklace, an electronic accessory, an electronic tattoo, a smart mirror, or a smart watch). Other examples of electronic devices include smart home appliances. Examples of smart home appliances may include a television, a digital video disc (DVD) player, an audio player, a refrigerator, an air conditioner, a vacuum cleaner, an oven, a microwave, a washing machine, a dryer, an air purifier, a set-top box, a home automation control panel, a security control panel, a TV box (such as SAMSUNG HOMESYNC, APPLETV, or GOOGLE TV), a smart speaker or speaker with an integrated digital assistant (such as SAMSUNG GALAXY HOME, APPLE HOMEPOD, or AMAZON ECHO), a game console (such as XBOX, PLAYSTATION, or NINTENDO), an electronic dictionary, an electronic key, a camera, or at least one of an electronic photo frame. Other examples of electronic devices include at least one of the following: various medical devices (such as various portable medical measuring devices (such as blood sugar measuring devices, heart rate measuring devices, or body temperature measuring devices), magnetic source angiography (MRA) devices, magnetic source imaging (MRI) devices, computed tomography (CT) devices, imaging devices, or ultrasound devices), navigation devices, global positioning system (GPS) receivers, event data recorders (EDRs), flight data recorders (FDRs), automotive infotainment devices, navigation electronic devices (such as navigation navigation devices or gyrocompasses), avionics equipment, security equipment, vehicle-mounted head units, industrial or household robots, automatic teller machines (ATMs), point-of-sale (POS) devices, or Internet of Things (IoT) devices (such as light bulbs, various sensors, electric or gas meters, sprinklers, fire alarms, thermostats, street lights, ovens, fitness equipment, hot water tanks, heaters, or boilers). Other examples of electronic devices include at least a portion of a piece of furniture or a building / structure, an electronic board, an electronic signature receiving device, a projector, or various measuring devices (such as devices for measuring water, electricity, gas, or electromagnetic waves). Note that according to various embodiments of the present disclosure, the electronic device may be one or a combination of the devices listed above. According to some embodiments of the present disclosure, the electronic device may be a flexible electronic device. The electronic devices disclosed herein are not limited to the devices listed above, and may include new electronic devices depending on technological developments.

[0016] In the following description, an electronic device is described with reference to the accompanying drawings according to various embodiments of the present disclosure. As used herein, the term "user" may refer to a person using an electronic device or another device (such as an artificial intelligence electronic device).

[0017] Definitions for certain other words and phrases may be provided throughout this patent document. Those of ordinary skill in the art should understand that in many, if not most instances, such definitions apply to prior, as well as future uses of such defined words and phrases.

[0018] Nothing in the description of this application should be read as implying that any particular element, step, or function is an essential element that must be included within the scope of the claims. The scope of the patented subject matter is limited solely by the claims. Furthermore, no claim is intended to invoke 35 USC 112(f), unless the exact words “means for” are followed by a participle. The use of any other terms in the claims, including but not limited to “mechanism,” “module,” “device,” “unit,” “component,” “element,” “member,” “apparatus,” “machine,” “system,” “processor,” or “controller,” are understood by applicants to refer to structures known to persons skilled in the relevant art and are not intended to invoke 35 U.S.C. 112(f). BRIEF DESCRIPTION OF THE DRAWINGS

[0019] For a more complete understanding of the present disclosure and its advantages, reference is now made to the following description taken in conjunction with the accompanying drawings, wherein like reference numerals represent like parts:

[0020] Figure 1 An example network configuration including electronic devices according to the present disclosure is shown;

[0021] Figure 2 An example system for efficient language model editing using a contextual prompt generator (CPG) according to the present disclosure is shown;

[0022] Figure 3 The method for training in accordance with the present disclosure is shown. Figure 2 Example process of using CPG in the system;

[0023] Figure 4 The method for editing the Figure 2 Example process of using CPG in the system;

[0024] Figure 5 The embodiment according to the present disclosure is shown in Figure 2 Additional details of an example of a CPG used in a system;

[0025] Figure 6 It shows that according to the present disclosure Figure 5 Additional details of an example of an edit encoder module that is part of a CPG;

[0026] Figure 7 It shows that according to the present disclosure Figure 5 Additional details of an example of an editorial choice module that is part of a CPG;

[0027] Figure 8 The method for editing the Figure 2 Another example process of using a CPG in a system of; and

[0028] Fig. 9 An example method for efficient language model editing using a contextual hint generator according to the present disclosure is shown. DETAILED DESCRIPTION

[0029] The following discussion is described with reference to the accompanying drawings. Figures 1 to 9 However, it should be understood that the present disclosure is not limited to these embodiments, and all changes and / or equivalents or replacements thereto also belong to the scope of the present disclosure.

[0030] As mentioned above, a large language model (LLM) is a neural network model that takes a piece of natural language as input and generates one or more words as output. LLMs can contain many parameters (e.g., billions of parameters) and can be trained with Internet-scale text data. The scale of the model and data can enhance the capabilities of LLMs and improve the quality of the output, making LLMs a viable option for commercial use. Various use cases for LLM applications have been widely discussed and explored in the industry, such as chatbots for customer service, intelligent assistants for software tools, question-answering for knowledge query, interfaces for interactive communication between people and devices, etc. In general, there is a high demand for LLM commercial deployment.

[0031] Commercial use of LLMs still faces several challenges, including those related to model maintenance, safety assurance, and quality control. For example, ongoing model maintenance may be required because certain facts about the world change over time. Therefore, LLMs with the latest facts need to be maintained. As a specific example, when a user asks for the name of the current US President, the model should reply "Donald Trump" if asked in 2020, and "Joe Biden" if asked in 2023. Some safety assurances mean that the output of LLMs should not cause any harm in the physical world. As a specific example, a user may ask for instructions on using cooking utensils, and LLMs should not instruct users to wrap food in tin foil when using a microwave (which is a fire hazard), even when the same practice works well with a conventional oven.

[0032] Regarding quality control, before releasing an LLM to the market, the LLM is usually thoroughly tested by a quality assurance department to identify erroneous outputs for the target business use case. These errors should be corrected in a controlled manner in order to minimize the probability of the error reoccurring. This error correction process is not trivial, as LLMs are complex probabilistic models that can give different responses with slightly different language inputs.

[0033] The use of model editing can address some of these challenges. However, current model editing strategies have shortcomings in several aspects that make them unsuitable for real-world deployment. For example, some current techniques experience a rapid drop in accuracy after being updated with a small amount of edited data. That is, the outputs for many irrelevant inputs are often changed to incorrect outputs. In addition, some current techniques do not predict the correct output with the required inputs. In addition, some current techniques introduce significant overhead in training and inference. Therefore, training LLMs can be extremely expensive, such as millions of dollars per model.

[0034] The present disclosure provides various techniques for performing efficient language model editing using a contextual hint generator. As described in more detail below, the disclosed systems and methods provide efficient and controllable model editing, thereby enabling large-scale LLM editing at lower cost and with improved accuracy. It should be noted that while some of the embodiments discussed below are described in the context of use in consumer electronic devices (such as smart phones), this is merely an example, and it should be understood that the principles of the present disclosure can be implemented in any number of other suitable contexts and using any suitable devices.

[0035] Figure 1 An example network configuration 100 including electronic devices according to the present disclosure is shown. Figure 1The embodiment of the network configuration 100 shown is for illustration only. Other embodiments of the network configuration 100 may be used without departing from the scope of the present disclosure.

[0036] According to an embodiment of the present disclosure, an electronic device 101 is included in a network configuration 100. The electronic device 101 may include at least one of a bus 110, a processor 120, a memory 130, an input / output (I / O) interface 150, a display 160, a communication interface 170, or a sensor 180. In some embodiments, the electronic device 101 may exclude at least one of these components, or may add at least one other component. The bus 110 includes circuits for connecting the components 120-180 to each other and for transmitting communications (such as control messages and / or data) between the components.

[0037] The processor 120 includes one or more processing devices, such as one or more microprocessors, microcontrollers, digital signal processors (DSPs), application specific integrated circuits (ASICs), or field programmable gate arrays (FPGAs). In some embodiments, the processor 120 includes one or more of a central processing unit (CPU), an application processor (AP), a communication processor (CP), or a graphics processor unit (GPU). The processor 120 is capable of performing control on at least one of the other components of the electronic device 101 and / or performing operations or data processing related to communication or other functions. As described in more detail below, the processor 120 can use a contextual hint generator to perform one or more operations for efficient language model editing.

[0038] The memory 130 may include volatile and / or non-volatile memory. For example, the memory 130 may store commands or data related to at least one other component of the electronic device 101. According to an embodiment of the present disclosure, the memory 130 may store software and / or programs 140. The programs 140 include, for example, a kernel 141, middleware 143, an application programming interface (API) 145, and / or an application program (or "application") 147. At least a portion of the kernel 141, the middleware 143, or the API 145 may be represented as an operating system (OS).

[0039] The kernel 141 may control or manage system resources (such as bus 110, processor 120, or memory 130) for executing operations or functions implemented in other programs (such as middleware 143, API 145, or application 147). The kernel 141 provides an interface that allows the middleware 143, API 145, or application 147 to access various components of the electronic device 101 to control or manage system resources. The application 147 may use a contextual prompt generator as described below to support one or more functions for efficient language model editing. These functions may be performed by a single application or multiple applications, each of which performs one or more of these functions. For example, the middleware 143 may be used as a relay to allow the API 145 or application 147 to communicate data with the kernel 141. Multiple applications 147 may be provided. The middleware 143 is capable of controlling work requests received from the application 147, such as by assigning a priority to at least one of the multiple applications 147 using the system resources (such as bus 110, processor 120, or memory 130) of the electronic device 101. The API 145 is an interface that allows the application 147 to control functions provided from the kernel 141 or the middleware 143. For example, the API 145 includes at least one interface or function (such as a command) for archive control, window control, image processing, or text control.

[0040] The I / O interface 150 serves as an interface that can, for example, transmit commands or data input from a user or other external devices to other components of the electronic device 101. The I / O interface 150 can also output commands or data received from other components of the electronic device 101 to the user or other external devices.

[0041] The display 160 includes, for example, a liquid crystal display (LCD), a light emitting diode (LED) display, an organic light emitting diode (OLED) display, a quantum dot light emitting diode (QLED) display, a microelectromechanical system (MEMS) display, or an electronic paper display. The display 160 may also be a depth perception display, such as a multi-focal display. The display 160 is capable of displaying, for example, various contents (such as text, images, videos, icons, or symbols) to the user. The display 160 may include a touch screen and may receive, for example, a touch, gesture, proximity, or hovering input using an electronic pen or a body part of the user.

[0042] For example, the communication interface 170 can establish communication between the electronic device 101 and an external electronic device (such as the first electronic device 102, the second electronic device 104, or the server 106). For example, the communication interface 170 can be connected to the network 162 or 164 through wireless or wired communication to communicate with the external electronic device. The communication interface 170 can be a wired or wireless transceiver or any other component for sending and receiving signals.

[0043] Wireless communication can use, for example, WiFi, long term evolution (LTE), long term evolution-advanced (LTE-A), fifth generation wireless system (5G), millimeter wave or 60GHz wireless communication, wireless USB, code division multiple access (CDMA), wideband code division multiple access (WCDMA), universal mobile telecommunications system (UMTS), wireless broadband (WiBro) or global system for mobile communications (GSM) as a communication protocol. Wired connection can include, for example, universal serial bus (USB), high definition multimedia interface (HDMI), recommended standard 232 (RS-232) or plain old telephone service (POTS) at least one. Network 162 or 164 includes at least one communication network, such as a computer network (such as a local area network (LAN) or a wide area network (WAN)), the Internet or a telephone network.

[0044] The electronic device 101 also includes one or more sensors 180, which can measure physical quantities or detect the activation state of the electronic device 101, and convert the measured or detected information into electrical signals. For example, one or more sensors 180 may include one or more cameras or other imaging sensors for capturing images of scenes. The sensor 180 may also include one or more buttons for touch input, gesture sensors, gyroscopes or gyroscope sensors, air pressure sensors, magnetic sensors or magnetometers, acceleration sensors or accelerometers, grip sensors, proximity sensors, color sensors (such as red, green, and blue (RGB) sensors), biophysical sensors, temperature sensors, humidity sensors, illumination sensors, ultraviolet (UV) sensors, electromyography (EMG) sensors, electroencephalography (EEG) sensors, electrocardiography (ECG) sensors, infrared (IR) sensors, ultrasonic sensors, iris sensors, or fingerprint sensors. The sensor 180 may also include an inertial measurement unit, which may include one or more accelerometers, gyroscopes, and other components. In addition, the sensor 180 may include a control circuit for controlling at least one of the sensors included here. Any of these sensors 180 may be located within the electronic device 101.

[0045] In some embodiments, the electronic device 101 may be a wearable device or a wearable device (such as an HMD) on which an electronic device may be mounted. For example, the electronic device 101 may represent an AR wearable device such as a headset with a display panel or smart glasses. In other embodiments, the first external electronic device 102 or the second external electronic device 104 may be a wearable device or a wearable device (such as an HMD) on which an electronic device may be mounted. In those other embodiments, when the electronic device 101 is mounted in the electronic device 102 (such as an HMD), the electronic device 101 may communicate with the electronic device 102 through the communication interface 170. The electronic device 101 may be directly connected to the electronic device 102 to communicate with the electronic device 102 without involving a separate network.

[0046] The first external electronic device 102 and the second external electronic device 104 and the server 106 may each be a device of the same or different type as the electronic device 101. According to certain embodiments of the present disclosure, the server 106 includes a group of one or more servers. In addition, according to certain embodiments of the present disclosure, all or some operations performed on the electronic device 101 may be performed on another or more other electronic devices (such as the electronic devices 102 and 104 or the server 106). In addition, according to certain embodiments of the present disclosure, when the electronic device 101 should automatically or upon request perform some functions or services, the electronic device 101 may request another device (such as the electronic devices 102 and 104 or the server 106) to perform at least some functions associated with it, instead of performing the function or service itself, or additionally performing the function or service. Other electronic devices (such as the electronic devices 102 and 104 or the server 106) are able to perform the requested function or additional function and transmit the result of the execution to the electronic device 101. The electronic device 101 may provide the requested function or service by processing the received result as is or additionally. To this end, for example, cloud computing, distributed computing, or client-server computing technology may be used. Although Figure 1 The electronic device 101 is shown to include a communication interface 170 that communicates with the external electronic device 104 or the server 106 via the network 162 or 164 , but according to some embodiments of the present disclosure, the electronic device 101 may operate independently without a separate communication function.

[0047] The server 106 may include components 110-180 (or a suitable subset thereof) that are the same or similar to the electronic device 101. The server 106 may support driving the electronic device 101 by performing at least one of the operations (or functions) implemented on the electronic device 101. For example, the server 106 may include a processing module or processor that may support the processor 120 implemented in the electronic device 101. As described in more detail below, the server 106 may perform one or more operations to support a technique for efficient language model editing using a contextual hint generator.

[0048] although Figure 1 An example of a network configuration 100 including an electronic device 101 is shown, but the Figure 1 Various changes may be made. For example, network configuration 100 may include any number of each component in any suitable arrangement. In general, computing and communication systems have a wide variety of configurations, and Figure 1 The scope of the present disclosure is not limited to any particular configuration. Figure 1 One operating environment is shown in which the various features disclosed in this patent document may be used, but these features may be used in any other suitable system.

[0049] Figure 2 An example system 200 for efficient language model editing using a contextual hint generator according to the present disclosure is shown. For ease of explanation, the system 200 is described as using the above Figure 1 The system 200 is implemented by one or more components of the network configuration 100 (such as the electronic device 101). However, this is merely an example, and the system 200 may be implemented using any other suitable device (such as the server 106) and in any other suitable system.

[0050] like Figure 2 As shown, using system 200, electronic device 101 receives input 202 of large language model (LLM) 204. In some embodiments, input 202 is text input provided by a user to electronic device 101, such as by speaking into a microphone of electronic device 101, typing on a touch screen of electronic device 101, or using any other suitable input mechanism. As a specific example, in some cases, input 202 represents a question to be answered by LLM 204.

[0051] The LLM 204 may represent a large pre-trained language model. In some embodiments, the LLM 204 includes an encoder-decoder or decoder-only architecture. Before providing the input 202 to the LLM 204, the electronic device 101 generates word units 206 from the input 202 and generates one or more word unit embeddings 208 from the word units 206. The word unit embeddings 208 are feature vectors, each of which represents a portion of the input 202. The electronic device 101 provides the word unit embeddings 208 as input to the LLM 204, and the LLM 204 generates a response 210 (such as a prediction) associated with the input 202.

[0052] In some embodiments, input 202 is a question whose answer can change over time. For example, input 202 can be "Who is the current US president?" If the question is asked in 2020, the correct answer is "Donald Trump". If the question is asked in 2023, the correct answer is "Joe Biden". In some cases, the pre-trained LLM 204 is static, which means that LLM 204 includes an existing knowledge base that does not change over time. Specifically, the weights of the encoder and decoder of LLM 204 do not change over time. In this case, LLM 204 can generate a correct response 210 to input 202 "Who is the current US president?" in one year, but provide an incorrect response 210 in another year (once the US president changes).

[0053] To address this problem, the system 200 also includes a contextual prompt generator (CPG) 212. The CPG 212 receives the input 202 and generates a prompt 214 based on the input 202. The prompt 214 is a single word that represents new or updated information that is not included in the existing knowledge of the LLM 204. Using the above example of "Who is the current US president?", the prompt 214 represents updated information about the current US president. After the CPG 212 generates the prompt 214 based on the input 202, the CPG 212 generates at least one prompt embedding 216. Similar to the word embedding 208, the prompt embedding 216 is a feature vector representing the prompt 214, thereby representing the new or updated information that is not included in the existing knowledge of the LLM 204. The prompt embedding 216 can have the same number of dimensions as the word embedding 208 (e.g., 1024 dimensions per word). The prompt embedding 216 and the word embedding 208 generally represent real-valued vectors. An example of such a vector might be [0.23, 0.15, -0.87, ... , 1.1].

[0054] In some embodiments, CPG 212 includes a neural network having one or more transformers, which are standard blocks for building neural networks. As described in more detail below, CPG 212 generates a prompt embedding 216 from a collection of edit embeddings representing edit descriptors, where the edit descriptors include new or updated information. CPG 212 is trained to generate prompt embedding 216 based on edit descriptors and input 202. Although the following description discusses a single prompt 214 and a single prompt embedding 216, it will be understood that CPG 212 can generate multiple prompts 214 and multiple prompt embeddings 216.

[0055] After the CPG 212 generates the prompt embedding 216, the electronic device 101 cascades the prompt embedding 216 and the word-unit embedding 208. For example, the electronic device 101 can cascade the prompt embedding 216 as a prefix at the beginning of the word-unit embedding 208. As another example, the electronic device 101 can cascade the prompt embedding 216 as a suffix at the end of the word-unit embedding 208. The combination of the prompt embedding 216 and the word-unit embedding 208 is provided to the LLM 204 as an input. The added prompt embedding 216 affects the inference of the LLM 204 and causes the LLM 204 to generate an updated (ideally correct) response 210, thereby achieving the effect of model editing. In other words, the prompt embedding 216 provides information about a portion of the edit descriptor associated with the input 202, so that the LLM 204 can provide an updated response 210. Here, the updated response 210 reflects new or updated information represented by the prompt embedding 216 but not present in the existing knowledge of the LLM 204.

[0056] although Figure 2 One example of a system 200 for efficient language model editing using a contextual hint generator and related details is shown, but may be used for Figure 2 For example, although CPG 212 and LLM 204 are described as involving a specific sequence of operations, Figure 2 The various operations described may overlap, occur in parallel, occur in a different order, or occur any number of times (including zero). Furthermore, Figure 2 The specific operations shown in are examples only and may be performed using other techniques. Figure 2 Each operation shown in .

[0057] Figure 3 An example training process 300 for training CPG 212 according to the present disclosure is shown. Training process 300 may be performed before system 200 is deployed to a final product, such as electronic device 101. For ease of explanation, training process 300 is described as using a computer other than electronic device 101. Figure 1The training process 300 may be implemented using one or more components of the network configuration 100 (such as the server 106). However, this is just an example, and the training process 300 may be implemented using any other suitable device (such as the electronic device 101) and in any other suitable system. It is also possible that the same device (such as the server 106) trains and uses the CPG 212.

[0058] like Figure 3 , CPG 212 is trained using a set of training edit embeddings 302, which can be prepared in any suitable manner. In some cases, the set of training edit embeddings 302 is prepared by a model developer. Training edit embeddings 302 represent example edit descriptors, which may include new or updated information not included in the prior knowledge of LLM 204. Training edit embeddings 302 are provided as input to LLM 204 and CPG 212.

[0059] The LLM 204 operates on the training edit embedding 302 to generate one or more outputs 304, which represent responses to the training edit embedding 302. The outputs 304 are compared to the language model training targets 306, and one or more losses are calculated. The one or more losses represent the difference between the outputs 304 and the training targets 306. Here, each of the training targets 306 and the calculated losses represent any suitable training targets and calculated losses for training the LLM, respectively. In some cases, each calculated loss is converted into a gradient, which is back-propagated to the CPG 212. Based on the gradients, one or more parameters (such as weighting parameters) are updated in the CPG 212. Note here that only the parameters in the CPG 212 can be updated, while the parameters of the LLM 204 can remain static.

[0060] although Figure 3 One example of a training process 300 for training the CPG 212 is shown, but may be used for Figure 3 For example, although the training process 300 is described as involving a specific sequence of operations, Figure 3 The various operations described may overlap, occur in parallel, occur in a different order, or occur any number of times (including zero).In addition, although the training process 300 is described above as being performed by the server 106, all or a portion of the training process 300 may be implemented by another device.

[0061] Figure 4An example editing process 400 for editing a CPG 212 according to the present disclosure is shown. In contrast to the training process 300 (which may be performed prior to deployment), the editing process 400 may be performed at various times, including after the system 200 is deployed to a final product. According to an embodiment, the editing process 400 may be performed using Figure 1 The editing process 400 is implemented by one or more components of the network configuration 100, such as the electronic device 101 or a device other than the electronic device 101 (such as the server 106). However, this is merely an example, and the editing process 400 can be implemented using any other suitable device and in any other suitable system.

[0062] like Figure 4 As shown in FIG, CPG 212 is edited using one or more edit inputs 402 provided by a user of the final product. Each of edit inputs 402 represents at least one edit descriptor that may include new or updated information not included in the existing knowledge of LLM 204. Edit inputs 402 are tokenized and encoded into embeddings that are provided as input to LLM 204 and CPG 212.

[0063] The LLM 204 operates on the edit input 402 to generate one or more outputs 404, which represent responses to the edit input 402. The outputs 404 are compared to a language model training target 406 (which can be the same as or similar to the language model training target 306), and one or more losses are calculated. The one or more losses represent the difference between the output 404 and the training target 406. Here, each of the training target 406 and the calculated losses represent any suitable training target and calculated loss for training the LLM, respectively. In some cases, each calculated loss is converted into a gradient, which is back-propagated to the CPG 212. Based on the gradient, one or more parameters (such as weighting parameters) are updated in the encoder of the CPG 212. Again, it is noted here that only the parameters in the CPG 212 can be updated, while the parameters of the LLM 204 can remain static.

[0064] although Figure 4 One example of an editing process 400 for editing a CPG 212 is shown, but may be used for Figure 4 For example, although the editing process 400 is described as involving a specific sequence of operations, Figure 4 The various operations described may overlap, occur in parallel, occur in a different order, or occur any number of times (including zero). In addition, although the editing process 400 is described above as being performed by the server 106 or the electronic device 101, all or part of the editing process 400 may be implemented by another device.

[0065] Figure 5 FIG. 2 shows additional details of an example of a CPG 212 according to the present disclosure. In this example, the CPG 212 does not require any model training during the editing process, such as Figure 4 As shown in Figure 5 , CPG 212 includes an edit encoder module 502, a storage 504 for encoded edits, and an edit selection module 506. Edit encoder module 502 receives one or more edits 508, such as from a user or model developer, and encodes edits 508. Storage 504 is used by CPG 212 to store the output of edit encoder module 502. In some embodiments, storage 504 represents a standard file system in volatile and / or non-volatile memory (such as memory 130). Depending on the embodiment, storage 504 can be on the device, in the cloud, or in any other suitable location. Edit selection module 506 compares input 202 and encoded edits 508 to generate prompts 214.

[0066] Figure 6 1 shows additional details of an example of an edit encoder module 502 according to the present disclosure. In this example, the edit encoder module 502 represents a trainable neural network module including multiple layers of transformer blocks. Figure 6 As shown, edit encoder module 502 receives edit 508 (in Figure 6 Identified as ). Each edit 508 includes an edit descriptor, such as text that describes the edit 508. The edit descriptors may be obtained from any suitable source, such as a model developer or one or more users. In some embodiments, each edit descriptor is formatted as a sentence or a question and answer, where the answer corresponds to the question. As a specific example, the answer to each question may be different from the corresponding answer to the same question in the prior knowledge of the LLM 204. For example, one edit 508 may include the text "Who is the current president of the USA? Joe Biden" and another edit 508 may include the text "When is tax day? October 16, 2023"

[0067] Edit 508 is provided as input to edit encoder module 502, which encodes edit 508 to generate corresponding edit embedding 602 (in Figure 6 Identified as ). Edit embedding 602 may represent a real-valued vector stored in storage 504. A representative example of edit embedding 602 may be the vector [0.3, 0.9, -2.1, 0, ... , -0.5]. However, each vector may have any suitable value and dimension. As described above, storage 504 may include a standard file system on a server, in the cloud, or on a user device.

[0068] Figure 7 1 shows additional details of an example of an edit selection module 506 according to the present disclosure. Figure 7 As shown, edit selection module 506 obtains a set of edit embeddings 602 from storage 504. In some embodiments, edit selection module 506 loads all or a subset of saved edit embeddings 602 from storage 504 into a temporary memory location. Edit selection module 506 also obtains input 202 from a user.

[0069] Edit selection module 506 includes an edit encoder 702 that encodes input 202 to generate an intermediate input embedding 704. In some embodiments, edit encoder 702 may represent edit encoder module 502, although this is not necessarily the case. Edit encoder 702 may represent any suitable language model encoder, such as a base encoder of BERT or LLM 204. The generated intermediate input embedding 704 may represent a real-valued vector having the same number of dimensions as edit embedding 602. A representative example of intermediate input embedding 704 may be the vector [-0.3, -0.1, 0.2, 0.5, ... , 0].

[0070] The edit selection module 506 also includes a crisscross attention layer 706 that receives the set of edit embeddings 602 and the intermediate input embeddings 704 and outputs the intermediate edit embeddings 708. For example, the crisscross attention layer 706 can generate the intermediate edit embeddings 708 by comparing the similarity between the intermediate input embeddings 704 and each of the edit embeddings 602. The crisscross attention layer 706 selects the edit embedding 602 that is most similar to the intermediate input embedding 704 and outputs the selected edit embedding 602 as the intermediate edit embedding 708. In some cases, the crisscross attention layer 706 represents a trainable neural network layer and can have any suitable architecture, such as the same or similar architecture as the crisscross attention layer in the standard transformer layer. Once the intermediate edit embedding 708 is determined using the crisscross attention layer 706, the edit selection module 506 applies a multi-layer perceptron (MLP) layer 710 to the intermediate edit embedding 708 to generate the hint embedding 216. In some cases, the MLP layer 710 represents a small trainable neural network with one or more linear layers.

[0071] Because the intermediate edit embedding 708 (and the corresponding hint embedding 216) can be within or outside the range of the input 202 (such as related or unrelated), the edit selection module 506 also determines the likelihood that the intermediate input embedding 708 is related to the input 202. For example, the edit selection module 506 can use a cascade operation 712 to cascade the intermediate input embedding 704 and the intermediate edit embedding 708, and apply another MLP layer 714 to the cascade result to generate a confidence value 716. In some cases, the MLP layer 714 represents a small trainable neural network with one or more linear layers. The confidence value 716 can represent a real value (such as 0.2 or 0.87) that represents the likelihood that the intermediate edit embedding 708 and the corresponding hint embedding 216 are related to the input 202.

[0072] The edit selection module 506 performs a gating operation 718 that compares the confidence value 716 to a specified threshold (such as 0.5) representing the minimum likelihood confidence that is being taken. If the confidence value 716 is greater than or equal to the threshold, the edit selection module 506 outputs the prompt embedding 216. Otherwise, the edit selection module 506 is not confident about the prompt embedding 216 and may output an empty prompt embedding 216a. The empty prompt embedding 216a allows the LLM 204 to generate the response 210 without interference from the CPG 212.

[0073] although Figures 5 to 7 An example of a CPG 212 and related details is shown, but may be Figures 5 to 7 For example, although the edit encoder module 502 and the edit selection module 506 are described as involving a specific sequence of operations, Figures 5 to 7 The various operations described may overlap, occur in parallel, occur in a different order, or occur any number of times (including zero). Furthermore, Figures 5 to 7 The specific operations shown in are examples only and may be performed using other techniques. Figures 5 to 7 Each operation shown in .

[0074] Figure 8 The method for editing the Figure 2 Another example process 800 of using the CPG 212 in the system 200. The editing process 800 may use Figure 1 The editing process 800 may be implemented using one or more components of the network configuration 100, such as the electronic device 101. However, this is merely an example, and the editing process 800 may be implemented using any other suitable device (such as the server 106) and in any other suitable system.

[0075] like Figure 8As shown in , CPG 212 is edited by receiving one or more edit inputs 802, which may be obtained from any suitable source, such as a user. Edit input 802 may represent one or more edit descriptors that include new or updated information not contained in the existing knowledge of LLM 204. Edit encoder module 502 receives edit input 802 and encodes edit input 802 into edit embeddings, which are stored as a set of updated edit embeddings in storage 504. In contrast to editing process 400, editing process 800 does not involve model training, which may make editing process 800 fast and computationally efficient, and therefore suitable for operation on a user device, such as a smart phone or tablet.

[0076] although Figure 8 Shows the Figure 2 Another example of a process 800 for using a CPG 212 in the system 200, but may be used for Figure 8 For example, although the editing process 800 is described as involving a specific sequence of operations, Figure 8 The various operations described may overlap, occur in parallel, occur in a different order, or occur any number of times (including zero). In addition, although the editing process 800 is described above as being performed by the electronic device 101 , all or part of the editing process 400 may be implemented by another device (such as the server 106 ).

[0077] It should be noted that Figures 2 to 8 The functions shown in or described above may be implemented in any suitable manner in the electronic device 101, 102, 104, the server 106, or other devices. For example, in some embodiments, one or more software applications or other software instructions executed by the processor 120 of the electronic device 101, 102, 104, the server 106, or other devices may be used to implement or support Figures 2 to 8 In other embodiments, dedicated hardware components may be used to implement or support at least some of the functions shown or described above. Figures 2 to 8 At least some of the functions shown or described above in . Generally, any suitable hardware or any suitable combination of hardware and software / firmware instructions may be used to implement Figures 2 to 8 In addition, Figures 2 to 8 The functions shown or described above may be performed by a single device or by multiple devices. For example, the server 106 may be used to train one or more components, and the server 106 may deploy one or more trained components to one or more other devices (such as the electronic device 101) for use.

[0078] Fig. 9An example method 900 for efficient language model editing using a contextual hint generator according to the present disclosure is shown. For ease of explanation, Fig. 9 The method 900 is described as using Figure 2 The system 200 shown Figure 1 The electronic device 101 shown is used for execution. However, Fig. 9 The method 900 shown in FIG. 9 may be used with any other suitable device or system.

[0079] like Fig. 9 As shown, in step 901, an input of the LLM is received from a user. This may include, for example, the electronic device 101 receiving an input 202 of the LLM 204 from a user. In step 903, one or more word element embeddings are generated based on the input. This may include, for example, the electronic device 101 generating one or more word element embeddings 208 based on the input 202. In step 905, one or more prompt embeddings are generated based on the input using the CPG. This may include, for example, the electronic device 101 generating one or more prompt embeddings 216 based on the input 202 using the CPG 212. The one or more prompt embeddings 216 represent new or updated information that is not included in the existing knowledge of the LLM 204.

[0080] At step 907, one or more word-gram embeddings and one or more prompt embeddings are provided to the LLM. This can include, for example, the electronic device 101 providing one or more word-gram embeddings 208 and one or more prompt embeddings 216 as inputs to the LLM 204. In some cases, the one or more word-gram embeddings 208 and the one or more prompt embeddings 216 can be cascaded. At step 909, a prediction is output based on the one or more word-gram embeddings and the one or more prompt embeddings using the LLM. This can include, for example, the electronic device 101 using the LLM 204 to output a predicted response 210 based on the one or more word-gram embeddings 208 and the one or more prompt embeddings 216. The prediction reflects new or updated information represented by the one or more prompt embeddings 216.

[0081] although Fig. 9 An example of a method 900 for efficient language model editing using a contextual hint generator is shown, but may be used for Fig. 9 For example, although shown as a series of steps, Fig. 9 The various steps in can overlap, occur in parallel, occur in a different order, or occur any number of times (including zero).

[0082] Although the present disclosure has been described with reference to various exemplary embodiments, various changes and modifications may be suggested to one skilled in the art. The present disclosure is intended to encompass such changes and modifications as fall within the scope of the appended claims.

Claims

1. A method comprising: receiving input from a user into a large language model (LLM); generating one or more word embeddings based on the input; generating, using a contextual cue generator (CPG), one or more cue embeddings based on the input, the one or more cue embeddings representing new or updated information not contained in existing knowledge of the LLM; providing the one or more word-gram embeddings and the one or more hint embeddings to the LLM; as well as Outputting a prediction based on the one or more word-gram embeddings and the one or more cue embeddings using the LLM, wherein the prediction reflects the new or updated information represented by the one or more cue embeddings.

2. The method according to claim 1, wherein: Generating the one or more prompt embeddings includes: obtaining a collection of edit embeddings representing edit descriptors, wherein the edit descriptors include the new or updated information; selecting an edit embedding from the set of edit embeddings that is most similar to the input; determining a likelihood that the selected edit embedding is associated with the input; and In response to the likelihood meeting or exceeding a threshold, generating the one or more hint embeddings based on the selected edit embedding.

3. The method according to claim 2, wherein: The set of edit embeddings is generated by an edit encoder of the CPG and is stored in memory after being generated by the CPG.

4. The method according to claim 3, wherein: Each edit descriptor includes a question and an answer corresponding to the question; as well as The answer for each edit descriptor is different from a corresponding answer to the same question in the prior knowledge of the LLM.

5. The method according to claim 2, wherein: The edited embedding that is most similar to the input is selected using criss-cross attention.

6. The method according to claim 1, wherein: The CPG is trained using a set of training edit embeddings representing edit descriptors, wherein the edit descriptors include the new or updated information.

7. The method according to claim 1, further comprising: editing the CPG in a user editing mode before receiving the input of the LLM from the user; Wherein, editing the CPG comprises: receiving one or more edit inputs, the one or more edit inputs comprising one or more edit descriptors, the one or more edit descriptors comprising the new or updated information; and An encoder of the CPG is updated based on the one or more editing inputs.

8. An electronic device comprising: At least one processing device configured to: receiving input from a user into a large language model (LLM); generating one or more word embeddings based on the input; generating, using a contextual cue generator (CPG), one or more cue embeddings based on the input, the one or more cue embeddings representing new or updated information not contained in existing knowledge of the LLM; providing the one or more word-gram embeddings and the one or more hint embeddings to the LLM; as well as Outputting a prediction based on the one or more word-gram embeddings and the one or more cue embeddings using the LLM, wherein the prediction reflects the new or updated information represented by the one or more cue embeddings.

9. The electronic device according to claim 8, wherein: To generate the one or more hint embeddings, the at least one processing device is configured to: obtaining a collection of edit embeddings representing edit descriptors, wherein the edit descriptors include the new or updated information; selecting an edit embedding from the set of edit embeddings that is most similar to the input; determining a likelihood that the selected edit embedding is associated with the input; and In response to the likelihood meeting or exceeding a threshold, generating the one or more hint embeddings based on the selected edit embedding.

10. The electronic device according to claim 9, wherein: The at least one processing device is configured to: generating the set of edit embeddings using an edit encoder of the CPG; and After the edit-embedded set is generated, the edit-embedded set is stored in a memory.

11. The electronic device according to claim 10, wherein: Each edit descriptor includes a question and an answer corresponding to the question; as well as The answer for each edit descriptor is different from a corresponding answer to the same question in the prior knowledge of the LLM.

12. The electronic device according to claim 9, wherein: The at least one processing device is configured to select the edit embedding that is most similar to the input using cross attention.

13. The electronic device according to claim 8, wherein: The at least one processing device is configured to train the CPG using a set of training edit embeddings representing edit descriptors, wherein the edit descriptors include the new or updated information.

14. The electronic device according to claim 8, wherein: The at least one processing device is further configured to: edit the CPG in a user editing mode before receiving the input of the LLM from the user; as well as In order to edit the CPG, the at least one processing device is configured to: receiving one or more edit inputs, the one or more edit inputs comprising one or more edit descriptors, the one or more edit descriptors comprising the new or updated information; as well as An encoder of the CPG is updated based on the one or more editing inputs.

15. A non-transitory machine-readable medium comprising instructions that, when executed, cause at least one processor of an electronic device to: receiving input from a user into a large language model (LLM); generating one or more word embeddings based on the input; generating, using a contextual cue generator (CPG), one or more cue embeddings based on the input, the one or more cue embeddings representing new or updated information not contained in existing knowledge of the LLM; providing the one or more word-gram embeddings and the one or more hint embeddings to the LLM; as well as Outputting a prediction based on the one or more word-gram embeddings and the one or more cue embeddings using the LLM, wherein the prediction reflects the new or updated information represented by the one or more cue embeddings.