Generating content items based on source document metadata using generative neural networks
By incorporating metadata from source electronic documents into a generative neural network, the problem of insufficient quality and relevance of generated content items is solved, achieving the generation of higher quality and more relevant content items.
Patent Information
- Application Number
- CN202511437042.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-10-08
- Filing Date
- 2025-10-09
- Publication Date
- 2026-01-13
AI Technical Summary
In existing technologies, generative neural networks fail to effectively utilize the metadata of the source electronic document when generating content items, resulting in insufficient quality and relevance of the generated content items.
By incorporating metadata from the source electronic document into the prompts of the generative neural network, and utilizing metadata such as creation date, publication date, and source, the performance of the generative neural network can be improved, thereby generating higher quality and more context-sensitive content items.
The generated content items are more informative and relevant to the context input, improving the user experience.
Smart Images

Figure CN121328482A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the use of generative neural networks to generate content items based on source document metadata. Background Technology
[0002] This specification relates to using neural networks to process input to generate content items. For example, content items may include text data, image data, video data, audio data, etc.
[0003] A neural network is a machine learning model that uses one or more layers of non-linear units to predict the output from a given input. In addition to the output layer, some neural networks also include one or more hidden layers. The output of each hidden layer serves as the input to the next layer in the network (i.e., another hidden layer or output layer). Each layer of the network generates its output from the received input based on the current values of its corresponding parameter set. Summary of the Invention
[0004] This specification describes a content item generation system implemented as a computer program on one or more computers in one or more locations, which uses a generative neural network to generate content items based on source document metadata.
[0005] According to one aspect, a method performed by one or more computers is provided. The method includes: receiving from a user a request to generate a content item using a generative neural network conditioned on contextual input, wherein the contextual input includes content derived from a source electronic document; obtaining metadata associated with the source electronic document; generating a prompt for the generative neural network based on the contextual input and the metadata associated with the source electronic document; processing the prompt using the generative neural network to generate the content item; and providing the content item to the user.
[0006] Generating the prompt for this generative neural network can include generating a prompt that includes the context input and the metadata.
[0007] Generating the prompt for the generative neural network may include generating a prompt that includes the context input, the metadata, and additional information, the additional information being generated based on the metadata associated with the source electronic document and the metadata associated with the generative neural network.
[0008] The metadata associated with the generative neural network may include the deadline of the generative neural network, which indicates the most recent release date of the data included in the training data used to train the generative neural network.
[0009] Obtaining metadata associated with the source electronic document may include receiving the metadata from an operating system running on the one or more computers.
[0010] Obtaining metadata associated with the source electronic document may include: performing a search in a document corpus to identify electronic documents related to the content; and using the metadata associated with the identified electronic documents as metadata associated with the source electronic document.
[0011] The content exported from the source electronic document may include content copied from the source electronic document.
[0012] According to another aspect, one or more computer-readable storage media encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform the operations described above.
[0013] According to another aspect, a system is provided, comprising: one or more computers; and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to perform corresponding operations in respect of the methods described above.
[0014] The subject matter described in this specification can be implemented in specific embodiments to achieve one or more of the following advantages.
[0015] Generally, content item generation systems can use generative neural networks to generate content items based on contextual inputs that provide the context for the content item. Typically, the contextual inputs provided to the system include content derived from (e.g., transmitted) a source electronic document. Metadata associated with the source electronic document is usually ignored and not incorporated into the content item generation process, even if it may contain additional information that could be used to improve the performance of the generative neural network to generate content items with improved quality or better relevance to the contextual inputs.
[0016] Using the techniques described in this specification, a content item generation system can incorporate metadata related to the source electronic document into prompts for processing by a generative neural network. This enables the network to generate higher-quality content items—for example, more informative and context-sensitive items. This ability to generate higher-quality content items also improves the user experience of the content item generation system.
[0017] Metadata obtained by the content item generation system using the techniques described in this specification can include information such as the source of the source electronic document, its creation date and time, and its publication date. Therefore, the metadata encodes additional information relevant to the generation task (e.g., information richer than the exported content itself or information that can attribute the exported content to a specific source), but this metadata is missing from the context input initially received by the system. Thus, content items generated by the system can be more informative and more relevant to the context input than content items generated solely based on the context input.
[0018] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of this subject matter will become apparent from the description, drawings, and claims. Attached Figure Description
[0019] Figure 1 This is a diagram of the example content item generation system.
[0020] Figure 2 It is a diagram of an example environment that includes a content item generation system.
[0021] Figure 3 This is an example illustration of a content item generation system that generates content items based on prompts including the exported content and metadata.
[0022] Figure 4 This is a flowchart of an example process for generating content items based on source document metadata using a generative neural network.
[0023] The same reference numerals and names in the various figures indicate the same elements. Detailed Implementation
[0024] Figure 1 This is a diagram of an example content item generation system 100. The content item generation system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations, wherein the systems, components and techniques described below can be implemented.
[0025] The content item generation system 100 is a system that uses a generative neural network 110 to generate content items 152, at least conditioned on the context input 104.
[0026] For example, the content item generation system 100 may receive contextual input 104 as part of or associated with a request 102 for content item 152, and in response, use a generative neural network 110 to process a prompt 108 including the contextual input 104 to generate content item 152.
[0027] The content item generation system 100 can generate any type of content item 152, such as text content items, image content items, video content items, audio content items, etc.
[0028] The content item 152 generated by the content item generation system 100 can be used in any of a variety of ways. For example, the system can provide the content item 152 to be presented to a user on a display device. As another example, the system can provide the content item 152 to another component in the system or a different system for further processing. As yet another example, the system can store the content item 152 in a data repository for some future purpose.
[0029] In some cases, the content item generation system 100 may be a text generation system that generates a sequence of text items, i.e., each content item 152 generated by the system is an output text sequence comprising a sequence of text lexicons from a vocabulary of text lexicons, which includes one or more of characters, subwords, words, punctuation marks, numbers, or other symbols appearing in natural or computer languages. For example, the system may generate a text sequence in response to contextual input 104 provided by a user of the system and provide the text sequence to the user who provided the contextual input 104.
[0030] For example, contextual input 104 could be an input text sequence, and the output sequence could be another text sequence, such as a translation of the input text sequence, completion of the input text sequence, interpretation of the input text sequence, a response to a question posed in the input sequence, or a text sequence about a topic specified by the input text sequence. As another example, contextual input 104 could be input other than text, such as an image, video, or audio, and the output sequence could be text describing the input.
[0031] As a specific example, the content item generation system 100 may be part of a dialogue system, and the context input 104 may include audio or text from the most recent conversation round submitted by the user during the dialogue, while the output text sequence is the next round in the conversation, for example, text or audio as a response to the most recent conversation round. Optionally, the context input 104 may also include one or more historical conversation rounds that occurred earlier in the conversation.
[0032] As another specific example, the content item generation system 100 may be part of a computer code generation system, and the context input 104 may be a textual description of the desired code segment or a computer code snippet in a programming language, and the output text sequence may be computer code, for example, a code snippet described by the context input or a code snippet in a computer program following the context input.
[0033] In some cases, the content item generation system 100 may be an image or video generation system that generates images or videos, each having multiple frames (where each frame is an image), by generating images (e.g., as a sequence of pixels or through an iterative denoising process). For example, the content item generation system 100 may generate images or videos conditional on contextual input 104 provided by the system's user, which includes a textual description of the image or video content.
[0034] In some cases, the content item generation system 100 may be an audio generation system that generates audio signals. For example, each content item 152 is an output audio sample that includes audio wave samples at each of a series of output time steps spanning a specified time window. For instance, the content item generation system 100 may generate audio conditionally based on contextual input 104 provided by the system's user, which includes a textual description of the audio content.
[0035] In these cases, output time steps can be arranged at regular intervals within a specified time window. The audio sample at a given output time step can be the amplitude value of an audio wave, or the amplitude value that has been compressed, compressed, or both. For example, the audio sample can be the original amplitude value or a mu-law compressed representation of the amplitude value.
[0036] In many scenarios, context input 104 includes content derived (e.g., transmitted) from source electronic document 106. An electronic document is data that presents a set of content. Examples of electronic documents include web pages, word processing documents, portable document format (PDF) documents, images, videos, audio, source code files, and feed sources. Native applications (e.g., "apps" and / or applications) such as those installed on mobile, tablet, or desktop computing devices are also examples of electronic documents.
[0037] Source electronic document 106 has associated metadata. Examples of such metadata include: the date and time of creation or last modification of source electronic document 106, the publication date of source electronic document 106, the title or file name of source electronic document 106, the owner or author (individual or organization) of source electronic document 106, the release or version number of source electronic document 106, the source of source electronic document 106, such as the network location of source electronic document 106, such as Uniform Resource Locator (URL), source hostname, source domain name, or source Internet Protocol (IP) address, etc.
[0038] The content exported from the source electronic document 106 also has associated metadata. Examples of such metadata include: the date and time the content was copied from the source electronic document 106, the location of the content within the source electronic document 106 (e.g., line number or byte offset within the source electronic document 106), the surrounding context of the content, the size of the content, and so on.
[0039] In these scenarios, the performance of the generative neural network 110 can be improved in many content generation applications (e.g., conditional text, image, video, or audio generation tasks) when it is provided not only with content derived from the source electronic document 106, but also with metadata associated with the source electronic document 106, the content, or both, as well as other possible metadata (e.g., metadata associated with the generative neural network 110 itself).
[0040] For a number of reasons, when metadata is provided to the generative neural network 110 as part of the prompt 108 for processing, it can improve the performance of the generative neural network 110 to facilitate the generation of higher quality content items, such as more informative content items 152 that are more relevant to the context input 104.
[0041] First, the metadata includes richer information than the exported content itself, and thus provides additional context for the content items to be generated by the generative neural network 110.
[0042] Secondly, metadata can attribute the exported content to a specific source, and thus facilitate the generation of content items as a more relevant response to the contextual input 104 that includes the exported content, since the generative neural network 110 can now access information about the source location and creation time of the exported content.
[0043] As an example, a user transmits a text fragment (e.g., a news article about a news event) from a webpage to contextual input 104, and contextual input 104 is an input text sequence that represents a question posed about the transmitted text fragment (e.g., a question about a news event) or another request made in reference to the transmitted text fragment (e.g., a request to summarize a news article). In various cases, when providing contextual input 104, the user can transmit the text fragment by using an input device through copy-paste, cut-paste, or inputting verbatim or paraphrased text.
[0044] In this example, metadata about when a webpage with a text fragment was created (and thus implies when a news event occurred) can enable the generative neural network 110 to generate more relevant content items, such as output sequences representing answers more relevant to the text fragment (e.g., answers more relevant to the news event) or more accurate summaries of the news article.
[0045] Assuming the text fragment was created after the knowledge deadline of the generative neural network 110, in some implementations, the generative neural network 110 can generate the output text sequence by using a search engine (or another external tool that can retrieve external data), and more specifically, it can generate the output text sequence based on the latest information included in the search engine results (or external data retrieved by another external tool).
[0046] In other words, the generative neural network 110 can avoid generating output text sequences by relying solely on stale information available to the generative neural network during training, and thus avoid generating incorrect or at least outdated and therefore unreliable output text sequences.
[0047] In this example, metadata about the source location of the webpage with the exported content (e.g., where the news article was found) can also enable the generative neural network 110 to generate more informative content items, such as an output sequence representing a more accurate factual answer to a question posed about a news event.
[0048] Assuming the metadata indicates that the text fragment originates from a source associated with a high level of factual accuracy, the generative neural network 110 can process the contextual input 104 including the text fragment to generate an output text sequence. Alternatively, assuming the metadata indicates that the text fragment originates from a source associated with a low level of factual accuracy, the generative neural network 110 can generate an output text sequence indicating that the text fragment lacks factual accuracy (e.g., “I don’t think this is accurate. Rather, here’s what I know about this…”).
[0049] As another example, a user copies an image from an image source (e.g., from a camera app, photo album app, or image / video processing app), and then pastes the copied image into context input 104. Context input 104 is a multimodal input sequence of both text and image, representing a question or request about the copied image. In this example, metadata about the image's capture time and / or location can enable the generative neural network 110 to generate more relevant content items.
[0050] For example, metadata can enable the generative neural network 110 to generate an output text sequence that represents an answer or another response (e.g., a text description) that is more relevant to the copied image than an output text sequence generated without utilizing metadata.
[0051] As another example, metadata can enable the generative neural network 110 to generate another image with higher quality (e.g., higher fidelity) compared to an image generated without utilizing metadata. For example, the other image could be a modified version of the copied image, such as a super-resolution image (with a higher resolution than the copied image), a repaired image (reconstructing any missing parts of the copied image), or an image that is the predicted next frame of the copied image.
[0052] As another example, a user transmits a fragment of a text report from a specific software tool to context input 104, and context input 104 represents a question or request made about the text fragment transmitted from the specific software tool.
[0053] For example, the software tool could be an integrated development environment (IDE) tool, and the text snippet could include source code snippets, such as incomplete source code snippets, buggy source code snippets, and so on. As another example, the software tool could be a compiler, and the text snippet could include error messages included as part of a bug report generated by the compiler after the compilation process of the source code has stopped due to an error that occurred in the source code.
[0054] In this example, metadata about a particular software tool (e.g., compiler version) and metadata about the source code (e.g., the owner or source of the source code) can enable the generative neural network 110 to generate more relevant content items, such as output sequences that are more relevant to the text fragment (e.g., more relevant answers to questions about why the source code fragment has a bug / failed to compile, or more accurate completion of an incomplete source code fragment).
[0055] To this end, after the user provides contextual input 104, which includes content exported (e.g., transferred) from the source electronic document 106, the content item generation system 100 collects metadata associated with the source electronic document 106, metadata associated with the exported content, or both, and may collect other metadata (e.g., metadata associated with the generative neural network 110) based on the contextual input 104, and then incorporates that metadata into the prompt 108 before processing the contextual input 104 using the generative neural network 110.
[0056] For example, prompt 108 may include (i) context input 104 (provided by the user) and one or more of the following: (ii) metadata associated with the source electronic document 106 (obtained by the system), (iii) metadata associated with the exported content (obtained by the system), or (iv) metadata associated with the generative neural network 110.
[0057] In this way, the prompt 108 contains richer information, including metadata that the user did not provide as part of or directly associated with the request 102 for content item 152, thus improving the performance of the generative neural network 110 to facilitate the generation of higher quality content item 152.
[0058] The generative neural network 110 can be any suitable generative neural network that has a generative neural network parameter set and can be used to generate content items including unimodal or multimodal data by processing the cue 108 according to the generative neural network parameter set.
[0059] In some implementations, the generative neural network 110 may have an architecture that allows it to more efficiently map a cue 108, including (i) context input 104 and (ii) metadata, to content items. As an example, the generative neural network 110 may include: a cue encoder subnetwork that processes the context input 104 to generate an embedding of the context input 104; a metadata encoder subnetwork that processes the metadata to generate an embedding of the metadata; and a core subnetwork that processes the embedding of the context input 104 and the embedding of the metadata to generate content items.
[0060] In some implementations, the generative neural network 110 may include a language model neural network, for example, as the core sub-network mentioned above or as another component of the generative neural network 110, which performs an autoregressive lexical generation process to autoregressively generate content items 152 across multiple time steps, for example, text lexical sequences, pixel lexical sequences, audio lexical sequences, multimodal lexical sequences (e.g., text and pixel lexical sequences, etc.), for example, by generating one lexical at each time step conditioned on any lexical that has been generated in a previous time step.
[0061] Language model neural networks can be any of a variety of Transformer-based neural network architectures, such as encoder-only Transformer architecture, encoder-decoder Transformer architecture, decoder-only Transformer architecture, other attention-based architectures, and so on.
[0062] Examples of neural networks for language models include those described in the following: Colin Raffel, et al., Exploring the limits of transfer learning with a unified text-to-text transformer. arXiv preprint arXiv:1910.10683, 2019; Daniel Adiwardana, et al., Towards a human-like open-domain chatbot. CoRR, abs / 2001.09977, 2020; Tom B Brown, et al., Language models are few-shot learners. arXiv preprint arXiv:2005.14165, 2020 (Tom B Brown et al., "Language Models as Few-Trial Learners", arXiv preprint arXiv:2005.14165, 2020); Aakanksha Chowdhery, et al., PaLM: Scaling Language Modeling with Pathways, arXiv preprint arXiv:2204.02311 (Aakanksha Chowdhery et al., "PaLM: Scaling Language Modeling with Pathways", arXiv preprint arXiv:2204.02311); Rohan Anil, et al., Palm 2 technical report. arXiv preprint arXiv:2305.10403, 2023 (Rohan Anil et al., "Palm 2 Technical Report", arXiv preprint arXiv:2305.10403, 2023); Borsos, Zalán, et al., Audiolm: a language modeling approach to audio generation.IEEE / ACM Transactions on Audio, Speech, and Language Processing (2023) (Borsos, Zalán et al., “Audiolm: A Language Modeling Approach for Audio Generation,” IEEE / ACM Transactions on Audio, Speech, and Language Processing (2023)); and Agostinelli, Andrea, et al., “Musiclm: Generating music from text.” arXiv preprint arXiv:2301.11325 (2023)
[0063] In some implementations, the generative neural network 110 may include a diffusion model neural network, for example, as a core sub-network mentioned above or as another component of the generative neural network 110, which performs a reverse diffusion process to iteratively generate content items 152, such as images, videos or audio, from random noise across multiple reverse diffusion steps.
[0064] For example, a diffusion model neural network can generate an image by performing a reverse diffusion process to produce a diffusion output that includes or otherwise specifies multiple color values of pixels in the image arranged in a specified order.
[0065] As another example, a diffusion model neural network can generate an image by performing a reverse diffusion process to produce a diffusion output that includes or otherwise specifies multiple words representing image patch embeddings of the image, which can then be processed by a decoder neural network to generate the image.
[0066] Examples of diffusion model neural networks include those described in the following: Chitwan Saharia, et al., Photorealistic text-to-image diffusion models with deep language understanding. Advances in Neural Information Processing Systems, 35:36479–36494, 2022; Aditya Ramesh, et al., Hierarchical text-conditional image generation with clip latents. arXivpreprint arXiv:2204.06125; and Robin Rombach, et al., High-resolution image synthesis with latent diffusion model, Proceedings of the IEEE / CVF conference. On computer vision and pattern recognition. 2022 (Robin Rombach et al., "High-resolution image synthesis using latent diffusion models", Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, 2022).
[0067] As another example, diffusion model neural networks can generate images or videos with multiple frames (where each frame is an image) by iteratively predicting masked lexical units in a discrete lexical space during the decoding process, for example, as in Huiwen Chang, et al., Muse: Text-to-image generation via masked generative transformers. arXiv preprint arXiv:2301.00704, 2023 and Huiwen Chang, et al., Maskgit: Masked generative image transformer. arXiv preprint arXiv:2202.04200, 2022. As described in arXiv:2202.04200, 2022.
[0068] Figure 2 It includes Figure 1 A diagram of an example environment 200 for a content item generation system 100 and a computing device 160. Additionally, environment 200 includes at least one of the following: a search engine 120, an operating system 130, or a training system 140.
[0069] In some implementations, environment 200 may include a network, such as a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof. When a network is included, the network connects content item generation system 100, computing device 160, and one or more of the following: search engine 120, operating system 130, or training system 140.
[0070] The content generation system 100 can operate in conjunction with an artificial intelligence software application 162 (or simply "AI application 162") installed on a computing device 160. Examples of computing devices 160 include personal computers, gaming devices, mobile communication devices, digital assistant devices, augmented reality (AR) devices, virtual reality (VR) devices, wearable devices, and other electronic devices.
[0071] Users can interact with the content item generation system 100 using the AI application 162. For example, users can use the AI application 162 to provide contextual input 104 as part of or associated with a request 102 for content item 152 through the input device of the computing device 160, and the generative neural network 110 can provide content item 152 to be presented to the user within the AI application on the display device of the computing device 160.
[0072] Examples of input devices include keyboards, mice, microphones, AR / VR input devices, touchscreens, and so on. Examples of output devices include monitors, screens, speakers, and so on. Output devices can be used to display images, text, video, and / or play audio to the user.
[0073] Temporary shift Figure 3 This is an example illustration 300 of a content item generation system 100 that generates content items based on prompts including exported content 105 and metadata 107. A user interacts with the content item generation system 100 using an AI application 162 to generate prompts 108, and a generative neural network 110 processes the prompts 108 to generate a content item 152, which can then be presented within the AI application 162.
[0074] Note 108 includes content exported from the source electronic document 106. For example, such as... Figure 3 As shown, the exported content 105 may include text or image data transferred (e.g., copied or moved) by the user from the source electronic document 106.
[0075] Typically, prompt 108 also includes a request made by the user with reference to the exported content 105. For example, when the exported content 105 includes text data, the request could be a request to translate, interpret, summarize, expand, analyze, or otherwise process the exported content 105. Alternatively, the request could be a request to generate some other modality of data (e.g., image data, video data, or audio data) conditioned on the exported content 105. As another example, when the exported content 105 includes image data, the request could be a request to generate text descriptions or some other description of the exported content 105.
[0076] In response to the user providing exported content 105, the content item generation system 100 obtains metadata 107 based on the exported content 105. To obtain the metadata 107, the content item generation system 100 may interact with one or more of the following: a search engine 120, an operating system 130, or a training system 140 included in the environment 200.
[0077] Search engine 120 can be any suitable search engine accessible to content item generation system 100 and capable of searching any appropriate document corpus (e.g., web pages, books, or other documents). For example, search engine 120 can be an internet search engine that searches for and returns results referencing electronic documents available on the internet. As another example, search engine 120 can be a different search engine searching a private document corpus (e.g., electronic documents available on an internal network or stored in a collection of one or more databases).
[0078] In an implementation where search engine 120 is included in environment 200, in response to computing device 160 providing contextual input 104 including the exported content to content item generation system 100, content item generation system 100 can use search engine 120 to perform a search in a document corpus based on the exported content to identify electronic documents related to the exported content, for example, electronic documents that meet a relevance threshold regarding the exported content. Then, content item generation system 100 can use the identified electronic documents as source electronic documents 106 and obtain metadata associated with source electronic documents 106.
[0079] When operating system 130 is included, it can run on computing device 160. The operating system provides an interface between the computing device's hardware (e.g., input / output devices and a processor that executes instructions retrieved from a computer-readable medium) and its software. The operating system provides a platform for the execution of various software applications on the computing device.
[0080] Software applications may include AI applications 162 as mentioned above, and one or more software applications provide transfer functionality. Examples of transfer functionality include copy commands, cut commands, paste commands, and so on. Examples of such software applications include word processing applications, spreadsheet applications, presentation applications, web browser applications, email applications, camera applications, photo album applications, image / video processing applications, and so on.
[0081] Examples of transmission may include content transmission between source electronic documents and software applications (e.g., between source electronic document 106 and an AI application that a user can use to interact with content item generation system 100), between two different electronic documents, between two different software applications, and so on.
[0082] A user of computing device 160 can select any content within a source electronic document 106 presented on a display device of computing device 160 within AI application 162 (or another software application), provide a transfer request to transfer the selected content to AI application 162 (optionally, via the system clipboard), and use AI application 162 to provide context input 104 to content item generation system 100 including the transferred content. For example, the transfer request may include a copy request to copy the selected content from source electronic document 106 to AI application 162, or a cut request to move the selected content from source electronic document 106 to AI application 162.
[0083] In an implementation in which the operating system 130 is included in the environment 200, in response to the computing device 160 providing the content item generation system 100 with context input 104 including the transmitted content, the content item generation system 100 may obtain, for example, metadata associated with the source electronic document 106, metadata associated with the transmitted content, or both, from the operating system 130 by issuing a data request to the operating system 130.
[0084] In some cases, the paste command may include an enhanced paste command that, in addition to the selected content, automatically transfers a predetermined set of content, such as that provided by a content provider. In these cases, the content item generation system 100 can still obtain metadata—for example, by issuing a data request to the operating system 130—which typically includes additional or different information compared to the information to be included in the predetermined content set. In fact, in some of these cases, the content item generation system 100 may remove the automatically transferred content, such that the context input includes the selected content but excludes the automatically transferred content.
[0085] The training system 140 is a system implemented as a computer program on one or more computers at one or more locations, which trains the generative neural network 110 to determine training values for a set of parameters of the generative neural network. That is, the generative neural network 110 has been trained by the training system 140 and configured to generate content items 152, including single-modality or multi-modality data, by processing cues 108 according to the training values of the generative neural network parameter set.
[0086] For example, the training system 140 can train the generative neural network 110 in two phases: a pre-training phase and a fine-tuning phase. In the pre-training phase, the training system 140 pre-trains the generative neural network 110 based on optimizing one or more unsupervised or self-supervised objective functions (e.g., maximum likelihood objective functions) on a large pre-training dataset.
[0087] Examples of large pre-trained datasets include: large datasets of text in one or more natural languages—for example, text publicly available from the Internet or another text corpus; large datasets of computer code in one or more programming languages—for example, computer code publicly available from the Internet or another code repository; large datasets of audio samples—for example, audio recordings or waveforms representing audio recordings; large datasets of images—where each image comprises an array of pixels; large datasets of videos—where each video comprises a sequence of time frames; or large multimodal datasets that include combinations of two or more of these datasets.
[0088] During the fine-tuning phase, the pre-trained generative neural network 110 is then tuned for the generation task through fine-tuning or another tuning technique (e.g., cue tuning or instruction tuning). The generation task can include any combination of one or more of the generation tasks mentioned above with other possible tasks. Examples of fine-tuning techniques include Supervised Fine-tuning (SFT), Human Feedback-Based Reinforcement Learning (RLHF), AI Feedback-Based Reinforcement Learning (RLAIF), etc., which use different training objectives, different fine-tuning datasets, or both.
[0089] Some implementations of the training system 140 may use low-rank tuning techniques or other techniques to achieve computationally efficient fine-tuning of the pre-trained generative neural network 110 by reducing the total number of parameter values that need to be learned during the fine-tuning phase.
[0090] The generative neural network 110, trained by the training system 140, has a knowledge cutoff date. The knowledge cutoff date represents the latest release date of the data included in the training data used to train the generative neural network 110. Therefore, the training data used by the training system 140 for training (e.g., a pre-training dataset or a fine-tuning dataset) includes only information up to the knowledge cutoff date, and not any latest information that only becomes available after the knowledge cutoff date.
[0091] In an implementation in which training system 140 is included in environment 200, content item generation system 100 can obtain metadata associated with generative neural network 110 from training system 140, including its knowledge expiration date and other possible information such as model version, model size, etc.
[0092] As part of the fine-tuning phase, in some implementations, the training system 140 can train the generative neural network 110 on a metadata-enhanced training dataset using any of the fine-tuning techniques mentioned above.
[0093] For example, when the generative neural network 110 is a language model neural network, the metadata-enhanced training dataset may include multiple training prompts. Each training prompt includes a training context input, which includes content derived from the source electronic document and training metadata, which includes metadata associated with the source electronic document.
[0094] As another example, when the generative neural network 110 is a diffusion model neural network, the metadata-enhanced training dataset may include multiple training content items, such as image content items, video content items, or audio content items. Each training content item is associated with training metadata, which includes metadata associated with the source electronic document from which the training content item was obtained.
[0095] Figure 4 This is a flowchart of an example process 400 for generating content items based on source document metadata using a generative neural network. For convenience, process 400 will be described as being executed by a system of one or more computers located in one or more locations. For example, a content item generation system appropriately programmed according to this specification (e.g., Figure 1 The content item generation system 100 described in the text is an executable process 400.
[0096] The system receives a request from a computing device to generate content items using a generative neural network conditioned on contextual input (step 402). The system can use the generative neural network to generate any type of content item, such as text data items, image data items, video content items, audio data items, etc.
[0097] Contextual input provides context for the content items to be generated by the generative neural network. Contextual input can include data provided by a user using an input device on a computing device. Specifically, contextual input includes content derived from a source electronic document. For example, contextual input can include text fragments (in natural or computer language), images (or blocks of images), videos (or frames of videos), or audio (or frames of audio) transmitted from the source electronic document.
[0098] The system obtains metadata associated with the source electronic document (step 404). For example, the system can obtain metadata by using a search engine to perform a search in a document corpus based on the exported content to identify electronic documents related to the exported content (e.g., electronic documents that meet a relevance threshold regarding the exported content). The system can then use the identified electronic document as the source electronic document and obtain the metadata associated with the source electronic document. As another example, the system can issue a data request to an operating system running on a computing device and obtain the metadata from the operating system.
[0099] Optionally, the system also obtains metadata associated with the exported content. For example, the system can similarly obtain the metadata from an operating system running on a computing device. Further optionally, the system also obtains metadata associated with the generative neural network. For example, the system can obtain the metadata from a training system that has trained the generative neural network or from a database that maintains metadata associated with the generative neural network.
[0100] The system generates a prompt for the generative neural network based on the context input, metadata associated with the source electronic document, and optionally metadata associated with the exported content or metadata associated with the generative neural network (step 406). Thus, the prompt includes not only the context input directly provided by the user but also additional information (metadata obtained by the system) not directly provided by the user.
[0101] In some implementations, the system can generate the prompt by concatenating the context input and the obtained metadata. For example, the prompt can take the following form: <context input(上下文输入)> <metadata(元数据)>, where "<context input>" represents the context input received from the user, the context input includes the content exported from the source electronic document, and " <metadata>"" indicates metadata associated with the source electronic document, metadata associated with the exported content, metadata associated with the generative neural network, or some combination thereof. In this example, <metadata>Can be arranged before <context input>.
[0102] As another example, the prompt can take the following form: <system prompt (system prompt)> <context input> <metadata>, in"<system prompt> "This can represent additional information generated by the system based on metadata." <systemprompt> 、 <metadata>and<context input> In this example, they can be arranged in different orders.
[0103] For example, the system can compare the knowledge deadline for generating the neural network with the publication date of the source electronic document and generate information indicating whether the knowledge deadline is earlier than the publication date as part of the system prompt.
[0104] As a concrete example of this situation, suppose the publication date of the source electronic document is "2024-May-01" and the knowledge cutoff date for the generative neural network is "2024-March-01". The information included in the system prompt can be formatted as a text sequence indicating the time difference: "the context is two months after knowledge cutoff".
[0105] As another example, the system can determine that the source electronic document was not included as part of a large pre-training dataset already used to pre-train the generative neural network. The source electronic document is also not included in the fine-tuning dataset used for subsequent fine-tuning of the generative neural network. In some cases, the system can make this determination after it has been determined that the specific content provider offering the source electronic document is not in the list of known content providers offering the data included in the pre-training and fine-tuning datasets.
[0106] In this example, the system can generate information as part of a system prompt indicating that the source electronic document was not included in the training data on which the generative neural network has been trained. In some cases, including such information will more likely lead the generative neural network to utilize a search engine (or another external tool) to generate content items more efficiently.
[0107] As a concrete example of this situation, suppose the system determines that the source electronic document was not included in the training data. The information included in the system prompt can be formatted as a text sequence that instructs the generative neural network to input contextual information not included in the training data used to train the neural network: "This information is from a set of data that was not used during the training."
[0108] In any example, the prompt may also include additional information, such as task-specific information or other system-generated information, such as a predefined set of system instructions.
[0109] Alternatively or alternatively,<system prompt> "This can refer to a predefined system prompt, which includes, for example, a predefined set of instructions on how to generate content items, a list of examples of content items to be generated, a list of search engine results obtained by using a search engine (or external data retrieved by another external tool) based on the content included in the context input, and so on."
[0110] In some implementations, the system may provide prompts to be presented to the user. If provided for presentation, the metadata may be presented to the user along with (e.g., together with) contextual input. Alternatively, the metadata may be presented separately from the contextual input, such as in a footnote. Optionally, the metadata may be displayed as a short summary that expands upon selection (e.g., double-click, hover, etc.) to show a more comprehensive presentation.
[0111] The system uses a generative neural network to process cues to generate content items (step 408). This generative neural network can generate any type of content item, such as text data items, image data items, video content items, audio data items, and so on.
[0112] In some implementations, the generative neural network has a generative neural network parameter set and can generate content items by processing prompts based on the generative neural network parameter set.
[0113] Generally, as mentioned above, including metadata in the prompt enables generative neural networks to generate more informative and context-relevant content items compared to those generated solely based on context input (even though the same set of generative neural network parameters is used).
[0114] In some implementations, the generative neural network has a generative neural network parameter set and multiple adaptation parameter sets corresponding to (or mapped to) multiple predetermined use cases. Each adaptation parameter set is an additional set of parameters that can be used in conjunction with the generative neural network parameter set to adapt the generative neural network to generate content items.
[0115] For example, a generative neural network may have: a first set of tuning parameters corresponding to a first use case, wherein the derived content included in the context input is after the knowledge deadline of the generative neural network; a second set of tuning parameters corresponding to a second use case, wherein the derived content included in the context input is before the knowledge deadline of the generative neural network, and so on.
[0116] As another example, a generative neural network may have: a first set of adaptation parameters corresponding to a first use case, wherein the content included in the context input is derived from a source electronic document provided by a first content provider; a second set of adaptation parameters corresponding to a second use case, wherein the content included in the context input is derived from a source electronic document provided by a second content provider, and so on.
[0117] In these implementations, the system may additionally utilize a classification neural network configured to categorize prompts into one of a plurality of predetermined use cases. For example, to generate content items, the system may first use the classification neural network to process metadata and optionally process contextual input to generate a classification output for a specific user case, and then use both (i) a set of generating neural network parameters and (ii) a specific set of adaptation parameters corresponding to that specific use case to process the prompts to generate content items.
[0118] The system provides content items to be presented to the user on the display device (step 410).
[0119] This specification uses the term "configured" in conjunction with system and computer program components. For a system of one or more computers to be configured to perform a specific operation or action, this means that the system has software, firmware, hardware, or a combination thereof installed thereon that causes the system to perform those operations or actions in operation. For one or more computer programs configured to perform a specific operation or action, this means that one or more programs include instructions that, when executed by a data processing device, cause that device to perform that operation or action.
[0120] Embodiments of the subject matter and functional operation described in this specification may be implemented in digital electronic circuit systems, in tangibly embodied computer software or firmware, in computer hardware (including the structures disclosed in this specification and their equivalents), or in one or more combinations thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, for example, one or more modules of computer program instructions encoded on a tangible, non-transitory storage medium for execution by a data processing device or for controlling the operation of a data processing device. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or one or more combinations thereof. Alternatively or additionally, program instructions may be encoded on artificially generated propagation signals (e.g., machine-generated electrical, optical, or electromagnetic signals) that are generated to encode information for transmission to a suitable receiver device for execution by the data processing device.
[0121] The term "data processing device" refers to data processing hardware and includes all kinds of devices, apparatuses, and machines for processing data, such as programmable processors, computers, or multiple processors or computers. The device may also be or further include a dedicated logic circuit system, such as a FPGA (Field-Programmable Gate Array) or an ASIC (Application-Specific Integrated Circuit). In addition to hardware, the device may optionally include code that creates an execution environment for computer programs, such as code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of these.
[0122] A computer program (which may also be referred to or described as a program, software, software application, app, module, software module, script, or code) can be written in any form of programming language, including compiled or interpreted languages or declarative or procedural languages, and can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but does not need to, correspond to a file in a file system. A program may be stored as a part of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), as a single file dedicated to the program in question, or as a plurality of coordinated files (e.g., a file storing portions of one or more modules, subroutines, or code). A computer program can be deployed to execute on a single computer or on multiple computers located at a single site or distributed across multiple sites and interconnected via a data communication network.
[0123] In this specification, the term "database" is used broadly to refer to any collection of data: data that does not need to be structured in any particular way, or does not need to be structured at all, and can be stored on storage devices in one or more locations. Thus, for example, an index database may include multiple collections of data, each of which can be organized and accessed differently.
[0124] Similarly, in this specification, the term "engine" is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines may be installed and run on the same one or more computers.
[0125] The processes and logic flows described in this specification can be executed by one or more programmable computers, which execute one or more computer programs to perform functions by manipulating input data and generating output. The processes and logic flows can also be executed by a dedicated logic circuit system (e.g., an FPGA or ASIC) or by a combination of a dedicated logic circuit system and one or more programmable computers.
[0126] A computer suitable for executing computer programs can be based on a general-purpose microprocessor or a special-purpose microprocessor or both, or any other type of central processing unit. Generally, the central processing unit receives instructions and data from read-only memory or random access memory or both. The basic components of a computer are the central processing unit for executing instructions and one or more memory devices for storing instructions and data. The central processing unit and memory may be supplemented by or incorporated into a special-purpose logic circuit system. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or operatively coupled to receive data from or transfer data to one or more mass storage devices or both. However, a computer need not have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive, to name a few.
[0127] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD ROMs and DVD-ROMs.
[0128] To provide interaction with the user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback, such as visual, auditory, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input. Furthermore, the computer can interact with the user by sending documents to and receiving documents from the device used by the user; for example, by sending a webpage to a web browser on the user's device in response to a request received from a web browser. Additionally, the computer can interact with the user by sending text messages or other forms of messages to a personal device (e.g., a smartphone running a messaging application) and receiving response messages from the user in response.
[0129] Data processing devices used to implement machine learning models may also include, for example, dedicated hardware accelerator units for handling the common and computationally intensive parts of machine learning training or production (i.e., inference, workloads).
[0130] Machine learning frameworks (such as TensorFlow or JAX) can be used to implement and deploy machine learning models.
[0131] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes back-end components (e.g., as a data server), or middleware components (e.g., an application server), or front-end components (e.g., a client computer having a graphical user interface, web browser, or app that a user can interact with through an implementation of the subject matter described in this specification), or any combination of one or more such back-end components, middleware components, or front-end components. Components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.
[0132] A computing system may include clients and servers. Clients and servers are generally geographically separated and typically interact via a communication network. The client-server relationship is established by computer programs executed on respective computers that establish a client-server relationship between them. In some embodiments, the server transmits data (e.g., HTML pages) to a user device, for example, for the purpose of displaying data to a user interacting with the device acting as a client and receiving user input from that user. Data generated at the user device, such as the result of user interaction, may be received at the server from the device.
[0133] While this specification contains numerous details of specific implementations, these details should not be construed as limiting the scope of any invention or the scope that may be claimed, but rather as descriptions of features that may be characteristic of particular embodiments of a particular invention. Certain features described in the context of individual embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from that claimed combination may be removed, and the claimed combination may involve sub-combinations or variations thereof.
[0134] Similarly, although operations are depicted in the accompanying drawings and described in a specific order in the claims, this should not be construed as requiring such operations to be performed in the specific order shown or in sequential order, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0135] Specific embodiments of this subject matter have been described. Other embodiments are within the scope of the appended claims. For example, the actions described in the claims can be performed in a different order and still achieve the desired result. As an example, the processes depicted in the drawings do not necessarily require a specific order or sequence shown to achieve the desired result. In some cases, multitasking and parallel processing can be advantageous.< / metadata> < / systemprompt> < / metadata> < / metadata> < / metadata>
Claims
1. A method performed by one or more computers, the method comprising: receiving, from a user, a request to generate a content item using a generative neural network conditioned on a contextual input, wherein the contextual input comprises content derived from a source electronic document; obtaining metadata associated with the source electronic document, wherein obtaining the metadata comprises receiving the metadata from an operating system running on the one or more computers; generating a prompt for the generative neural network based on the contextual input and the metadata associated with the source electronic document; processing the prompt using the generative neural network to generate the content item; and providing the content item for presentation to the user.
2. The method of claim 1, wherein generating the prompt for the generative neural network comprises: generating the prompt comprising the contextual input and the metadata.
3. The method of claim 1, wherein obtaining the metadata associated with the source electronic document comprises: performing a search in a corpus of documents to identify electronic documents related to the content; and using metadata associated with the identified electronic documents as the metadata associated with the source electronic document.
4. The method of claim 1, wherein the metadata associated with the generative neural network comprises a cutoff date for the generative neural network, the cutoff date representing a most recent publication date of data included in training data used to train the generative neural network.
5. The method of claim 1, wherein the content derived from the source electronic document comprises content copied from the source electronic document.
6. The method of any one of claims 1-5, wherein the generative neural network comprises a multi-modal neural network.
7. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform the method of any one of claims 1-6.
8. A non-transitory computer storage medium storing instructions that when executed by one or more computers cause the one or more computers to perform the method of any one of claims 1-6.
9. A method performed by one or more computers, the method comprising: receiving, from a user, a request to generate a content item using a generative neural network conditioned on a contextual input, wherein the contextual input comprises content derived from a source electronic document; obtaining metadata associated with the source electronic document; generating a prompt for the generative neural network based on the contextual input and the metadata associated with the source electronic document, wherein the prompt comprises the contextual input, the metadata, and additional information generated based on the metadata associated with the source electronic document and metadata associated with the generative neural network; processing the prompt using the generative neural network to generate the content item; and providing the content item for presentation to the user.
10. The method of claim 9, wherein the metadata associated with the generative neural network comprises a cutoff date for the generative neural network, the cutoff date representing a most recently published date for data included in training data used to train the generative neural network.
11. The method of claim 9, wherein the content derived from the source electronic document comprises content copied from the source electronic document.
12. The method of any one of claims 9 to 11, wherein the generative neural network comprises a multi-modal neural network.
13. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform the methods of any one of claims 9 to 12.
14. A non-transitory computer storage medium storing instructions that when executed by one or more computers cause the one or more computers to perform the methods of any one of claims 9 to 12.
Citation Information
Patent Citations
Generating images using sequences of generative neural networks
CN117561549A
Generating output sequence with inline evidence using linguistic model neural network
CN118715523A
Creating content items
US20200090220A1
Generative ai inferred prompt outpainting
US20240273670A1
Confidence enhancement for responses by document-based large language models
US20240296279A1