Knowledge Injection Model for Generative Commonsense Reasoning

By using the encoder-decoder model and knowledge injection model in the generative common sense reasoning technology, combining the generation of prototypes and scaling factors, the problem of generating unreasonable descriptions when processing concept sets in vacuum is solved, and a more reasonable and logical description generation is achieved.

CN116438529BActive Publication Date: 2025-06-03MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080107084.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-12
Publication Date
2025-06-03
Estimated Expiration
2040-11-12

AI Technical Summary

Technical Problem

Existing generative common sense reasoning techniques may generate unreasonable or meaningless descriptions when processing concept sets in a vacuum, making it difficult to prioritize certain concept combinations.

Method used

The encoder-decoder model is adopted in combination with the knowledge injection model, and by generating prototypes and combining them with the concept set, scaling factors and position indicators are generated to reduce the possibility of distorted model output of the prototype input token and adapt to the scene bias introduced by the generated prototype.

Benefits of technology

Improve the rationality and logic of generated descriptions, and enable more credible descriptions to be generated in the absence of additional context, reducing the user's cognitive and knowledge burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116438529B_ABST
    Figure CN116438529B_ABST
Patent Text Reader

Abstract

A knowledge injection model for generative commonsense reasoning. In an example, an encoder-decoder model is used to generate a model output (204), a reasonable description of a concept set. Prototypes (218) are generated from in-domain or out-of-domain knowledge corpora, and the prototypes are further used as inputs (202) to the encoder-decoder model. The concept input tokens and the prototype input tokens are scaled to limit potential biases that may be introduced by the prototypes (218). Additionally, position indicators are generated for each input token, and these position indicators indicate the relative position of each input token compared to other input tokens. In this way, when decoding the scaled and encoded input tokens, the decoder (214) can be more adaptable to the scenario biases introduced by the prototypes (218) when generating the model output (204). Therefore, when generating the model output (204), the encoder-decoder model does not need to rely solely on the concept set.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND OF THE INVENTION

[0001] Concept sets can be processed according to generative commonsense reasoning techniques to generate reasonable descriptions based on the concepts. However, processing concepts in a vacuum may not be sufficient to produce reasonable descriptions. Instead, in at least some cases, the resulting model output may be illogical or meaningless.

[0002] Embodiments have been described for these and other general considerations. Moreover, although relatively specific problems have been discussed, it should be understood that the embodiments should not be limited to solving the specific problems identified in the background. SUMMARY OF THE INVENTION

[0003] Aspects of the present disclosure relate to a knowledge injection model for generative commonsense reasoning. In an example, an encoder-decoder model is used to generate a model output (e.g., a reasonable description or descriptive sentence) based on an input that includes a concept set. A prototype is generated based on the concept set, and the prototype is also used as an input to the encoder-decoder model. The prototype can be generated from one or more in-domain and / or out-of-domain knowledge corpora. A scaling engine scales the input concept input tokens and the prototype input tokens to reduce the likelihood that the prototype input tokens that overlap with the concept input tokens distort the model output. For example, if a prototype input token is likely to be helpful in generation, the norm of the encoder output state associated with the prototype input token may be increased, while the norm may be decreased when there is a conflict between the prototype input token and the concept input token.

[0004] Additionally, a position indicator is generated for each input token, which provides an indication of the relative position of each corresponding input token compared to other input tokens. In this way, when decoding the scaled and encoded input tokens, the decoder can be more adaptable to the scenario bias introduced by the generated prototype when generating the model output. Thus, when generating the model output, the encoder-decoder model does not need to rely solely on the concept set, but can further incorporate the prototype generated from the knowledge corpus based on the techniques of immediate scaling and position indicators.

[0005] The present summary is provided to introduce some concepts in a simplified form that will be further described in the detailed description below. The present summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0006] References

[0007] The following publications are hereby incorporated by reference in their entirety:

[0008] 1. "An Enhanced Knowledge Injection Model for Commonsense Generation" paper (12 pages) (copy attached).

[0009] 2. Bill Yuchen Lin, Ming Shen, Wangchunshu Zhou, Pei Zhou, Chandra Bhagavatula, Yejin Choi, and Xiang Ren. 2019b. Commongen: A constrained text generation challenge for generative commonsense reasoning. CoRR, abs / 1911.03705.

[0010] 3. Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019. Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Non-limiting and non-exhaustive examples are described with reference to the following figures.

[0012] Figure 1 An overview of an example system in which the knowledge injection model described herein can be utilized is shown.

[0013] Figure 2 An overview of an example framework for generative commonsense reasoning according to the disclosed knowledge injection model is shown.

[0014] Figure 3 An overview of an example method for processing a set of concepts according to the disclosed knowledge injection model for generative commonsense reasoning is shown.

[0015] Figure 4 A block diagram of an example physical component of a computing device with which aspects of the present disclosure may be practiced is shown.

[0016] Figure 5A and Figure 5B is a simplified block diagram of a mobile computing device in which aspects of the present disclosure may be practiced.

[0017] Figure 6 is a simplified block diagram of a distributed computing system in which aspects of the present disclosure may be practiced.

[0018] Figure 7 illustrates a tablet computing device for performing one or more aspects of the present disclosure. Detailed Description

[0019] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof, and in which are shown by way of illustration specific embodiments or examples. These aspects may be combined, other aspects may be utilized, and structural changes may be made without departing from the present disclosure. The embodiments may be practiced as a method, system, or apparatus. Thus, the embodiments may take the form of a hardware implementation, a fully software implementation, or an implementation combining software and hardware aspects. Accordingly, the following detailed description should not be taken in a limiting sense, and the scope of the present disclosure is defined by the appended claims and their equivalents.

[0020] In an example, generative commonsense reasoning is used to generate plausible descriptions from a set of concepts. Compared to the set of concepts, the generated descriptions may enable improved data retrieval such that a larger amount and / or more accurate dataset responsive to a user query is identified. As another example, the generated descriptions may be more readily understandable by the user, or may be used as an alternative to requesting additional information from the user, thereby reducing the user's cognitive and knowledge burden and also reducing the time the user needs to spend inputting information. For example, descriptive sentences may be generated for an image based on a relevant set of tags (e.g., which may be generated using computer vision techniques and / or provided by the user). As another example, target content may be provided for a set of concepts, such that a descriptive title and / or descriptive summary may be generated for the target content. The descriptive title and / or summary may be used to identify the target content associated with the user, or as another example, a search query from the user may be used to generate a descriptive string for identifying such target content. Thus, it will be understood that generative reasoning and the related aspects described herein have applicability in a variety of contexts.

[0021] Examples of generative commonsense reasoning include, but are not limited to: Situations With Adversarial Generations (SWAG), CommonsenseQA, and CommonGen. For example, SWAG infers possible subsequent events based on a given textual description of an event. As another example, CommonsenseQA focuses on commonsense questions by describing relationships between concepts from a semantic network such as ConceptNet. Different from the discriminative tasks performed by SWAG and CommonsenseQA, CommonGen is an example of a computational generation ability trained based on background commonsense knowledge. Thus, it will be understood that aspects of the present disclosure apply to any generative commonsense reasoning context among various generative commonsense reasoning contexts.

[0022] For example, given a set of concepts such as "dog", "Frisbee", "catch", "throw", a reasonable description that may result is "the dog catches the Frisbee when the boy throws it". However, processing the set of concepts in a vacuum (e.g., lacking additional context) may result in an unreasonable description. For example, the generated description may be "two dogs are throwing Frisbees to each other". Thus, in the absence of additional context (e.g., dogs typically catch Frisbees or dogs cannot throw Frisbees), generative commonsense reasoning may not be able to prioritize certain concept combinations and may produce descriptions that are implausible or do not make logical sense.

[0023] Accordingly, aspects of the present disclosure relate to a knowledge injection model for generative commonsense reasoning. As an example, prototypes are generated from an in-domain knowledge corpus and / or an out-of-domain knowledge corpus based on a set of concepts. The prototypes are combined with the set of concepts to generate an input (e.g., including input tokens), and the input is processed using a pre-trained model. A scaling factor is assigned to each input token encoded by the model. In the example, the scaling factor is generated to reduce the attention weights of certain input tokens associated with the prototype, thereby reducing the likelihood that prototype input tokens overlapping with the concept input tokens receive distorted attention weights. Additionally, the surrounding input tokens can describe how the concepts interact. Thus, a position indicator can be generated for each input token, which provides an indication of the relative position of the input token compared to other input tokens. Accordingly, the decoder is more adaptable to the scenario bias introduced by the generated prototype when processing the encoded tokens in view of the position indicator to generate the model output.

[0024] Examples of using an encoder-decoder model, such as BART, are described herein. The encoder-decoder model can include one or more encoder layers, where each encoder layer is composed of a self-attention network and a feed-forward network. The encoder layer can also include an encoder-decoder attention mechanism between the self-attention network and the feed-forward network. Similarly, the encoder-decoder model can include one or more decoder layers, where each decoder layer can include a self-attention network and a feed-forward network. Although examples are discussed in the context of using the BART encoder-decoder model, it will be understood that any of a variety of other generative models (e.g., including an encoder and a decoder) can be used.

[0025] The following provides a set of example equations associated with the encoder-decoder attention mechanism. An input token set (e.g., including a set of concepts and associated generated prototypes) can be encoded into a hidden state or an encoder output sequence

[0026]

[0027]

[0028]

[0029]

[0030] In the example equation above, d u is input into the decoder, while h v is the output from the encoder. Additionally, x represents the x-th attention head, and are the trainable parameters of the query, key, and value, d is the size of the hidden state sequence, d k is the attention head dimension, and LN is the layer normalization function.

[0031] A concept set can be in any of a variety of forms. For example, a concept set can be received from a user or generated from sentences. In some examples, the concept set is a set of keywords received from a user or generated from metadata, or other instances. Accordingly, a prototype is generated based on the concept set. As used herein, the prototype includes background knowledge to improve the model output from an encoder-decoder model. The prototype can be a sentence or search snippet associated with search results in response to a user's search query, or other instances. The prototype can be generated from an in-domain and / or out-of-domain knowledge corpus. For example, if the concept set is related to common scenarios, example in-domain and out-of-domain knowledge corpora include but are not limited to VaTex (Wang et al., 2019), SNLI (Bowman et al., 2015), Activity (Krishna et al., 2017), or the training set of commonGen.

[0032] However, such an in-domain corpus may be difficult to generalize to other domains. Thus, an out-of-domain knowledge corpus (e.g., Wikipedia, websites, or social networks) (e.g., as an alternative or addition to the in-domain knowledge corpus) can be used to generate prototypes. One or more information retrieval techniques can be used to generate prototypes from the knowledge corpus, such as keyword search, exact or inexact matching techniques, or graph search techniques of an ontology graph database, etc., or other instances.

[0033] The generated prototype can be combined with the concept set to generate the input to the encoder-decoder model. For example, the input such that the input tokens can be However, there may be an overlap between the prototype input tokens and the concept input tokens, thus causing more attention weights to be given to certain input tokens and, in some examples, additional noise may be introduced.

[0034] Accordingly, a scaling engine can generate a scaling factor for each input token. In some examples, the scaling factor can be used as an alternative to using a simple hard mask that omits concept input tokens that are not also presented as prototype input tokens. If a prototype input token may be helpful for generation, then the scaling engine can increase the norm of the encoder output state associated with the prototype input token (e.g., h v in the above equation), while when there is a conflict between the prototype input token and the concept input token, the scaling engine can decrease the norm of the associated encoder output state. An example set of equations that can be used by the scaling engine is provided below.

[0035] Λ = Sigmoid(W 2 ReLU(W 1 h v +b1 ) + b 2 )

[0036] h v = h v ⊙(2 × Λ)

[0037] In this example, are trainable parameters used to adjust the scaling engine. In some cases, the parameters can be initialized to N(0, var), where var is a small value so that the scaling factors generated by the scaling engine do not significantly harm the operation of the encoder-decoder model.

[0038] Additionally, the prototype input tokens co-occurring with the output tokens of the encoder-decoder model may be more important than other tokens when generating the model output. Therefore, the encoder classification task can be used to enable the scaling engine to determine which tokens should appear in the generated output. An example loss function is shown below that the scaling engine of the encoder can use to perform such classification.

[0039]

[0040] In the example above, is an indicator function such that if then Or alternatively, when

[0041] As described above, in addition to (or, in some examples, as an alternative to) utilizing the scaling engine, position indicators can be generated to inform the decoder of the positions of the input tokens. Such position indicators can enable the decoder to more effectively identify and incorporate the scene biases that may be introduced through the prototype. For example, the position indicator for a given token can be determined based on its proximity to the concept input token.

[0042] For example, the conceptual input tokens within the input can each receive a value of "0", while the prototype input tokens can receive a value of "1" or greater. For a set of conceptual input tokens "dog" and "thrown", a prototype token including "the Frisbee was thrown to the dog" (which can alternatively be represented as a list including each token) can receive the position indicators 4, 3, 2, 1, 2, 2, 1. In this example, both "to" and "the" receive a position indicator of "2" because they are both close to prototype tokens that are also conceptual input tokens (e.g., "thrown" and "dog" respectively). Thus, the position indicators can be determined based on the minimum proximity to the conceptual input tokens. In an alternative, the second "the" would instead receive a position token of "3" associated with "thrown", rather than the position token of "2" associated with "dog" discussed earlier.

[0043] Thus, the generated set of position indicators can be incorporated into the encoder-decoder attention mechanism according to the following set of example equations. As shown, the above technique for generating position indicators for a given input token is implemented as a function D(s v ) and E D is the embedding of those distance values in D.

[0044] ED(h v ) = E D (D(s v ))

[0045]

[0046] Thus, incorporating ED(h v ) into the attention equation as shown enables the decoder to process the encoder output h vWhen applicable, it can incorporate associated location indicators to better learn the effective scenario deviations generated from the generated prototypes. For example, applying generative commonsense reasoning to the concept set "ear", "feel", "pain", "pierce" in a vacuum may produce an output similar to "I can feel the pain in my ears and feel the pierce in my neck from the piercing". However, incorporating the prototype of "if you pierce your hand, you also feel pain" injects additional knowledge into the processing performed by the encoder-decoder model, enabling the model to include scenario deviations when processing this concept set. Thus, the generated output can be replaced with "one feels the pain of having an ear pierced".

[0047] It will be understood that aspects of the present disclosure can be used during the generation phase (e.g., a pre-trained encoder-decoder model) and / or during the training phase. For example, the loss function can include as described above. The loss function can further incorporate the defined below to maximize the and log-likelihood of given

[0048]

[0049]

[0050] In the above example, t k is the k-th token in and t<k is the first (k - 1) tokens in and During model training, λ can be used to balance

[0051] It will be understood that aspects of the present disclosure are applicable in various contexts. For example, the disclosed knowledge injection model can be used in generative commonsense reasoning scenarios for generating descriptions based on concept sets. For example, computer vision techniques or user-submitted tags can be used to generate a tag set for an image, enabling a descriptive sentence of the image to be generated accordingly. The descriptive sentence can be provided to a client computing device as an alternative text label associated with the image.

[0052] As another example, target content can be provided for a concept set such that a descriptive title and / or descriptive summary can be generated for the target content. The target content can be provided to a user device based on a query from the user device that matches the descriptive title and / or descriptive summary. As a further example, a descriptive query can be generated from a concept set of a user query (e.g., as a search query string), where the user query is received from a user device. Target content can be identified on the descriptive query. Thus, the disclosed techniques can enable improved target content identification and distribution, thereby enabling identification of relevant content and presenting it to a customer that otherwise may not have been determined to be responsive to a user query. Additionally, the disclosed aspects can improve the associated user experience because a user does not need to provide as much information to a computer system, thereby reducing the user's cognitive and knowledge burden and also reducing the amount of time the user needs to spend entering information. Instead, generative commonsense reasoning techniques and associated knowledge injection models are used to supplement the amount of information used in order to generate a more complete representation of concepts that might have been provided by a user.

[0053] Figure 1 An overview of an example system 100 in which the knowledge injection model described herein can be utilized is shown. As shown, system 100 includes a server device 102, client devices 104, 106, a network 108, and an out-of-domain data source 110. In the example, the server device 102, the out-of-domain data source 110, and the client devices 104 and 106 communicate using the network 108, which can include a local area network, a wireless network, or the Internet, or any combination thereof, or other instances.

[0054] The server device 102 and the out-of-domain data source 110 can each be any computing device among a variety of computing devices, including but not limited to a server computing device or a set of computing devices that make up a distributed computing device. Similarly, each of the client devices 104 and 106 can be any computing device among a variety of computing devices, including but not limited to a mobile computing device, a laptop computing device, a tablet computing device, or a desktop computing device. It will be understood that while system 100 is shown as including one server device 102, two client devices 104 and 106, and one out-of-domain data source 110, any number of such elements can be used in other examples. Additionally, the functions described herein with respect to the server device 102, the client devices 104 and 106, and the out-of-domain data source 110 can be distributed among any number of different computing devices in any configuration among various configurations in other examples or otherwise implemented thereon. For example, client device 104 can include an out-of-domain data source similar to out-of-domain data source 110, which can be used as a knowledge corpus from which prototypes can be generated in accordance with the aspects disclosed herein.

[0055] Client device 104 is shown as including a client application 118, which can be any of a variety of applications, such as a web application executed in a web browser, a native application, or a combination thereof. For example, a user of client device 104 can use client application 118 to navigate to a website associated with server device 102, through which a set of concepts is provided. Similarly, client device 106 is shown as including a client application 120. Aspects of client device 106 are similar to aspects of client device 104 and thus need not be re-described in detail below.

[0056] As an example, client application 118 can display a website where a user can enter a query to search for content. The query can be transmitted to server device 102, which can extract a set of concepts from the query. The generative inference engine 112 can generate prototypes based on the set of concepts (e.g., from the in-domain data store 114, the out-of-domain data store 116, and / or the out-of-domain data source 110). Then, the generative inference engine 112 can generate a model output based on the input including the set of concepts and the generated prototypes. The model output can be used to identify target content associated with the user's query, which can be transmitted to client device 104 and presented by client application 118 along with search results in response to the user's search query. It will be understood that in other examples, it is not necessary to receive the set of concepts as a search query. For example, client application 118 can use an application programming interface (API) to provide the set of concepts to server device 102 and receive the model output and / or other associated processing results generated by the generative inference engine 112.

[0057] As another example, client application 118 can enable a user to enter a set of keywords associated with target content, which can be provided to server device 102 for processing according to the aspects described herein. The generative inference engine 112 can process the input and generate one or more model outputs that include a descriptive title and / or a descriptive summary of the target content associated with the set of concepts. In an example, the target content, the descriptive title, and / or the descriptive summary can be stored by server device 102 for subsequent use (e.g., to provide target content associated with search results in response to a user's search query). As another example, the set of concepts and the generated model outputs can be received and transmitted via an API, respectively. Thus, it will be understood that the disclosed aspects can be implemented according to any of a variety of paradigms (e.g., as a service via an API, according to a client / server approach, or locally to the client device, or other examples).

[0058] The server device 102 includes a generative inference engine 112, an in-domain data repository 114, and an out-of-domain data repository 116. The generative inference engine 112 processes a set of concepts to generate prototypes. Prototypes can be generated based on a knowledge corpus, which can be stored in or otherwise accessed from the in-domain data repository 114, the out-of-domain data repository 116, and / or the out-of-domain data source 110. For example, the out-of-domain data source 110 can be a third-party data source, such as a social network or an online knowledge base (e.g., an online encyclopedia or a knowledge base website), or other examples. In some instances, in-domain or out-of-domain data can be accessed or otherwise received from a client device. Thus, the knowledge corpus need not be limited to the server device 102. One or more information retrieval techniques can be used to generate prototypes from the knowledge corpus, such as keyword search, exact or inexact matching techniques, or graph search techniques of an ontology graph database, or other examples.

[0059] In an example, the generative inference engine 112 processes the set of concepts in combination with the generated prototypes to generate a model output according to aspects of the present disclosure. The concepts and prototypes form an input that includes the input tokens described herein. Example concepts include, but are not limited to, words, topics, or phrases. Thus, returning to the example above, concepts can be extracted from a search query based on word boundaries or based on identifying one or more topics therein, or other examples. The model output generated by the generative inference engine 112 can take any of a variety of forms. For example, the generative inference engine 112 can generate one or more sentences (e.g., a descriptive title or a descriptive summary in the example above), or can use the model output to subsequently identify relevant content (e.g., the target content in the example above). While example concepts and the resulting model output are described herein, it will be understood that any of a variety of other inputs and outputs can be used according to the techniques described herein.

[0060] Figure 2 An overview of an example framework 200 for generative commonsense reasoning according to the disclosed knowledge injection model is shown. As shown by the dashed box, the framework 200 can be implemented by Figure 1 the generative inference engine 112 therein. In an example, the framework 200 is based on an encoder-decoder model, such as BART.

[0061] The input 202 is a set of concepts, which in some examples can be received from a client device (such as Figure 1 the client devices 104 or 106 therein). The group embedding 206 includes a set of input tokens based on the input 202, which are shown as a set of concepts 216 and prototypes 218. For example, the prototype 218 can be generated by Figure 1The generative inference engine 112 within is generated based on an in-domain and / or out-of-domain knowledge corpus. In an example, for the concept and prototype a group embedding 206 can be generated according to the following example equation, where E B is the original BART embedding function.

[0062]

[0063] As shown, the group embedding 206 is processed by the encoder 208. For example, each encoder layer of the encoder 208 can consist of a self-attention network and a feed-forward network. The encoder layer can also include an encoder-decoder attention mechanism between the self-attention network and the feed-forward network. The scaling engine 210 further assigns a scaling factor to each input token of the concept set 216 and the prototype 218. As described above, if a prototype input token may contribute to generation, then the scaling engine 210 can increase the norm of the encoder output state associated with the prototype input token of the prototype 218. Conversely, when there is a conflict between the prototype input token of the prototype 218 and the concept input token of the concept set 216, the scaling engine 210 can decrease the norm of the associated encoder output state.

[0064] The position indicator generator 212 generates a position indicator for each input token of the input 202. Such a position indicator can enable the decoder 214 to more effectively identify and incorporate the scene bias that may be introduced by the prototype 218. As an example, the position indicator for a given token can be determined based on its proximity to an input token that is the same or similar to the concept.

[0065] The decoder 214 can include one or more decoder layers, where each decoder layer can include a self-attention network and a feed-forward network. In an example, for the encoded group embedding generated by the encoder 208, the decoder 214 generates the model output 204 based on the scaling factor generated by the scaling engine 210 and the position indicator generated by the position indicator generator 212. As described above, the scaling engine 210 ensures that the input tokens of the concept set 216 do not receive distorted attention due to potential overlap with the prototype 218. Additionally, since the decoder 214 incorporates the position indicator generated by the position indicator generator 212, the decoder 214 is more effective at incorporating the scene bias generated by the generated prototype compared to processing the concept set alone.

[0066] Figure 3 An overview of an example method 300 for processing a concept set according to the disclosed knowledge injection model for generative commonsense reasoning is shown. In an example, aspects of the method 300 are performed by the generative inference engine, such as Figure 1 and Figure 2The generative inference engine 112 in Figure 1 client devices 104 or 106 in

[0067] The process proceeds to operation 304, where a prototype is generated based on the concept set. In an example, the prototype is generated from an in-domain and / or out-of-domain knowledge corpus, which can be accessed from an out-of-domain data source (e.g., Figure 1 the out-of-domain data source 110 in

[0068] or stored by an in-domain data repository (e.g., the in-domain data repository 114) or an out-of-domain data repository (e.g., the out-of-domain data repository 116). One or more information retrieval techniques can be used to generate the prototype from the knowledge corpus, such as keyword search, exact or approximate matching techniques, or graph search techniques of an ontology graph database, or other examples.

[0069] In operation 306, the concept set and the generated prototype are treated as inputs to an encoder-decoder model and encoded accordingly. For example, aspects of operation 306 can be performed by an encoder, such as Figure 2 the encoder 208 in

[0070] The process proceeds to operation 308, where the encoder output is scaled based on the concept set and the generated prototypes. In an example, aspects of operation 308 are performed by a scaling engine, such as the scaling engine 210 in Figure 2 . For example, if a prototype input token may contribute to generation (e.g., it can be determined whether the prototype input token is the same as or similar to a concept), the norm of the encoder output state associated with the prototype input token can be increased in operation 308. Conversely, when there is a conflict between the prototype input token and the concept input token, the norm of the associated encoder output state may be decreased. In some examples, operation 308 also includes performing an encoder classification task, which determines which encoded tokens may appear in the model output, as described above. The determined encoded tokens can be prioritized and scaled accordingly. In an example, operations 306 and 308 are performed iteratively for each layer of the encoder.

[0071] The process proceeds to operation 310 of generating a position indicator. In an example, aspects of operation 310 are performed by a position indicator generator, such as the position indicator generator 212 in Figure 2 . As described above, position indicators can be generated for concept input tokens and prototype input tokens. The position indicator for a given token can be determined based on its proximity to the concept input token. The concept input token can be assigned a position indicator of "0", while the prototype token can receive a value of "1" or more. For example, if the prototype input token is the same as or similar to the concept input token, an indicator value of "1" can be used, such that the indicator values of other input tokens can increase accordingly with the increase in distance. It will be understood that while the example is described as linearly increasing the position indicator with the increase in distance from the proximity input token (the same as or similar to the concept input token), other techniques can be used. For example, the position indicator can be multiplied or exponentially scaled, or scaled according to any of its various mathematical formulas.

[0072] At operation 312, the scaled encoded output is decoded based on the generated position indicator. In an example, aspects of operation 312 are performed by a decoder, such as the decoder 214 in Figure 2 . As described above, the decoder can include one or more decoder layers, where each decoder layer can include a self-attention network and a feed-forward network. For example, the model output can be generated word by word, while referring to the scaled representation generated by the encoder in combination with the generated position indicator. For example, the model output can be generated one word at a time (e.g., from left to right).

[0073] The process proceeds to operation 314, where the generated model output is provided. In an example, the model output is provided via an API such that another application, process, and / or computing device can use the model output accordingly. For example, the model output can subsequently be used as a descriptive query to better identify content (and / or target content) as compared to using only a search query. As another example, operation 314 can include storing the generated model output (e.g., as a descriptive summary or title associated with the target content). Thus, it will be understood that the generated model output can be used in any of a variety of scenarios. Method 300 terminates at operation 314.

[0074] Although method 300 is shown as occurring sequentially, it will be understood that these aspects need not be performed in the order shown in method 300 and, in some examples, can be performed concurrently. As an example, operation 310 need not be performed after operations 306 and 308, but in some examples, can occur concurrently with at least one of operations 306 and 308, instead.

[0075] Figures 4 - 7 and the associated description provide a discussion of various operating environments in which aspects of the present disclosure can be practiced. However, the devices and systems shown and discussed are for purposes of example and illustration and do not limit the numerous computing device configurations that can be used to practice aspects of the present disclosure described herein. Figures 4 - 7 The systems shown and discussed are for purposes of example and illustration and do not limit the numerous computing device configurations that can be used to practice aspects of the present disclosure described herein.

[0076] Figure 4 is a block diagram showing the physical components (e.g., hardware) of a computing device 400 with which aspects of the present disclosure can be practiced. The computing device components described below can be applicable to the computing devices described above, including the devices 102, 104, and 106 in Figure 1 In a basic configuration, computing device 400 can include at least one processing unit 402 and system memory 404. Depending on the configuration and type of the computing device, system memory 404 can include, but is not limited to, volatile memory (e.g., random access memory), non-volatile memory (e.g., read-only memory), flash memory, or any combination of such memories.

[0077] System memory 404 can include an operating system 405 and one or more program modules 406 suitable for running software applications 420, such as one or more components supported by the systems described herein. As an example, system memory 404 can include a scaling engine 424 and a location indicator generator 426. For example, operating system 405 can be suitable for controlling the operation of computing device 400.

[0078] In addition, embodiments of the present disclosure may be practiced in conjunction with a graphics library, other operating systems, or any other application, and are not limited to any particular application or system. This basic configuration is shown in Figure 4 by those components within dashed line 408. Computing device 400 may have additional features or functionality. For example, computing device 400 may also include additional data storage devices (removable and / or non-removable) (e.g., magnetic disks, optical disks, or magnetic tapes). Such additional memory is shown in Figure 4 by removable storage device 409 and non-removable storage device 410.

[0079] As described above, multiple program modules and data files may be stored in system memory 404. When executed on processing unit 402, program modules 406 (e.g., applications 420) may perform processing including but not limited to aspects described herein. Other program modules that may be used in accordance with aspects of the present disclosure may include email and contact applications, word processing applications, spreadsheet applications, database applications, slide presentation applications, drawing or computer-aided applications, etc.

[0080] In addition, embodiments of the present disclosure may be practiced in a circuit including discrete electronic components, a packaged or integrated electronic chip containing logic gates, a circuit utilizing a microprocessor, or on a single chip containing electronic elements or a microprocessor. For example, embodiments of the present invention may be practiced via a system-on-a-chip (SOC), where Figure 4 each or many of the components shown therein may be integrated onto a single integrated circuit. Such SOC devices may include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all of which are integrated (or "burned") onto a chip substrate as a single integrated circuit. When operating via an SOC, functions related to the performance of the client switching protocol described herein may be operated via specific application logic integrated with other components of computing device 400 on a single integrated circuit (chip). Embodiments of the present disclosure may also be implemented using other technologies capable of performing logical operations (e.g., AND, OR, and NOT), including but not limited to mechanical, optical, fluidic, and quantum technologies. In addition, embodiments of the present invention may be practiced in a general-purpose computer or any other circuit or system.

[0081] The computing device 400 may also have one or more input devices 412, such as a keyboard, mouse, pen, voice or speech input device, touch or swipe input device, etc. (Multiple) output devices 414, such as a display, speaker, printer, etc., may also be included. The above devices are examples, and other devices may also be used. The computing device 400 may include one or more communication connections 416 that allow communication with other computing devices 450. Examples of suitable communication connections 416 include, but are not limited to, radio frequency (RF) transmitter, receiver, and / or transceiver circuits; universal serial bus (USB), parallel, and / or serial ports.

[0082] The term computer-readable medium as used herein may include computer storage media. Computer storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, or program modules. System memory 404, removable storage device 409, and non-removable storage device 410 are all examples of computer storage media (e.g., memory storage devices). Computer storage media may include RAM, ROM, electrically erasable programmable read-only memory (EEPROM), flash memory, or other storage technologies, CD-ROM, digital versatile disk (DVD), or other optical memory, magnetic cassettes, tapes, disk storage, or other magnetic storage devices, or any other article of manufacture that can be used to store information and can be accessed by the computing device 400. Any such computer storage media may be part of the computing device 400. Computer storage media does not include carrier waves or other propagated or modulated data signals.

[0083] Communication media may be embodied by computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transmission mechanism, and includes any information delivery media. The term "modulated data signal" may describe a signal having one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media may include wired media, such as a wired network or direct wired connection, and wireless media, such as acoustic, radio frequency (RF), infrared, and other wireless media.

[0084] Figure 5A and Figure 5B A mobile computing device 500 is shown, for example, a mobile phone, smartphone, wearable computer (such as a smartwatch), tablet computer, laptop computer, etc., which can be used to practice embodiments of the present disclosure. In some aspects, the client may be a mobile computing device. Referring to Figure 5A, which shows one aspect of the mobile computing device 500 for implementing aspects. In a basic configuration, the mobile computing device 500 is a handheld computer with both input elements and output elements. The mobile computing device 500 generally includes a display 505 that allows a user to input information into the mobile computing device 500 and one or more input buttons 510. The display 505 of the mobile computing device 500 can also be used as an input device (e.g., a touchscreen display).

[0085] If an optional side input element 515 is included, it allows for further user input. The side input element 515 can be a rotary switch, a button, or any other type of manual input element. In an alternative aspect, the mobile computing device 500 can include more or fewer input elements. For example, the display 505 may not be a touchscreen in some embodiments.

[0086] In yet another alternative embodiment, the mobile computing device 500 is a portable telephone system, such as a cellular phone. The mobile computing device 500 can also include an optional keyboard 535. The optional keyboard 535 can be a physical keyboard or a "soft" keyboard generated on a touchscreen display.

[0087] In various embodiments, the output elements include a display 505 for displaying a graphical user interface (GUI), a visual indicator 520 (e.g., a light-emitting diode), and / or an audio transducer 525 (e.g., a speaker). In some aspects, the mobile computing device 500 includes a vibration sensor for providing tactile feedback to the user. In another aspect, the mobile computing device 500 includes input and / or output ports, such as an audio input (e.g., a microphone jack) for sending signals to or receiving signals from an external device, an audio output (e.g., a headphone jack), and a video output (e.g., an HDMI port).

[0088] Figure 5B is a block diagram showing the architecture of one aspect of the mobile computing device. That is, the mobile computing device 500 can incorporate a system (e.g., an architecture) 502 to implement some aspects. In one embodiment, the system 502 is implemented as a "smartphone" capable of running one or more applications (e.g., a browser, an email client, a calendar, a contact manager, a messaging client, a game, and a media client / player). In some aspects, the system 502 is integrated as a computing device, such as an integrated personal digital assistant (PDA) and a wireless phone.

[0089] One or more applications 566 can be loaded into the memory 562 and run on or associated with the operating system 564. Examples of applications include a phone dialer, an email program, a personal information management (PIM) program, a word processing program, a spreadsheet program, an Internet browser program, a messaging program, and the like. The system 502 also includes a non-volatile storage area 568 within the memory 562. The non-volatile storage area 568 can be used to store persistent information that should not be lost when the system 502 is powered off. The applications 566 can use and store information in the non-volatile storage area 568, such as emails or other messages used by an email application, and so on. A synchronization application (not shown) also resides on the system 502 and is programmed to interact with a corresponding synchronization application residing on a host to keep the information stored in the non-volatile storage area 568 synchronized with the corresponding information stored on the host. It should be understood that other applications can be loaded into the memory 562 and run on the mobile computing device 500 described herein (e.g., a search engine, an extractor module, a relevance ranking module, an answer scoring module, etc.).

[0090] The system 502 has a power supply 570, which can be implemented as one or more batteries. The power supply 570 can also include an external power source, such as an AC adapter or a power dock for supplementing or charging the battery.

[0091] The system 502 can also include a radio interface layer 572, which performs the functions of transmitting and receiving radio frequency communications. The radio interface layer 572 facilitates a wireless connection between the system 502 and the "outside world" via a communication carrier or service provider. Transmissions in and out of the radio interface layer 572 are under the control of the operating system 564. In other words, communications received by the radio interface layer 572 can be propagated to the applications 566 via the operating system 564, and vice versa.

[0092] The visual indicator 520 can be used to provide visual notifications, and / or the audio interface 574 can be used to generate audible notifications via the audio transducer 525. In the illustrated embodiment, the visual indicator 520 is a light-emitting diode (LED), and the audio sensor 525 is a speaker. These devices can be directly coupled to the power source 570 so that when activated, they remain on for a duration specified by the notification mechanism, even if the processor 560 and other components can be turned off to conserve battery power. The LED can be programmed to remain on indefinitely until the user takes an action indicating the powered-on state of the device. The audio interface 574 is used to provide sound signals to the user and receive sound signals from the user. For example, in addition to being coupled to the audio transducer 525, the audio interface 574 can also be coupled to a microphone to receive sound input, such as for facilitating a telephone conversation. According to an embodiment of the present disclosure, the microphone can also be used as an audio sensor to facilitate control of the notifications, as described below. The system 502 can also include a video interface 576 that enables operation of the vehicle-mounted camera 530 to record still images, video streams, and the like.

[0093] The mobile computing device 500 implementing the system 502 can have additional features or functionality. For example, the mobile computing device 500 can also include additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or magnetic tapes. Such additional memory is Figure 5B shown by the non-volatile storage area 568.

[0094] Data / information generated or captured by the mobile computing device 500 and stored by the system 502 can be stored on the local mobile computing device 500 as described above, or the data can be stored on any number of storage media that can be accessed by the device via the radio interface layer 572 or via a wired connection between the mobile computing device 500 and a separate computing device associated with the mobile computing device 500, e.g., a server computer in a distributed computing network such as the Internet. It should be understood that such data / information can be accessed by the mobile computing device 500 via the radio interface layer 572 or via a distributed computing network. Similarly, such data / information can be easily transmitted between computing devices for storage and use according to well-known data / information transmission and storage means, including email and collaborative data / information sharing systems.

[0095] Figure 6Illustrates an aspect of a system architecture for processing data received from a remote source on a computing system, as described above, where the remote source is, for example, a personal computer 604, a tablet computing device 606, or a mobile computing device 608. The content displayed on the server device 602 can be stored in different communication channels or other storage types. For example, various documents can be stored using a directory service 622, a portal website 624, a mailbox service 626, an instant messaging repository 628, or a social networking site 630.

[0096] A client communicating with the server device 602 can use a prototype generation engine 620 (e.g., performing aspects similar to operation 304 of method 300 in Figure 3 ), and / or a generative inference engine 621 can be used by the server device 602. The server device 602 can provide data to and receive data from client computing devices such as a personal computer 604, a tablet computing device 606, and / or a mobile computing device 608 (e.g., a smart phone) via a network 615. As an example, the computer system described above can be implemented in a personal computer 604, a tablet computing device 606, and / or a mobile computing device 608 (e.g., a smart phone). In addition to receiving graphical data, any of these embodiments of the computing device can obtain content from a memory 616, and the graphical data can be pre-processed on a graphical source system or post-processed on the receiving computing system.

[0097] Figure 7 An exemplary tablet computing device 700 is shown that can execute one or more aspects disclosed herein. Additionally, the aspects and functions described herein can operate on a distributed system (e.g., a cloud-based computing system), where application functions, memory, data storage and retrieval, and various processing functions can operate remotely from each other via a distributed computing network such as the Internet or an intranet. Various types of user interfaces and information can be displayed via an in-vehicle computing device display or via a remote display unit associated with one or more computing devices. For example, various types of user interfaces and information can be displayed on and interacted with a wall onto which they are projected. Interactions with the numerous computing systems available for practicing embodiments of the present invention include keystroke input, touchscreen input, voice or other audio input, gesture input, where the relevant computing device is equipped with detection (e.g., camera) functions for capturing and interpreting user gestures for controlling the functions of the computing device, and so on.

[0098] This disclosure relates to systems and methods for generating model outputs based on a set of concepts according to examples provided at least in the following sections:

[0099] (A1) In one aspect, some embodiments include a system (e.g., 400, 500), the system comprising: at least one processor (e.g., 402, 560, 561); and a memory (e.g., 404, 562), the memory storing instructions that, when executed by the at least one processor, cause the system to perform a set of operations (e.g., Figure 3 ), the set of operations including: receiving (e.g., 302) an indication including a search query (e.g., 202) from a computing device (e.g., 104, 106); obtaining (e.g., 304) prototypes (e.g., 218) for a set of concepts (e.g., 216) associated with the search query based on a knowledge corpus (e.g., 110, 114, 116); encoding (e.g., 208, 306) an input based on the set of concepts and the obtained prototypes, the input including one or more concept input tokens for the set of concepts and one or more prototype input tokens for the obtained prototypes; scaling (e.g., 210, 308) the encoded input to reduce a first norm of the encoded output state for a first prototype input token that is similar to a first concept input token of the concept input tokens; generating (e.g., 212, 310) a set of position indicators for the input tokens of the input; decoding (e.g., 214, 312) the scaled and encoded output based on the set of position indicators to generate a model output; identifying target content based on the generated model output (e.g., 204, 314); and providing (e.g., 314) the identified target content to the computing device in response to the received indication.

[0100] (A2) In some embodiments of A1, the prototypes (e.g., 218) of the set of concepts (e.g., 216) are obtained (e.g., 302) based on search results in response to the received search query.

[0101] (A3) In some embodiments of A1 - A2, generating (e.g., 212, 310) the set of position indicators includes: for each input token: when the input token is a concept input token (e.g., 216), generating a position indicator of a first value; when the input token is a prototype input token similar to a concept input token (e.g., 218), generating a position indicator of a second value that is greater than the first value; and when the input token is a prototype input token not similar to a concept input token, generating a position indicator of a third value that is greater than the position indicator value of the closest prototype input token similar to the concept input token.

[0102] (A4) In some embodiments of A1 - A3, the third value is linearly determined based on the distance to the closest prototype input token that is similar to the concept input token.

[0103] (A5) In some embodiments of A1 - A4, the search results in response to the received search query are retrieved (e.g., 304) from the knowledge corpus (e.g., 110, 114, 116).

[0104] (A6) In some embodiments of A1 - A5, the knowledge corpus is determined from a knowledge corpus set (e.g., 110, 114, 116) based on the received search query.

[0105] (A7) In some embodiments of A1 - A6, the knowledge corpus is one of an in - domain knowledge corpus (e.g., 114) or an out - of - domain knowledge corpus (e.g., 110, 116).

[0106] (B1) In another aspect, some embodiments include a system (e.g., 400, 500) that includes: at least one processor (e.g., 402, 560, 561); and a memory (e.g., 404, 562) that stores instructions which, when executed by the at least one processor, cause the system to perform a set of operations (e.g., Figure 3 ) that include: receiving (e.g., 302) a request (e.g., 202) that includes a set of concepts (e.g., 216); generating (e.g., 304) prototypes (e.g., 218) for the set of concepts based on a knowledge corpus (e.g., 110, 114, 116); encoding (e.g., 208, 306) an input that includes a set of input tokens, where the set of input tokens includes concept input tokens of the set of concepts and prototype input tokens of the prototypes; generating (e.g., 212, 310) a set of position indicators for the input tokens of the input, where each position indicator indicates the relative distance of an input token to the closest input token that is similar to a concept input token; decoding (e.g., 214, 312) the encoded output based on the set of position indicators to generate a model output (e.g., 204); and providing (e.g., 314) the generated model output in response to the request.

[0107] (B2) In some embodiments of B1, the set of operations further includes: scaling (210, 308) the encoded input to reduce a first norm of the encoded output state for a first prototype input token that is similar to a first concept input token of the concept input token.

[0108] (B3) In some embodiments of B1 - B2, the knowledge corpus is one of an in - domain knowledge corpus or an out - of - domain knowledge corpus.

[0109] In a further aspect, some embodiments include a method (e.g., Figure 3 ) for generating a model output (e.g., 204) based on a concept set (e.g., 202), the method comprising: generating (e.g., 304) a prototype (e.g., 218) for the concept set (e.g., 216) based on a knowledge corpus (e.g., 110, 114, 116); encoding an input comprising an input token set, wherein the input token set includes concept input tokens of the concept set and prototype input tokens of the prototype; scaling the encoded input (e.g., 210, 308) to reduce a first norm of an encoded output state for a first prototype input token that is similar to a first concept input token of the concept input tokens; generating (e.g., 212, 310) a set of position indicators for the input tokens of the input; and decoding the scaled and encoded output based on the set of position indicators (e.g., 214, 312) to generate a model output.

[0110] (C2) In some embodiments of C1, the method further comprises: receiving (202, 302) an indication comprising a search query from a computing device; generating (302) the concept set (e.g., 216) based on the search query; and identifying target content based on the generated model output (204, 314); and in response to the indication, providing (e.g., 314) the identified target content.

[0111] (C3) In some embodiments of C1 - C2, the method further comprises: receiving (202, 302) from a computing device a concept set as a keyword associated with target content; and storing the model output (e.g., 204, 314) as one of a descriptive title or a descriptive summary associated with the target content.

[0112] (C4) In some embodiments of C1 - C3, the knowledge corpus (e.g., 110, 114, 116) is one of an in - domain knowledge corpus (e.g., 114) or an out - of - domain knowledge corpus (e.g., 110, 116).

[0113] (C5) In some embodiments of C1 - C4, the knowledge corpus is determined from a set of knowledge corpora (e.g., 110, 114, 116) based on the concept set.

[0114] For example, aspects of the present disclosure have been described above with reference to block diagrams and / or operational descriptions of methods, systems, and computer program products. The functions / actions recited in the block diagrams may not occur in any order shown in any flowchart. For example, depending on the functionality / acts involved, two blocks shown in succession may actually be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order.

[0115] The descriptions and illustrations provided in this application of one or more aspects are not intended to limit or in any way restrict the scope of the disclosure as claimed. The aspects, embodiments, and details provided in this application are considered sufficient to convey ownership and enable others to make and use the best mode of the disclosure as claimed. The disclosure as claimed should not be construed as limited to any aspect, example, or detail provided in this application. Various features (structural and methodical) are intended to be selectively included or omitted, whether shown and described in combination or separately, to yield embodiments having a particular set of features. After the description and illustration of this application have been provided, those skilled in the art may envision variations, modifications, and alternative aspects that fall within the spirit of the broader aspects of the general inventive concept embodied in this application, which do not depart from the broader scope of the disclosure as claimed.

Claims

1. A system, comprising: at least one processor; and a memory storing instructions which, when executed by the at least one processor, cause the system to perform a set of operations, the set of operations including: receiving an indication including a search query from a computing device; obtaining prototypes for the concept set from a knowledge corpus based on a set of concepts associated with the search query; encoding an input based on the concept set and the obtained prototypes, the input including one or more concept input tokens for the concept set and one or more prototype input tokens for the obtained prototypes; scaling the encoded input by generating a scaling factor to assign to corresponding input tokens encoded by a model to reduce a first norm of an encoded output state for a first prototype input token, the first prototype input token being similar to a first concept input token of the concept input tokens; generating a set of position indicators for input tokens of the input, wherein when an input token is a prototype input token, the corresponding position indicator indicates a relative distance of the prototype input token to the closest prototype input token, the closest prototype input token being similar to a concept input token; decoding the scaled and encoded input based on the set of position indicators to generate a model output; identifying target content based on the generated model output; and providing the identified target content to the computing device in response to the received indication.

2. The system according to claim 1, wherein the prototypes for the concept set are obtained based on search results in response to the received search query.

3. The system according to claim 1, wherein generating the set of position indicators comprises: for each input token: when the input token is a concept input token, generating a position indicator of a first value; when the input token is a prototype input token similar to a concept input token, generating a position indicator of a second value, the second value being greater than the first value; and when the input token is a prototype input token not similar to a concept input token, generating a position indicator of a third value, the third value being greater than a position indicator value of the closest prototype input token similar to the concept input token.

4. The system according to claim 3, wherein the third value is linearly determined based on a distance to the closest prototype input token similar to the concept input token.

5. The system according to claim 2, wherein the search results in response to the received search query are retrieved from the knowledge corpus.

6. The system according to claim 5, wherein the knowledge corpus is determined from a set of knowledge corpora based on the received search query.

7. The system according to claim 1, wherein the knowledge corpus is one of an in-domain knowledge corpus or an out-of-domain knowledge corpus.

8. A system, comprising: at least one processor; and a memory storing instructions which, when executed by the at least one processor, cause the system to perform a set of operations, the set of operations including: Receive a request that includes a concept set; Based on the concept set, generate a prototype for the concept set from a knowledge corpus; Encode an input that includes an input token set, where the input token set includes concept input tokens of the concept set and prototype input tokens of the prototype; Generate a set of position indicators for the input tokens of the input, where, when an input token is a prototype input token, the corresponding position indicator indicates the relative distance of the prototype input token to the closest prototype input token that is similar to a concept input token; Based on the set of position indicators, decode the encoded input to generate a model output; and In response to the request, provide the generated model output.

9. The system according to claim 8, wherein the set of operations further includes: Scale the encoded input by generating a scaling factor to be assigned to the corresponding input tokens encoded by the model to reduce a first norm of an encoded output state of a first prototype input token that is similar to a first concept input token of the concept input tokens.

10. The system according to claim 8, wherein the knowledge corpus is one of an in-domain knowledge corpus or an out-of-domain knowledge corpus.

11. A method for generating a model output based on a concept set, including: Based on the concept set, generate a prototype for the concept set from a knowledge corpus; Encode an input that includes an input token set, where the input token set includes concept input tokens of the concept set and prototype input tokens of the prototype; Scale the encoded input by generating a scaling factor to be assigned to the corresponding input tokens encoded by the model to reduce a first norm of an encoded output state of a first prototype input token that is similar to a first concept input token of the concept input tokens; Generate a set of position indicators for the input tokens of the input, where, when an input token is a prototype input token, the corresponding position indicator indicates the relative distance of the prototype input token to the closest prototype input token that is similar to a concept input token; and Based on the set of position indicators, decode the scaled and encoded input to generate a model output.

12. The method according to claim 11, further including: Receive an indication that includes a search query from a computing device; Based on the search query, generate the concept set; and Based on the generated model output, identify target content; and In response to the indication, provide the identified target content.

13. The method according to claim 11, further including: Receive a concept set from a computing device as a keyword associated with target content; and Store the model output as one of a descriptive title or a descriptive summary associated with the target content.

14. The method according to claim 11, wherein the knowledge corpus is one of an in-domain knowledge corpus or an out-of-domain knowledge corpus.

15. The method according to claim 14, wherein the knowledge corpus is determined from a knowledge corpus set based on the concept set.

Citation Information

Patent Citations

  • Systems and methods for semantic search and extraction of related concepts from clinical documents

    CN107408156A

  • Network service semanteme register system and method thereof

    CN1464426A