Text summary in information system based on personalized priori knowledge

By identifying user requests and prior knowledge, a personalized document or entity summary is generated, which solves the problem of inaccurate summaries in the existing technology and realizes summary generation that better meets user needs.

CN120604228APending Publication Date: 2025-09-05MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480009301.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-02-27
Filing Date
2024-02-20
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing document summarization systems are unable to provide accurate document or entity summaries based on users' personalized prior knowledge and context, often resulting in over-inclusion or under-inclusion of summaries.

Method used

By receiving user requests, identifying the user's context and prior knowledge, separating document or entity information into segments, generating semantic embeddings, and leveraging a personal knowledge system to generate targeted summaries based on the user's personalized knowledge.

Benefits of technology

It provides a document or entity summary that meets user needs, improves the accuracy of the summary and user experience, and avoids the problem of excessive or insufficient summary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120604228A_ABST
    Figure CN120604228A_ABST
Patent Text Reader

Abstract

Examples of the present disclosure describe systems and methods for providing a text summary based on personalized priori knowledge. In an example, a user request for a summary of a document or entity is received. One or more documents associated with an entity are separated into segments, and a semantic embedding is created for each segment. Semantic embedding, the context of the user request, and related information available in a previous knowledge base of the requesting user are provided as input to a personal knowledge system. Based on the input, the personal knowledge system outputs an indication of the segment that should be summarized. The indicated segment is provided to the summary system. A summary system generates a document summary or an entity summary and provides the summary to a user.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] A document summary provides a user with a brief overview of the content in a document. In most cases, the document summary is based on the entire document or on a time-dependent portion of the document (e.g., modifications within the last three days). However, given a particular user's knowledge of the subject matter in the document and that particular user's recent document access history, such document summaries are often over-inclusive and / or under-inclusive.

[0002] With respect to the above and other overall considerations, various technical solutions disclosed herein have been proposed. In addition, although relatively specific problems may be described, it should be understood that these examples should not be limited to solving the specific problems mentioned in the background technology section or other locations of this disclosure. Summary of the Invention

[0003] Examples of the present disclosure describe systems and methods for providing text summaries based on personalized prior knowledge. In the examples, a user request for a document summary is received at a processing system. The document is segmented into segments, and a semantic embedding is created for each segment. The semantic embedding, the context of the user's request, and relevant information available in the user's prior knowledge base are provided as input to a personal knowledge system. Based on the input, the personal knowledge system outputs an indication of the segments in the document that should be summarized. The indicated segments are provided to a summarization system, which generates one or more summaries of the document based on the indicated segments and provides the one or more summaries of the document to the user.

[0004] In another example of the present disclosure, a user request for an entity summary is received at a processing system. Documents and information associated with the entity are identified and segmented into segments. Semantic embeddings are created for each segment. The semantic embeddings, the context of the user request, and relevant information available from the user's prior knowledge base are provided as input to a personal knowledge system. Based on the input, the personal knowledge system outputs an indication of the segments in the document and the information that should be summarized. The indicated segments are provided to a summarization system, which generates a summary of the entity based on the indicated segments and provides the summary of the entity to the user.

[0005] This summary is provided to introduce a selection of concepts that are further described below in a simplified form in the detailed description. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Additional aspects, features, and / or advantages of the examples will be set forth in part in the description that follows and, in part, will be obvious from the description, or may be learned through practice of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] Examples are described with reference to the following figures.

[0007] Figure 1 An overview of an example system for providing text summaries based on personalized prior knowledge is shown.

[0008] Figure 2 A method for providing a summary of a text document based on a user's personalized prior knowledge is shown.

[0009] Figure 3 A method for providing text entity summaries based on a user's personalized prior knowledge is shown.

[0010] Figure 4 is a block diagram illustrating example physical components of a computing device for practicing aspects of the present disclosure.

[0011] Figure 5 is a simplified block diagram of an example distributed computing system for practicing aspects of the present disclosure. DETAILED DESCRIPTION

[0012] In many systems, a document summary is provided to the user so that the user can access a brief overview of the document content before opening or requesting the document. As used herein, the term "document" includes files, electronic communications (e.g., email messages, calendar invitations, instant messages, text messages, social media posts, phone call data, and voicemail data), video data, and image data, etc. In some cases, the document summary is based on the entire document. For example, the document summary can include a one-sentence summary of each paragraph in the document, regardless of the length or complexity of each paragraph. Alternatively, the document summary can summarize the topics detected in the document or describe the overall topic of the document. In other cases, the document summary is based on the time-related portion of the document. For example, the document summary can give priority to modifications made in the most recent time period (e.g., the last three days) or modifications to the document made by a specific entity (e.g., user, group, or organization). However, for most users, such document summaries are typically over-inclusive and / or under-inclusive. For example, users with extensive or up-to-date knowledge of a topic often perceive document summaries as over-inclusive (e.g., providing information the user already knows), whereas users with limited or non-up-to-date knowledge of a topic perceive document summaries as under-inclusive (e.g., omitting information that would be helpful to the user).

[0013] The present disclosure solves the above-mentioned problems of document summarization technology by providing a solution for providing text document summaries and text entity summaries based on the user's personalized prior knowledge. In an embodiment of the present disclosure, a user request for a summary of a document or a summary of an entity is received at an information processing system. For example, a user can specifically request a summary (e.g., by selecting a "summary" button / option or entering a voice command for the summary). In other examples, the user request is implicit. For example, a user can access a document in a content feed or while browsing; a user can receive an electronic communication that includes an attached document, references a related document, or identifies another user (e.g., a sender, another recipient, or a related user); a user can perform a hover event (or other interaction event) in association with a document, entity, or corresponding link; or a user can compose an electronic communication to one or more recipients. In this example, each event (e.g., document access, electronic communication sending / receiving, or interaction event) can be interpreted as an implicit request for a summary of the corresponding document or entity.

[0014] Upon receiving a user request for a summary, identifying a context for the user request. The context can be based on information or settings associated with the user, the document, the user's device, and / or members of the user's social circle. Non-exhaustive examples of information included in the context include the user's average reading speed, the maximum amount of time the user has to read a document summary, previous summary consumption information (e.g., the average attention span of the user spent on similar topics at similar times of the day or week or in similar locations), the user's current location, the user's travel pattern, the length of the document, the number of words or sentences in the document, the complexity of the document (e.g., structure, semantics, or comprehension complexity), the topics in the document, whether a user preference has been expressed regarding new information or information retention, whether a user preference has been expressed regarding the length or scope of the document summary, whether a user preference has been expressed for displaying or describing images and / or videos, information / capabilities of the user's device (e.g., brand / model, hardware / software version, display size or resolution, and touch screen capabilities), the identities or roles of members of the user's social circle, the user's communication behavior with other users (e.g., communication frequency, communication length, and communication topics), commonalities between users, and one or more events (e.g., previous, current, or upcoming events).

[0015] If the received user request is for a summary of a document, the document is divided into a plurality of segments. In an example, each segment corresponds to a chapter, paragraph, sentence, image, or another document object of the document. A semantic embedding is generated for each segment. As used herein, a semantic embedding refers to a representation of a segment created by converting high-dimensional data into low-dimensional data in a manner that the high-dimensional data and the low-dimensional data remain semantically similar. In one example, each semantic embedding is represented as a feature vector (e.g., an n-dimensional vector representing numerical features of one or more objects).

[0016] If a user request is received for a summary of an entity, information related to the entity is collected from one or more data stores. Non-exhaustive examples of information related to the entity include documents associated with the entity, public news and events associated with the entity, social circle information of the entity, mailing lists including the entity, role or title information of the entity, project information associated with the entity, social media information of the entity, and user signals associated with the entity (e.g., detected events). In some examples, a semantic embedding is generated for the information related to the entity.

[0017] The semantic embedding, the context of the user's request, and information associated with the user's existing knowledge base are provided as input to a personal knowledge system personalized for the user. Non-exhaustive examples of information associated with the user's existing knowledge base include the user's areas of interest or expertise, the user's interest or knowledge level in various knowledge areas or with various documents, the user's document interactions (e.g., documents or document parts created, read, modified, commented on, received, sent, or quoted by the user), the user's electronic communications, text or voice data from meetings and events attended by the user, video and image data associated with the user or interacted with by the user, and other user signals associated with the user (e.g., detected events). Based on this input, the personal knowledge system outputs an indication of the semantic embeddings to be summarized for the user. In the example, the indications are ranked in order of importance to the user. These indications can also indicate the scope of summary for each semantic embedding or indicate the degree to which the information associated with each semantic embedding should be summarized.

[0018] The indication of semantic embedding is used to generate a summary of information associated with the semantic embedding. As an example, a document can be summarized so that information known to the user or not of interest to the user is omitted or briefly summarized at a high conceptual level, while information unknown to the user or of interest to the user is more comprehensively summarized at a conceptual level based on the user's level of knowledge about the information. As another example, a summary for an entity is generated so that entity information that is of interest to the user, temporally relevant (e.g., an upcoming meeting or visit date) or particularly noteworthy (e.g., an important achievement, an event with an important public scope, or an anniversary) is prioritized (e.g., ranked accordingly). In such an example, the presentation of the summary can be based on the context of the user's request. For example, the context can affect the length of the summary, the highlighting or emphasis applied to the summary, the output mode of the summary (e.g., delivered on a screen as text or delivered as voice by a digital assistant), the order in which the summary is presented, or whether additional (e.g., supplementary) information is provided together with the summary.

[0019] Therefore, the present disclosure provides multiple technical benefits and improvements over previous document summarization solutions, including: using the user's personalized prior knowledge to determine the applicability and complexity of the document summary to be generated for the user, using the user's personalized prior knowledge to generate entity summaries for the user, and using the context of the user's request to determine the optimal presentation length and format of the summary.

[0020] Figure 1 An overview of an example system for providing a text summary based on personalized prior knowledge is shown. The example system can correspond to an enterprise system that includes one or more entities, stores multiple documents specific to one or more entities or owned by one or more entities, and can access multiple documents outside the boundaries of the enterprise system (e.g., public documents or documents of other enterprise systems). As presented, system 100 is a combination of interdependent components that interact to form an integrated whole. The components of system 100 can be hardware components or software components (e.g., applications, application programming interfaces (APIs), modules, virtual machines, or runtime libraries) implemented on and / or executed by hardware components of system 100. In one example, the components of the system disclosed herein are implemented on a single computing device. In another example, the components of the system disclosed herein are distributed across multiple computing devices and / or computing environments.

[0021] exist Figure 1 In FIG, system 100 includes user devices 102A, 102B, and 102C (collectively referred to as "user devices 102"), a network 104, and a service environment 106. Figure 1Although depicted as including a specific combination of computing environments and devices, the size and configuration of the devices and computing environments described herein may vary and may include more than Figure 1 Furthermore, although the description will be in the context of personalized text summaries of documents and entities Figure 1 and examples in subsequent figures, but these examples also apply to summaries based on prior knowledge of multiple users, summaries of events (and other types of activities), and summaries provided in non-text formats (e.g., voice-based formats or video-based formats).

[0022] The user device 102 is configured to detect and / or collect input data from one or more users or user devices. In some examples, the input data corresponds to user interactions with one or more software applications or services implemented or accessible by the user device 102. In other examples, the input data corresponds to automatic interactions with the software applications or services, such as automatically (e.g., non-manually) executing a script or set of commands at a scheduled time or in response to a predetermined event. The user interaction or automatic interaction can be related to the execution of a user activity corresponding to a task, project, data request, etc. The input data includes, for example, audio input, touch input, text-based input, gesture input, image input, and / or corresponding user signals. One or more sensor components of the user device 102 (e.g., a keyboard, microphone, pointing / selection tool, touch-based sensor, optical / magnetic sensor, geolocation sensor, and accelerometer) are used to detect and / or collect the input data. Examples of user device(s) 102 include personal computers (PCs), mobile devices (e.g., smartphones, tablets, laptops, personal digital assistants (PDAs)), wearable devices (e.g., smart watches, smart glasses, fitness trackers, smart clothing, body-mounted devices, head-mounted displays), as well as gaming consoles or devices and Internet of Things (IoT) devices.

[0023] (Multiple) user devices 102 may include or otherwise access various applications and services. Non-exhaustive examples of applications include word processing applications, spreadsheet applications, presentation applications, document reader software, social media software / platforms, search engines, media software / platforms, multimedia player software, content design software / tools, and database applications. Applications and services enable users to access and / or interact with one or more types of content, such as text, audio, images, video, animation, and multimedia (e.g., a combination of text, audio, images, video, and / or animation).

[0024] User device(s) 102 provide the received input data to service environment 106. In some examples, user device(s) 102 transmit the input data to service environment 106 using network 104. Examples of network 104 include a private local area network (PAN), a local area network (LAN), a wide area network (WAN), and the like. Although network 104 is depicted as a single network, it is contemplated that network 104 may represent several networks of similar or different types. In other examples, input data is provided to service environment 106 without using network 104. Service environment 106 provides user devices 102 with access to various computing services and resources (e.g., applications, devices, storage, processing power, networking, analytics, intelligence). In some examples, service environment 106 is implemented in a cloud-based or server-based environment using one or more computing devices, such as server devices (e.g., web servers, file servers, application servers, database servers), edge computing devices (e.g., routers, switches, firewalls, multiplexers), personal computers (PCs), virtual appliances, and mobile devices. In other examples, the service environment 106 is implemented in a local environment (e.g., a home or office) using such a computing device. The computing device may include one or more sensor components, as discussed with respect to the user device(s) 102. In other examples, the service environment 106 is implemented locally on the user device(s) 102. The service environment 106 may include many hardware and / or software components and may be controlled by one or more distributed computing models / services (e.g., Infrastructure as a Service (IaaS), Platform as a Service (PaaS), Software as a Service (SaaS), Function as a Service (FaaS)).

[0025] exist Figure 1 , the service environment 106 includes a signal detector 108, a user knowledge base 110, an embedding engine 112, a personal knowledge system 114, and a summarization engine 116. The signal detector 108 detects or receives user signals associated with input data generated by the user device 102. The signal detector 108 can use an event listener function or process (or similar functions) to detect user signals. The listener function or process is programmed to react to input or signals indicating the occurrence of a particular event by calling an event handler. In an example, at least a portion of the user signals generated by the user device 102 relates to a user request (e.g., an implicit user request or an explicit user request) for a summary of documents and / or entities. For example, a user of (multiple) user devices 102 can execute a search query that generates search results that include a collection of documents. Executing the search query can be considered an implicit user request to generate a summary for one or more documents in the search results.

[0026] In some examples, signal detector 108 identifies the context of a detected user request for a summary. In other examples, the context is identified by an alternative component of system 100. The context can be based on information associated with the user, the document (or document collection), the user's device, or another entity. Such information can be based on current user behavior or location, previous user behavior, user preferences or settings configured for the user, or a specified user goal for the user request. As an example, signal detector 108 can determine that the context of a user who has submitted a user request for a summary of a document indicates that the user is currently commuting to a destination, the user will arrive at the destination in approximately ten minutes, the user is an expert on several topics in the document, the average reading speed of the document, the average reading speed of the user, and the type of device from which the user signal is received. The context can also indicate that the user typically focuses on each summary for approximately thirty seconds, the user typically focuses on the first part of the summary (e.g., the first 50% of the lines in the summary), the user has expressed an interest in learning about one or more specific topics, and the user typically prefers to receive low-concept-level summaries on topics in which the user is not an expert.

[0027] The user knowledge base 110 stores user data related to previously collected user signals and the user's existing user knowledge (collectively referred to as "knowledge data"). The user knowledge base 110 can be represented or stored in a data store, such as a database, a table, or a similar storage device or system. The user knowledge base 110 aggregates user data from various data sources (e.g., the signal detector 108, the user's application data, a data store external to the system 100), and stores the aggregated user data in one or more data structures. As an example, the knowledge base 110 includes a personalized knowledge graph (or any other type of ontology-based data structure) for each user who is a member of the system 100. In this example, the personalized knowledge graph includes objects (e.g., documents, document parts, entities) interacted with by the user, relationships between objects, and metadata associated with the objects and relationships (e.g., creation date, modification date, most recent interaction date, document properties, and entity properties). The personalized knowledge graph may also include weights or scores assigned to objects based on various factors, such as the date / time the object was added to the personalized knowledge graph, the user's expertise or familiarity with the object, the user's interest in the object, recent modifications to the object, the frequency of the user's interaction with the object, the total number of interactions the user has with the object, the user's most recent interactions with the object, the type of interaction(s) with the object, and the like.

[0028] The embedding engine 112 retrieves a document or object from a data source such as (multiple) user devices 102, a user knowledge base 110, or another data repository of the system 100. In some examples, the format of the document or object is a format other than text, such as an audio format or a video format. In such an example, the embedding engine 112 can convert the audio data or the audio portion of the video data into a text format. The embedding engine 112 uses, for example, a data parsing tool to separate the document or object into one or more segments. A semantic embedding is then generated for each segment using an embedding model, such as Bidirectional Encoder Transformer Representation (BERT), Sentence BERT (SBERT), Principal Component Analysis (PCA), Singular Value Decomposition (SVD), and Word2Vec. In an example, the embedding model decomposes the segment into one or more feature vectors.

[0029] The personal knowledge system 114 identifies the semantic embedding to be summarized. In an example, the personal knowledge system 114 is a machine learning (ML) system that is personalized (e.g., trained) for the user using information in or related to the user knowledge base 110 and / or other user-specific data of the user. Non-exhaustive examples of ML techniques implemented by the personal knowledge system 114 include neural networks (e.g., generative adversarial networks (GANs), recurrent neural networks (RNNs), convolutional neural networks (CNNs), and autoencoders), ensemble methods (e.g., random forests and gradient boosting), classification models (e.g., support vector machines (SVMs), k-nearest neighbor models, and decision trees), and regression models (e.g., linear regression models and logistic regression models).

[0030] The personal knowledge system 114 receives input in the form of semantic embeddings from the signal detector 108, receives a user request context from the signal detector 108 (or from an alternative component of the system 100), and / or receives knowledge information from the user knowledge base 110. In an example, the personal knowledge system 114 compares the semantic embeddings with the knowledge information. The comparison may include dynamically (e.g., in response to detecting a user request) creating a semantic embedding for the knowledge information. Alternatively, previously generated semantic embeddings may be stored in the knowledge information. In some examples, the comparison is performed by calculating a similarity metric (e.g., a cosine similarity metric or a Euclidean distance metric) between the semantic embeddings and the knowledge information. The similarity metric indicates the similarity of content or topic between the semantic embeddings and the knowledge information. In other examples, the comparison is performed using statistical techniques such as maximum mean difference (MMD). MMD represents the distance between distributions as the distance between the average embeddings of features.

[0031] The results of comparing the semantic embeddings to the knowledge information are evaluated based on the context of the user request. For example, based on the context of the user request, the personal knowledge system 114 may prioritize semantic embeddings that summarize new information about a topic in which the user has expertise or experience, new information about a topic in which the user has limited knowledge or experience, or known information about a topic in which the user has expertise or extensive experience (e.g., information already known to the user). Based on the evaluation of the comparison, the personal knowledge system 114 identifies the semantic embeddings to be summarized for the user who submitted the user request. Identifying the semantic embeddings may include generating an indication, such as a unique identifier (e.g., an index value or a hash value), for each semantic embedding. Alternatively, identifying the semantic embeddings may include collecting the semantic embeddings or fragments corresponding to the semantic embeddings.

[0032] The personal knowledge system 114 also determines the scope of summarization for each semantic embedding or the degree to which information associated with each semantic embedding should be summarized based on the context of the user request. For example, the personal knowledge system 114 may determine that semantic embeddings related to a particular topic will be summarized at a high concept level (or a low concept level), that the semantic embedding will be limited to a specific number of sentences or words, or that the semantic embedding will include supplementary information (e.g., links to additional content). The personal knowledge system 114 also determines the presentation order and / or output mode of the semantic embeddings based on the context of the user request. As an example, the personal knowledge system 114 may assign an order to the semantic embeddings so that segments with the highest or lowest similarity metrics are presented before other segments. As another example, the personal knowledge system 114 may determine whether to deliver the semantic embeddings via voice or text based on the user's current traffic mode or activity.

[0033] The summarization engine 116 receives an indication of the semantic embedding to be summarized and corresponding summarization instructions (e.g., summarization scope, summarization amount, presentation order, output mode) from the personal knowledge system 114. In an example, the summarization engine 116 implements an ML language model that generates human-like text from semantic embeddings, such as Generative Pre-Trained Transformer 3 (GPT-3), Language Model from Conversational Applications (LaMDA), and Big Science Large-Scale Open Science Open Access Multilingual Language Model (BLOOM). The summarization engine 116 generates one or more summaries for the semantic embedding according to the summarization instructions. For example, the summary may include segments with various summarization scopes and arranged in a specific segment presentation order. The summarization engine 116 provides the summary to the user device 102 to fulfill the user request for the summary. For example, the summary may be provided to one or more user devices 102 of the user who submitted the user request. The summary may then be displayed on or by the user device(s) 102 in response to the user request.

[0034] Having described systems that may be employed by embodiments disclosed herein, methods that may be performed by such systems are now provided.While methods 200-300 are described in the context of system 100, execution of methods 200-300 is not limited to such an example.

[0035] Figure 2 A method 200 for providing a summary of a text document based on personalized prior knowledge of a user is shown. The method 200 can be performed by one or more components of a computing environment, such as the service environment 106. The method 200 begins at operation 202, where a user request for a document summary is received. In an example, the user request is determined by a user signal monitoring mechanism (e.g., the signal detector 108) in response to detecting a user signal of a user device (e.g., user device(s) 102). The user signal can correspond to an explicit request by the user to retrieve a document summary or an implicit request by the user to retrieve a document summary. As an example, in response to a user performing a hover event on a hyperlink or icon of a document, the signal monitoring mechanism receives or creates a set of commands (e.g., computing instructions) for generating a summary of the document.

[0036] At operation 204, the context of the user request is identified. In an example, the context of the user request is identified using information associated with the user submitting the user request (e.g., behavioral tendencies, configuration settings / user preferences, and user knowledge goals), information associated with the user device of the user (e.g., geographic location data, destination data, and user device capabilities), and information associated with the document for which the document summary is requested (e.g., required or average reading time, reading complexity, and topics discussed in the document). Such information can be collected from data sources stored in the computing environment and / or the user device, such as user profiles, event logs, application data, device sensor data, documents, and data stores including metadata or indexes of documents.

[0037] In operation 206, a semantic embedding is generated for the document for which the document summary is requested. In an example, a document (or a copy of a document) is retrieved or accessed from a data source (e.g., any other data store accessible to a user device or computing environment). For example, an identifier (e.g., a document name, a uniform resource locator (URL), or a uniform resource name (RON)) of the document is requested from the user device, or an identifier of the document is identified from the context of the user request. The identifier is used to locate or access the document. A segmentation mechanism such as embedding engine 112 separates the document into one or more segments (e.g., chapters, paragraphs, sentences, or document objects). A semantic embedding is generated for each segment of the document. In an example, each semantic embedding represents a feature vector that includes semantic information for the corresponding segment. The feature vector may include a numerical value representing the semantic information.

[0038] In operation 208, the semantic embedding of the document and the context of the user request are provided to a knowledge system that has been personalized for the user, such as the personal knowledge system 114. In some examples, the knowledge system also receives knowledge information of the user who submitted the user request. The knowledge information includes previously collected user signals and the user's existing user knowledge. The knowledge information is retrieved from a data store such as the user knowledge base 110. The knowledge system compares the semantic embedding with the knowledge information to determine the similarity (or difference) between the semantic embedding and the knowledge information. The similarity (or difference) identifies whether the user understands or has prior experience with the information (e.g., topic or fact) represented by the semantic embedding. As an example, if a semantic embedding is generated for a user for a topic in which the user is an expert, the knowledge system's comparison will indicate a close similarity (or match) between the semantic embedding and the knowledge information.

[0039] At operation 210, a summary determination is generated for the document. In an example, the knowledge system evaluates the similarity (or difference) between the semantic embedding and the knowledge information based on the context of the user's request to determine the summary scope of the document summary. The summary scope identifies the information in each semantic embedding that should be summarized and the degree to which the information should be summarized. As an example, the summary scope may indicate that topics for which the user is an expert should be summarized at a high conceptual level using a single sentence of no more than twenty words, while topics for which the user has limited knowledge expertise should be summarized at a low conceptual level using multiple sentences that collectively include no more than five hundred words. The knowledge system also determines presentation options for the document summary (e.g., segment presentation order and document summary output mode) based on the context of the user's request. As an example, for a user currently attending an in-person meeting based on the user's calendar event (identified by the context of the user's request), the document summary can be arranged so that topics in the document that are related to the topic in the meeting are provided at the top (e.g., beginning) of the document summary and are presented emphatically (e.g., highlighted, bolded, or italicized). The summary determination includes the summary scope and presentation options.

[0040] In operation 212, a document summary is generated for the document. In an example, the summary determination is provided to a summary generation mechanism, such as the summary engine 116. The summary generation mechanism generates a text document summary for the document based on the summary determination. In some examples, the summary generation mechanism retrieves or accesses the document, as discussed above in operation 206. The summary determination is then applied to the document to generate the document summary. For example, the summary determination may include summary instructions (e.g., calculation instructions) for converting the document into a document summary. In other examples, the document summary is generated using semantic embeddings and the summary determination. For example, the summary determination may be applied to the segments corresponding to each semantic embedding. Alternatively, the summary determination may be used to generate a modified semantic embedding that is used to recreate the corresponding modified segments. The modified segments are then used to generate the document summary. After the document summary is generated, the document summary is provided to the user upon fulfillment of the user request.

[0041] Figure 3 A method 300 for providing text entity summaries based on a user's personalized prior knowledge is shown. The method 300 may be performed by one or more components of a computing environment, such as the service environment 106. The method 300 begins at operation 302, where a user request for an entity summary is received. In the example, as in Figure 2 A user request is determined as described in operation 202. For example, in response to a user receiving an introductory email message from an unknown entity (e.g., an entity with which the user has not previously exchanged electronic communications), a command set for generating a profile for the entity is received or created.

[0042] At operation 304, the context of the user request is identified. In an example, the context of the user request is identified using information associated with the user submitting the user request, information associated with the user device of the user, and information associated with the entity for which the entity profile is being requested (e.g., biographical data of the entity, documents and items associated with the entity, public news and events associated with the entity, social media information of the entity, and user signals associated with the entity). Such information can be collected from data sources internal to the computing environment (e.g., user profiles, knowledge graphs, event logs, application data, device sensor data, and electronic communications) and data sources external to the computing environment (e.g., public documents, search engines, or a separate computing environment of the entity).

[0043] At operation 306, a semantic embedding is generated for the information associated with the entity. In an example, the information associated with the entity is provided to a segmentation mechanism, such as the embedding engine 112. The segmentation mechanism separates the information associated with the entity into segments, such as Figure 2As described in operation 206 of . For example, an entity's resume, recent (e.g., the past three months) social media posts, calendar events, and knowledge graphs are each separated into corresponding segments. Semantic embeddings are then generated for each segment. In some examples, before separating the information associated with the entity into segments, the segmentation mechanism determines the importance of each content item (e.g., documents, images, calendar events, electronic communications) in the information associated with the entity. The importance of each content item can be based on, for example, the creation date or modification date of the content item, the amount of data in the content item, the user's level of interest in each content item (e.g., explicit or implicit indication by the user), and / or the impact of the content item (e.g., the audience range of the content item). The importance of each content item can be expressed using an importance value, such as a numerical value (e.g., 90 out of 100) or as a text-based value (e.g., "low importance," "medium importance," or "high importance"). The segmentation mechanism can then assign the importance value to the content item. Alternatively, the content item can include a pre-existing importance value. In either scenario, the segmentation mechanism determines whether to separate the content item into segments based on the importance value. For example, if the importance value of a content item is above a predefined threshold, the content item is separated into segments.

[0044] At operation 308, the semantic embedding of the entity and the context of the user request are provided to a knowledge system that has been personalized for the user, such as the personal knowledge system 114. In some examples, the knowledge system also receives knowledge information of the user who submitted the user request and compares the semantic embedding with the knowledge information, such as Figure 2 The similarity (or difference) between the semantic embedding and the knowledge information identifies whether the user is aware of the entity or the information (eg, topic or fact) represented by the semantic embedding of the entity or has prior experience with the entity or the information.

[0045] At operation 310, a summary determination is generated for the entity. In an example, the knowledge system evaluates the similarity (or difference) between the semantic embeddings and the knowledge information based on the context of the user request to determine the summary scope of the entity summary. The summary scope identifies the information in each semantic embedding that should be summarized and the degree to which the information should be summarized. As an example, the summary scope may indicate that the user and the entity share a common acquaintance who is a close friend of the user. Thus, the summary scope may indicate that the summary of the common acquaintance should be extremely brief (e.g., no more than ten words) or omitted. As another example, the summary scope may indicate that the entity will be speaking at a conference on a topic about which the user has limited or no knowledge. Thus, the summary scope may summarize the topic at a low concept level that includes broken down acronyms, definitions of various terms, and / or links to additional illustrative references. The knowledge system also determines presentation options for the entity summary based on the context of the user request, such as Figure 2Operation 210 is as described above.

[0046] At operation 312, an entity summary is generated for the document. In an example, the summary determination is provided to a summary generation mechanism, such as the summary engine 116. The summary generation mechanism generates a textual entity summary for the entity based on the summary determination. For example, the entity summary can be a report that organizes snippets corresponding to the semantic embeddings into various categories, such as background facts about the entity, commonalities and / or differences between users and the entity, recommended talking points, etc. In some examples, the entity summary includes one or more images of the entity and / or other entities associated with the entity. For example, the entity summary can include an organizational chart that includes images of the entity and the entity team (e.g., direct reports, peers, and managers) and information about the entity and the entity team. The entity summary can also include social media images and information related to the user. In some examples, the entity summary includes or provides actionable tasks associated with the entity. For example, the entity summary can include a link to register for an event that the entity will attend. Alternatively, the entity summary can cause an email service to provide a preformatted email that includes the entity's contact information and / or automatically generated text related to one or more content items in the information associated with the entity.

[0047] Figure 4-5 and the associated descriptions provide a discussion of various operating environments in which aspects of the present disclosure may be practiced. Figure 4-5 The devices and systems shown and discussed are for purposes of example and explanation, and as will be appreciated, numerous configurations of computing devices may be utilized to practice the aspects of the present disclosure described herein.

[0048] Figure 4 4 is a block diagram illustrating the physical components (e.g., hardware) of a computing device 400 with which aspects of the present disclosure may be practiced. The computing device components described below may be applicable to the computing devices and systems described above. In a basic configuration, the computing device 400 includes at least one processing unit 402 and system memory 404. Depending on the configuration and type of the computing device, the system memory 404 may include volatile storage (e.g., random access memory (RAM)), non-volatile storage (e.g., read-only memory (ROM)), flash memory, or any combination of these memories.

[0049] System memory 404 includes an operating system 405 and one or more program modules 406 suitable for running software applications 420, such as one or more components supported by the system described herein. Operating system 405 may, for example, be suitable for controlling the operation of computing device 400.

[0050] Furthermore, embodiments of the present disclosure may be practiced in conjunction with a gallery, other operating systems, or any other application, and are not limited to any particular application or system. Figure 4 408. The computing device 400 may have additional features or functionality. For example, the computing device 400 may also include additional data storage devices (removable and / or non-removable), such as magnetic or optical disks. Such additional storage Figure 4 4 is shown by a removable storage device 407 and a non-removable storage device 410.

[0051] As described above, a number of program modules and data files can be stored in system memory 404. When executed on processing unit 402, program modules 406 (e.g., applications 420) can perform processes including aspects as described herein. Other program modules that can be used in accordance with aspects of the present disclosure can include email and contact applications, word processing applications, spreadsheet applications, database applications, slide presentation applications, drawing or computer-assisted applications, and the like.

[0052] Furthermore, embodiments of the present disclosure may be practiced on circuits comprising discrete electronic components, packaged or integrated electronic chips containing logic gates, circuits utilizing a microprocessor, or a single chip containing electronic components or a microprocessor. For example, embodiments of the present disclosure may be practiced via a system on a chip (SOC) wherein Figure 4 Each or many of the components shown can be integrated onto a single integrated circuit. Such a SOC device may include one or more processing units, a graphics unit, a communication unit, a system virtualization unit, and various application functions, all of which are integrated (or "burned") onto a chip substrate as a single integrated circuit. When operated via the SOC, the functionality described herein regarding the ability to switch protocols for the client can be operated via dedicated logic integrated with other components of the computing device 400 on a single integrated circuit (chip). Embodiments of the present disclosure may also be practiced using other technologies capable of performing logical operations, such as AND, OR, and NOT, including mechanical, optical, fluidic, and quantum technologies. In addition, embodiments of the present disclosure may be practiced within a general-purpose computer or in any other circuit or system.

[0053] The computing device 400 may also have one or more input devices 412, such as a keyboard, a mouse, a pen, an audio or voice input device, a touch or slide input device, and the like. Output devices 414 may also be included, such as a display, speakers, a printer, and the like. The above devices are examples, and other devices may be used. The computing device 400 may include one or more communication connections 416 that allow communication with other computing devices 450. Examples of suitable communication connections 416 include radio frequency (RF) transmitters, receivers, and / or transceiver circuits; universal serial bus (USB), parallel, and / or serial ports.

[0054] The term computer-readable media as used herein may include computer storage media. Computer storage media may include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures or program modules). System memory 404, removable storage device 407 and non-removable storage device 410 are all examples of computer storage media (e.g., memory storage). Computer storage media may include RAM, ROM, electrically erasable ROM (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other product that can be used to store information and can be accessed by computing device 400. Any such computer storage media may be part of computing device 400. Computer storage media does not include carrier waves or other propagated or modulated data signals.

[0055] Communication media may be embodied by computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and includes any information delivery media. The term "modulated data signal" may describe a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, communication media may include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared, and other wireless media.

[0056] Figure 5 One aspect of the architecture of a system for processing data received at a computing system from a remote source (e.g., a personal computer 504, a tablet computing device 506, or a mobile computing device 508) is shown, as described above. The content displayed at the server device 502 can be stored in different communication channels or other storage types. For example, a directory service 522, a web portal 524, a mailbox service 526, an instant messaging storage 528, or a social networking site 530 can be used to store various documents.

[0057] The input evaluation service 520 can be employed by a client communicating with the server device 502, and / or the input evaluation service 520 can be employed by the server device 502. The server device 502 can provide data to and from client computing devices such as a personal computer 504, a tablet computing device 506, and / or a mobile computing device 508 (e.g., a smartphone) via a network 515. As an example, the computer system described above can be embodied in the personal computer 504, the tablet computing device 506, and / or the mobile computing device 508 (e.g., a smartphone). In addition to receiving graphics data that can be pre-processed at a graphics originating system or post-processed at a receiving computing system, any of these embodiments of the computing device can also obtain content from storage 516.

[0058] As will be understood from the foregoing disclosure, one example of the present disclosure relates to a system comprising: a processor; and a memory coupled to the processor, the memory comprising computer-executable instructions that, when executed by the processor, perform operations. The operations include: receiving a user request for a summary of a document from a user; identifying a context for the user request; segmenting the document into segments; generating semantic embeddings for the segments; providing the semantic embeddings and the context to a knowledge system personalized for the user; generating, by the knowledge system, a summary determination of the document based on the semantic embeddings and the context; generating a summary of the document based on the summary determination; and providing the summary of the document to the user.

[0059] Another example of the present disclosure relates to a computer-implemented method. The method includes: receiving a user request for a summary of a document from a user; identifying a context of the user request; separating the document into segments; generating semantic embeddings of the segments; providing the semantic embeddings and context to a knowledge system personalized for the user; generating, by the knowledge system, a summary determination of the document based on the semantic embeddings and context; generating a summary of the document based on the summary determination; and providing the summary of the document to the user. Another example of the present disclosure relates to a computing environment, including: a processor; and a memory coupled to the processor, the memory including computer-executable instructions that, when executed by the processor, perform operations. The operations include: receiving a user request for a summary of a document from a user; identifying a context of the user request; generating a semantic embedding of the document; providing the user's semantic embedding, context, and knowledge information to the knowledge system, the knowledge information identifying the user's knowledge level of one or more knowledge areas; generating, by the knowledge system, a summary determination of the document based on the semantic embedding, context, and knowledge information; generating a summary of the document based on the summary determination; and providing the summary of the document to the user.

[0060] For example, various aspects of the present disclosure are described above with reference to block diagrams and / or operational diagrams of methods, systems, and computer program products according to various aspects of the present disclosure. The functions / actions indicated in the blocks may not occur in the order shown in any flowchart. For example, depending on the functions / actions involved, two blocks shown in succession may actually be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order.

[0061] The description and explanation of one or more aspects provided in this application are not intended to limit or restrict the scope of this disclosure in any way. The aspects, examples and details provided in this application are considered to be sufficient to convey the best mode of having and enabling others to make and use the disclosed claimed invention. The disclosed claimed invention should not be interpreted as being limited to any aspect, example or details provided in this application. No matter whether shown and described in combination or individually, various features (structure and method) are intended to be selectively included or omitted to produce an embodiment with a specific feature set. The description and explanation of this application have been provided, and those skilled in the art can envision variations, modifications and alternative aspects within the spirit of the broader aspects of the overall inventive concept embodied in this application, which variations, modifications and alternative aspects do not depart from the broader scope of the disclosed claimed invention.

Claims

1. A system comprising: processing unit; as well as a memory coupled to the processing unit, the memory comprising computer-executable instructions that, when executed, perform operations comprising: receiving a user request for a summary of the document from a user; identifying the context of the user request; dividing the document into segments; generating a semantic embedding of the segment; providing the semantic embedding and the context to a knowledge system personalized for the user; generating, by the knowledge system, a summary determination of the document based on the semantic embedding and the context; generating the summary of the document based on the summary determination; and The summary of the document is provided to the user.

2. The system according to claim 1, wherein: The user request is an implicit user request determined based on a user signal received from a user device of the user.

3. The system according to claim 2, wherein: The implicit user request corresponds to: receiving the document via electronic communication; receiving a semantic reference to said document; or Hover events on links or icons of said document.

4. The system according to claim 1, wherein: The context of the user request is identified using information associated with the user and at least one of: information associated with a user device of the user; or Information associated with the document.

5. The system according to claim 4, wherein: The information associated with the user includes at least one of the following: The user's behavioral tendencies; configuration settings or user preferences associated with the user; or A knowledge goal indicated by the user.

6. The system according to claim 4, wherein: The information associated with the user equipment of the user includes at least one of the following: geographic location data of the user; the user's destination data; or The capabilities of the user equipment.

7. The system according to claim 4, wherein: The information associated with the document includes at least one of the following: the average reading time of said document; the reading complexity of the document; or Topics discussed in said document.

8. The system of claim 4, wherein: The system is an enterprise system and the user is a member of the enterprise system; and The context is identified using data sources of the enterprise system.

9. The system according to claim 1, wherein: The fragments correspond to: the sections of the document; a paragraph from the document; or A sentence from the document in question.

10. The system according to claim 1, wherein: The semantic embedding represents a feature vector including semantic information of the corresponding segment.

11. The system according to claim 1, wherein: As part of providing the semantic embedding and the context to the knowledge system, the user's knowledge information is further provided to the knowledge system.

12. The system according to claim 11, wherein The knowledge information includes at least one of the following: User signals previously collected for the user; or Existing user knowledge of the user.

13. The system according to claim 11, wherein: The knowledge system compares the semantic embedding with the knowledge information to determine similarity between the semantic embedding and the knowledge information.

14. The system according to claim 13, wherein: The similarity between the semantic embedding and the knowledge information indicates whether the user is aware of a topic or fact represented by the semantic embedding.

15. A method comprising: receiving a user request for a summary of the document from a user; identifying the context of the user request; dividing the document into segments; generating a semantic embedding of the segment; providing the semantic embedding and the context to a knowledge system personalized for the user; generating, by the knowledge system, a summary determination of the document based on the semantic embedding and the context; generating the summary of the document based on the summary determination; as well as The summary of the document is provided to the user.

16. The method according to claim 15, wherein Generating the summary determination includes: Determining a summary scope of the summary of the document, the summary scope indicating: Information to be summarized in the semantic embedding; and A value of a degree of summarization to be performed on the information to be summarized in the semantic embedding.

17. The method according to claim 16, wherein Generating the summary determination further comprises: Determining a presentation option for the summary of the document, the presentation option indicating at least one of: the presentation order for the segments; or An output mode for the summary of the document.

18. The method according to claim 15, wherein Generating the summary of the document includes: identifying the segment corresponding to the semantic embedding; applying the summarization determination to the segment to generate a modified segment; and The summary of the document is generated using the modified segment.

19. The method according to claim 15, wherein Generating the summary of the document includes: access to the document or a copy of the document; and The summarization determination is applied to the document or a copy of the document, wherein the summarization determination includes instructions for converting the document or the copy of the document into a summary of the document.

20. A computing environment comprising: processing unit; as well as a memory coupled to the processing unit, the memory comprising computer-executable instructions that, when executed, perform operations comprising: receiving a user request for a summary of the document from a user; identifying the context of the user request; generating a semantic embedding of the document; providing the semantic embedding, the context, and knowledge information of the user to a knowledge system, wherein the knowledge information identifies the user's knowledge level in one or more knowledge domains; generating, by the knowledge system, a summary determination of the document based on the semantic embedding, the context, and the knowledge information; generating the summary of the document based on the summarization determination; and The summary of the document is provided to the user.