Computer-Implemented Method, System, and Computer Program Product (Recognizing and Reconstructing Previously Presented Information)
The system uses NLP to identify and differentiate previously consumed content within machine-encoded files, improving computing efficiency and storage by reducing the processing and storage of duplicate information.
Patent Information
- Application Number
- JP2021200925
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-11
- Filing Date
- 2021-12-10
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2041-12-10
Smart Images

Figure 0007786843000001 
Figure 0007786843000002 
Figure 0007786843000003
Abstract
Description
[Technical Field]
[0001] FIELD OF THE DISCLOSURE This disclosure relates to electronically processing and presenting information to a user, and more particularly to distinguishing between newly presented information and previously presented information to a user. [Background technology]
[0002] The advent of the Internet and other modern means of electronic communication has greatly expanded the amount of information available to individual information consumers and the manner in which it is received. Numerous news articles and other types of information can be frequently published and widely distributed electronically on a variety of communication and computing platforms. Such information can be transmitted electronically via wired and wireless connections over a variety of electronic communication networks and consumed using a variety of electronic devices. Summary of the Invention [Problem to be solved by the invention]
[0003] The present disclosure aims to provide a technology that can distinguish between information that is newly presented to a user and information that has been presented in the past. [Means for solving the problem]
[0004] In one or more embodiments, a method includes performing natural language processing (NLP) on a machine-encoded file to determine one or more file segment classifiers. The machine-encoded file can be obtained via a data communications network. The method includes detecting previously consumed content in the machine-encoded file that corresponds to the one or more file segment classifiers by performing a computer database search against a user-specific consumption corpus database for NLP elements identified in electronic content previously presented to the user that match NLP elements extracted from the machine-encoded file. The method includes modifying an electronic presentation structure for electronically presenting the machine-encoded file to the user based on the detection. The electronic presentation structure is modified to distinguish portions of the machine-encoded file that correspond to previously consumed content from other portions of the machine-encoded file.
[0005] In one or more embodiments, a method includes analyzing electronic media content embedded in an electronic presentation structure presented to a user. The method includes detecting portions of the electronic media content previously consumed by the user based on the analysis. Based on the detection, the method modifies the electronic presentation structure to distinguish the portions of the electronic media content previously consumed by the user from other portions of the electronic media content.
[0006] In one or more embodiments, a method includes detecting one or more revisions made to a machine-encoded file after a user retrieves the previous version of the machine-encoded file over a data communications network. The detection includes comparing natural language processing (NLP) elements identified in the electronic content of the previous version of the machine-encoded file with NLP elements identified in the electronic content of the revised version of the machine-encoded file. Based on the detection of the one or more revisions, the method includes presenting the machine-encoded file using a separate electronic presentation structure that indicates which portions are newly added or revised.
[0007] In one or more embodiments, a system includes a processor configured to initiate operations. The operations include performing natural language processing (NLP) on a machine-encoded file to determine one or more file segment classifiers. The machine-encoded file can be obtained over a data communications network. The operations include detecting previously consumed content in the machine-encoded file that corresponds to the one or more file segment classifiers by performing a computer database search against a user-specific consumption corpus database for NLP elements identified in electronic content previously presented to the user that match NLP elements extracted from the machine-encoded file. The operations include modifying an electronic presentation structure for electronically presenting the machine-encoded file to the user based on the detection. The electronic presentation structure is modified to distinguish portions of the machine-encoded file that correspond to previously consumed content from other portions of the machine-encoded file.
[0008] In one or more embodiments, a computer program product includes one or more computer-readable storage media having instructions stored thereon. The instructions are executable by a processor to cause the processor to initiate operations. The operations include performing natural language processing (NLP) on a machine-encoded file to determine one or more file segment classifiers. The machine-encoded file can be obtained over a data communications network. The operations include detecting previously consumed content in the machine-encoded file that corresponds to the one or more file segment classifiers by performing a computer database search against a user-specific consumption corpus database for NLP elements identified in electronic content previously presented to the user that match NLP elements extracted from the machine-encoded file. The operations include modifying an electronic presentation structure for electronically presenting the machine-encoded file to the user based on the detection. The electronic presentation structure is modified to distinguish portions of the machine-encoded file that correspond to previously consumed content from other portions of the machine-encoded file.
[0009] This Summary is intended to introduce some concepts only and does not identify key or essential features of the claimed subject matter. Other features of the inventive structure will be apparent from the accompanying drawings and the detailed description that follows. [Brief explanation of the drawings]
[0010] Examples of configurations of the present invention are illustrated in the accompanying drawings. However, the drawings should not be construed as limiting the configuration of the present invention to only the particular implementations shown. Various aspects and advantages will become apparent upon reading the following detailed description and upon reference to the drawings.
[0011] [Figure 1] FIG. 1 illustrates an example computing environment in which one or more information awareness presentation systems may be used, according to one embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of an information recognition and presentation system according to one embodiment. [Figure 3] FIG. 3 is a diagram illustrating some operational features of an exemplary information awareness presentation system, according to one embodiment. [Figure 4] FIG. 4 illustrates an example electronic presentation structure and a modified electronic presentation structure, according to one embodiment. [Figure 5] FIG. 5 is a flowchart illustrating an example method for electronically recognizing and presenting information, according to one embodiment. [Figure 6] FIG. 6 is a flowchart illustrating an example method for electronically recognizing and presenting information, according to one embodiment. [Figure 7] FIG. 7 is a flow chart illustrating an example of a method for electronically recognizing and presenting information relative to a particular location of a user, according to one embodiment. [Figure 8] FIG. 8 is a flow chart illustrating an example of a method for detecting updates to a piece of information previously electronically presented to a user, according to one embodiment. [Figure 9] FIG. 9 is a diagram illustrating a cloud computing environment, according to one embodiment. [Figure 10]FIG. 10 is a diagram illustrating abstraction model layers, according to one embodiment. [Figure 11] FIG. 11 is a diagram illustrating a cloud computing node, according to one embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of a portable device, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0012] While the present disclosure concludes with claims defining novel features, it is believed that the various features described within this disclosure will be better understood by considering the description in conjunction with the drawings. The processes, machines, manufactures, and variations thereof described herein are provided for illustrative purposes. The specific structural and functional details described within this disclosure should not be construed as limiting, but merely as a basis for the claims and as a representative basis for teaching those skilled in the art to employ various of the described features in nearly all appropriately detailed structures. Furthermore, the terms and phrases used within this disclosure are not limiting, but are intended to describe the described features in an understandable manner.
[0013] The present disclosure relates to electronically processing and presenting information to a user, and more particularly, to distinguishing between information newly presented to a user and information previously consumed by the user. Electronic communication technologies and information sources are constantly expanding, enabling access to ever greater amounts of information. Despite the benefits of rapid access to vast amounts of information, the volume of information can reach a point of diminishing returns for individual information consumers. For example, a current event may generate multiple news articles. However, some of these articles simply repeat the same facts. As a result, consumers must sift through a significant amount of information just to discover which portions of the information are already consumed by the consumer and which portions are new.
[0014] One aspect of the systems, methods, and computer program products disclosed herein relates to automatically discovering information a user has already consumed, thereby reducing the need for the user to reread, relisten, or rewatch information previously electronically presented to the user. Another aspect relates to restructuring an electronic presentation structure to contrast, compress, or eliminate portions of information previously presented to the user as electronic content. As defined herein, "electronic content" is any information or data presented in machine-encoded files that can be electronically stored and / or processed by a computer system. Information presented as electronic content includes news articles, educational materials, academic journals, front-page articles, entertainment materials, and the like. The machine-encoded files for presenting electronic content can be formatted as electronic documents, audio files, video files, or any combination thereof. As defined herein, an "electronic presentation structure" refers to any text, image, audio, or video embedded in a machine-encoded file that is received or generated electronically and presented to a user using an electronic device, which may be, for example, a desktop or laptop computer, a tablet, a smartphone, or similar device capable of receiving text, images, audio, or video, or a combination thereof, via a wired or wireless or both connection.
[0015] The systems, methods, and computer program products disclosed herein detect distinct portions of a given amount of electronic content that a user has previously consumed. As defined herein, a user "consumes" electronic content by retrieving the electronic content embedded in an electronic document, audio file, or video file over a data communications network. Alternatively, in a system that allows users to tag content as they read, listen to, or watch it, a user "consumes" electronic content by tagging an article to indicate that they have read, listened to, or watched the content. As a user consumes electronic content in various formats (e.g., text, audio, video) from various sources, the consumed electronic content is added to a database consisting of a user-specific consumption corpus. The consumed electronic content is then compared with subsequent presentations to prevent repeated presentations of previously consumed electronic content.
[0016] The various machine learning and natural language processing techniques disclosed herein, each in a different configuration, determine one or more file segment classifiers within a machine-encoded file that classify distinct segments in the newly presented electronic content and correspond to distinct topics identified by topic analysis. For each identified topic, the corresponding segment in the electronic content can be compared to previously consumed electronic content to determine whether any of the topical segments in the newly presented electronic content are identical to or sufficiently similar to previously consumed content related to the same topic.
[0017] Based on the detection, an electronic presentation structure for presenting the electronic content in the machine-encoded file is modified. That is, the electronic presentation structure is modified to distinguish portions of the electronic content in the machine-encoded file that correspond to previously consumed portions from other portions of the file. This allows a user to avoid re-reading, re-listening, or re-watching information that was previously presented to the user as electronic content. For example, in a text presentation of electronic content, the electronic presentation structure can be annotated by highlighting or rendering the text in a different color or style to distinguish newly presented content from previously consumed content. In an audio presentation, for example, audio portions associated with previously presented content can be deleted, skipped, or fast-forwarded. Similarly, in a video presentation, for example, video portions associated with previously presented content can be skipped or fast-forwarded. The electronic presentation structure can be restructured through reformatting and / or annotation to distinguish between previously consumed electronic content and newly presented electronic content. Optionally, this reconstruction may include annotating the electronic content to indicate portions of the electronic content that have been previously consumed.
[0018] The arrangements described herein relate to computer technology and provide improvements to existing computer technology. One such improvement relates to reducing the waste of computing resources. The amount of information disseminated as electronic content continues to expand. As the amount of electronic content increases, the expenditure of computing resources to electronically store and process the information embedded therein also increases. The arrangements described herein improve the processing efficiency of computer systems and reduce data storage requirements by reducing the processing and storage of electronic content that merely duplicates information already consumed by a user.
[0019] For example, with respect to any computing device (e.g., desktop computer, laptop, tablet, smartwatch), the disclosed configurations improve computing technology by avoiding the unnecessary use of computing resources to process and present duplicate electronic content. Reducing the need to process repetitive electronic content increases the processing efficiency of the computing device. Also, avoiding data storage of consumed electronic content correspondingly increases the available storage capacity on the computing device. Thus, computer and device hardware can operate more efficiently by avoiding the retrieval, storage, and processing of information already consumed by the user.
[0020] Further aspects of the embodiments described within the present disclosure are described in detail below with reference to the drawings. Note that for simplicity and clarity of illustration, elements in the figures may not be drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. Furthermore, where considered appropriate, the same reference numerals may be repeated throughout the figures to indicate corresponding, similar, or like features.
[0021] FIG. 1 illustrates a computing environment 100. The illustrated computing environment 100 includes a computing node 102, one or more informational nodes 104, and devices 106a, 106b, and 106c, communicatively connected via a data communications network 108. The computing node 102 and the information node 104 may each include one or more of the hardware components described with reference to the exemplary computing node 1100, including computer system 1112 (FIG. 11). The devices 106a-106c may comprise a computer system (e.g., a desktop computer) including memory, one or more processors, and other hardware components, such as those described with reference to computer system 1112 (FIG. 11). One or more of the devices 106a-106c may be a portable device (laptop computer, tablet, smartphone), such as the exemplary portable device 1200 (FIG. 12). Note that while three devices are illustrated, many more of these devices are typically present in an environment such as computing environment 100. In operation, users of devices 106a-106c electronically access electronic content embedded in machine-encoded files, such as electronic documents, audio files, video files, and various other media, hosted by information node 104 and accessed via a wired or wireless connection to data communications network 108. The various media carry a variety of electronic content, including news articles, educational materials, tabloid articles, entertainment materials, and the like. Data communications network 108 typically includes the Internet but may also include wide area networks (WANs), local area networks (LANs), and various other data communications networks, through which electronic documents, audio, video, and various other media may be accessed and retrieved.
[0022] The amount of information available to users of devices 106a-106c from information node 104 via data communications network 108 can be substantial. Given the amount of information available to device users electronically as electronic content from information node 104 via data communications network 108, the information electronically presented to the device user may overlap in whole or in part with previously presented information. For example, a tablet user may want to catch up on news about a current event by accessing multiple articles produced by different news organizations. The user may have been following developments in the event for some time, and some articles may simply repeat facts the user has already read. As another example, a computer user attending an online university may be assigned several educational videos to watch. Some or all of the educational videos assigned to the user may contain portions that cover the same or substantially similar content.
[0023] In these and other situations, informational recognition and presentation (IRP) systems according to various embodiments described herein can relieve device users from having to unnecessarily spend time and effort sifting through various forms of electronic content (e.g., text, audio, video) and separating newly presented information from previously presented information.
[0024] The IRP system analyzes electronic media content embedded in an electronic presentation structure presented to a user. Based on the analysis, the IRP system detects portions of the electronic media content previously consumed by the user. Based on the detection of one or more portions previously consumed by the user, the IRP system modifies the electronic presentation structure to distinguish the electronic media content previously consumed by the user from other portions of the electronic media content.
[0025] The IRP system also improves the efficiency of the underlying computer system or device hardware by reducing the need to store, retrieve, and process electronic content already consumed by a user. When a user electronically retrieves electronic content as a machine-encoded file, the IRP system identifies distinct portions of the machine-encoded file that correspond to previously consumed content.
[0026] In some embodiments disclosed herein, the IRP system uses natural language processing (NLP) or other machine learning to identify distinct portions, or segments, in machine-encoded files. Segments are identified by file segment classifiers, each corresponding to a particular topic of information presented as electronic content. The IRP system detects previously consumed content within the machine-encoded files that corresponds to one or more file segment classifiers. The previously consumed content is detected by searching a computer database for NLP elements identified in electronic content previously presented to the user that match NLP elements extracted from the machine-encoded files. This search is performed against a user-specific consumption corpus database, which additionally supports the machine learning performed in the IRP system.
[0027] Based on the detection, the IRP system modifies an electronic presentation structure (e.g., an electronic document, an audio file, a video file) for electronically presenting the machine-encoded file to a user. The electronic presentation structure can be modified to distinguish portions of the machine-encoded file that correspond to previously consumed content from other portions of the machine-encoded file.
[0028] In the case of an electronic presentation structure comprised of text, the IRP system may modify the electronic presentation structure by annotating the text to distinguish portions of the text corresponding to portions of the electronic media content previously consumed by the user from other portions of the text. Additionally or alternatively, in some embodiments, the IRP system may modify one or more formatting tags to change the format of the text corresponding to portions of the electronic media content previously consumed by the user from other portions of the text. In other embodiments, the IRP system may also or alternatively extract portions of the text corresponding to portions of the electronic media content previously consumed by the user. Optionally, the IRP system may substitute a link for the extracted portion. This link may electronically link the user to a database where the extracted portion is electronically stored. The IRP system may issue instructions directing the user to link to the database if the user desires to add the extracted portion back into the electronic presentation structure. Optionally, the IRP system may also communicate to the user an indicator of when or how long ago the user consumed the extracted portion.
[0029] In the case of an electronic presentation structure comprised of audio or video files, the IRP system may modify the electronic presentation structure by increasing the speed of audio or video rendering for portions of the electronic media content previously consumed by the user. Additionally or alternatively, in other embodiments, the IRP system may shorten the audio or video to remove portions of the electronic media content previously consumed by the user. The IRP system may render the audio or video at a normal speed for portions determined not previously consumed by the user.
[0030] As shown, in the computing environment 100, the IRP system 110 is implemented as a cloud-based system integrated with or operably connected to the computing node 102. The computing node 102 may be a cloud-based server. The IRP system 110 may be implemented in hardware, software, or a combination of both. By performing the operations and functions described herein, the IRP system 110 compares a newly presented electronic presentation structure (e.g., electronic document, audio, video) to a user with previously presented data 112 previously presented to the user and electronically stored in a database accessible to the IRP system 110. Based on the comparison, the IRP system 110 distinguishes portions of information previously consumed by the user from new information electronically presented to the user. The IRP system 110 implemented as a cloud-based system is accessed via a data communications network 108 by a user using a device 106a or other electronic device communicatively connected to the data communications network 108.
[0031] In another configuration, the IRP system 114 is a local IRP system implemented on the device 106b. The IRP system 114 can perform the same or similar operations and functions as the IRP system 110, but is implemented on the device 106b rather than on a cloud-based server like the IRP system 110. For example, the IRP system 114 can be a computer application downloaded to the device 106b (e.g., a desktop, laptop, tablet, smartphone). Optionally, the IRP system 114 can also locally store past presentation data (e.g., electronic documents, audio, video) in performing the operations and functions described herein. Alternatively, the IRP system 114 can access the independently stored past presentation data 112 via the data communications network 108. The device 106c can be another device used by the same user. The IRP system 114 operating on device 106b can be operatively connected to device 106c to perform the same operations and functions with respect to the electronic presentation structure presented to the user on device 106c.
[0032] FIG. 2 illustrates in more detail some example components of an IRP system 200 according to one embodiment. The IRP system 200 is implemented with separate software modules. As shown, these modules include an analyzer 202, a selector 204, a comparator 206, and a reconstructor 208. The modules that make up the IRP system 200 process machine-encoded files acquired electronically by a user. The machine-encoded files may contain electronic content presented to the user as electronic documents, audio renderings, videos, or other electronic presentation structures. The IRP system 200 operates by processing textual data embedded in the machine-encoded files. Thus, the textual data may be contained in electronic documents (e.g., web pages, online news articles, messages) or may be converted from audio (e.g., pre-recorded audio, audio portions of videos) by a speech-to-text processor (not shown).
[0033] In various embodiments, the IRP system 200 integrates various elements of machine learning. The analyzer 202 identifies topics within articles based on topic analysis of electronic content embedded in machine-encoded files. The analyzer 202 performs natural language processing to parse the electronic content into textual components, evaluate the meaning of the text, and resolve inherent ambiguities. The natural language processing by the analyzer 202 may include semantic analysis and / or word-sense disambiguation based on tokenization, lemmatization, and / or part-of-speech tagging, or a combination thereof.
[0034] The analyzer 202 employs natural language understanding using deep learning. In some embodiments described herein, deep learning is performed by a multi-hidden layer neural network. Using the multi-hidden layer neural network, the analyzer 202 performs text analysis to extract keywords, concepts, categories, relations, sentiments, emotions, syntax, and other natural language elements from given text. As defined herein, a "natural language processing (NLP) element" is any object detected in text or derived (explicitly or implicitly) from text based on natural language processing that at least partially explains the natural language meaning of a particular segment of text. As described below, NLP elements include keywords, concepts, relations, and the like, in addition to semantic and syntactic objects. In some embodiments, the analyzer 202 performs text analysis to detect keyword frequency and extract frequently occurring phrases related to each other from unstructured text. When analyzing the machine-encoded file 210, the analyzer 202 performs text analysis in conjunction with a machine learning model trained to identify one or more topics that correspond to distinct segments 212 in the machine-encoded file 210.
[0035] In some embodiments, analyzer 202 performs NLP to determine one or more file segment classifiers contained within machine-encoded file 210. Each file segment classifier corresponds to a distinct portion, or segment, of machine-encoded file 210 that relates to a particular topic of the information embedded in machine-encoded file 210 as electronic content. The information contained in distinct segments 212 may cover multiple topics.
[0036] In some embodiments, the analyzer 202 performs topic modeling, an unsupervised machine learning technique, to determine topics corresponding to the file segment classifiers. By performing topic modeling, the analyzer 202 groups texts based on the words they contain and the correlations identified between them. In some embodiments, the analyzer 202 performs topic modeling using latent semantic analysis, which groups texts based on the frequency with which the same words appear in each text, based on a distributional hypothesis. In other embodiments, the analyzer 202 performs topic modeling using latent Dirichlet analysis, a generative statistical model that maps texts to topics and groups the texts so that the words in each text are captured (or nearly captured) by the topics.
[0037] Topic modeling based on unsupervised machine learning is suitable for discovering topics of interest to a user that are not explicitly specified by the user. However, if the user specifies topics of interest, the analyzer 202 can perform supervised machine learning to train a classification model that classifies distinct pieces of information based on the predetermined topics. For example, as a user consumes (e.g., reads, listens to, or watches) electronic presentations (e.g., documents, audio, or videos) of information (electronic content), the user can label or tag the electronic presentations according to user-identified topics. The labeled electronic presentations can be added to a corpus of labeled training examples and stored electronically. In some embodiments, the analyzer 202 implements a deep learning neural network classification model trained with the user-labeled electronic presentations to classify distinct pieces of electronic content according to topics.
[0038] Thus, in some embodiments, the IRP system 200 offers users the option of having the system learn through unstructured machine learning (e.g., topic modeling) to discover topics of interest to them. Alternatively, users have the option of explicitly specifying topics of interest by labeling electronic presentations as they consume the information conveyed by the electronic presentations. By labeling electronic presentations, users build a corpus of labeled training examples that are used to train a classification model (e.g., a deep learning neural network) that identifies topics of particular interest to the user. Additionally, regardless of whether a particular topic is of particular interest to the user, the IRP system 200 in various embodiments described herein detects and identifies previously consumed content, allowing users to skip or otherwise avoid previously consumed content.
[0039] Based on the topics determined by the analyzer 202, the selector 204 retrieves electronically stored topic-specific data. This data structure may include NLP elements generated based on natural language processing of the user's previously acquired electronic files 214. The data structure includes a user-specific consumption corpus electronically stored in a consumption corpus database 216. This includes previously read, listened to, or viewed electronic content stored electronically in a relational or non-relational database. The data structure corresponds to text extracted from electronic documents or text converted from speech by a speech-to-text processor. The data structure may also include a timestamp indicating the date the corresponding information was presented to the user.
[0040] The selector 204 performs a computer database search against the consumed corpus database 216 to search for NLP elements identified within the electronic content in the previously presented files 214 and identify NLP elements that match NLP elements extracted from the machine-encoded files 210. Based on the matches, the selector 204 can select topic-specific segments in the previously presented electronic content for comparison with the topic-specific distinct segments 212 in the machine-encoded files 210.
[0041] Optionally, the selector 204 may, based on a user's instruction, select only files from the previously presented files 214 that were presented to the user within a specified time period. In this manner, electronic comparison of information for a particular user is based on both the topic and the time period over which the user consumed information. For example, the user may specify that newly presented information be compared to information consumed within the last week, last month, last year, or any other time period. This may also reflect the user's level of interest in a particular topic.
[0042] For each of one or more topics identified by the analyzer 202, the comparator 206 compares the distinct segments 212 with corresponding topic-specific content contained in articles electronically stored in a consumer corpus database 216 and retrieved by the selector 204 based on the topics identified by the analyzer 202. The comparator 206 compares the NLP elements extracted from the distinct segments 212 by the analyzer 202 with NLP elements extracted from a previously retrieved file 214 electronically stored in the consumer corpus database 216.
[0043] The analyzer 202 performs text analysis on distinct segments 212 in the retrieved machine-encoded file 210 and previously consumed text with topics corresponding to those of the machine-encoded file 210 to extract NLP elements. Based on natural language understanding, the analyzer 202 identifies and extracts NLP elements, such as names, places, and events. The analyzer 202 also extracts keywords. The analyzer 202 also extracts concepts. Concepts are broad characterizations of topics or ideas expressed in the text. For example, if the machine-encoded file 210 consists of research papers on deep learning, the analyzer 202 can extract the concept of "artificial intelligence" even if the words "artificial" and "intelligence" are not explicitly used in the text. The analyzer 202 also extracts relationships between NLP elements. For example, if the machine-encoded file 210 consists of news articles about the Nobel Prize, including the NLP elements "Nobel Prize in Physics" and the name "Andrea M. Ghez," the analyzer 202 can connect these NLP elements with the relationship "awarded to."
[0044] The analyzer 202 also classifies NLP elements. In some embodiments, the analyzer 202 generates a classification hierarchy of text topics (e.g., finance / financial news, computer / operating system / hardware). The analyzer 202 can also classify text based on the NLP elements of sentiments (e.g., positive, negative, neutral) and / or emotions (e.g., happy, sad, anger, frustration).
[0045] Based on a topic-based comparison of the NLP elements of the machine-encoded file 210 with the NLP elements of the previously retrieved file 214, the comparator 206 identifies which distinct segments 212 in the machine-encoded file 210 contain previously consumed information. A distinct segment in the machine-encoded file 210 may repeat, word for word, a portion of the previously retrieved file related to the same topic, in which case the comparator 206 identifies the distinct segment in the machine-encoded file 210 as previously consumed content. However, in other situations, a distinct segment in the machine-encoded file 210 may not be word for word identical to a segment in the previously retrieved file related to the same topic, but the similarity between the segments may be sufficient to allow the comparator 206 to identify the distinct segment in the machine-encoded file 210 as previously consumed content.
[0046] The similarity can be determined by the comparator 206 performing a similarity analysis that compares corresponding NLP elements extracted from each file. For example, the comparator 206 can determine the percentage of NLP elements that are common to both files. Based on that percentage, the comparator 206 can determine the probability that each segment is an overlap. Sufficient similarity for the comparator 206 to identify distinct segments in the machine-encoded file 210 as previously consumed content can be based on whether the probability exceeds a certain level (e.g., 80 percent) predetermined by the user.
[0047] Thus, the IRP system 200 can identify previously consumed content appearing in a newly presented machine-encoded file based on topic analysis and natural language understanding, even if a side-by-side comparison of topic-specific segments does not reveal that the previously acquired content covers substantially the same subject matter but is presented differently. For example, a newly acquired financial news article may consist of text written differently from a previously acquired news article. However, a comparison of NLP elements, such as keywords, terms (e.g., names, dates, numbers), concepts (e.g., monetary policy), or relationships (e.g., monetary policy and interest rates), or combinations thereof, may reveal a high percentage (e.g., 90%) of the same NLP elements appearing in both articles, indicating that the later article substantially overlaps with the earlier article.
[0048] In some embodiments, the IRP system 200 generates a similarity metric, such as the percentage of identical or common NLP elements, to determine whether one or more of the distinct segments 212 are composed of information the user has previously consumed. The IRP system 200 can also include an offset to the similarity metric, for example, based on sentiment and / or emotion. A similarity metric offset is a negating factor that reduces the similarity metric between two or more articles that are similar except for differences in sentiment and / or emotion. The basic idea is that such articles may state common facts but offer contrasting perspectives on the same set of facts. For example, two articles consisting of editorials on the same topic are likely to have some similarity in terms of NLP elements. However, if the articles have different perspectives, the sentiment and / or emotion in the articles are likely to differ. Therefore, the IRP system 200 reduces the similarity and concludes that the content in a later-acquired article is not previously consumed based on differences in sentiment and / or emotion. Sentiments and emotions, along with other NLP elements, can be stored electronically in a consumption corpus.
[0049] The cooperative operations and functions performed by the analyzer 202, selector 204, and comparator 206 can be performed on text obtained directly from any medium (e.g., electronic documents, audio, video) or through a speech-to-text process. Similarly, a user-specific consumption corpus aggregates data structures corresponding to text obtained from different media. Therefore, newly presented information in one medium can be compared with previously presented information in a completely different medium. For example, newly presented information in an electronic presentation structure consisting of electronic documents can be compared with previously presented information in other electronic presentation structures, including not only other electronic documents but also audio and video. In this way, the IRP system 200 can compare the content of information electronically presented to a user across media. For example, a user may begin listening to a novel through an audiobook but stop at some point. In this case, the user can resume the story by reading the electronic text on a tablet, where the electronic content has been annotated to indicate the point where the user stopped listening to the audiobook. In this way, the IRP system 200 can present the portion of the novel that was heard in the audiobook, and the user can resume reading the newly presented portion of the novel from the tablet.
[0050] In other embodiments, the IRP system 200 may additionally or alternatively identify distinct portions of newly presented information that correspond to previously presented information based on metadata corresponding to the electronic document, audio, video, or other electronic presentation structure presented to the user. For example, metadata corresponding to an electronic document may indicate the version or publication date of the electronic document. If the electronic document is subsequently revised, for example, by updating old information (e.g., proper nouns, dates, or numbers) or adding new information, the comparison unit 206 may identify changes between the two versions by comparing metadata corresponding to the earlier and later published versions. In yet other embodiments, the IRP system 200 may additionally or alternatively capture URL hyperlinks for online publications consumed by the user and store the content in a relational or non-relational database, thereby enabling subsequent comparison of ongoing changes to the information contained in the content as it is updated.
[0051] The reconstructor 208 modifies the electronic presentation structure based on the comparator 206's detection of distinct portions of the presented information that correspond to previously presented information. For information presented as text, the reconstructor 208 can annotate the text of the electronic document to contrast previously presented portions of the information with newly presented portions to the user. In some embodiments, the reconstructor 208 highlights the text with a different color, font, formatting (e.g., bold, italics, strikethrough), background, or combination thereof to contrast newly presented portions with previously presented portions. In some embodiments, the reconstructor 208 shortens the text, removing portions identified as previously presented to the user.
[0052] In some embodiments, the reconstructor 208 modifies the audio or video electronic presentation structure by shortening the structure and removing distinct portions corresponding to previously presented information. In other embodiments, the reconstructor 208 alters the presentation speed of portions of the audio or video that contain previously presented information. The reconstructor 208 can alter the presentation speed of the audio or video to fast-forward or skip portions previously heard or viewed so that the user does not have to hear or view those portions again. The reconstructor 208 renders the audio or video at normal speed for portions identified as newly presented portions.
[0053] For example, a user interested in a topic may read electronically acquired electronic documents, watch videos, and listen to audio podcasts. As the user's understanding of the topic deepens, the user may be more interested in consuming new information rather than rereading, rewatching, or relistening to previously consumed content. The IRP system 200 continuously updates the consumption corpus database 216 as the user consumes content related to the same topic. If the user listens to a podcast related to the topic, the reconstructor 208 may advance the audio (e.g., in seven-minute increments) and begin the audio at the point where new content is introduced. If the user watches a video related to the topic, the reconstructor 208 may fast-forward the video (e.g., at twice the normal speed) for the portion corresponding to the previously consumed content and resume presentation of the video at the point where the newly presented content begins. When a user retrieves electronic documents on this topic, the reconstructor 208 annotates information that overlaps with previously consumed information (e.g., other electronic texts, previous audio or video), and highlights, annotates, or otherwise distinguishes newly presented information from previously consumed information.
[0054] In some embodiments, the IRP system 200 recognizes metadata associated with documents, audio, or video. The metadata indicates that the information presented in the content is related to the specific location (e.g., GPS coordinates) of the device that electronically presents the information. For example, a user may read wirelessly captured electronic documents using a tablet or other device capable of receiving and processing wireless signals while traveling through multiple locations in a factory. At each location, the user may be required to review a set of usage or safety documents, regardless of whether they have previously read them at other locations. The goal is to ensure that the user reviews the information presented each time the user visits a particular location. Based on the metadata recognized by the IRP system 200, the IRP system 200 can highlight specific portions of the corresponding electronic documents at each location to indicate the specific portions of information that need to be read. Device location data (e.g., GPS coordinates) can be included in content requests sent by devices when requesting content. For example, a plug-in can include such data. The IRP system 200 can use a comparison of the device location data with other data associated with the requested location-related content when modifying the electronic presentation structure.
[0055] In other embodiments, the IRP system 200 tracks the context in which information is presented in various media (e.g., text, audio, video) to treat previously presented information differently depending on the context in which the information is presented. A user can impose limitations on modifying the electronic presentation structure by considering the article's contextual basis, determined from keywords, concepts, relationships, or other NLP elements extracted by the IRP system 200 as described above. As defined herein, a "contextual basis" is a topic, subtopic, or concept determined by natural language processing that characterizes an article in whole or in part based on an overarching theme. For example, a user may want to track information consumed in an article about the economic impact of COVID-19, but may not want to read repetitive facts in a new article about the economic impact of COVID-19. However, the user may want to read repetitions of the same facts when presented in the context of a new pick, such as in an article about Major League Baseball. The IRP system 200 modifies the article's electronic presentation structure if the article's contextual basis is economic impact, but not if the contextual basis is Major League Baseball. Thus, a user can use machine-encoded instructions to specify that the reconstructor 208 should reconstruct an electronic presentation structure for a topic when presented in one context (e.g., economic impact) but not when presented in another context (e.g., professional sports).
[0056] In yet another embodiment, a user can use machine-encoded instructions to specify that the IRP system 200 electronically tag information with a purpose indicator, a date indicator, or both, to indicate when and why the user consumed the information. For example, a user engaged in online learning can specify that the information consumed on November 3rd was for a particular assignment. If the same information is subsequently presented in a newly generated electronic presentation structure, the reconstructor 208 can annotate the text or otherwise inform the user when and / or why the same information was previously consumed. For information conveyed via audio or video, the reconstructor 208 can also generate a recording that indicates to the user the audio portions already heard or the video portions already viewed. Optionally, for example, in online learning, the IRP system 200 can generate credits to be awarded to the user each time the user consumes a new assignment if the corresponding information has not previously been consumed by the user.
[0057] In yet another embodiment, the IRP system 200 can track the amount of previously presented information contained in an electronic presentation structure. The IRP system 200 can generate an indicator of the proportion of information contained in an article that has been previously consumed. For example, an electronic presentation structure composed of text can include an annotation indicating the proportion of the text contained in the structure that corresponds to previously presented information. For example, in response to a user retrieving an online article over a data communications network, the IRP system 200 can communicate an electronic message to the user with a corresponding "relationship score" indicating the portion of the article that has been previously consumed. Based on the relationship score, the IRP system 200 can query the user as to whether they wish to view the online article or cancel their request to access the article.
[0058] Similarly, an electronic presentation structure comprised of audio or video may include a pre-recorded indicator that tells the user what percentage of the audio or video the user has previously consumed, thus allowing the user to ascertain whether the article, audio, or video contains a sufficient amount of newly presented information to merit the user's time to read, listen to, or watch the electronic presentation structure.
[0059] FIG. 3 illustrates some operational aspects 300 of an IRP system. According to one embodiment, the IRP system is implemented as a software module or suite of modules that function as an add-on to a web browser. At block 302, a user installs a web plug-in for the IRP system. In various embodiments, this plug-in can function as part of a cloud-based IRP system, such as IRP system 110 (FIG. 1), an IRP system, such as IRP system 114 (FIG. 1), or both. At block 304, the IRP system captures a URL hyperlink to electronic content 306 in response to the user accessing and obtaining a machine-encoded file containing electronic content 306 over a data communications network. The machine-encoded file can be formatted as an electronic document, an audio file, or a video file. At block 308, the IRP system captures electronic content 306. At block 310, a data structure 312, along with the URL hyperlink, is electronically stored in a database containing a consumption corpus 314. The data structure 312 consists of the URL hyperlink and the electronic content of the machine-encoded file. At block 316, the IRP system performs one or more of the various NLP and relevance analyses described above to determine whether any portions of the electronic content 306 have been previously consumed by the user.
[0060] This determination may be based on matching NLP elements extracted from the electronic content 306 with NLP elements identified within the electronic content stored in the consumption corpus 314. If the IRP system is cloud-based, the IRP system may act as a proxy, and the request for the electronic content 306 may be forwarded to the proxy via a plug-in. The proxy then performs the NLP and relevance analysis. Optionally, at block 318, the IRP system generates a relevance score and attaches it to the URL hyperlink. The relevance score indicates the percentage of the electronic content 306 determined to have been previously consumed by the user. The IRP system may send a message indicating the relevance score and ask the user whether they wish to load the electronic content 306. In response to determining that a portion of the electronic content 306 has been previously consumed by the user, at block 320, an electronic presentation structure for presenting the electronic content 306 to the user is reconstructed.
[0061] FIG. 4 illustrates an example of a reconstructed electronic presentation structure comprised of a text document 400. Based on the above-described operational aspects 300 performed by the IRP system, portions of the text document 400 are identified as corresponding to information previously consumed by a user. The text document 400 is reconstructed as an electronic presentation structure 402. The previously consumed portions are contrasted with newly presented portions of electronic content. The electronic presentation structure 402, which is a modified or reconstructed version of the text document 400, includes distinct portions 404, 406, 408, and 410. As shown, each portion 404-410 is uniquely highlighted with distinct shadings to indicate previously consumed portions of information. The highlight colors can correspond to different topics within the text document, and the different shadings of the highlights can indicate the length of time since the corresponding portion (by topic) was consumed. For example, the text document 400 may be a news article about a new virus. When a user electronically requests the text document 400 over a data communications network, the IRP system identifies previously consumed portions. The separate portions 404, 406, 408, 410 can be highlighted in different shades to indicate previously consumed portions and the topics to which each relates.
[0062] Figure 5 is a flow chart illustrating a method 500 for electronically recognizing and presenting information, according to one embodiment. Method 500 may be performed by the same or similar systems as those described with reference to Figures 1 and 2. At block 502, the system analyzes electronic media content embedded in an electronic presentation structure to be presented to a user. The electronic media content may be presented as an electronic document, an audio recording, a video recording, a multimedia presentation, a CD-ROM, online content, etc., and is typically presented in digital form, but may include analog data in some embodiments.
[0063] Based on the analysis, the system detects portions of the electronic media content previously consumed by the user at block 504. Based on the detection, the system modifies the electronic presentation structure at block 506. The electronic presentation structure is modified to distinguish portions of the electronic media content previously consumed by the user from other portions of the content.
[0064] If the electronic presentation structure is comprised of audio or video files, the system may modify the electronic presentation structure by increasing the rendering speed of the audio or video. By increasing the speed of portions of the electronic media content previously consumed by the user, the user can quickly advance through the previously consumed portions of the electronic media content without having to reread or rewatch those portions. Alternatively or additionally, in other embodiments, the system may shorten the audio or video to remove portions of the electronic media content previously consumed by the user.
[0065] If the electronic presentation structure is composed of text, the system may modify the electronic presentation structure by annotating the text to distinguish portions of the text that correspond to portions of the electronic media content that the user has previously consumed from portions of the text that the user has not previously consumed. In other embodiments, the system may alternatively or additionally modify the electronic presentation structure by reformatting one or more formatting tags to distinguish portions of the text that correspond to portions of the electronic media content that the user has previously consumed from portions of the text that the user has not previously consumed.
[0066] In yet another embodiment, the system may alternatively or additionally modify the electronic presentation structure by extracting text portions corresponding to portions of the electronic media content previously consumed by the user. Optionally, the system may substitute a link for the extracted portion. This link may electronically connect the user (e.g., via a data communications network) to a database in which the extracted portion is electronically stored. The system may communicate the link to the user, along with instructions for the user to link to the database if the user wishes to add the extracted portion back to the electronic presentation structure. A notification provided by the system to the user may indicate that the extracted portion was previously consumed by the user and may ask whether the user wishes to consume the extracted portion again. If so, the user may do so using a link provided by the system. The notification may indicate when or how long ago the user consumed the extracted portion.
[0067] Figure 6 is a flow chart illustrating a method 600 for electronically recognizing and presenting information, according to one embodiment. Method 600 can be performed by the same or similar systems as those described with reference to Figures 1 and 2. At block 602, the system performs natural language processing (NLP) on a machine-encoded file to identify one or more file segment classifiers. The machine-encoded file can be obtained by a user over a data communications network.
[0068] At block 604, the system locates previously consumed content in the machine-encoded files that corresponds to one or more file segment classifiers by performing a computer database search for NLP elements identified in electronic content previously presented to the user that match the NLP elements extracted from the machine-encoded files. This search is performed against a user-specific consumption corpus database.
[0069] Based on the detection, the system modifies an electronic presentation structure for electronically presenting the machine-encoded file to the user at block 606. The electronic presentation structure is modified to distinguish portions of the machine-encoded file that correspond to previously consumed content from other portions of the machine-encoded file.
[0070] In some embodiments, the system modifies the electronic presentation structure in response to determining the contextual basis of the one or more file segment classifiers. The one or more file segment classifiers can correspond to one or more topics of interest to the user. In some embodiments, the system performs unsupervised machine learning based topic modeling to automatically discover topics of interest to the user. The system uses topic modeling to discover topics of interest to the user based on multiple machine-encoded files previously acquired by the user over a data communications network.
[0071] In another embodiment, the one or more file segment classifiers correspond to one or more topics of interest to the user that the system identifies through topic classification based on supervised machine learning. The system identifies the topics based on classification performed by a topic classifier trained through machine learning using a consumption corpus consisting of electronic content selected by the user.
[0072] In yet another embodiment, the previously presented electronic content includes one or more machine-encoded files previously obtained over a data communications network within a user-specified time period.
[0073] In yet another embodiment, the system modifies the electronic presentation structure by annotating the electronic presentation structure with a relevance score. The system generates the relevance score by determining the percentage of the machine-encoded file that consists of previously consumed content.
[0074] 7 is a flowchart illustrating an exemplary method 700 for electronically recognizing and presenting information relative to a particular location of a user, according to one embodiment. Method 700 can be performed by the same or similar system as described with reference to FIGS. 1 and 2. At block 702, the system electronically tracks the location of a device for electronically receiving and presenting an electronic presentation structure to a user. At block 704, the system modifies the electronic presentation structure based on the device's location.
[0075] FIG. 8 is a flowchart illustrating an exemplary method 800 for detecting updates to portions of information previously presented to a user, according to one embodiment. Method 800 can be performed by the same or similar system as described with reference to FIGS. 1 and 2. At block 802, the system detects one or more revisions made to the machine-encoded file since the user acquired the file. At block 804, in response to detecting one or more revisions, the system presents the machine-encoded file using a different electronic presentation structure that indicates which portions are newly added or revised. In some embodiments, the system can automatically timestamp and record versions on links to sites where users access the machine-encoded file. When a new version of a machine-encoded file is generated, the system can detect the revision by performing topic analysis and natural language processing described herein to discover whether the newly generated version associated with the link contains portions that the user has not previously consumed. Based on detecting portions that have not previously been consumed, the system can communicate with the user. The system can indicate to the user which portions of the revised version are newly presented, allowing the user to access the newly presented information if desired, without having to re-consume portions they have already consumed.
[0076] Although this disclosure includes detailed descriptions of cloud computing, implementation of the teachings described herein is not limited to cloud computing environments. Rather, embodiments of the invention can be practiced in conjunction with any other type of computing environment now known or later developed.
[0077] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with the service provider. This cloud model may include at least five characteristics, at least three service models, and at least four implementation models.
[0078] The characteristics are as follows:
[0079] On-Demand Self-Service: Cloud consumers can unilaterally provision computing capacity, such as server time or network storage, automatically as needed, without the need for human interaction with the service provider.
[0080] Broad network access: Computing power is available over the network and can be accessed through standard mechanisms, facilitating use by heterogeneous thin or thick client platforms (e.g., cell phones, laptops, PDAs).
[0081] Resource Pooling: Computing resources from a provider are pooled and offered to multiple consumers using a multi-tenant model. Various physical and virtual resources are dynamically allocated and reallocated based on demand. Consumers generally have no control or knowledge of the exact location of the resources they are provided with, so there is a sense of location independence. However, consumers may be able to identify their location at a higher level of abstraction (e.g., country, state, data center).
[0082] Rapid Elasticity: Computing capacity can be provisioned quickly and elastically, sometimes automatically, to instantly scale out and quickly released to instantly scale in. To the consumer, the computing power available to provision often appears unlimited, and can be purchased at any time and in any quantity.
[0083] Metered Services: Cloud systems leverage measurement capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, active user accounts) to automatically control and optimize resource usage. Resource usage can be monitored, controlled, and reported to provide transparency to both providers and consumers of utilized services.
[0084] The service model is as follows:
[0085] Software as a Service (SaaS): The functionality offered to the consumer is the availability of a provider's applications running on a cloud infrastructure. The applications can be accessed from a variety of client devices through a thin client interface such as a web browser (e.g., webmail). The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even individual application functions, except for limited user-specific application configuration settings.
[0086] Platform as a Service (PaaS): The capability offered to consumers is to deploy applications they create or acquire onto a cloud infrastructure using programming languages and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but does have control over the deployed applications and, in some cases, the configuration of their hosting environment.
[0087] Infrastructure as a Service (IaaS): The functionality provided to consumers is the provisioning of processors, storage, networking, and other basic computing resources on which the consumer can deploy and run any software, including operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure, but has control over the operating systems, storage, and deployed applications, and in some cases partial control over some network components (e.g., host firewalls).
[0088] The deployment model is as follows:
[0089] Private Cloud: This cloud infrastructure is dedicated to a specific organization and can be managed by that organization or a third party, and can exist on-premise or off-premise.
[0090] Community Cloud: This cloud infrastructure is shared by multiple organizations to support a specific community with common concerns (e.g., mission, security requirements, policies, and compliance). This cloud infrastructure can be managed by those organizations or a third party and can exist on-premises or off-premises.
[0091] Public cloud: This cloud infrastructure is available to the general public or large industry organizations and is owned by an organization that sells cloud services.
[0092] Hybrid cloud: This cloud infrastructure combines two or more cloud models (private, community, or public), each of which retains its inherent nuances but is bound by standards or specific technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).
[0093] A cloud computing environment is a service-oriented environment that emphasizes statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is an infrastructure that includes a network of interconnected nodes.
[0094] FIG. 9 illustrates an exemplary cloud computing environment 900. As illustrated, the cloud computing environment 900 includes one or more cloud computing nodes 910, with which local computing devices used by cloud consumers (e.g., PDA or cell phone 940a, desktop computer 940b, laptop computer 940c, or automobile computer system 940n, or combinations thereof) can communicate. The computing nodes 910 can communicate with each other. The nodes 910 can be physically or virtually grouped (not shown) in one or more networks, such as, for example, a private, community, public, or hybrid cloud, or combinations thereof, as described above. This enables the cloud computing environment 900 to provide infrastructure, platform, or software as a service, or combinations thereof, for which cloud consumers are not required to maintain resources on their local computing devices. It should be understood that the types of computing devices 940a-n illustrated in FIG. 9 are merely exemplary, and that the computing nodes 910 and the cloud computing environment 900 can communicate with any type of electronic device via any type of network or network-addressable connection (e.g., using a web browser), or both.
[0095] Next, a set of functional abstraction layers provided by the cloud computing environment 900 (FIG. 9) is shown in FIG. 10. It should be understood in advance that the components, layers, and functions shown in FIG. 10 are merely examples, and embodiments of the present invention are not limited thereto. As shown, the following layers and corresponding functions are provided:
[0096] Hardware and software layer 1060 includes hardware and software components. Examples of hardware components include mainframe 1061, reduced instruction set computer (RISC) architecture-based server 1062, server 1063, blade server 1064, storage device 1065, and network and network components 1066. In some embodiments, software components include network application server software 1067 and database software 1068.
[0097] The virtualization layer 1070 provides an abstraction layer from which virtual entities such as virtual servers 1071, virtual storage 1072, virtual networks including virtual private networks 1073, virtual applications and operating systems 1074, and virtual clients 1075 can be provided.
[0098] By way of example, the management layer 1080 may provide the following functionality: Resource provisioning 1081 enables dynamic procurement of computing and other resources utilized to execute tasks within the cloud computing environment. Metering and pricing 1082 enables cost tracking as resources are utilized within the cloud computing environment and billing or invoicing for the consumption of these resources. By way of example, these resources may include application software licenses. Security enables identification and verification of cloud consumers and tasks, as well as protection for data and other resources. User portal 1083 provides consumers and system administrators with access to the cloud computing environment. Service level management 1084 enables allocation and management of cloud computing resources so that requested service levels are met. Service level agreement (SLA) planning and fulfillment 1085 enables advance arrangement and procurement of anticipated future cloud computing resources required in accordance with SLAs.
[0099] The workload layer 1090 provides examples of functionality that can be utilized in a cloud computing environment. Examples of workloads and functionality that can be provided from this layer include mapping and navigation 1091, software development and lifecycle management 1092, virtual classroom instruction delivery 1093, data analytics processing 1094, transaction processing 1095, and IRP systems 1096.
[0100] 11 is a schematic diagram of an example computing node 1100. In one or more embodiments, computing node 1100 is an example of a suitable cloud computing node. Note that computing node 1100 is not intended to suggest any limitation as to the scope of use or functionality of the embodiments of the invention described herein. Computing node 1100 may perform any of the functions described within this disclosure.
[0101] Computing node 1100 includes a computer system 1112. The computer system 1112 is operational with numerous other general purpose or special purpose computing system environments or configurations. Some well-known computing systems, environments, or configurations, or combinations thereof, suitable for use with the computer system 1112 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices.
[0102] The computer system 1112 may be described in the general context of computer system-executable instructions, such as program modules, executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, etc. that perform particular tasks or implement particular data types. The computer system 1112 may be practiced in a distributed cloud computing environment where tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules may be stored in both local and remote computer system storage media, including memory storage devices.
[0103] As shown in Figure 11, computer system 1112 is depicted as a general-purpose computing device. Example components of computer system 1112 include one or more processors 1116, memory 1128, and a bus 1118 connecting various system components, including memory 1128, to processor 1116. As defined herein, a "processor" means at least one hardware circuit configured to execute instructions. A hardware circuit may be an integrated circuit. Example processors include a CPU, an array processor, a vector processor, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic array (PLA), an application-specific integrated circuit (ASIC), a programmable logic circuit, a controller, etc.
[0104] Executing computer program instructions by a processor includes running the program. As defined herein, "run" includes a sequence of actions or events carried out by a processor in accordance with one or more machine-readable instructions. As defined herein, "running" means the active performing of actions or events by a processor. As used herein, the terms "run," "running," "execute," and "executing" are used interchangeably.
[0105] Bus 1118 represents any one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port (AGP), and a processor or local bus using any of a variety of bus architectures. By way of example, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, a Peripheral Component Interconnect (PCI) bus, and a PCI Express (PCIe) bus.
[0106] Computer system 1112 typically includes a variety of computer system readable media. Such media can be any available media that can be accessed by computer system 1112 and can include both volatile and nonvolatile media, removable and non-removable media.
[0107] The memory 1128 may include a computer system-readable medium such as a volatile memory, such as RAM 1130 or cache memory 1132, or both. The computer 1112 may also include other removable / non-removable and volatile / non-volatile computer system-readable media. As an example, the storage system 1134 may be provided for reading from and writing to a non-removable, non-volatile magnetic medium and / or a solid-state drive (not shown, commonly referred to as a "hard drive"). Also, although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a floppy disk) and an optical disk drive for reading from and writing to a removable, non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these examples, each may be connected to the bus 1118 by one or more data medium interfaces. As further shown and described below, the memory 1128 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of embodiments of the present invention.
[0108] As an example, a program / utility 1140 having a set (at least one) of program modules 1142 can be stored in memory 1128, as can an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, can include an implementation of a network environment. The program modules 1142 generally perform the functions and / or methods of embodiments of the invention described herein. For example, one or more of the program modules can include an IRP system 996 or portions thereof.
[0109] The programs / utilities 1140 are executable by the processor 1116. The programs / utilities 1140, and the data items used, generated, or manipulated, or a combination thereof, by the computer system 1112, are functional data structures that, when utilized by the computer system 1112, provide functionality to the system. A "data structure," as defined in this disclosure, is a physical implementation in physical memory of the data organization of a data model. As such, a data structure is formed by specific electrical or magnetic structural elements in memory. The data structure provides a physical organization to data stored in memory for use by application programs executed by the processor.
[0110] The computer system 1112 may also communicate with one or more external devices 1114, such as a keyboard, pointing device, display 1124, one or more devices that enable a user to interact with the computer system 1112, or any device (e.g., a network card, modem, etc.) that enables the computer system 1112 to communicate with one or more other computer devices, or a combination thereof. Such communication may occur via an input / output (I / O) interface 1122. The computer system 1112 may also communicate with one or more networks (e.g., a local area network (LAN), a general wide area network (WAN), or a public network (e.g., the Internet), or a combination thereof) via a network adapter 1120. As shown, the network adapter 1120 may communicate with other components of the computer system 1112 via a bus 1118. Note that other hardware and / or software components (not shown) may be used with the computer system 1112. Examples include microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archive storage systems.
[0111] It should be noted that while computing node 1100 is illustrated as an example of a cloud computing node, computer systems using the same or similar architecture as described with reference to FIG. 11 may be used in non-cloud computing implementations to perform various operations described herein. In this regard, the example embodiments described herein are not intended to be limited to cloud computing environments. Computing node 1100 is an example of a data processing system. As defined herein, "data processing system" means one or more hardware systems configured to process data, each including at least one processor programmed to initiate operations and memory.
[0112] Computing node 1100 is an example of computer hardware. Computing node 1100 may include fewer components than those shown in FIG. 11 or may include additional components not shown, depending on the specific type of device and / or system being implemented. The specific operating system and / or applications included on the node may vary depending on the type of device and / or system, as may the types of I / O devices included on the node. Additionally, one or more of the illustrated components may be incorporated into or otherwise form part of other components. For example, a processor may include at least some memory.
[0113] Computing node 1100 is also an example of a server. As defined herein, a "server" refers to a data processing system configured to share services with one or more other data processing systems. As defined herein, a "client device" refers to a data processing system that requests shared services from a server and with which a user interacts directly. Examples of client devices include workstations, desktop computers, computer terminals, mobile computers, laptop computers, netbook computers, tablet computers, smartphones, PDAs, smartwatches, smart glasses, gaming consoles, set-top boxes, and smart televisions. In one or more embodiments, the various user devices described herein may be client devices. Network infrastructure such as routers, firewalls, switches, and access points are not client devices as defined herein.
[0114] 12 illustrates an exemplary portable device 1200 in accordance with one or more embodiments described herein. The portable device 1200 may include a memory 1202, one or more processors 1204 (e.g., an image processor, a digital signal processor, a data processor), and an interface circuit 1206.
[0115] In one embodiment, the memory 1202, the processor 1204, or the interface circuit 1206, or a combination thereof, are implemented as separate components. In other embodiments, the memory 1202, the processor 1204, or the interface circuit 1206, or a combination thereof, are integrated into one or more integrated circuits. The various components in the portable device 1200 can be connected, for example, by one or more communication buses or signal lines (e.g., interconnects and / or wires). In one embodiment, the memory 1202 can be connected to the interface circuit 1206 via a memory interface (not shown).
[0116] Sensors, devices, subsystems, and / or input / output (I / O) devices can be connected to interface circuit 1206 to facilitate the functions and / or operations described herein, including generating sensor data. Various sensors, devices, subsystems, and / or input / output (I / O) devices can be connected to interface circuit 1206 directly or through one or more I / O controllers (not shown).
[0117] For example, a position sensor 1210, a light sensor 1212, and a proximity sensor 1214 can be connected to the interface circuit 1206 to facilitate orientation, lighting, and proximity functions, respectively, of the portable device 1200. A position sensor 1210 (e.g., a GPS receiver and / or a GPS processor) can be connected to the interface circuit 1206 to provide geographic location sensor data. An electronic magnetometer 1218 (e.g., an integrated circuit chip) can be connected to the interface circuit 1206 to provide sensor data usable for determining the direction of magnetic north for directional navigation purposes. An accelerometer 1220 can be connected to the interface circuit 1206 to provide sensor data usable for determining the device's change in velocity and direction of movement in three dimensions. An altimeter 1222 (e.g., an integrated circuit) can be connected to the interface circuit 1206 to provide sensor data usable for determining altitude. A voice recorder 1224 can be connected to the interface circuit 1206 to store recorded speech.
[0118] The camera subsystem 1226 can be connected to an optical sensor 1228. The optical sensor 1228 can be implemented using any of a variety of technologies. Examples of the optical sensor 1228 include CCD, CMOS optical sensors, etc. The camera subsystem 1226 and the optical sensor 1228 can be used to facilitate camera functions such as recording images or video clips, or both (hereinafter "image data"). In one embodiment, the image data is a subset of the sensor data.
[0119] The communication functionality may be facilitated by one or more wireless communication subsystems 1230. The wireless communication subsystems 1230 may include radio frequency transceivers, optical (e.g., infrared) transceivers, etc. The specific design and implementation of the wireless communication subsystems 1230 may depend on the specific type of portable device 1200 being implemented and / or the communication network in which the portable device 1200 operates.
[0120] By way of example, wireless communication subsystem 1230 may be designed to operate over one or more mobile networks (e.g., GSM, GPRS, EDGE), Wi-Fi networks including WiMax networks, short-range wireless networks (e.g., Bluetooth networks), or any combination thereof. Wireless communication subsystem 1230 may implement a hosting protocol to allow portable device 1200 to be configured as a base station for other wireless devices.
[0121] The audio subsystem 1232 can be connected to a speaker 1234 and a microphone 1236 to facilitate voice-enabled functions such as voice recognition, voice replication, digital recording, voice processing, and telephony. The audio subsystem 1232 can generate audio-type sensor data. In one or more embodiments, the microphone 1236 can be utilized as a respiratory sensor.
[0122] I / O devices 1238 can be connected to interface circuit 1206. Examples of I / O devices 1238 include, for example, a display device, a touch-sensitive display device, a trackpad, a keyboard, a pointing device, a communication port (e.g., a USB port), a network adapter, a button, or other physical control means. A touch-sensitive device, such as a display screen or pad, or both, is configured to detect contact, movement, contact disruption, etc. using any of a variety of touch sensor technologies. Examples of touch sensor technologies include, for example, capacitive, resistive, infrared, surface acoustic wave technology, other proximity sensor arrays, or other elements for identifying one or more points of contact with the touch-sensitive device. One or more of I / O devices 1238 may be adapted to control functions of sensors, subsystems, etc. of portable device 1200.
[0123] Portable device 1200 further includes a power source 1240. Power source 1240 can provide power to various elements of portable device 1200. In one embodiment, power source 1240 is implemented as one or more batteries. The batteries can be implemented using any of a variety of battery technologies, whether disposable (e.g., replaceable) or rechargeable. In another embodiment, power source 1240 is configured to obtain power from an external power source and provide power (e.g., DC power) to each element of portable device 1200. In the case of a rechargeable battery, power source 1240 can further include circuitry that can charge the battery when connected to an external power source.
[0124] Memory 1202 may include random access memory (e.g., volatile memory) and / or non-volatile memory, such as one or more magnetic disk storage devices, one or more optical storage devices, flash memory, etc. Memory 1202 may store an operating system 1252, such as LINUX, UNIX, a mobile operating system, an embedded operating system, etc. Operating system 1252 may include instructions for handling system services and for performing hardware-dependent tasks.
[0125] The memory 1202 can store other program code 1254. Examples of the other program code 1254 include instructions facilitating communication with one or more additional devices, one or more computers, or one or more servers, or a combination thereof; instructions facilitating graphic user interface processing; processing instructions facilitating sensor-related functions, telephone-related functions, electronic messaging-related functions, web browsing-related functions, media processing-related functions, GPS and navigation-related functions, security functions, camera-related functions including webcam and / or web video functions, etc. The memory 1202 can store applications including the IRP program code 1256. The memory 1202 can also store one or more other applications 1258.
[0126] The various instructions and / or program code described herein are exemplary only and not limiting. The program code may be implemented as separate software programs, procedures, or modules. Memory 1202 may include additional instructions or fewer instructions. Furthermore, various functions of portable device 1200 may be implemented in hardware and / or software, including one or more signal processing circuits and / or application specific integrated circuits.
[0127] The program code stored in memory 1202 and the data used, generated, and / or manipulated by portable device 1200 are functional data structures that, when utilized as part of a device, provide functionality to the device. Further examples of functional data structures include, for example, sensor data, data obtained by user input, data obtained by querying external data sources, baseline information, etc. A "data structure" is a physical implementation in physical memory of the data organization of a data model. As such, a data structure is formed by specific electrical or magnetic structural elements in memory. A data structure adds physical organization to data stored in memory for use by a processor.
[0128] In some embodiments, the various sensors and / or subsystems described with respect to portable device 1200 may be separate devices connected or communicatively linked to portable device 1200 via wired or wireless connections. For example, one or more (or all) of position sensor 1210, light sensor 1212, proximity sensor 1214, gyroscope 1216, magnetometer 1218, accelerometer 1220, altimeter 1222, voice recorder 1224, camera subsystem 1226, audio subsystem 1232, etc. may be implemented as separate systems or subsystems operatively connected to portable device 1200 via I / O devices 1238 and / or wireless communication subsystem 1230.
[0129] Portable device 1200 may include fewer components than those shown in FIG. 12 or may include additional components not shown, depending on the particular type of system in which it is implemented. Furthermore, the particular operating system or applications or other program code, or combinations thereof, included in the device may also vary depending on the type of system. Furthermore, one or more of the illustrated components may be incorporated into or otherwise form part of other components. For example, a processor may include at least some memory.
[0130] Portable device 1200 is illustrative only and not limiting of the present invention. Devices and / or systems configured to perform the operations described herein may have architectures different from that shown in FIG. 12. The architecture may be a simplified version of portable device 1200 and may include a processor and memory storing instructions. The architecture may include one or more sensors described herein. Portable device 1200 or a similar system may collect data using various sensors on the device or sensors connected thereto. However, portable device 1200 may include fewer sensors or other additional sensors. In this disclosure, data generated by a sensor is referred to as "sensor data."
[0131] Implementations of the portable device 1200 include, for example, a smartphone or other mobile device or telephone, a wearable computing device (e.g., a smartwatch), a dedicated medical device, or other suitable handheld, wearable, or easily portable electronic device capable of sensing and processing signals and data sensed by sensors. It should be noted that embodiments may be deployed as a standalone device or as multiple devices in a distributed client / server network system. For example, in some embodiments, a smartwatch may be operably connected to a mobile device (e.g., a smartphone), which may or may not be configured to interact with a remote server and / or computing system.
[0132] Although the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention, some definitions are provided below that apply throughout the specification.
[0133] As used herein, the singular forms "a", "an" and "the" include the plural forms unless the context clearly indicates otherwise.
[0134] As used herein, "another" means at least a second or more.
[0135] As used herein, the phrases "at least one," "one or more," and "and / or" are open-ended expressions that, in their operation, refer to both conjunctions and disjunctions (unless expressly stated otherwise). For example, the phrases "at least one of A, B, and C," "at least one of A, B, or C," "one or more of A, B, and C," "one or more of A, B, or C," and "A, B, or C, or combinations thereof" each refer to A only, B only, C only, both A and B, both A and C, both B and C, or all of A, B, and C.
[0136] As used herein, "automatically" means without user intervention.
[0137] In this specification, the terms "includes / inducing" and / or "comprise / comprising" specify the presence of stated features, integers, steps, operations, elements or components or combinations thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components or groups thereof or combinations thereof.
[0138] As used herein, "if" means "in response to" or "responsive to," depending on the context. Thus, the phrase "if it is determined" may be interpreted as meaning "in response to determining" or "responsive to determining," depending on the context. Similarly, the phrase "if [a stated condition or event] is detected" may be interpreted as meaning "upon detecting [the stated condition or event]," "in response to detecting [the stated condition or event]," or "responsive to detecting [the stated condition or event]," depending on the context.
[0139] As used herein, the term "one embodiment," "an embodiment," "in one or more embodiments," "in particular embodiments," or similar language means that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment described within this disclosure. Thus, the appearances of these phrases and / or similar language throughout this disclosure may (but do not necessarily) all refer to the same embodiment.
[0140] As used herein, the phrases "in response to" and "responsive to" mean to respond or react readily to an action or event. Thus, when a second action is performed "in response to" or "in response to" a first action, there is a causal relationship between the occurrence of the first action and the occurrence of the second action. The phrases "in response to" and "responsive to" express that causal relationship.
[0141] As used herein, "real time" means a level of processing responsiveness that is perceived by a user or system as being sufficiently immediate with respect to the execution of a particular process or decision, or that allows a processor to keep up with some external process.
[0142] As used herein, the term "substantially" means that the described characteristics, parameters, or values need not be achieved exactly, and deviations or variations, including, for example, tolerances, measurement errors, measurement accuracy limits, and other factors known to those skilled in the art, may occur to an extent that does not interfere with the effects inherently obtained by the characteristics.
[0143] As used herein, "user" and "individual" each mean a human being.
[0144] In this specification, terms such as first, second, etc. may be used to describe various elements. These terms are used only to distinguish one element from another, unless otherwise specified or clear from the context. Therefore, these elements are not limited by these terms.
[0145] The present invention may be a system, method, or computer program product, or combination thereof, integrated at any possible level of technical detail. The computer program product may include a computer-readable storage medium having stored thereon computer-readable program instructions for causing a processor to carry out aspects of the present invention.
[0146] A computer-readable storage medium may be a tangible device capable of retaining and storing instructions for use by an instruction execution device. The computer-readable storage medium may be, by way of example, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or a suitable combination thereof. More specific examples of computer-readable storage media include portable computer diskettes, hard disks, RAM, ROM, EPROM (or flash memory), SRAM, CD-ROMs, DVDs, memory sticks, floppy disks, mechanically encoded devices having instructions recorded on punch cards or ridge-in-groove structures, or the like, and suitable combinations thereof. Computer-readable storage devices, as used herein, should not be construed as ephemeral signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., light pulses passing through a fiber optic cable), or electrical signals transmitted over wires.
[0147] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computer / processing device. Alternatively, they can be downloaded to an external computer or external storage device via a network (e.g., the Internet, a LAN, a WAN, or a wireless network, or a combination thereof). The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, or edge servers, or a combination thereof. A network adapter card or network interface within each computer / processing device receives the computer-readable program instructions from the network and transfers the computer-readable program instructions to a computer-readable storage medium in the respective computer / processing device for storage.
[0148] The computer-readable program instructions for carrying out the operations of the present invention can be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for integrated circuits, or source or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk and C++, and procedural programming languages such as the "C" programming language and similar programming languages. The computer-readable program instructions can execute entirely on the user's computer as a stand-alone software package, partially on the user's computer, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a LAN or WAN, or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider). In some embodiments, electronic circuitry, including, for example, programmable logic circuits, field programmable gate arrays (FPGAs), programmable logic arrays (PLAs), can execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to customize the electronic circuitry for carrying out aspects of the present invention.
[0149] Aspects of the present invention are described herein with reference to flowchart and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. Each block of the flowchart and / or block diagram, and combinations of blocks in the flowchart and / or block diagram, can be implemented by computer-readable program instructions.
[0150] The computer-readable program instructions may be provided to a processor of a computer or other programmable data processing apparatus to produce a machine, whereby the instructions, executed by the processor of such computer or other programmable data processing apparatus, form means for performing the functions / acts identified in one or more blocks of the flowcharts and / or block diagrams. The computer-readable program instructions may also be stored on a computer-readable storage medium capable of instructing a computer, programmable data processing apparatus, or other device, or combination thereof, to function in a particular manner, whereby the computer-readable storage medium having instructions stored thereon constitutes an article of manufacture including instructions for performing aspects of the functions / acts identified in one or more blocks of the flowcharts and / or block diagrams.
[0151] Computer-readable program instructions may also be loaded into a computer, other programmable device, or other device and a series of operational steps executed on the computer, other programmable device, or other device to create a computer-implemented process, whereby the instructions executing on the computer, other programmable device, or other device perform the functions / operations identified in one or more blocks in the flowcharts and / or block diagrams.
[0152] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions, including one or more executable instructions for performing a specific logical function. In some implementations, the functions depicted in the blocks may be performed in an order different from that depicted in the figures. For example, two blocks shown in succession may actually be accomplished as a single step, may be executed simultaneously or substantially simultaneously, may be executed in a partially or fully overlapping manner, or may even be executed in reverse order, depending on the functionality involved. Note that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs a specific function or operation, or by a combination of dedicated hardware and computer instructions.
[0153] While various embodiments of the present invention have been described by way of example, they are not intended to be exhaustive or limited to these embodiments. It will be apparent to those skilled in the art that many modifications and variations are possible without departing from the scope and spirit of the described embodiments. The terms used herein have been selected to best explain the principles, practical applications, or technical improvements to commercially recognized technologies of the embodiments, or to enable those skilled in the art to understand the embodiments disclosed herein.
Claims
1. 1. A computer-implemented method comprising: performing natural language processing (NLP) on machine-encoded files obtained over a data communications network to determine one or more file segment classifiers; locating previously consumed content in the machine-encoded file that corresponds to the one or more file segment classifiers by performing a computer database search against a user-specific consumption corpus database for NLP elements identified in electronic content previously presented to the user that match NLP elements extracted from the machine-encoded file; based on said detecting, modifying an electronic presentation structure for electronically presenting said machine-encoded file to a user so as to distinguish portions of said machine-encoded file corresponding to said previously consumed content from other portions of said machine-encoded file; 11. A computer-implemented method comprising:
2. modifying the electronic presentation structure in response to determining a contextual basis of the one or more file segment classifiers; The computer-implemented method of claim 1 .
3. electronically tracking the location of a device used to electronically receive and present the electronic presentation structure to the user; modifying the electronic presentation structure based on the location of the device. The computer-implemented method of claim 1 .
4. subsequently detecting one or more revisions made to the machine-encoded file since the user obtained the machine-encoded file; and in response to detecting the one or more revisions, presenting the machine-encoded file using a separate electronic presentation structure that indicates which portions are newly added or revised. The computer-implemented method of claim 1 .
5. the one or more file segment classifiers correspond to one or more topics of interest to the user discovered by performing unsupervised machine learning based topic modeling; the topic modeling discovers the topics of interest to the user based on a plurality of machine-encoded files acquired by the user via the data communications network; The computer-implemented method of claim 1 .
6. the one or more file segment classifiers correspond to one or more topics of interest to the user as identified by performing a topic classification based on supervised machine learning; the topic classification identifies topics specified by a user prior to retrieving the machine-encoded file; The computer-implemented method of claim 1 .
7. the previously submitted electronic content includes one or more machine-encoded files previously obtained over a data communications network within a user-specified time period; The computer-implemented method of claim 1 .
8. 1. A computer-implemented method comprising: analyzing electronic media content embedded in an electronic presentation structure presented to a user by comparing it to other electronic media content previously consumed by said user; detecting portions of the electronic media content that correspond to information previously consumed by the user based on the analysis; modifying the electronic presentation structure based on detecting one or more portions of the electronic media content corresponding to information previously consumed by the user to distinguish portions of the electronic media content corresponding to information previously consumed by the user from other portions of the electronic media content; 11. A computer-implemented method comprising:
9. Based on the electronic presentation structure including audio or video, modifying includes: Increasing the audio or video rendering speed of portions of the electronic media content that correspond to information previously consumed by the user; or shortening the audio or video to remove portions of the electronic media content corresponding to information previously consumed by the user; or doing both of the above, 9. The computer-implemented method of claim 8.
10. Based on the electronic presentation structure including text, modifying includes: annotating the text to distinguish portions of the text that correspond to information previously consumed by the user of the electronic media content from other portions of the text; modifying one or more formatting tags to change the formatting of the text corresponding to information previously consumed by the user of the electronic media content from other portions of the text; or extracting portions of the text corresponding to information of the electronic media content previously consumed by the user; 9. The computer-implemented method of claim 8.
11. The extracting step comprises: replacing the extracted portion with a link to a database in which the extracted portion is electronically stored; issuing instructions directing a user link to said database when said user desires to return said extracted portion to said electronic presentation structure; The computer-implemented method of claim 10.
12. 1. A computer-implemented method comprising: Detecting one or more revisions made to the machine-encoded file since a previous version of the machine-encoded file was obtained by a user over a data communications network, the detection comprising comparing natural language processing (NLP) elements identified in the electronic content of the previous version of the machine-encoded file with NLP elements identified in the electronic content of the revised version of the machine-encoded file; presenting the machine-encoded file using a separate electronic presentation structure that indicates which portions are newly added or revised based on detecting one or more revisions; 11. A computer-implemented method comprising:
13. 1. A system including a processor, The processor: performing natural language processing (NLP) on machine-encoded files obtained over a data communications network to determine one or more file segment classifiers; locating previously consumed content in the machine-encoded file that corresponds to the one or more file segment classifiers by performing a computer database search against a user-specific consumption corpus database for NLP elements identified in electronic content previously presented to the user that match NLP elements extracted from the machine-encoded file; based on said detecting, modifying an electronic presentation structure for electronically presenting said machine-encoded file to a user so as to distinguish portions of said machine-encoded file corresponding to said previously consumed content from other portions of said machine-encoded file; The system is configured to initiate an action that includes:
14. modifying the electronic presentation structure in response to determining a contextual basis of the one or more file segment classifiers; The system of claim 13.
15. The processor: electronically tracking the location of a device used to electronically receive and present the electronic presentation structure to the user; and modifying the electronic presentation structure based on the location of the device. The system of claim 13.
16. The processor: subsequently detecting one or more revisions made to the machine-encoded file since the user obtained the machine-encoded file; and in response to detecting the one or more revisions, initiating further operations including presenting the machine-encoded file using a separate electronic presentation structure that indicates which portions are newly added or revised. The system of claim 13.
17. the one or more file segment classifiers correspond to one or more topics of interest to the user discovered by performing unsupervised machine learning based topic modeling; the topic modeling discovers the topics of interest to the user based on a plurality of machine-encoded files acquired by the user via the data communications network; The system of claim 13.
18. the one or more file segment classifiers correspond to one or more topics of interest to the user as identified by performing a topic classification based on supervised machine learning; the topic classification identifies topics specified by a user prior to retrieving the machine-encoded file; The system of claim 13.
19. A computer program that causes a computer to execute the computer-implemented method described in any one of claims 1 to 12.
Citation Information
Patent Citations
Method and system for document management and storage medium having processing program stored thereon
JP2002024211A
Television broadcast receiving and output apparatus, and program
JP2008154126A
Electronic book device, object display method and program
JP2012194771A
Information processor and control program for the same
JP2013012246A
Digital content browsing support device, browsing support method, and program
JP2016066115A