Recognition and reconstruction of previously presented material

Through machine learning and natural language processing technology, it automatically identifies and distinguishes new and old electronic information, solves the problem of users' repeated consumption under information overload, and improves the efficiency of computing devices and storage utilization.

CN114625699BActive Publication Date: 2025-09-23INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111434346.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-12-11
Filing Date
2021-11-29
Publication Date
2025-09-23
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

When faced with a large amount of electronic information, users find it difficult to quickly distinguish new information from previously consumed information, resulting in repeated reading, listening or re-watching, wasting time and computing resources.

Method used

Through machine learning and natural language processing technologies, the electronic content consumed by users is analyzed, newly presented information is identified and distinguished from previously consumed information, and the electronic presentation structure is modified to highlight or skip repetitive parts, such as distinguishing new and old content through different colors, styles, speeds, etc.

Benefits of technology

It reduces users' need to process and store duplicate information, improves the processing efficiency and storage capacity of computing devices, and reduces the burden of users' repeated consumption of information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114625699B_ABST
    Figure CN114625699B_ABST
Patent Text Reader

Abstract

The information identification and reconstruction includes analyzing electronic media content embedded in an electronic presentation structure presented to a user and, based on the analysis, detecting portions of the electronic media content previously consumed by the user. The method includes modifying the electronic presentation structure based on the detection to distinguish the electronic media content previously consumed by the user from other portions of the electronic media content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to electronically processing and presenting information to a user, and more particularly, to distinguishing information newly presented to a user from previously presented information.

[0002] The advent of the Internet and other modern forms of electronic communication has greatly expanded the amount and modes of information available to individual consumers. Numerous news articles and a wide variety of other types of information are frequently published and widely distributed electronically across a variety of communication and computing platforms. This information can be transmitted electronically via various types of electronic communication networks, both wired and wireless, and consumed using a variety of electronic devices. Background Art

[0003] In one or more embodiments, a method includes performing natural language processing (NLP) on a machine-coded file to determine one or more file segmentation classifiers. The machine-coded file can be retrieved via a data communications network. The method includes detecting previously consumed content within the machine-coded file that corresponds to the one or more file segmentation classifiers by performing a computer database search to find NLP elements identified within electronic content previously presented to a user that match NLP elements extracted from the machine-coded file, wherein the search is performed against a user-specific consumption corpus database. The method includes modifying an electronic presentation structure for electronically presenting the machine-coded file to the user based on the detection. The electronic presentation structure is modified to distinguish portions of the machine-coded file that correspond to the previously consumed content from other portions of the machine-coded file.

[0004] In one or more embodiments, a method includes analyzing electronic media content embedded in an electronic presentation structure presented to a user. The method includes detecting, based on the analysis, a portion of the electronic media content previously consumed by the user. The method includes modifying the electronic presentation structure, based on the detection, to distinguish the electronic media content previously consumed by the user from other portions of the electronic media content.

[0005] In one or more embodiments, a method includes detecting one or more revisions to a machine-coded file subsequent to an earlier retrieval of an earlier version of the machine-coded file by a user via a data communications network. The detecting includes comparing natural language processing (NLP) elements identified within electronic content of the earlier version of the machine-coded file with NLP elements identified within electronic content of a revised version of the machine-coded file. The method includes, based on detecting the one or more revisions, presenting the machine-coded file using another electronic presentation structure that indicates which portions were newly added or revised.

[0006] In one or more embodiments, a system includes a processor configured to initiate operations. The operations include performing natural language processing (NLP) on a machine-coded file to determine one or more file segmentation classifiers. The machine-coded file can be retrieved via a data communication network. The operations include detecting previously consumed content corresponding to the one or more file segmentation classifiers within the machine-coded file by performing a computer database search to find NLP elements identified within electronic content previously presented to a user that match NLP elements extracted from the machine-coded file, wherein the search is performed on a user-specific consumption corpus database. The operations include modifying an electronic presentation structure for electronically presenting the machine-coded file to the user based on the detection. The electronic presentation structure is modified to distinguish a portion of the machine-coded file corresponding to the previously consumed content from other portions of the machine-coded file.

[0007] In one or more embodiments, a computer program product includes one or more computer-readable storage media having instructions stored thereon. The instructions are executable by a processor to initiate operations. The operations include performing natural language processing (NLP) on a machine-coded file to determine one or more file segmentation classifiers. The machine-coded file can be retrieved via a data communications network. The operations include detecting previously consumed content corresponding to the one or more file segmentation classifiers within the machine-coded file by performing a computer database search to find NLP elements identified within electronic content previously presented to a user that match NLP elements extracted from the machine-coded file, wherein the search is performed on a user-specific consumption corpus database. The operations include modifying an electronic presentation structure for electronically presenting the machine-coded file to the user based on the detection. The electronic presentation structure is modified to distinguish portions of the machine-coded file corresponding to the previously consumed content from other portions of the machine-coded file.

[0008] This summary is provided only to introduce certain concepts and is not intended to identify any key or essential features of the claimed subject matter.Other features of the invention will become apparent from the accompanying drawings and the detailed description that follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The arrangement of the present invention is illustrated by way of example in the accompanying drawings. However, the drawings should not be construed as limiting the arrangement of the present invention to the specific embodiments shown. Various aspects and advantages will become apparent upon reading the following detailed description and upon reference to the accompanying drawings:

[0010] Figure 1 depicts an example computing environment in which one or more information identification and presentation systems are used according to embodiments;

[0011] Figure 2 An exemplary information identification and presentation system according to an embodiment is shown;

[0012] Figure 3 Certain operational features of an exemplary information identification and presentation system according to an embodiment are shown;

[0013] Figure 4 Described are example electronic presentation structures and modified electronic presentation structures according to embodiments;

[0014] Figure 5 is a flow chart of an example method of electronically identifying and presenting information according to an embodiment;

[0015] Figure 6 is a flow chart of an example method of electronically identifying and presenting information according to an embodiment;

[0016] Figure 7 is a flow chart of an example method of electronically identifying and presenting information in the context of a user's specific location, according to an embodiment;

[0017] Figure 8 is a flow chart of an example method of detecting an update to a portion of information previously presented electronically to a user, according to an embodiment;

[0018] Figure 9 depicts a cloud computing environment according to an embodiment;

[0019] Figure 10 depicts an abstract model layer according to an embodiment;

[0020] Figure 11 depicts a cloud computing node according to an embodiment; and

[0021] Figure 12 An example portable device according to an embodiment is depicted. DETAILED DESCRIPTION

[0022] Although the present disclosure concludes with claims defining novel features, it is believed that the various features described within the present disclosure will be better understood by considering the description in conjunction with the accompanying drawings. The processes, machines, manufactures, and any variations thereof described herein are provided for illustrative purposes. The specific structural and functional details described in this disclosure are not to be construed as limiting, but merely as a basis for the claims and as a representative basis for teaching one skilled in the art to variously employ the features described in virtually any appropriate detailed structure. Furthermore, the terms and phrases used in this disclosure are not intended to be limiting, but rather to provide an understandable description of the features described.

[0023] The present disclosure relates to electronically processing and presenting information to users, and more particularly, to distinguishing between information newly presented to a user and information previously consumed by the user. The seemingly ever-expanding array of electronic communication technologies and information sources provides greater access to ever-larger amounts of information. Despite the benefits of rapidly accessing vast amounts of information, the volume of information can reach a point of diminishing returns for individual consumers of the information. For example, a developing current event may prompt multiple news articles. However, some of these many articles may simply be a repetition of the same facts. Consequently, an individual must laboriously read through considerable amounts of material simply to discover which portions of the information the individual has already consumed and which portions are new.

[0024] One aspect of the systems, methods, and computer program products disclosed herein is the automatic discovery of information a user has already consumed, thereby alleviating the need for the user to reread, relisten to, or rewatch material previously presented to the user electronically. Another aspect is the reconstruction of an electronic presentation structure that contrasts, compresses, or eliminates portions of information previously presented to the user as electronic content. As defined herein, "electronic content" is any information or data presented in a machine-encoded file that can be electronically stored and / or processed by a computer system. Information presented as electronic content can include news articles, educational materials, academic journals, human interest stories, entertainment material, and the like. The machine-encoded files used to present electronic content can be formatted as electronic documents, audio files, video files, or any combination thereof. As defined herein, an "electronic presentation structure" is any text, image, audio, or video embedded in a machine-encoded file that is electronically received or generated and presented to a user using an electronic device. The electronic device can be, for example, a desktop or laptop computer, a tablet computer, a smartphone, or a similar device capable of receiving text, images, audio, and / or video via a wired and / or wireless connection.

[0025] The systems, methods, and computer program products disclosed herein detect different portions of electronic content previously consumed by a user within a given volume of electronic content. As defined herein, a user "consumes" electronic content by retrieving the electronic content embedded in an electronic document, audio file, or video file via a data communications network, or, in the context of a system that enables a user to tag content as they read, listen to, or view it, by tagging an article to indicate that the user has read, listened to, or viewed the electronic content. As a user consumes electronic content from various sources in different formats (e.g., text, audio, video), the consumed electronic content is added to a database comprising a user-specific consumption corpus for comparison with subsequently presented information, thereby avoiding repeated presentation of already consumed electronic content.

[0026] Various machine learning and natural language processing techniques disclosed herein determine one or more file segment classifiers within a machine-encoded file in various arrangements. The one or more file segment classifiers classify different segments of newly presented electronic content corresponding to different topics identified by topic analysis. For each identified topic, the corresponding segment of the electronic content can be compared with previously consumed electronic content to determine whether any topic-specific segment of the newly presented electronic content is identical or sufficiently similar to previously consumed content belonging to the same topic.

[0027] Based on this detection, the electronic presentation structure for presenting the electronic content in the machine-coded file is modified. The electronic presentation structure is modified to distinguish the portion of the machine-coded file corresponding to the previously consumed portion of the electronic content from the other portions of the machine-coded file. This enables the user to avoid re-reading, re-listening to, or re-watching information previously presented to the user as electronic content. For example, for a textual presentation of electronic content, the electronic presentation structure can be annotated by highlighting or presenting the text in a different color or style to distinguish the newly presented content from the previously consumed content. In an audio presentation, for example, the audio portion related to the previously presented content can be quickly eliminated, skipped, or accelerated through. In a video-based presentation, for example, the portion of the video related to the previously presented content can also be skipped or accelerated. By reformatting and / or annotating, the electronic presentation structure can be reconstructed to distinguish the previously consumed electronic content from the newly presented electronic content. Optionally, the reconstruction can include annotating the electronic content to indicate the portion of the previously consumed electronic content.

[0028] The arrangements described herein relate to computer technology and provide improvements to existing computer technology. One such improvement involves mitigating the waste of computing resources. An ever-increasing amount of information is being disseminated as electronic content. The greater the volume of electronic content, the greater the expense of computing resources required to electronically store and process the information embedded therein. The arrangements described herein enhance computer system processing efficiency and reduce data storage requirements by reducing the processing and storage of electronic content that merely replicates information already consumed by users.

[0029] For example, for any computing device (e.g., desktop computer, laptop computer, tablet computer, smart watch), these arrangements improve computer technology by avoiding unnecessary use of computing resources when processing and presenting repeated electronic content. Reducing the need to process repeated electronic content enhances the processing efficiency of the computing device. By avoiding the storage of data of electronic content that has already been consumed, the available storage capacity of the computing device is correspondingly enhanced. Thus, by avoiding the retrieval, storage, and processing of information that the user has already consumed, the computer and device hardware performs more efficiently.

[0030] Other aspects of the embodiments described in this disclosure are described in more detail with reference to the following figures. For simplicity and clarity of illustration, the elements shown in the figures are not necessarily drawn to scale. For example, the dimensions of some elements may be exaggerated relative to other elements for clarity. In addition, where deemed appropriate, reference numerals have been repeated in the figures to indicate corresponding, similar, or analogous features.

[0031] Figure 1 A computing environment 100 is depicted. The computing environment 100 illustratively includes a computing node 102, one or more information nodes 104, and devices 106a, 106b, and 106c communicatively coupled via a data communications network 108. The computing node 102 and the information node(s) 104 may each include a computer system 1112 ( Figure 11 ) of the example computing node 1100. Devices 106a-106c may include a computer system (e.g., a desktop computer) having memory, one or more processors, and a processor such as reference computer system 1112 ( Figure 11 ) other hardware components such as those described in the accompanying drawings. One or more of the devices 106a-106c may be portable devices (laptops, tablets, smartphones), such as the example portable device 1200 ( Figure 12 ). Although three devices are shown, environments such as computing environment 100 typically include many more such devices. In operation, users of devices 106a-106c electronically access electronic content embedded in machine-encoded files, such as electronic documents, audio files, video files, and various other media hosted by information node 104 and accessed via a wired or wireless connection to data communication network 108. The various media convey various types of electronic content, including news articles, educational materials, human interest stories, entertainment materials, etc. Data communication network 108 typically includes the Internet, but may include wide area networks (WANs), local area networks (LANs), and various other data communication networks through which electronic documents, audio, video, and various other media can be accessed and retrieved.

[0032] The amount of information available from information nodes 104 to users of devices 106a-106c via data communications network 108 can be substantial. Given the amount of information that a device user can electronically retrieve as electronic content from information nodes 104 via data communications network 108, sometimes the information electronically presented to the device user is a repetition of all or part of previously presented information. For example, a tablet user may wish to catch up on news of current events by accessing several articles whose text is generated by different news organizations. The user may have been closely following developments related to the event for some time, and some of the articles may merely reiterate facts that the user has already read. In another context, for example, a computer user attending an online university may be assigned several instructional videos for viewing. Some or all of the instructional videos assigned to the user may include portions covering the same or substantially similar material.

[0033] In these and other situations, an information identification and presentation (IRP) system according to various embodiments described herein can relieve device users from needlessly expending time and effort sorting through various forms of electronic content (e.g., text, audio, video) to separate newly presented information from previously presented information.

[0034] The IRP system analyzes electronic media content embedded in an electronic presentation structure presented to a user. Based on the analysis, the IRP system detects portions of the electronic media content previously consumed by the user. Based on detecting the one or more portions previously consumed by the user, the IRP system modifies the electronic presentation structure to distinguish the electronic media content previously consumed by the user from other portions of the electronic media content.

[0035] The IRP system also improves the efficiency of the underlying computer system or device hardware by reducing the need to store, retrieve, and process electronic content that a user has already consumed. When a user electronically retrieves electronic content as a machine-encoded file, the IRP system identifies the different portions of the machine-encoded file that correspond to previously consumed content.

[0036] In certain embodiments disclosed herein, an IRP system implements natural language processing (NLP) and other machine learning aspects in identifying different sections or segments of a machine-encoded file, these segments being identified by file segment classifiers, each of which corresponds to a particular subject matter of information presented as electronic content. The IRP system detects previously consumed content within the machine-encoded file that corresponds to one or more file segment classifiers. The previously consumed content is detected by performing a computer database search for NLP elements identified in electronic content previously presented to the user that match NLP elements extracted from the machine-encoded file. The search is performed against a user-specific consumption corpus database, which additionally supports the machine learning implemented by the IRP system.

[0037] Based on the detection, the IRP system modifies an electronic presentation structure (e.g., an electronic document, an audio file, a video file) for electronically presenting the machine-encoded file to the user. The electronic presentation structure may be modified to distinguish portions of the machine-encoded file corresponding to previously consumed content from other portions of the machine-encoded file.

[0038] For an electronic presentation structure that includes text, the IRP system can modify the electronic presentation structure by annotating the text to distinguish the portion of the text corresponding to the portion of the electronic media content previously consumed by the user from the rest of the text. Additionally or alternatively, in certain embodiments, the IRP system can modify one or more formatting tags to change the format of the text corresponding to the portion of the electronic media content previously consumed by the user from the rest of the text. In other embodiments, the IRP system can additionally or alternatively extract the portion of the text corresponding to the portion of the electronic media content previously consumed by the user. Alternatively, the IRP system can replace the extracted portion with a link. The link can electronically link the user to a database in which the extracted portion is electronically stored. If the user wishes to restore the extracted portion to the electronic presentation structure, the IRP system can provide instructions instructing the user to link to the database. The IRP system can also optionally transmit to the user an indicator indicating when the user consumed the extracted portion or how long ago the extracted portion was consumed.

[0039] For electronic presentation structures that include audio or video files, the IRP system can modify the electronic presentation structure by accelerating the rendering speed of the audio or video for portions of the electronic media content previously consumed by the user. Additionally or alternatively, the IRP system in other embodiments can compress the audio or video to eliminate portions of the electronic media content previously consumed by the user. The IRP system can render the audio or video at a normal speed for portions determined not to have been previously consumed by the user.

[0040] Illustratively, within computing environment 100, IRP system 110 is implemented as a cloud-based system integrated into or operatively coupled to computing node 102 (which may be a cloud-based server). IRP system 110 may be implemented in hardware, software, or a combination of both. In performing the operations and functions described herein, IRP system 110 compares electronic presentation structures (e.g., electronic documents, audio, video) newly presented to a user with previous presentations 112 previously presented to the user and electronically stored in a database accessible to IRP system 110. Based on this comparison, IRP system 110 distinguishes portions of the information previously consumed by the user from new information electronically presented to the user. IRP system 110, implemented as a cloud-based system, is accessed by a user via data communications network 108 using device 106a or other electronic device communicatively coupled to data communications network 108.

[0041] In an alternative arrangement, IRP system 114 is a localized IRP system implemented on device 106b. IRP system 114 can perform the same or similar operations and functions as IRP system 110, but is implemented on device 106b rather than on a cloud-based server as IRP system 110. For example, IRP system 114 can comprise a computer application downloaded to device 106b (e.g., a desktop, laptop, tablet, smartphone). Optionally, while performing the operations and functions described herein, IRP system 114 can also locally store previous presentations (e.g., electronic documents, audio, video). IRP system 114 can alternatively access separately stored previous presentations 112 via data communications network 108. Device 106c can be another device used by the same user. IRP system 114 operating on device 106b can be operatively coupled to device 106c to perform the same operations and functions related to the electronic presentation structure presented to the user on device 106c.

[0042] Figure 2 Certain example components of an IRP system 200 according to an embodiment are depicted in greater detail. IRP system 200 is implemented in various software modules. Illustratively, the modules include an analyzer 202, a selector 204, a comparator 206, and a reconstructor 208. The modules comprising IRP system 200 electronically process machine-encoded files retrieved by a user. Machine-encoded files may include electronic content presented to the user as electronic documents, audio renderings, videos, or other electronic presentation structures. Operationally, IRP system 200 processes textual data embedded in machine-encoded files. The textual data may thus be contained in electronic documents (e.g., web pages, online news articles, messages), or converted from speech (e.g., pre-recorded audio, the audio portion of a video) by a speech-to-text processor (not shown).

[0043] In various embodiments, IRP system 200 integrates various aspects of machine learning. Analyzer 202 identifies topics within an article based on topic analysis of electronic content embedded in machine-encoded documents. Analyzer 202 implements natural language processing to parse electronic content into its constituent textual elements, assess the meaning of the text, and resolve inherent ambiguity. The natural language processing performed by analyzer 202 may include semantic analysis and / or word sense disambiguation based on tokenization, lemmatization, and / or part-of-speech tagging.

[0044] Analyzer 202 uses deep learning to combine natural language understanding. In certain embodiments described herein, deep learning is implemented in a multi-hidden layer neural network. Using the multi-hidden layer neural network, analyzer 202 performs text analysis to extract keywords, concepts, categories, relationships, sentiment, syntax, and other natural language elements from a given text. As defined herein, a "natural language processing (NLP) element" is any object that is detected in or derived from text (explicitly or implicitly) based on natural language processing and that at least partially explains the natural language meaning of a particular segment of text. As described below, NLP elements include keywords, concepts, relationships, etc., as well as semantic and syntactic objects. In certain embodiments, analyzer 202 performs text analysis to detect keyword frequencies and extract phrases that frequently appear in conjunction with each other from unstructured text. Text analysis is performed by analyzer 202 in conjunction with a machine learning model in analyzing machine-coded file 210, which is trained to identify one or more topics corresponding to different segments 212 of machine-coded file 210.

[0045] In some embodiments, analyzer 202 performs NLP to determine one or more file segment classifiers contained within machine coded file 210. Each file segment classifier corresponds to a different portion or segment of machine coded file 210 that pertains to a particular subject matter of information embedded as electronic content within machine coded file 210. The information contained in different segments 212 can cover multiple subjects.

[0046] In some embodiments, analyzer 202 implements topic modeling, a kind of unsupervised machine learning technique, to determine the theme corresponding to the file segmentation classifier.To implement topic modeling, analyzer 202 groups texts based on the words contained in each text and the observed correlations between them.In certain embodiments, analyzer 202 uses latent semantic analysis to implement topic modeling, and latent semantic analysis groups texts based on the frequency of occurrence of the same words in each text based on the distribution hypothesis.In other embodiments, analyzer 202 uses latent Dirichlet analysis to implement topic modeling, and described latent Dirichlet analysis is a generation statistical model that maps texts to themes and groups texts so that the words in each text are captured (or almost captured) by the themes.

[0047] Topic modeling based on unsupervised machine learning is well suited for discovering topics that a user is interested in, even though the user has not explicitly specified them. However, if a user specifies topics of interest, the analyzer 202 can implement supervised machine learning to train a classification model to classify different parts of the information based on predetermined topics. For example, when a user consumes (e.g., reads, listens to, watches) an electronic presentation (e.g., a document, audio, video) of information (electronic content), the user can tag or annotate the electronic presentation according to the topic identified by the user. The tagged electronic presentation can be added to a corpus of tagged training examples and stored electronically. In some embodiments, the analyzer 202 implements a deep learning neural network classification model that is trained using the electronic presentation tagged by the user to classify different parts of the electronic content according to the topic.

[0048] Thus, in some embodiments, the IRP system 200 provides users with the option of having the system learn through unstructured machine learning (e.g., topic modeling) to discover topics of interest to the user. Alternatively, the user has the option of explicitly specifying topics of interest by tagging the electronic presentation as the user consumes information conveyed in the presentation. By tagging the electronic presentation, the user builds a corpus of labeled training examples for training a classification model (e.g., a deep learning neural network) to identify topics of particular interest to the user. Furthermore, regardless of whether the user is particularly interested in any particular topic, the IRP system 200 in the various embodiments described herein detects and identifies previously consumed content, thereby enabling the user to skip or otherwise avoid content that has already been consumed.

[0049] Based on the topic determined by analyzer 202, selector 204 retrieves electronically stored topic-specific data. The structure of the data may include NLP elements generated based on natural language processing of electronic documents 214 previously retrieved by the user. The data structure includes a user-specific consumption corpus electronically stored in consumption corpus database 216—i.e., electronic content previously read, heard, or viewed—electronically stored in a relational or non-relational database. The data structure corresponds to text extracted from an electronic document or text converted from audio by a speech-to-text processor. The data structure may also include a timestamp indicating the date the corresponding information was presented to the user.

[0050] The selector 204 performs a computer database search of the consumption corpus database 216 for NLP elements identified within the electronic content in the previously rendered file 214 to identify NLP elements that match the NLP elements extracted from the machine-encoded file 210. Based on the match, the selector 204 may select a topic-specific segment of the previously rendered electronic content for comparison with the different topic-specific segment 212 of the machine-encoded file 210.

[0051] Alternatively, based on user-provided instructions, selector 204 may select from previously presented files 214 only those files that were previously presented to the user within a specific time interval. Thus, for a particular user, the electronic comparison of information is based on both subject matter and the time period during which the user consumed information. For example, a user may specify that newly presented information be compared with information consumed within the previous week, month, year, or any other time period, which may also reflect the user's interest in certain topics.

[0052] For each of the one or more topics determined by the analyzer 202, the comparator 206 compares the different segments 212 with corresponding topic-specific content contained in the article electronically stored in the consumer corpus database 216 and retrieved by the selector 204 based on the topic identified by the analyzer 202. The comparator 206 compares the NLP elements extracted by the analyzer 202 from the different segments 212 with NLP elements extracted from the previously retrieved file 214 electronically stored in the consumer corpus database 216.

[0053] NLP elements are extracted by analyzer 202, which performs text analysis on different segments 212 of the retrieved machine-encoded file 210 and previously consumed text whose subject matter corresponds to the subject matter of machine-encoded file 210. Based on natural language understanding, analyzer 202 identifies and extracts NLP elements such as names, places, and events. Analyzer 202 also extracts keywords. Analyzer 202 also extracts concepts. Concepts are high-level representations of themes and ideas expressed in the text. For example, if machine-encoded file 210 includes a research paper on deep learning, analyzer 202 may extract the concept "artificial intelligence," even if neither "artificial" nor "intelligence" is explicitly stated in the text. Analyzer 202 also extracts relationships between NLP elements. For example, if machine-encoded file 210 includes a news article about the Nobel Prize, which includes the NLP elements "Nobel Prize in Physics" and the name "Andrea M. Ghez," analyzer 202 may connect the NLP elements via the relationship "awarded."

[0054] Analyzer 202 also categorizes NLP elements. In some embodiments, analyzer 202 generates a classification hierarchy of text topics (e.g., finance / financial news, computers / operating systems / hardware). Analyzer 202 can also categorize text based on NLP elements of sentiment (e.g., positive, negative, neutral) and / or emotion (e.g., joy, sadness, anger, frustration).

[0055] Based on a topic-based comparison of the NLP elements of the machine-encoded file 210 and the NLP elements of the previously retrieved file 214, the comparator 206 identifies which of the different segments 212 of the machine-encoded file 210 include previously consumed information. The different segments of the machine-encoded file 210 may verbatim restate portions related to the same topic identified in the previously retrieved file, in which event the different segments of the machine-encoded file 210 are identified by the comparator 206 as previously consumed content. However, in other cases, although there may not be verbatim identity between the different segments of the machine-encoded file 210 and the previously retrieved file segments pertaining to the same topic, the similarity between the segments may be sufficient for the comparator 206 to identify the different segments of the machine-encoded file 210 as previously consumed content.

[0056] Similarity can be determined based on a similarity analysis performed by comparator 206, which compares corresponding NLP elements extracted from the respective files. For example, the percentage of NLP elements shared by two files can be determined by comparator 206. Based on this percentage, comparator 206 can determine the probability that the respective segments are duplicates. The similarity sufficient for comparator 206 to identify different segments of machine-encoded file 210 as previously consumed content can be based on a probability exceeding a certain level predetermined by the user (e.g., 80%).

[0057] Thus, it is noted that, based on topic analysis and natural language understanding, the IRP system 200 is able to identify previously consumed content that appears in a newly presented machine-encoded document, even though a side-by-side comparison of topic-specific snippets may not reveal that the previously retrieved content covers substantially the same material, albeit presented in a different manner. For example, a newly retrieved financial news article may include text that is written differently than a previously retrieved news article. However, a comparison of NLP elements such as keywords or terms (e.g., names, dates, numerical values) and / or concepts (e.g., monetary policy) and / or relationships (e.g., monetary policy and interest rates) may reveal that a high percentage (e.g., 90%) of the same NLP elements appear in both articles, making the later article essentially a duplicate of the earlier retrieved article.

[0058] In certain embodiments, the IRP system 200 generates a similarity metric, such as the percentage of identical or common NLP elements, to determine whether one or more of the different segments 212 include information previously consumed by the user. The IRP system 200 may also include a bias in the similarity metric based on, for example, sentiment and / or emotion. A similarity metric bias is a negative factor that reduces the similarity metric between two or more articles that are similar in other ways except for differences in sentiment and / or emotion. The underlying concept is that, despite recounting common facts, such articles can offer contrasting perspectives on the same set of facts. For example, two articles containing editorial opinions on the same subject may contain similarities with respect to NLP elements. However, if each takes a different perspective, the sentiment and / or emotion of each will likely differ. Therefore, the IRP system 200 discounts the similarity and, based on the difference in sentiment and / or emotion, infers that the material in the later retrieved article was not previously consumed. Sentiment and emotion, along with other NLP elements, can be electronically stored in a consumption corpus.

[0059] The operations and functions collaboratively performed by analyzer 202, selector 204, and comparator 206 can be performed on text derived directly from any medium (e.g., electronic documents, audio, video) or derived through speech-to-text processing. Similarly, a user-specific consumption corpus merges data structures corresponding to text derived from different media. Thus, information newly presented in one medium can be compared with information previously presented in a completely different medium. For example, information newly presented in an electronic presentation structure including an electronic document can be compared with information previously presented in another electronic presentation structure including audio or video and other electronic documents. The IRP system 200 thus enables mixed-media content comparison of information electronically presented to a user. For example, a user who begins listening to a novel presented as an audio book and stops at a certain point can resume the storyline by reading the electronic text on a tablet computer, where the electronic content is annotated to indicate the point at which the user stopped listening to the audio book. Thus, the IRP system 200 can indicate the portion of the novel being listened to in the audio book and enable the user to continue reading the newly presented portion of the novel from the tablet computer.

[0060] In other embodiments, IRP system 200 additionally or alternatively identifies portions of newly presented information that differ from previously presented information based on metadata corresponding to an electronic document, audio, video, or other electronic presentation structure presented to the user. For example, metadata corresponding to an electronic document may indicate the document's version or publication date. If the electronic document is subsequently revised, for example by updating old information (e.g., appropriate names, dates, values) or adding new information, a comparison of the metadata corresponding to the earlier and later published versions may prompt comparator 204 to determine changes between the two versions. In still other embodiments, IRP system 200 additionally or alternatively captures URL hyperlinks to online publications consumed by users, storing the content in a relational or non-relational database for subsequent comparison of successive changes to the information contained therein with each content update.

[0061] Reconstructor 208 modifies the electronic presentation structure based on comparator 206 detecting different portions of the information being presented that correspond to previously presented information. For information presented as text, reconstructor 208 can annotate the text of the electronic document to contrast those portions of the previously presented information with the portions newly presented to the user. In some embodiments, reconstructor 208 highlights the text using a different color, font, and / or stylistic variation (e.g., bold, italics, strikethrough), and / or background that contrasts the newly presented portion with the previously presented portion. In some embodiments, reconstructor 208 compresses the text, eliminating portions identified as having been previously presented to the user.

[0062] In some embodiments, the structure of the electronic presentation, including the audio or video, is modified by the reconstructor 208, which compresses the structure to eliminate different portions corresponding to previously presented information. In other embodiments, the reconstructor 208 changes the presentation speed of the audio or video portion containing the previously presented information. The reconstructor 208 can change the speed of the presentation so that the audio or video is accelerated or skips portions that were previously heard or seen and do not need to be heard or seen again by the user. The reconstructor 208 causes the audio or video to render the portion determined to be newly presented at normal speed.

[0063] For example, a user interested in a particular topic may engage in reading electronically retrieved electronic documents, watching videos, and listening to audio podcasts. As the user's understanding of the topic grows, the user desires to consume new material rather than rereading, revisiting, or listening to previously consumed content. As the user consumes more content on the same topic, the IRP system 200 continuously updates the consumption corpus database 216. When the user listens to a podcast on the topic, the reconstructor 208 advances the audio (e.g., by 7 minutes) to begin at the point where the new material is introduced. When watching a video about the subject, the reconstructor 208 can speed up the video presentation (e.g., to 2x normal speed) through the portion corresponding to the previously consumed content, and then slow the video presentation back to normal speed at the point where the newly presented content begins. When the user retrieves an electronic document on the topic, the reconstructor 208 annotates information that is duplicated with previously consumed information (e.g., from other electronic texts, previously consumed audio, or previously consumed video) and highlights, annotates, or otherwise distinguishes the newly presented information from the previously consumed information.

[0064] In certain embodiments, the IRP system 200 identifies metadata associated with a document, audio, or video, where the metadata indicates that the information presented therein is related to a specific location (e.g., GPS coordinates) of the device rendering the electronic presentation of the information. For example, a user may move between multiple sites on a factory floor, accessing electronic documents captured wirelessly using a tablet or other device capable of receiving and processing wireless signals. The user may need to review a set of usage and safety documents at each site, regardless of whether the information has been previously accessed at another site. The goal is to ensure that the presented information is reviewed each time the user visits a particular site. Based on the metadata identified by the IRP system 200, the IRP system 200 at each site may highlight specific portions of the corresponding electronic document to indicate the specific portion containing the information required for reading. Device location data (e.g., GPS coordinates) may be included in the content request sent by the device when requesting content. For example, a plug-in may include such data. The IRP system 200 may use the device location data to compare with any other data associated with the requested content, where the requested content is location-dependent, to modify the electronic presentation structure.

[0065] In other embodiments, the IRP system 200 tracks the context in which information is presented in various media (e.g., documents, audio, video) so that previously presented information is treated differently by the system based on the context in which the information is presented. As described above, given the contextual basis of an article, the user is able to limit modifications to the electronic presentation structure, which contextual basis is determined based on keywords, concepts, relationships, or other NLP elements extracted by the IRP system 200. As defined herein, a "contextual basis" is a topic, subtopic, or concept determined by natural language processing that characterizes an article based in whole or in part on an overall topic. For example, a user may want to track information consumed in articles about the economic impact of COVID-19, but does not want to read repeated facts about the economic impact of COVID-19 in a new article, however, if presented in the context of a new topic, such as an article about major league baseball, the user may still want to see the same facts repeated. If the contextual basis of the article is the economic impact, the IPR system 200 modifies the electronic presentation structure of the article, but if the contextual basis is major league baseball, the IPR system 200 does not modify the electronic presentation structure of the article. Thus, a user may specify in machine-coded instructions that the reconstructor 208 reconstruct electronic presentation structures related to a topic if presented in one context (eg, economic impact) but not in another context (eg, professional sports).

[0066] In yet other embodiments, a user can specify, using machine-coded instructions, that the IRP system 200 electronically tag information with a purpose and / or date indicator to indicate when the user consumed the information and why. For example, a user participating in online learning can specify that the material consumed on November 3rd was for the purpose of a specific task. When the same information is newly presented in a subsequently generated electronic presentation structure, the reconstructor 208 can annotate the text or otherwise inform the user when and / or why the same information was previously consumed. The reconstructor 208 for information transmitted in audio or video can also generate a record that indicates to the user which audio portion has been heard or which video portion has been viewed. Alternatively, in the context of online learning, for example, the IRP system 200 can generate credits assigned to users because each new assignment is consumed if the corresponding material has not been previously consumed by the user.

[0067] In yet other embodiments, IRP system 200 may track the amount of information contained in previously presented electronic presentation structures. IRP system 200 may generate a metric indicating the percentage of information contained in previously consumed articles. For example, an electronic presentation structure containing text may include a note indicating the percentage of text contained in the electronic presentation structure that corresponds to previously presented information. For example, in response to a user retrieving an online article via a data communications network, IRP system 200 may transmit an electronic message to the user specifying a corresponding "relationship score" indicating the portion of the article that was previously consumed. IRP system 200 may inquire, given the relationship score, whether the user wishes to view the online article or to deny the request to access the article.

[0068] Similarly, an electronic presentation structure including audio or video can include a pre-recorded indicator that informs the user of the percentage of the audio or video that the user has previously consumed. Thus, the user can decide whether the article, audio, or video contains enough new presentation information to be worthwhile to spend time reading, listening to, or viewing the electronic presentation structure.

[0069] Figure 3 Certain operational aspects of an IRP system 300 are depicted, which, according to one embodiment, is implemented in a software module or module suite that functions as an add-on to a web browser. At block 302, a user installs an IRP system web plug-in. In various embodiments, the plug-in can be used as a web browser for a cloud-based IRP system (e.g., IRP system 110). Figure 1 )) and / or IPR system (e.g., IPR system 114 ( Figure 1 )). At block 304, the IRP system captures a URL hyperlink to electronic content 306 in response to a user accessing and retrieving a machine-coded file containing electronic content 306 via a data communications network. The machine-coded file may be formatted as an electronic document, an audio file, or a video file. The IRP system captures electronic content 306 at block 308 and electronically stores a data structure 312 of the electronic content, including the URL hyperlink and the machine-coded file, along with the URL hyperlink, in a database comprising a consumption corpus 314 at block 310. At block 316, the IRP system performs one or more of the various types of NLP and related analyses described above to determine whether any portion of electronic content 306 has been previously consumed by the user.

[0070] This determination can be based on matching NLP elements extracted from the electronic content 306 with NLP elements identified in the electronic content stored in the consumption corpus 314. If the IRP system is cloud-based, the IRP system can act as a proxy, and the request for the electronic content 306 can be forwarded to the proxy via a plug-in, which performs NLP and related analysis. Optionally, at box 318, the IRP system generates a relationship score and attaches the relationship score to the URL hyperlink, the relationship score indicating the proportion of the electronic content 306 that is determined to have been previously consumed by the user. The IRP system can send a message indicating the relationship score and ask the user whether he or she wishes to load the electronic content 306. In response to determining that some portion of the electronic content 306 has been previously consumed by the user, at box 320, the electronic presentation structure for presenting the electronic content 306 to the user is reconstructed.

[0071] Additional references Figure 4 , describes an example reconstruction of an electronic presentation structure including a text document 400. Based on the aforementioned operational aspects 300 performed by the IRP system, portions within text document 400 are identified as corresponding to information previously consumed by a user. Text document 400 is reconstructed into an electronic presentation structure 402, which compares the previously consumed portions with portions of electronic content newly presented to the user. Electronic presentation structure 402, i.e., the modification or reconstruction of text document 400, includes different portions 404, 406, 408, and 410. For example, each portion 404-410 is highlighted using a different shading to indicate the previously consumed portion. The highlighting colors can correspond to different topics within the text document, and the differences in highlight shading can indicate the length of time since the corresponding portion (by topic) was consumed. For example, text document 400 can be a news article about a new virus. In response to a user submitting an electronic request for text document 400 via a data communications network, the IRP system determines the previously consumed portion. Different sections 404, 406, 408, and 410 may be highlighted with different shading to indicate the previously consumed sections and the subject matter to which each belongs.

[0072] Figure 5 is a flow chart of a method 500 for electronically identifying and presenting information according to an embodiment. The method 500 may be performed by referring to Figure 1 and 2 At block 502, the system analyzes electronic media content embedded in an electronic presentation structure for presentation to a user. The electronic media content may be presented as an electronic document, an audio recording, a video recording, a multimedia presentation, a CD-ROM, as online content, etc., and is typically presented in a digital format, but may include analog data in some embodiments.

[0073] Based on the analysis, the system detects portions of the electronic media content previously consumed by the user at block 504. Based on the detection, the system modifies the electronic presentation structure at block 506. The electronic presentation structure is modified to distinguish the electronic media content previously consumed by the user from other portions of the electronic media content.

[0074] If the electronic presentation structure includes an audio or video file, the system can modify the electronic presentation structure by speeding up the speed at which the audio or video is presented. Accelerating the speed for portions of the electronic media content previously consumed by the user allows the user to quickly move through the electronic media content without having to reread or re-view the previously consumed portions of the electronic media content. Alternatively or additionally, in other embodiments, the system can compress the audio or video to eliminate portions of the electronic media content previously consumed by the user.

[0075] If the electronic presentation structure includes text, the system can modify the electronic presentation structure by annotating the text to distinguish portions of the text corresponding to portions of the electronic media content previously consumed by the user from portions of the text not previously consumed by the user. In other embodiments, the system can alternatively or additionally modify the electronic presentation structure by reformatting one or more formatting tags to distinguish portions of the text corresponding to portions of the electronic media content previously consumed by the user from portions of the text not previously consumed by the user.

[0076] In other embodiments, the system can alternatively or additionally modify the electronic presentation structure by extracting a text portion corresponding to a portion of the electronic media content previously consumed by the user. Alternatively, the system can replace the extracted portion with a link. The link can connect the user electronically (e.g., via a data communications network) to a database in which the extracted portion is electronically stored. If the user wishes to restore the extracted portion to the electronic presentation structure, the system can transmit the link to the user together with an instruction that instructs the user to link to the database. A notification provided to the user by the system can indicate that the extracted portion was previously consumed by the user and ask the user whether or not to consume the extracted portion again, in which case the user can do so using a link provided by the system. The notification can indicate when the user consumed the extracted portion or how long ago the extracted portion was consumed.

[0077] Figure 6 is a flow chart of a method 600 for electronically identifying and presenting information according to an embodiment. The method 600 may be performed by referring to Figure 1 and 2 At block 602, the system performs natural language processing (NLP) on the machine-encoded document to determine one or more document segment classifiers. The machine-encoded document may be retrieved by a user via a data communication network.

[0078] At block 604, the system detects previously consumed content in the machine-encoded file that corresponds to one or more document segmentation classifiers by performing a computer database search for NLP elements identified in electronic content previously presented to the user that match NLP elements extracted from the machine-encoded file. The search is performed against a user-specific consumption corpus database.

[0079] Based on the detection, the system modifies the electronic presentation structure to electronically present the machine-encoded file to the user at block 606. The electronic presentation structure is modified to distinguish portions of the machine-encoded file corresponding to previously consumed content from other portions of the machine-encoded file.

[0080] In some embodiments, the system modifies the electronic presentation structure in response to determining a contextual basis for one or more file segmentation classifiers. The one or more file segmentation classifiers may correspond to one or more topics of interest to the user. In some embodiments, the system automatically discovers topics of interest by performing unsupervised machine learning topic modeling. Using topic modeling, the system discovers topics of interest to the user based on a plurality of machine-encoded files previously retrieved by the user via a data communication network.

[0081] In other embodiments, the one or more document segment classifiers correspond to one or more topics of interest to the user identified by the system by performing supervised machine learning topic classification. The system identifies topics based on classification performed by a topic classifier trained using machine learning using a consumed corpus comprising electronic content selected by the user.

[0082] In other embodiments, the previously presented electronic content includes one or more machine-encoded files previously retrieved via a data communications network within a user-specified time interval.

[0083] In other embodiments, the system modifies the electronic presentation structure by annotating the electronic presentation structure with a relationship score.The system generates the relationship score by determining a proportion of the machine-encoded file that includes previously consumed content.

[0084] Figure 7 is a flow chart of an example method 700 for electronically identifying and presenting information in the context of a user's specific location, according to an embodiment. The method 700 may be performed by referring to Figure 1 and 2 In block 702, the system electronically tracks the location of a device used to receive and present an electronic presentation structure to a user. In block 704, the system modifies the electronic presentation structure based on the location of the device.

[0085] Figure 8is a flow chart of an example method 800 for detecting an update to a portion of information previously presented to a user, according to an embodiment. Figure 1 and 2 The system may be performed by the same or similar system described herein. In box 802, the system detects one or more modifications to the machine coded file subsequent to an earlier search for the machine coded file by the user. In box 804, the system responds to detecting one or more revisions by presenting the machine coded file using another electronic presentation structure that indicates which portions were newly added or modified. In some embodiments, the system may automatically timestamp links to sites from which users access machine coded files and record versions. If a new version of the machine coded file is generated, the system may detect the revision by performing the subject analysis and natural language processing described herein to discover whether the newly generated version associated with the link contains portions that the user has not previously consumed. Based on detecting the previously unconsumed portions, the system may communicate with the user. The system may indicate to the user which portions of the revised version are newly presented, thereby enabling the user to access the newly presented information as the user wishes without having to reconsume portions that have already been consumed.

[0086] It is expressly noted that although the present disclosure includes detailed descriptions about cloud computing, implementation of the teachings recited herein is not limited to a cloud computing environment. Rather, embodiments of the present invention can be implemented in conjunction with any other type of computing environment now known or later developed.

[0087] Cloud computing is a service delivery model for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be quickly provisioned and released with minimal management effort or interaction with the service provider. The cloud model can include at least five characteristics, at least three service models, and at least four deployment models.

[0088] Features are as follows:

[0089] On-demand self-service: Cloud consumers can automatically and unilaterally provision computing capabilities, such as server time and network storage, as needed without manual interaction with the service provider.

[0090] Broad Network Access: Capabilities are available over the network and accessed through standard mechanisms that facilitate use by heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0091] Resource pooling: Providers' computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated based on demand. This is location-agnostic in the sense that consumers typically do not control or know the exact location of the provided resources, but are able to specify the location at a higher level of abstraction (e.g., country, state, or data center).

[0092] Rapid Elasticity: Capacity that can be quickly and elastically (in some cases automatically) provisioned to quickly scale out and quickly released to quickly scale in. To the consumer, the capacity available for provisioning generally appears unlimited and can be purchased in any quantity at any time.

[0093] Metered Services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both providers and consumers of the utilized services.

[0094] The service model is as follows:

[0095] Software as a Service (SaaS): The ability provided to consumers is to use the provider's applications running on a cloud infrastructure. Applications are accessed from a variety of client devices through a thin-client interface such as a web browser (e.g., web-based email). Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0096] Platform as a Service (PaaS): The capability provided to consumers is to deploy consumer-created or acquired applications onto cloud infrastructure. These applications are built using programming languages ​​and tools supported by the provider. Consumers do not manage or control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but do have control over the deployed applications and possibly the configuration of the application hosting environment.

[0097] Infrastructure as a Service (IaaS): The capabilities provided to consumers are processing, storage, networking, and other basic computing resources on which consumers can deploy and run arbitrary software, including operating systems and applications. Consumers do not manage or control the underlying cloud infrastructure, but do have control over the operating system, storage, deployed applications, and possibly limited control over selected networking components (e.g., host firewalls).

[0098] The deployment model is as follows:

[0099] Private cloud: Cloud infrastructure is run for the sole use of an organization. It can be managed by that organization or a third party and can exist inside or outside a building.

[0100] Community cloud: Cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., mission, security requirements, policies, and compliance considerations). It can be managed by the organization or a third party and can exist on-premises or off-premises.

[0101] Public cloud: Cloud infrastructure is available to the general public or large industry groups and is owned by the organization that sells cloud services.

[0102] Hybrid cloud: A cloud infrastructure is a combination of two or more clouds (private, community, or public) that remain unique entities but are bound together by standardized or proprietary technologies that enable data and application portability (e.g., cloud bursting for load balancing between clouds).

[0103] The cloud computing environment is service-oriented, with a focus on statelessness, low coupling, modularity, and semantic interoperability. At the core of cloud computing is the infrastructure consisting of a network of interconnected nodes.

[0104] Now refer to Figure 9 , depicts an illustrative cloud computing environment 900. As shown, the cloud computing environment 900 includes one or more cloud computing nodes 910, and local computing devices used by cloud consumers can communicate with one or more cloud computing nodes 910, and the local computing devices, such as personal digital assistants (PDAs) or cellular phones 940a, desktop computers 940b, laptop computers 940c and / or automobile computer systems 940n. The computing nodes 910 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above. This allows the cloud computing environment 900 to provide infrastructure, platforms, and / or software as services for which cloud consumers do not need to maintain resources on local computing devices. It should be understood that Figure 9 The types of computing devices 940a-n shown in are intended for illustration only, and computing node 910 and cloud computing environment 900 can communicate with any type of computerized device over any type of network and / or network-addressable connection (eg, using a web browser).

[0105] Now refer to Figure 10 , showing a cloud computing environment 900 ( Figure 9 ) provides a set of functional abstraction layers. It should be understood in advance that Figure 10 The components, layers, and functions shown in are intended to be illustrative only, and embodiments of the present invention are not limited thereto. As depicted, the following layers and corresponding functions are provided:

[0106] The hardware and software layer 1060 includes hardware and software components. Examples of hardware components include mainframes 1061; servers based on RISC (Reduced Instruction Set Computer) architecture 1062; servers 1063; blade servers 1064; storage devices 1065; and network and networking components 1066. In some embodiments, software components include web application server software 1067 and database software 1068.

[0107] The virtualization layer 1070 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 1071 ; virtual storage 1072 ; virtual networks 1073 , including virtual private networks; virtual applications and operating systems 1074 ; and virtual clients 1075 .

[0108] In one example, the management layer 1080 may provide the functionality described below. Resource provisioning 1081 provides dynamic procurement of computing resources and other resources used to perform tasks within a cloud computing environment. Metering and pricing 1082 provides cost tracking when utilizing resources in a cloud computing environment, as well as accounting or invoicing for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. User portal 1083 provides access to the cloud computing environment for consumers and system administrators. Service level management 1084 provides cloud computing resource allocation and management so that required service levels are met. Service level agreement (SLA) planning and fulfillment 1085 provides pre-scheduling and procurement of cloud computing resources, where future demand is anticipated based on the SLA.

[0109] The workload layer 1090 provides examples of functionality that can utilize a cloud computing environment. Examples of workloads and functionality that can be provided from this layer include: mapping and navigation 1091; software development and lifecycle management 1092; virtual classroom education delivery 1093; data analytics processing 1094; transaction processing 1095; and IPR systems 1096.

[0110] Figure 11 A schematic diagram of an example of a computing node 1100 is shown. In one or more embodiments, computing node 1100 is an example of a suitable cloud computing node. Computing node 1100 is not intended to limit the scope of use or functionality of the embodiments of the present invention described herein. Computing node 1100 is capable of performing any of the functions described in this disclosure.

[0111] Computing node 1100 includes a computer system 1112 that can operate with numerous other general-purpose or special-purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations suitable for use with computer system 1112 include, but are not limited to, personal computer systems, server computer systems, thin clients, fat clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems or devices.

[0112] Computer system 1112 can be described in the general context of computer system-executable instructions, such as program modules, executed by the computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. Computer system 1112 can be practiced in a distributed cloud computing environment, where tasks are performed by remote processing devices linked through a communications network. In a distributed cloud computing environment, program modules can be located in both local and remote computer system storage media, including memory storage devices.

[0113] like Figure 11 As shown, computer system 1112 is shown in the form of a general-purpose computing device. Components of computer system 1112 may include, but are not limited to, one or more processors 1116, memory 1128, and a bus 1118 that couples various system components, including memory 1128, to processor 1116. As defined herein, a "processor" means at least one hardware circuit configured to execute instructions. The hardware circuit may be an integrated circuit. Examples of processors include, but are not limited to, a central processing unit (CPU), an array processor, a vector processor, a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic array (PLA), an application-specific integrated circuit (ASIC), a programmable logic circuit, and a controller.

[0114] Execution of instructions of a computer program by a processor includes executing or running the program. As defined herein, "run" and "execute" include a series of actions or events performed by a processor in accordance with one or more machine-readable instructions. "Run" and "execute," as defined herein, refer to the active performance of actions or events by a processor. The terms run, execute, and execute are used synonymously herein.

[0115] Bus 1118 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example only and not limitation, these architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, a Peripheral Component Interconnect (PCI) bus, and a PCI Express (PCIe) bus.

[0116] Computer system 1112 typically includes a variety of computer system readable media. Such media can be any available media that can be accessed by computer system 1112 and can include volatile and nonvolatile media, removable and non-removable media.

[0117] Memory 1128 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 1130 and / or cache memory 1132. Computer system 1112 may also include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example, storage system 1134 may be provided for reading from and writing to non-removable, non-volatile magnetic media and / or solid-state drives (not shown and typically referred to as "hard drives"). Although not shown, magnetic disk drives for reading from and writing to removable, non-volatile magnetic disks (e.g., "floppy disks") may be provided, as well as optical disk drives for reading from and writing to removable, non-volatile optical disks such as CD-ROMs, DVD-ROMs, or other optical media. In such cases, each may be connected to bus 1118 via one or more data media interfaces. As will be further depicted and described below, memory 1128 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of embodiments of the present invention.

[0118] By way of example and not limitation, a program / utility 1140 having a set (at least one) of program modules 1142, as well as an operating system, one or more application programs, other program modules, and program data, may be stored in memory 1128. Each of the operating system, one or more application programs, other program modules, and program data, or some combination thereof, may include an implementation of a networked environment. Program modules 1142 generally perform the functions and / or methods of embodiments of the present invention as described herein. For example, one or more program modules may include an IPR system 996 or portions thereof.

[0119] Program / utility 1140 is executable by processor 1116. Program / utility 1140 and any data items used, generated, and / or manipulated by computer system 1112 are functional data structures that impart functionality when used by computer system 1112. As defined in this disclosure, a "data structure" is a physical implementation of the data organization of a data model within physical memory. As such, a data structure is formed by specific electrical or magnetic structural elements in memory. A data structure imposes a physical organization on data stored in memory, as used by an application program executed using a processor.

[0120] Computer system 1112 can also communicate with one or more external devices 1114, such as a keyboard, pointing device, display 1124, and the like; one or more devices that enable a user to interact with computer system 1112; and / or any device that enables computer system 1112 to communicate with one or more other computing devices (e.g., a network card, modem, etc.). Such communication can occur via input / output (I / O) interface 1122. Computer system 1112 can also communicate with one or more networks, such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet), via network adapter 1120. As depicted, network adapter 1120 communicates with the other components of computer system 1112 via bus 1118. It should be understood that, although not shown, other hardware and / or software components may be used in conjunction with computer system 1112. Examples include, but are not limited to, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems.

[0121] Although computing node 1100 is used to illustrate an example of a cloud computing node, it should be understood that the use of computing node 1100 in conjunction with Figure 11 Computer systems with the same or similar architecture as described herein can be used in non-cloud computing implementations to perform the various operations described herein. In this regard, the example embodiments described herein are not intended to be limited to cloud computing environments. Compute node 1100 is an example of a data processing system. As defined herein, a "data processing system" refers to one or more hardware systems configured to process data, each hardware system including at least one processor and memory programmed to initiate operations.

[0122] Computing node 1100 is an example of computer hardware. Computing node 1100 may include Figure 11The illustrated components may include fewer components than shown or include additional components not shown, depending on the specific type of device and / or system implemented. The specific operating system and / or applications included may vary depending on the device and / or system type, and the type of I / O devices included may also vary. In addition, one or more illustrative components may be incorporated into another component or otherwise form a part of another component. For example, a processor may include at least some memory.

[0123] Computing node 1100 is also an example of a server. As defined herein, a "server" represents a data processing system configured to share services with one or more other data processing systems. As defined herein, a "client device" represents a data processing system that requests shared services from a server, and a user interacts directly with it. Examples of client devices include, but are not limited to, workstations, desktop computers, computer terminals, mobile computers, laptop computers, netbook computers, tablet computers, smart phones, personal digital assistants, smart watches, smart glasses, gaming devices, set-top boxes, smart TVs, etc. In one or more embodiments, the various user devices described herein can be client devices. Network infrastructure such as routers, firewalls, switches, access points, etc. is not a client device because the term "client device" is defined here.

[0124] Figure 12 An example portable device 1200 is shown in accordance with one or more embodiments described in this disclosure. The portable device 1200 may include a memory 1202, one or more processors 1204 (eg, an image processor, a digital signal processor, a data processor), and an interface circuit 1206.

[0125] In one aspect, the memory 1202, the processor(s) 1204, and / or the interface circuitry 1206 are implemented as separate components. In another aspect, the memory 1202, the processor(s) 1204, and / or the interface circuitry 1206 are integrated into one or more integrated circuits. The various components of the portable device 1200 can be coupled, for example, via one or more communication buses or signal lines (e.g., interconnects and / or wires). In one aspect, the memory 1202 can be coupled to the interface circuitry 1206 via a memory interface (not shown).

[0126] Sensors, devices, subsystems, and / or input / output (I / O) devices can be coupled to the interface circuit 1206 to facilitate the functions and / or operations described herein, including the generation of sensor data. Various sensors, devices, subsystems, and / or I / O devices can be coupled to the interface circuit 1206 directly or through one or more intermediate I / O controllers (not shown).

[0127] For example, a position sensor 1210, a light sensor 1212, and a proximity sensor 1214 can be coupled to the interface circuit 1206 to facilitate the orientation, lighting, and proximity functions of the portable device 1200, respectively. A position sensor 1210 (e.g., a GPS receiver and / or processor) can be connected to the interface circuit 1206 to provide geolocation sensor data. An electronic magnetometer 1218 (e.g., an integrated circuit chip) can be connected to the interface circuit 1206 to provide sensor data that can be used to determine the direction of magnetic north for directional navigation. An accelerometer 1220 can be connected to the interface circuit 1206 to provide sensor data that can be used to determine the speed and direction of movement of the device in three dimensions. An altimeter 1222 (e.g., an integrated circuit) can be connected to the interface circuit 1206 to provide sensor data that can be used to determine altitude. A voice recorder 1224 can be connected to the interface circuit 1206 to store recorded speech.

[0128] The camera subsystem 1226 can be coupled to an optical sensor 1228. The optical sensor 1228 can be implemented using any of a variety of technologies. Examples of the optical sensor 1228 include a charge coupled device (CCD), a complementary metal oxide semiconductor (CMOS) optical sensor, and the like. The camera subsystem 1226 and the optical sensor 1228 can be used to facilitate camera functions, such as recording images and / or video clips (hereinafter referred to as "image data"). In one aspect, the image data is a subset of the sensor data.

[0129] Communication functions may be facilitated through one or more wireless communication subsystems 1230. The wireless communication subsystems 1230 may include radio frequency receivers and transmitters, optical (e.g., infrared) receivers and transmitters, etc. The specific design and implementation of the wireless communication subsystem(s) 1230 may depend on the specific type of portable device 1200 implemented and / or the communication network(s) over which the portable device 1200 is intended to operate.

[0130] For illustrative purposes, the wireless communication subsystem 1230 may be designed to operate on one or more mobile networks (e.g., GSM, GPRS, EDGE), Wi-Fi networks that may include WiMax networks, short-range wireless networks (e.g., Bluetooth networks), and / or any combination thereof. The wireless communication subsystem(s) 1230 may implement a host protocol so that the portable device 1200 may be configured as a base station for other wireless devices.

[0131] The audio subsystem 1232 can be coupled to a speaker 1234 and a microphone 1236 to facilitate voice-enabled functions such as voice recognition, voice replication, digital recording, audio processing, and telephony. The audio subsystem 1232 can generate audio-type sensor data. In one or more embodiments, the microphone 1236 can be used as a respirator sensor.

[0132] I / O devices 1238 may be coupled to interface circuitry 1206. Examples of I / O devices 1238 include, for example, a display device, a touch-sensitive display device, a trackpad, a keyboard, a pointing device, a communication port (e.g., a USB port), a network adapter, buttons or other physical controls, and the like. Touch-sensitive devices such as display screens and / or pads are configured to detect contact, movement, interruption of contact, and the like using any of a variety of touch-sensitive technologies. For example, example touch-sensitive technologies include capacitive, resistive, infrared, and surface acoustic wave technologies, other proximity sensor arrays or other elements for determining one or more points of contact with a touch-sensitive device, and the like. One or more of I / O devices 1238 may be adapted to control functionality of sensors, subsystems, and the like of portable device 1200.

[0133] The portable device 1200 also includes a power supply 1240. The power supply 1240 is capable of providing power to various components of the portable device 1200. In one embodiment, the power supply 1240 is implemented as one or more batteries. The batteries can be implemented using any of a variety of different battery technologies, whether disposable (e.g., replaceable) or rechargeable. In another embodiment, the power supply 1240 is configured to obtain electrical power from an external source and provide power (e.g., DC power) to the components of the portable device 1200. In the case of rechargeable batteries, the power supply 1240 may also include circuitry capable of charging the one or more batteries when coupled to an external power source.

[0134] The memory 1202 may include random access memory (e.g., volatile memory) and / or non-volatile memory, such as one or more magnetic disk storage devices, one or more optical storage devices, flash memory, etc. The memory 1202 may store an operating system 1252, such as LINUX, UNIX, a mobile operating system, an embedded operating system, etc. The operating system 1252 may include instructions for handling system services and for performing hardware-dependent tasks.

[0135] Memory 1202 may store other program code 1254. Examples of other program code 1254 may include instructions for facilitating communication with one or more additional devices, one or more computers, and / or one or more servers; graphical user interface processing; processing instructions for facilitating sensor-related functionality; telephony-related functionality; electronic messaging-related functionality; web browsing-related functionality; media processing-related functionality; GPS and navigation-related functionality; security functionality; camera-related functionality, including web camera and / or web video functionality; and the like. Memory 1202 may store applications including IPR program code 1256. Memory 1202 may also store one or more other applications 1258.

[0136] The various types of instructions and / or program codes described are provided for purposes of illustration and not limitation. The program code can be implemented as a separate software program, process, or module. The memory 1202 can include additional instructions or fewer instructions. In addition, the various functions of the portable device 1200 can be implemented in hardware and / or software, including implementation with one or more signal processing and / or application specific integrated circuits.

[0137] The program code stored in memory 1202 and any data used, generated, and / or manipulated by portable device 1200 are functional data structures that, when used as part of the device, impart functionality to the device. Further examples of functional data structures include, for example, sensor data, data obtained via user input, data obtained via querying external data sources, baseline information, and the like. The term "data structure" refers to the physical implementation of the data organization of a data model within physical memory. Thus, a data structure is formed by specific electrical or magnetic structural elements within the memory. The data structure imposes a physical organization on the data stored in the memory for use by the processor.

[0138] In some embodiments, one or more of the various sensors and / or subsystems described with reference to the portable device 1200 may be separate devices coupled or communicatively linked to the portable device 1200 via a wired or wireless connection. For example, one or more (or all) of the position sensor 1210, light sensor 1212, proximity sensor 1214, gyroscope 1216, magnetometer 1218, accelerometer 1220, altimeter 1222, voice recorder 1224, camera subsystem 1226, audio subsystem 1232, etc. may be implemented as separate systems or subsystems operatively coupled to the portable device 1200 via the I / O device 1238 and / or the wireless communication subsystem 1230.

[0139] The portable device 1200 may include fewer components than shown, or may include components other than Figure 12Additional components beyond those shown may be included, depending on the specific type of system being implemented. In addition, the specific operating system and / or applications and / or other program code included may also vary depending on the type of system. Furthermore, one or more of the illustrative components may be incorporated into another component or otherwise form part of another component. For example, a processor may include at least some memory.

[0140] The portable device 1200 is provided for purposes of illustration and not limitation. Devices and / or systems configured to perform the operations described herein may have Figure 12 12. The portable device 1200 may be a system that can be configured to perform tasks in a manner different from that shown in the preceding text. The system may be a simplified version of the portable device 1200 and may include a processor and a memory for storing instructions. The system may include one or more sensors as described herein. The portable device 1200 or a similar system may use various sensors of the device or sensors coupled thereto to collect data. However, it should be understood that the portable device 1200 may include fewer sensors or other additional sensors. In this disclosure, data generated by sensors is referred to as "sensor data."

[0141] Example implementations of the portable device 1200 include, for example, a smartphone or other mobile device or phone, a wearable computing device (e.g., a smartwatch), a dedicated medical device, or other suitable handheld, wearable, or comfortably portable electronic device capable of sensing and processing signals and data detected by sensors. It will be understood that various embodiments can be deployed as a standalone device or as multiple devices in a distributed client-server networked system. For example, in some embodiments, a smartwatch can be operably coupled to a mobile device (e.g., a smartphone). The mobile device may or may not be configured to interact with a remote server and / or computer system.

[0142] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. Nevertheless, several definitions that apply throughout this document will now be presented.

[0143] As defined herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0144] As defined herein, "another" means at least a second or more.

[0145] As defined herein, "at least one," "one or more," and "and / or" are open-ended expressions that are conjunctive and disjunctive in operation unless expressly stated otherwise. For example, each of the expressions "at least one of A, B, and C," "at least one of A, B, or C," "one or more of A, B, and C," "one or more of A, B, or C," and "A, B, and / or C" means A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.

[0146] As defined herein, "automatically" means without user intervention.

[0147] As defined herein, “includes,” “comprising,” and / or “containing” specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0148] As defined herein, "if" means "in response to" or "in response to," depending on the context. Thus, the phrase "if it is determined" can be interpreted to mean "in response to determining" or "in response to determining," depending on the context. Similarly, the phrase "if [the condition or event] is detected" can be interpreted to mean "upon detection of [the condition or event]," "in response to detecting [the condition or event]," or "in response to detecting [the condition or event]," depending on the context.

[0149] As defined herein, "one embodiment," "an embodiment," "in one or more embodiments," "in a specific embodiment," or similar language means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment described within the present disclosure. Thus, appearances of the above phrases and / or similar language throughout this disclosure may, but do not necessarily, all refer to the same embodiment.

[0150] As defined herein, the phrases "in response to" and "respond" indicate a tendency to respond or react to an action or event. Thus, if a second action is performed "in response to" or "in response to" a first action, there is a causal relationship between the occurrence of the first action and the occurrence of the second action. The phrases "in response to" and "respond" indicate a causal relationship.

[0151] As defined herein, "real time" means a level of processing responsiveness that is perceived by a user or system as sufficiently immediate for a particular process or determination to be made, or that enables the processor to keep up with some external process.

[0152] As defined herein, "substantially" means that the stated characteristic, parameter or value need not be achieved precisely, but deviations or changes, including, for example, tolerances, measurement errors, measurement precision limitations and other factors known to those skilled in the art, may occur in amounts that do not negate the effect that the characteristic is intended to provide.

[0153] As defined herein, "User" and "Individual" each refers to a human being.

[0154] The terms first, second, etc. may be used herein to describe various elements. These elements should not be limited by these terms, as these terms are only used to distinguish one element from another, unless otherwise stated or the context clearly indicates otherwise.

[0155] The present invention may be a system, method and / or computer program product at any possible level of technical detail integration. The computer program product may include a computer-readable storage medium (or multiple media) having computer-readable program instructions thereon, the computer-readable program instructions being used to cause a processor to perform various aspects of the present invention.

[0156] A computer-readable storage medium can be a tangible device that can retain and store instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to: an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punch card or a raised structure in a groove on which instructions are recorded, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium should not be interpreted as a temporary signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted by a wire.

[0157] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a corresponding computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in a computer-readable storage medium within the corresponding computing / processing device.

[0158] The computer-readable program instructions for performing the operations of the present invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, configuration data for an integrated circuit, or source code or object code written in any combination of one or more programming languages ​​(including object-oriented programming languages, such as Smalltalk, C++, etc.) and procedural programming languages ​​(such as "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to perform various aspects of the present invention, an electronic circuit comprising, for example, a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA) may execute the computer-readable program instructions to personalize the electronic circuit by utilizing the state information of the computer-readable program instructions.

[0159] Aspects of the present invention are described herein with reference to the flowcharts and / or block diagrams of the methods, apparatus (systems) and computer program products according to embodiments of the present invention. It will be understood that each block of the flowcharts and / or block diagrams and the combination of blocks in the flowcharts and / or block diagrams can be implemented by computer-readable program instructions.

[0160] These computer-readable program instructions can be provided to a processor of a computer or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device create a device for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can direct the computer, programmable data processing device and / or other equipment to operate in a specific manner, so that the computer-readable storage medium having the instructions stored therein includes an article of manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0161] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable apparatus, or other device to implement the functions / actions specified in one or more boxes of the flowchart and / or block diagram.

[0162] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the system, method and computer program product according to various embodiments of the present invention. In this regard, each block in the flow chart or block diagram can represent a module, segment or part of an instruction, which includes one or more executable instructions for realizing the specified logical function. In some alternative embodiments, the functions noted in the frame may not occur in the order noted in the figure. For example, the two frames shown in succession can actually be implemented as a step, simultaneously, substantially simultaneously, in a manner that overlaps part or all of the time, or these frames can sometimes be performed in reverse order, depending on the functions involved. It will also be noted that each block of the block diagram and / or flow chart illustration and the combination of the blocks in the block diagram and / or flow chart illustration can be implemented by a dedicated hardware-based system that performs a specified function or action or performs a combination of dedicated hardware and computer instructions.

[0163] The description of various embodiments of the present invention has been provided for the purpose of illustration and is not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, practical applications, or improvements over existing technologies in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A computer-implemented method comprising: performing natural language processing (NLP) on a machine-encoded document to determine one or more document segment classifiers, wherein the machine-encoded document is retrieved via a data communications network; detecting previously consumed content within the machine-coded file that corresponds to the one or more document segment classifiers by performing a computer database search for NLP elements identified within electronic content previously presented to the user that match NLP elements extracted from the machine-coded file, wherein the search is performed against a user-specific consumption corpus database; and Based on the detection, an electronic presentation structure for electronically presenting the machine encoded file to a user is modified, wherein the electronic presentation structure is modified to distinguish a portion of the machine encoded file corresponding to the previously consumed content from other portions of the machine encoded file.

2. The method according to claim 1, wherein Modifying the electronic presentation structure is performed in response to determining a contextual basis of the one or more document segment classifiers.

3. The method according to claim 1, further comprising: electronically tracking the location of a device used to receive the electronic presentation structure and electronically present it to a user; as well as The electronic presentation structure is modified based on the location of the device.

4. The method according to claim 1, further comprising: subsequently detecting one or more revisions to the machine-coded file after a user retrieves the machine-coded file; as well as In response to detecting one or more revisions, the machine-encoded file is rendered using another electronic rendering structure indicating which portions are newly added or revised.

5. The method according to claim 1, wherein The one or more file segmentation classifiers correspond to one or more topics of interest to the user discovered by performing unsupervised machine learning topic modeling, and wherein the topic modeling discovers the topics of interest to the user based on a plurality of machine-encoded files retrieved by the user via the data communications network.

6. The method of claim 1 , wherein the one or more document segment classifiers correspond to one or more topics of interest to the user identified by performing supervised machine learning topic classification, and wherein, The subject classification identifies subjects specified by a user prior to retrieving the machine-encoded document.

7. The method according to claim 1, wherein The previously presented electronic content includes one or more machine-encoded files previously retrieved via the data communications network within a user-specified time interval.

8. The method according to claim 1, wherein The electronic presentation structure includes text, images, audio, or video embedded in the machine-coded file.

9. The method according to claim 8, wherein For electronic presentation structures that include audio or video, the modifications include: Accelerating the audio or video rendering speed for portions of the electronic content previously consumed by the user; and / or The audio or video is compressed to eliminate portions of the electronic content that were previously consumed by the user.

10. The method according to claim 8, wherein For the electronic presentation structure including text, the modification includes: annotating the text to distinguish portions of the text corresponding to portions of the electronic content previously consumed by the user from other portions of the text; modifying one or more formatting tags to change the format of the text corresponding to the portion of the electronic content previously consumed by the user from other portions of the text; or A portion of the text is extracted that corresponds to a portion of the electronic content that was previously consumed by the user.

11. The method according to claim 10, wherein: The extracting includes replacing the extracted portion with a link pointing to a database in which the extracted portion is electronically stored, and providing instructions instructing the user to link to the database to restore the extracted portion to the electronic presentation structure.

12. The method according to claim 1, further comprising: detecting one or more revisions to a machine-coded file subsequent to an earlier retrieval of an earlier version of the machine-coded file by a user via a data communications network, wherein the detecting comprises comparing natural language processing (NLP) elements identified within electronic content of the earlier version of the machine-coded file with NLP elements identified within electronic content of the revised version of the machine-coded file; as well as Based on detecting one or more revisions, the machine-encoded file is rendered using another electronic rendering structure indicating which portions were newly added or revised.

13. A computer system comprising: Memory; as well as A processor in communication with the memory, the processor being configured to perform the steps of the method according to any one of claims 1 to 12.

14. A computer program product comprising program instructions executable by a processor to cause the processor to initiate the operation of the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Method and apparatus for incorporating visual deltas for new documents based on previous consumption

    US20100095203A1

  • Controlling Start Times at which Skippable Video Advertisements Begin Playback in a Digital Medium Environment

    US20170195746A1