Dynamically generated content understanding system

By dynamically building the system, the problems of low efficiency and low accuracy of composite content item parsing in the prior art are solved, and efficient and accurate content parsing and performance optimization are achieved.

CN113196276BActive Publication Date: 2025-06-06MICROSOFT TECHNOLOGY LICENSING LLC

Patent Information

Application Number
CN201980082436.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-12-14
Filing Date
2019-12-05
Publication Date
2025-06-06
Estimated Expiration
2039-12-05

AI Technical Summary

Technical Problem

When existing computing systems deal with composite content items, it is difficult to effectively select and deploy appropriate parsers, resulting in low parsing efficiency and low accuracy.

Method used

Dynamically build the system, monitor parser usage and user feedback, analyze data based on the type and history of content items, and dynamically select and deploy the appropriate parsers to achieve efficient content parsing.

Benefits of technology

Improve the efficiency and accuracy of content parsing, meet different tenants' diverse needs for parser performance characteristics, and reduce user latency and waste of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113196276B_ABST
    Figure CN113196276B_ABST
Patent Text Reader

Abstract

A content item is received and analyzed to identify any different types of parsers that can be used to parse the content item based on a previous user-selected parser. One or more parsers are selected based on a content type in the content item and based on a previous user-selected parser. The selected parsers are built in a server environment and controlled to parse the content item.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Computing systems are currently in widespread use. Some such computing systems include server-based systems that host or serve content-related functions.

[0002] Some of these types of systems handle different types of content for different tenants. By way of example, some such systems may include productivity and email systems, document management systems, search engines and data mining systems, music and photo cataloging systems, and a variety of other systems. The content items handled and processed by these types of systems may include a variety of different types of content. For example, the content may be text-based documents, images, videos, voice or other sound recordings, compressed archive files, business line documents (e.g., quotes, opportunity records, etc.).

[0003] In order for search functions and intelligence functions to work properly on these content items, the content items are typically processed by a content understanding system that attempts to generate summaries and annotations. The annotations may include information such as special characteristics or formatting corresponding to the content item. The operation of generating summaries and annotations corresponding to the content item is typically referred to as "parsing", and this is performed by a logical item called a "parser".

[0004] Some content items are also called "compound items" or "complex items." In these content items, there may be multiple different content types in a single item. For example, a word processing document may include embedded images. In these cases, multiple different parsers may sometimes be required to generate acceptable summaries and annotations corresponding to the content item.

[0005] The above discussion is provided for general background information only and is not intended to be used as an aid in determining the scope of the claimed subject matter. Summary of the invention

[0006] A content item is received and analyzed to identify any different types of parsers that can be used to parse the content item based on a previously selected parser. One or more parsers are selected based on a content type in the content item and based on a previously user-selected parser. The selected parser is built in a server environment and controlled to parse the content item.

[0007] This summary is provided to introduce some concepts in a simplified form, which will be further described in the detailed description below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to help determine the scope of the claimed subject matter. The claimed subject matter is not limited to implementations that solve any or all of the shortcomings pointed out in the background technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 is a block diagram of one example of a computing system architecture in which one or more content understanding systems process content items.

[0009] Figure 2 is a block diagram illustrating one example of a dynamic build system.

[0010] Figure 3 is a flow chart illustrating one example of the operation of the detection and monitoring subsystem for training a dynamically built model.

[0011] Figure 4A and Figure 4B (collectively referred to herein as FIG. 4 ) shows an example of Figure 2 A flowchart of one example of the operation of a dynamic construction system in dynamically building a content understanding system to parse a received content item is shown in FIG.

[0012] Figure 5 Yes Description Figure 1 0 is a block diagram of an example of an architecture shown in , which is deployed in a cloud computing architecture.

[0013] Figure 6 is a block diagram illustrating one example of a computing environment that may be used in the architecture shown in the previous figures. DETAILED DESCRIPTION

[0014] Figure 1 1 is a block diagram of one example of a computing system architecture 100, which illustrates a plurality of different content server systems 102-104, each serving a plurality of tenants 106, 108, 110, and 112. Systems 102 and 104 may be connected to each other (and / or to their corresponding tenants) via a network 114. Network 114 may be any of a variety of different types of networks (e.g., a local area network, a wide area network, a cellular communication network, a near field communication network) or any of a variety of other types of networks or combinations of networks.

[0015] The content server systems 102-104 may be any of a variety of different types of content server systems (e.g., email systems, productivity systems, document sharing and management systems, search systems, data mining systems, line of business systems, music and image cataloging systems, or a variety of other systems). The content server system 102 is shown as having one or more processors or servers 116, a content understanding system 118, a data repository 120, and the content server system 102 may have a variety of other system functions 122. Similarly, the content server system 104 is shown as having one or more processors or servers 124, a data repository 126, a content understanding system 128, and the content server system 104 may also have a variety of other server system functions 130. The content server systems 102-104 may be similar or different. For the purposes of this discussion, it is assumed that the content server systems 102-104 are similar, so only the content server system 102 is described more fully below.

[0016] In one example, tenants 106-108 access or interact with documents (or content items) stored and processed by content server system 102. System 102 may also serve or host other functions. Content items accessed or provided by tenants 106-108 are processed by content understanding system 118. Content understanding system 118 illustratively generates summaries and annotations corresponding to each content item. In doing so, content understanding system 118 may operate on binary representations of content items to generate full-text summaries, as well as annotations that may identify the type of content, its format, etc. The parsed content items (and their corresponding summaries and annotations) may be stored in data repository 120 so that the content items may be searched or otherwise intelligently processed by systems such as search engines and / or other data or document processing systems.

[0017] Therefore, the type of parsing function used by the content understanding system 118 will vary based on the type of content item it is processing. For example, if the content understanding system 118 is processing a relatively large number of text image files (e.g., word processing documents or PDF files), the content understanding system 118 may use a text parser to parse the documents. On the other hand, if the PDF file is generated by splicing images from a scanner, the content understanding system 118 may use an optical character recognition parser to generate text summaries and annotations. If the content understanding system 118 is processing photographic images, it may use a different type of parser to generate text summaries and annotations.

[0018] The processing performed when executing a parser in the content understanding system 118 may involve relatively heavy computational tasks in terms of central processing unit (CPU) usage, memory and / or hard disk usage, etc. Moreover, depending on the specific parser capabilities of the content understanding system 118, the time it takes for the parser to parse a content item and the accuracy with which the item is parsed can vary greatly.

[0019] Additionally, the performance characteristics of the parsers desired by different tenants 106-108 may also vary depending on different scenarios. Some tenant computing systems may have requirements for user latency, memory / CPU usage, tenant preferences for parser types, etc. Users of tenants 106-108 may also identify preferences for different parser results, parsers, etc.

[0020] Thus, in one example (and also as Figure 1 ), the dynamic build system 132 is deployed in the architecture 100. The dynamic build system 132 is deployed in the architecture 100. Figure 1 104. The dynamic build system 132 is shown as being separate from both the content server system 102 and the content server system 104. Thus, the dynamic build system 132 may be a separately hosted service or system accessed by the systems 102-104. In another example, the dynamic build system 132 may be located on the content server system 102 and accessed by the content server system 104, or the dynamic build system 132 may be located on the content server system 104 and accessed by the content server system 102. In yet another example, a separate dynamic build system 132 may be located on each of the content server systems 102-104.

[0021] Briefly, by way of overview, the dynamic build system 132 illustratively monitors parser usage, tenant user input or feedback, and parser usage characteristics relative to different content items processed by the content server system 102, and identifies different parser functions that should be deployed to the content understanding system 118 based on how the tenants 106-108 use the parsing functions, etc. In the event that the dynamic build system 132 also serves the content service system 104, the dynamic build system 132 can perform the same operations as the system 104. The dynamic build system 132 can also pre-deploy parsing functions to different systems 102-104 based on current, historical, and / or predicted parser usage.

[0022] In addition to pre-deploying parser functions to different content understanding systems, the dynamic build system 132 can also receive content items to be parsed and dynamically build parsers in the corresponding content understanding system 118 in which the content items are to be parsed. As described in more detail below, the dynamic build system 132 illustratively processes the received content items to identify one or more parsers that can be used to parse the content items. The dynamic build system 132 then generates a parser graph that identifies the sequence of parsers to be used (where more than one parser is to be used, for example, in the case of complex content items) and builds or deploys these parsers in the appropriate sequence on the corresponding content understanding system 118 in which the parsing function is to be performed. In doing so, the dynamic build system 132 can take into account a variety of information, such as: what content types are contained in the content items; parsers that have been historically used for these content types; parsers identified by users; the degree of similarity of the content item (whose type is unknown) to other content items; feedback from tenants about parser preferences, latency or other user-defined performance preferences, and a variety of other information. In this way, when parsing capabilities have not been pre-deployed to the content understanding system, parsing capabilities can be dynamically built during runtime. Parsing capabilities can be built during runtime, taking into account historical parsing operations and user feedback preferences in order to provide highly accurate summaries and corresponding annotations.

[0023] Figure 2 is a block diagram illustrating one example of dynamic build system 132 in greater detail. Figure 2 Shown, in one example, the dynamic build system 132 illustratively includes one or more processors or servers 134, an orchestration subsystem 136, a data repository 138, a detection and monitoring subsystem 140, a parser tracking subsystem 142 that tracks available parsers 143, a model generation logic unit 144, a dynamic build model 146 (which may include a parser graph builder subsystem 150), a deployment subsystem 148, and the dynamic build system 132 may include a variety of other items 152. Before describing the overall operation of the system 132 in more detail, a brief description of some of the items in the system 132 and their operation will first be provided.

[0024] The orchestration system 136 illustratively determines which content understanding system 118-128 will perform parsing operations on the received content item. The detection and monitoring subsystem 140 detects various different types of information, which the model generation logic unit 144 can use to generate a dynamic build model 146. The dynamic build model 146 can include a subsystem 150 (as described below) and illustratively generates an output based on the received content item or based on a call for an output that identifies the parsing function that should be deployed in certain systems.

[0025] The detection and monitoring subsystem 140 may include a content type to location identifier logic unit 154, a performance data detection logic unit 156, a user feedback detection logic unit 158, a user selection criteria detection logic unit 160, a data aggregation / correlation logic unit 161, and the detection and monitoring subsystem 140 may include other items 162. In one example, the various resolvers may have reporting logic units deployed therein so that the resolvers can report back the type of content item they are being called to resolve, the geographic location they are located in, the data center they are located in, the server ID (which identifies the server on which the resolver is running), the tenant ID (which identifies the tenant for which the content item is being resolved), and the like.

[0026] Thus, the content type to location identifier logic 154 identifies the relationship between the type of content being parsed and its location (with respect to its geographic location, data center, server, tenant, etc.). The performance data detection logic 156 detects performance data corresponding to the parsing operations performed by the different parsers. For example, the performance data detection logic 156 may detect latency, CPU usage, memory and hard disk usage, etc. for different parsing operations.

[0027] The user feedback detection logic unit 158 ​​illustratively generates a representation of a user interface that can be presented to a user of a particular tenant to provide feedback on a parsing operation. For example, the results of two different parsers operating on the same content item can be provided to the user, and the user can be asked to select the parser they prefer. The user can be asked to rate the parser output for satisfaction. The user can be asked to select a parser for a particular type of content, or the user can be asked to provide feedback in a variety of other ways.

[0028] The user selection criteria detection logic unit 160 may also illustratively provide a user input mechanism by which a user of a particular tenant may identify the selection criteria that he or she wishes to use when selecting a parser to parse various content items. The selection criteria may allow the user to identify things such as latency, accuracy, performance (e.g., CPU and memory usage), etc. The user may identify which of these criteria are most important so that the dynamically constructed model 146 may take these criteria into account when identifying parsers to be deployed and used in the content understanding system for parsing content items.

[0029] The data aggregation / correlation logic 161 may aggregate data received by the subsystem 140. The data aggregation / correlation logic 161 may aggregate the data through a parsing operation or otherwise. The data aggregation / correlation logic 161 may then illustratively correlate the aggregated data with a parser, location, or otherwise.

[0030] The detection and monitoring subsystem 140 may detect and monitor a variety of other information that may also be used. This is indicated by block 162 .

[0031] The parser tracking subsystem 142 illustratively includes a registration logic unit 164, a parser access logic unit 166, and the parser tracking subsystem 142 may include various other items 168. The registration logic unit 164 allows a developer to register or otherwise provide an indication that a parser is available. The registration logic unit 164 may provide the type of content it parses, as well as various other data that identifies a particular parser. The parser access logic unit 166 illustratively accesses a list of parsers 143 or identification data for available parsers 143 (parsers that have been identified to the registration logic unit 164) and provides this information to the dynamic construction model 146 and the parser graph builder subsystem 150.

[0032] The model generation logic unit 144 receives information and other items from the detection and monitoring system 140, and in one example, the model generation logic unit 144 runs a machine learning system to generate and improve the dynamic construction model 146. In another example, the dynamic construction model 146 can also be generated based on a set of rules (e.g., business rules) and a model generation algorithm. The dynamic construction model 146 can be any type of machine-generated model that is used to identify parsers to be used in different scenarios. The model 146 can generate an output indicating one or more parsers to be used during runtime, or the model 146 can generate the output to pre-deploy the parsing function to different locations before performing the parsing or intermittently during the operation of the system 118, 128. In doing so, the model 146 can use the parser graph builder subsystem 150 and / or other logic units 151.

[0033] The parser graph builder subsystem 150 illustratively receives information or content items to be parsed from the subsystem 140 during runtime and analyzes the information content items. Based on the analysis, the parser graph builder subsystem 150 accesses the available parsers through the parser access logic unit 166 and identifies which specific parsers should be used and in what sequence these parsers should be used to parse the content item (when the content item is a complex content item containing two or more types of content). Therefore, the subsystem 150 illustratively includes a content header / extension processor logic unit 170, a content classifier logic unit 172, a historical usage logic unit 174, a complex content analysis logic unit 176, an other content analysis logic unit 178, a parser selection logic unit 180, a graph generator logic unit 181, and the subsystem 150 may include a variety of other items 182. The content header / extension processor logic unit 172 analyzes the header information and the extension information for the content item to identify the type of content being processed. The complex content analysis logic 176 identifies whether the content item is a complex item, which means that the content item has more than one type of content (for which more than one type of parser may be required). Such examples may include a spreadsheet document with attached or embedded images, a word processing document with embedded audio files or images, etc.

[0034] When the type of content cannot be identified by logic units 170 and 176, content classifier logic unit 172 analyzes the binary representation of the content item and classifies it into one of a plurality of different predefined categories corresponding to the previously parsed content item. For example, when content classifier logic unit 172 receives an unknown content item, the binary representation of the content item may have characteristics very similar to a word processing document. The binary representation of the content item may have characteristics very similar to an optical image or a photograph. The binary representation of the content item may have characteristics representing a complex document including a plurality of different types of content. In this case, content classifier logic unit 172 (which may be a machine learning classifier) ​​classifies the content item into predefined categories. Each predefined category may have one or more corresponding parsers corresponding thereto, so that once a category is identified, the corresponding parser or set of parsers that will be used to parse the content item will also be known.

[0035] The historical usage logic 174 may identify parsers that have been used in the past to parse this type of content. The historical usage logic 174 may identify a parser selected by a user or a parser selected based on user-defined selection criteria or otherwise.

[0036] Once the logic has been executed on the content item, the parser selection logic 180 selects one or more parsers, and the graph generator logic 181 arranges the selected parsers in a sequence corresponding to the sequence in which the parsers will operate on the content item. It will be noted that the parser selection logic 180 illustratively selects a parser that has been identified as being most suitable for the type of content. For example, if the type of content is an image, the parser selection logic 180 will ask the processing and analysis logic to identify the type of image (e.g., whether it is a photograph or an image of a text document (e.g., a PDF file)). A parser based on optical character recognition may be suitable for images of text documents, but not for different types of visual images (e.g., photographs). Therefore, the parser selection logic 180 will distinguish the type of image in order to select a parser to be used when parsing the content item.

[0037] The graph output by the logic unit 181 or the selected parser output by the logic unit 180 can be provided to and used by the deployment subsystem 148. The deployment subsystem 148 illustratively includes a dynamic model access logic unit 184, a pre-deployment logic unit 186, a runtime deployment logic unit 188, a post-processing deployment logic unit 190, and the deployment subsystem 148 may include other items 192. The logic unit 184 illustratively accesses the parsers identified by the parser graph builder subsystem 150 to identify the parsers that should be deployed to different content understanding systems 118-128. This can be done before receiving the items to be parsed or at runtime. For example, the logic unit 184 can identify that a particular tenant is showing particularly heavy usage of a certain type of parser function. Currently, the parser function may also not be located in a geographic location or data center close to the tenant using the function. In such a scenario, the pre-deployment logic unit 186 can deploy the identified parser function (or parser 143) to a data center or geographic location close to the tenant in order to reduce the latency of the tenant. The pre-deployment logic unit 186 can deploy the parser functionality to adapt to workloads in different server environments, or for other reasons, the pre-deployment logic unit 186 can also adapt the output of the dynamic construction model 146.

[0038] The runtime deployment logic 188 may also use the output from the dynamically built model 146 (e.g., the graph output by the graph generator logic 188) to perform runtime deployment of parsers to the content understanding systems 118-128 based on content items that have been received and are to be parsed. The post-processing deployment logic 190 may also build post-processing logic at the content understanding systems 118-128 based on the specific needs of the tenant, based on a user-selected parser or user-defined selection criteria, based on historically observed processing, or for other reasons.

[0039] Figure 3 is a flow chart illustrating one example of the operation of the detection and monitoring subsystem 140 in detecting and monitoring information that may be used by the model generation logic 144 , the parser graph builder subsystem 150 , and / or the deployment subsystem 148 .

[0040] The detection and monitoring subsystem 140 detects the type of content parsed at different locations. Figure 3 In one example, the parsers in the content understanding systems 118, 128 detect and send indications of the type of data they are being asked to parse. This is indicated by box 196. The data may identify the type of content 198, the specific parser used (as indicated by box 200), the geographic location where the parsing is performed (as indicated by box 202), the identification of the data center where the parsing is performed (as indicated by box 204), a server identifier identifying the server performing the parsing (as indicated by box 206), a tenant identifier identifying the tenant for which the parsing is performed (as indicated by box 208), and this may also include a variety of other items. These other items are indicated by box 210.

[0041] The performance data detector logic 156 then detects performance data for various parsing operations. This is indicated by block 212. The performance data may include an indication of parsing accuracy, as indicated by block 214. The performance data may include CPU usage 216, memory usage 218, disk usage 220, network capacity 222, network transmission delay 224, and the performance data may include a variety of other items 226.

[0042] The user feedback detection logic 158 can detect any user feedback 228 provided. The user feedback 228 can include preference data indicating a user preference for one parser over another, which is indicated by box 230. The preference data can include a user rating that rates the user's satisfaction with the parsing results or operation of a particular parser, as indicated by box 232. The user selection criteria detection logic 160 can detect selection criteria entered by the user and to be used to select a particular parser. The selection criteria can include performance data or other items. The selection criteria is indicated by box 234. The user feedback can also include a variety of other items, and this is indicated by box 236.

[0043] The detection and monitoring subsystem 140 illustratively detects changes in data over time. This is indicated by box 238. For example, the detection and monitoring subsystem 140 can detect cyclic or periodic changes, such as a particular tenant processing a particular type of content more frequently at different times of the day, year, etc. This is indicated by box 240. The detection and monitoring subsystem 140 can detect trends that indicate, for example, the types of content that different tenants process more frequently over time, the accuracy or performance of a parser as it changes over time, or a variety of other trends. Box 242 indicates detecting trends over time. The system 140 can also detect changes over time in a variety of other ways, and this is indicated by box 244.

[0044] The data aggregation and correlation logic 161 may then perform data aggregation and correlation on the data and control the data repository 138 to store the data, the aggregation and correlation, and other related information. The data may be aggregated over time, aggregated through a parsing operation, or may be combined in other ways. The data may be associated with a parser, a location, a user, etc. Block 246 indicates performing data aggregation and correlation, and block 248 indicates controlling the data repository 138 to store the data.

[0045] The model generation logic unit 144 then uses a machine learning algorithm to access the data stored in the data repository 138 and trains the dynamic build model 146 based on the data, aggregations, correlations, etc. This is indicated by box 250. Then, as indicated by Figure 3 As indicated by block 252 in the flowchart of , the dynamic build model 246 is output. In this way, the dynamic build model uses machine learning to identify the models that should be pre-deployed to different environments (e.g., Figure 1 The model generation logic unit 144 can also receive the document to be identified (or the analysis results of the content item) and generate an output to the deployment system 148, which indicates which parsers should be deployed to which locations to parse the content item.

[0046] Figure 4A and Figure 4B(collectively referred to herein as FIG. 4 ) shows a flowchart illustrating an example of the operation of the dynamic build system 132 when dynamically building a parser or parsing capability at a particular location. It is first assumed that the dynamic build model access logic unit 184 accesses the dynamic build model 146, and the pre-deployment logic unit 186 uses the model 146 to perform pre-deployment of certain parsing capabilities (or parsers) to different content understanding systems in the architecture 100. This can be based on historical usage of the parser, the types of content most commonly used by various tenants, user input such as parser selection and selection criteria input, and so on. Box 260 in the flowchart of FIG. 4 indicates accessing the dynamic build model to perform pre-deployment, and box 262 indicates using the pre-deployment logic unit 186 to pre-deploy the parsing functionality to the content understanding system in the architecture 100 based on the output of the model 146.

[0047] At a certain moment, the orchestration subsystem 126 will receive the content item to be parsed. The orchestration subsystem 126 will identify the specific content understanding system that will be used to parse the content item. The orchestration subsystem 126 also provides the content item to the dynamic construction model 146. Box 264 indicates that the content item to be parsed is received. Then, the model 146 identifies one or more parsers for parsing the content item. In doing so, the parser graph builder subsystem 150 processes the content item to obtain parser-related data, which will be used to select one or more parsers that will be used to parse the content item. This is indicated by box 266. It can operate on the binary representation of the content item. This information can indicate the type of content in the item. This is indicated by box 268. The complex content analysis logic unit 176 can identify whether the content item is a complex content item (meaning that the content item includes more than one type of content to be parsed). This is indicated by box 270. The content header / extension processor logic unit 170 illustratively analyzes the header and extension information about the content item. This information can indicate the type of content in the item. This is indicated by box 272. The historical usage logic 174 identifies parsers that have been used historically to parse this type of content. This is indicated by box 274. The content classifier logic 172 may classify the content item into one of a plurality of predefined categories based on characteristics of the binary representation of the content item (e.g., when the content item cannot be determined by other analysis). This is indicated by box 276.

[0048] The content may also be processed in a variety of other ways. For example, the other content analysis logic unit 178 may identify the most important content type in the complex content item to be parsed. By way of example, if the content item is a spreadsheet, but it has a single image attached, the spreadsheet may be the most important content type and may be marked as such. Identifying the most important content type in the complex content item is indicated by box 278 in the flowchart of Figure 4. The content item may be processed in a variety of other ways to identify data relevant to the parser, and this is indicated by box 280.

[0049] Then, the parser selection logic unit 180 identifies or selects one or more parsers to be used to parse the document. Box 282 indicates the selection of a parser. The parser selection logic unit 180 may select a parser based on a previously user-selected parser selected for different types of content. This is indicated by box 281. The parser selection logic unit 180 may select based on a selection criterion that the user prioritizes. For example, the user may indicate that speed has a higher priority than CPU usage. This is just one example and is indicated by box 283. It may use the parser access logic unit 166 to access available parsers. This is indicated by box 284. As indicated by box 286, it identifies multiple parsers for complex items. It may identify parsers based on the categories into which the classifier logic unit 172 classifies the content item. This is indicated by box 288. It may select the best parser in the case where there may be multiple parsers. In the case where multiple parsers are available for parsing the target file, the specific content of the target file may be identified in order to select the best parser. This is indicated by box 290. Parsers may also be selected in a variety of other ways, and this is indicated by box 292.

[0050] Once the parser selection logic unit 180 has selected a set of parsers, the graph generator logic unit 181 identifies the sequence of parsers to be operated on the complex content item. Box 294 indicates whether more than one parser has been selected, and box 296 indicates the sequence of parsers to be applied. In one example, the parser sequence will be arranged so that the most important content type is parsed first. This is indicated by box 298. Moreover, some content may not be parsed for delay, performance or other reasons. By way of example, only the most important content type can be parsed, and summaries and annotations are generated based on the content type, while for the content item, the remaining content types are not parsed. Box 300 in the flowchart of Figure 4 indicates that the parsing of certain content is omitted. Identifying the sequence of parsers to be applied can also be done in other ways, and this is indicated by box 302.

[0051] The runtime deployment logic 188 then dynamically builds the identified parsers (which were identified by the graph generator logic 181) in the environment in which they are to be built. This is indicated by block 304 in the flowchart of Figure 4. For example, if the orchestration subsystem 136 has identified that the content understanding system 118 is to perform a parsing, and the dynamic build model 146 and / or subsystem 150 has identified a set of parsers to be used for that parsing, the runtime deployment logic 188 deploys those parsers to the content understanding system 118 (or builds upon them).

[0052] The orchestration subsystem 136 then forwards the received content item to the constructed parser to parse it. The parser then outputs the results of the parsing. Box 306 indicates that the received content item is sent to the environment to be parsed, and box 308 indicates that the content item is parsed and the parser results are output. In one example, the parser results include a full text summary 310, a set of annotations 312 (some examples of which have been described above), and the parser results may include a variety of other items 314.

[0053] Then, the parser that is running or has just run detects the monitoring data monitored by the subsystem 140 and sends it to the subsystem 140. This is indicated by block 316.

[0054] Thus, it can be seen that parsing functionality can be pre-deployed based on machine learning model outputs that take into account historical parsing operations, parsing selection criteria, user-selected parsers or selection criteria, latency information, performance information, and a variety of other information when pre-deploying parsing capabilities to different environments or server systems. Moreover, during runtime, specific parsers are selected and used to parse content items in a dynamic manner, such that the desired parsing functionality is deployed during runtime to perform efficient parsing of multiple different types of content items.

[0055] It will be noted that the above discussion has described various systems, components and / or logical units. It will be appreciated that such systems, components and / or logical units can be composed of hardware items (e.g., processors and associated memory or other processing components, some of which are described below) that perform the functions associated with those systems, components and / or logical units. Additionally, as described below, systems, components and / or logical units can be composed of software that is loaded into the memory and then executed by a processor or server or other computing component. Systems, components and / or logical units can also be composed of different combinations of hardware, software, firmware, etc., some examples of which are described below. These are only some examples of different structures that can be used to form the systems, components and / or logical units described above. Other structures can also be used.

[0056] This discussion has referred to processors and servers. In one example, processors and servers include computer processors with associated memory and timing circuits (not shown separately). They are functional parts of the systems or devices to which they belong and by which they are activated, and facilitate the functions of other components or items in these systems.

[0057] Moreover, many user interface displays have been discussed. These user interface displays can take a variety of different forms, and a variety of different user-actuated input mechanisms can be provided thereon. For example, the user-actuated input mechanism can be a text box, a check box, an icon, a link, a drop-down menu, a search box, etc. They can also be actuated in a variety of different ways. For example, a click device (for example, a tracking ball or a mouse) can be used to actuate them. Hardware buttons, switches, joysticks or keyboards, thumb switches or thumb pads, etc. can be used to actuate them. Virtual keyboards or other virtual actuators can also be used to actuate them. Additionally, in the case where the screen on which they are displayed is a touch-sensitive screen, touch gestures can be used to actuate them. And, in the case where the device displaying them has a voice recognition component, voice commands can be used to actuate them.

[0058] Many data repositories have also been discussed. It is to be noted that each of these data repositories can be divided into multiple data repositories. All of them can be local to the system accessing them, all of them can be remote, or some can be local and others can be remote. All of these configurations are contemplated herein.

[0059] Moreover, these figures show multiple blocks with functions attributed to each block. It will be noted that fewer blocks can be used, so the functions are performed by fewer components. Moreover, more blocks can be used with functions distributed between more components.

[0060] Figure 5 yes Figure 1, except that its elements are arranged in a cloud computing architecture 500. Cloud computing provides computing, software, data access and storage services that do not require end users to understand the physical location or configuration of the system that delivers the services. In various embodiments, cloud computing uses appropriate protocols to deliver services over a wide area network such as the Internet. For example, a cloud computing provider delivers an application over a wide area network and can access it through a web browser or any other computing component. The software or components of architecture 100 and the corresponding data can be stored on a server at a remote location. The computing resources in a cloud computing environment can be integrated at a remote data center location, or they can also be dispersed. Cloud computing infrastructure can deliver services through a shared data center, even if they appear as a single access point to the user. Therefore, the components and functions described herein can be provided from a service provider at a remote location using a cloud computing architecture. Alternatively, these components and functions can be provided from a conventional server, or these components and functions can be installed directly or otherwise on a client device.

[0061] This description is intended to include both public and private cloud computing. Cloud computing (both public and private) provides substantially seamless pooling of resources and reduces the need to manage and configure underlying hardware infrastructure.

[0062] Public clouds are managed by the vendor and typically use the same infrastructure to support multiple consumers. Also, in contrast to private clouds, public clouds free the end user from the burden of managing hardware. Private clouds can be managed by the organization itself and typically do not share infrastructure with other organizations. The organization still maintains the hardware to some extent, e.g., installation and repairs, etc.

[0063] exist Figure 5 In the example shown in , some items are related to Figure 1 The items shown in FIG. 1 are similar and they have similar reference numerals. Figure 5 Specifically shown are systems 102, 104, and 132 that can be located in a cloud 502 (which can be public, private, or a combination of some public and others private). Thus, tenants 106, 108, 110, 112 access these systems through the cloud 502 using user devices.

[0064] Figure 5 Another example of a cloud architecture is also depicted. Figure 5It is shown that it is also contemplated that some elements of the architecture 100 may be provided in the cloud 502, while other elements are not provided in the cloud 502. By way of example, the data repositories 120, 126, 138 may be provided outside the cloud 502 and accessed through the cloud 502. In another example, the dynamic build system 132 (or other items) may be external to the cloud 502. Regardless of where they are located, they may be directly accessed by the tenants 106, 108, 110, 112 or the systems 102, 104 over a network (wide area network or local area network), they may be hosted at a remote site by a service, or they may be provided as a service through the cloud, or accessed through a connection service residing in the cloud. All of these architectures are contemplated herein.

[0065] It will also be noted that the architecture 100 or a portion thereof can be disposed on a variety of different devices. Some of these devices include servers, desktop computers, laptop computers, tablet computers or other mobile devices, such as palmtop computers, mobile phones, smart phones, multimedia players, personal digital assistants, etc.

[0066] Figure 6 is an example of a computing environment in which, for example, architecture 100 or a portion thereof may be deployed. Figure 6 , an example system for implementing some embodiments includes a general purpose computing device in the form of a computer 810. Components of the computer 810 may include, but are not limited to, a processing unit 820 (which may include a processor or server from the previous figures), a system memory 830, and a system bus 821 that couples various system components including the system memory to the processing unit 820. The system bus 821 may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus (which is also known as a mezzanine bus). With regard to Figure 1 The memory and programs described can be deployed in Figure 6 in the corresponding part of .

[0067] Computer 810 typically includes various computer-readable media. Computer-readable media can be any available media that can be accessed by computer 810, and include volatile media and non-volatile media, removable media and non-removable media. By way of example and not limitation, computer-readable media can include computer storage media and communication media. Computer storage media is different from (and does not include) modulated data signals or carrier waves. Computer storage media include hardware storage media, which include volatile media and non-volatile media, removable media and non-removable media implemented in any method or technology, which are used to store information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other storage technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage devices, cassettes, tapes, magnetic disk storage devices or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by computer 810. Communication media typically embody computer-readable instructions, data structures, program modules or other data in a transmission mechanism, and include any information delivery media. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information into the signal. By way of example and not limitation, communication media includes wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, RF, infrared and other wireless media. Combinations of any of the above media should also be included within the scope of computer-readable media.

[0068] The system memory 830 includes computer storage media in the form of volatile and / or nonvolatile memory, such as read-only memory (ROM) 831 and random access memory (RAM) 832. The basic input / output system 833 (BIOS), which contains the basic routines that help transfer information between elements within the computer 810, such as during startup, is typically stored in ROM 831. RAM 832 typically contains data and / or program modules that are immediately accessible and / or currently being operated on by the processing unit 820. By way of example and not limitation, Figure 6 Operating system 834 , application programs 835 , other program modules 836 , and program data 837 are shown.

[0069] The computer 810 may also include other removable / non-removable volatile / non-volatile computer storage media. By way of example only, Figure 6A hard disk drive 841 is shown that reads from or writes to a non-removable non-volatile magnetic medium, and an optical disk drive 855 that reads from or writes to a removable non-volatile optical disk 856 (e.g., a CD ROM or other optical media). Other removable / non-removable, volatile / non-volatile computer storage media that can be used in the exemplary operating environment include, but are not limited to, magnetic tape cassettes, flash memory cards, digital versatile disks, digital video tapes, solid-state RAM, solid-state ROM, etc. The hard disk drive 841 is typically connected to the system bus 821 via a non-removable memory interface such as interface 840, and the optical disk drive 855 is typically connected to the system bus 821 via a removable memory interface such as interface 850.

[0070] Alternatively or additionally, the functions described herein may be performed at least in part by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field programmable gate arrays (FPGAs), program-specific integrated circuits (ASICs), program-specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), etc.

[0071] discussed above and in Figure 6 The drives and their associated computer storage media shown in FIG. 8 provide storage of computer readable instructions, data structures, program modules and other data for the computer 810. Figure 6 844, application programs 845, other program modules 846, and program data 847. Note that these components may be the same as or different from operating system 834, application programs 835, other program modules 836, and program data 837. Operating system 844, application programs 845, other program modules 846, and program data 847 are given different reference numbers here to illustrate that they are at least different copies.

[0072] A user can enter commands and information into the computer 810 through input devices such as a keyboard 862, a microphone 863, and a pointing device 861 (e.g., a mouse, trackball, or touch pad). Other input devices (not shown) may include a joystick, a game pad, a satellite dish, a scanner, and the like. These and other input devices are typically connected to the processing unit 820 through a user input interface 860 coupled to the system bus, but may be connected through other interfaces and bus structures (e.g., a parallel port, a game port, or a universal serial bus (USB)). A visual display 891 or other type of display device is also connected to the system bus 821 via an interface such as a video interface 890. In addition to a monitor, the computer may also include other peripheral output devices such as speakers 897 and a printer 896, which may be connected through an output peripheral interface 895.

[0073] The computer 810 operates in a networked environment using logical connections to one or more remote computers, such as remote computer 880. The remote computer 880 may be a personal computer, handheld device, server, router, network PC, peer device or other public network node, and typically includes many or all of the elements described above relative to the computer 810. Figure 6 The logical connections depicted in the diagram include a local area network (LAN) 871 and a wide area network (WAN) 873, but other networks may also be included. Such networking environments are common in offices, enterprise-wide computer networks, intranets, and the Internet.

[0074] When the computer 810 is used in a LAN networking environment, the computer 810 is connected to the LAN 871 through a network interface or adapter 870. When the computer 810 is used in a WAN networking environment, the computer 810 typically includes a modem 872 or other module for establishing communications over the WAN 873 (e.g., the Internet). The modem 872, which may be internal or external, may be connected to the system bus 821 via the user input interface 860 or other appropriate mechanism. In a networking environment, program modules depicted relative to the computer 810 or portions thereof may be stored in the remote memory storage device. By way of example and not limitation, Figure 6 Remote application programs 885 are shown residing on remote computer 880. It will be appreciated that the network connections shown are exemplary and other means of establishing a communications link between the computers may be used.

[0075] It should also be noted that the different examples described herein can be combined in different ways. That is, parts of one or more examples can be combined with parts of one or more other examples. All of this is contemplated herein.

[0076] Example 1 is a computing system comprising:

[0077] a detection and monitoring logic unit that detects, from a content understanding system, data associated with a parser that parses content items to generate a textual summary of each of the content items, the parser-associated data indicating a type of content parsed by the content understanding system, and a parser identifier indicating a user-selected parser used to parse the content item;

[0078] a resolver tracking subsystem, which identifies available resolvers;

[0079] a parser selection logic unit for: receiving a content item to be parsed; accessing a parser tracking system; and automatically selecting a parser from among available parsers for parsing the content item to be parsed based on the parser identifier; and

[0080] A deployment subsystem that generates control signals to deploy the selected parser.

[0081] Example 2 is the computing system of any or all of the previous examples, and also includes:

[0082] A user selection criteria detection logic unit is configured to detect a priority input indicating a prioritized selection criteria, and a parser selection logic unit is configured to select a parser based on the prioritized selection criteria.

[0083] Example 3 is the computing system of any or all of the previous examples, wherein the detection and monitoring logic unit includes:

[0084] A user feedback detection logic unit is configured to detect a separate parser identifier for each of a plurality of different corresponding content types, each separate parser identifier identifying a user-selected parser for parsing the corresponding content type.

[0085] Example 4 is a computing system of any or all of the previous examples, wherein the received content item to be parsed includes a complex content item having multiple different content types, and wherein the parser selection logic unit selects multiple different parsers based on parser identifiers corresponding to the multiple different content types.

[0086] Example 5 is the computing system of any or all of the previous examples, wherein the parser selection logic unit selects the plurality of different parsers based on prioritized selection criteria.

[0087] Example 6 is the computing system of any or all of the previous examples, and also includes:

[0088] A content classifier logic unit is configured to classify the content item to be parsed into predefined content type categories based on characteristics of the content item to be parsed, each predefined category having a corresponding parser identifier.

[0089] Example 7 is the computing system of any or all of the previous examples, wherein the parser selection logic unit is configured to select a parser based on a parser identifier corresponding to a predefined category into which the content classifier logic unit classifies the content item to be parsed.

[0090] Example 8 is the computing system of any or all of the previous examples, wherein the detection and monitoring logic unit includes:

[0091] A performance data detection logic unit is configured to detect parser performance data indicative of parser latency and resource usage.

[0092] Example 9 is the computing system of any or all of the previous examples, and also includes:

[0093] A model generation logic unit is configured to generate a dynamically built model based on data related to the parser and based on the parser performance data, the dynamically built model being configured to output a parser deployment signal indicating the parser identified by the model to be deployed to the location.

[0094] Example 10 is the computing system of any or all of the previous examples, and also includes:

[0095] A pre-deployment logic unit is configured to deploy the parser identified by the model to a location.

[0096] Example 11 is a computer-implemented method comprising:

[0097] detecting, from a content understanding system, data associated with a parser that parses content items to generate a textual summary for each of the content items, the parser-associated data indicating a type of content parsed by the content understanding system, and the parser identifier indicating a user-selected parser for parsing the type of content;

[0098] receiving a content item to be parsed;

[0099] Access the collection of available parsers;

[0100] automatically selecting, based on the parser identifier, a parser among the available parsers for parsing the content item to be parsed; and

[0101] A control signal is generated to deploy the selected parser to parse the received content item to be parsed.

[0102] Example 12 is the computer-implemented method of any or all of the previous examples, and also includes:

[0103] A priority input is detected indicating a user-prioritized selection criterion, wherein automatically selecting a parser includes selecting a parser based on the prioritized selection criterion.

[0104] Example 13 is a computer-implemented method of any or all of the previous examples, wherein detecting data associated with the parser comprises:

[0105] An individual parser identifier is detected for each of a plurality of different corresponding content types, each individual parser identifier identifying a user-selected parser for parsing the corresponding content type.

[0106] Example 14 is the computer-implemented method of any or all of the previous examples, wherein the received content item to be parsed comprises a complex content item having a plurality of different content types, and wherein automatically selecting a parser comprises:

[0107] A plurality of different parsers are selected based on parser identifiers corresponding to the plurality of different content types.

[0108] Example 15 is a computer-implemented method of any or all of the previous examples, wherein automatically selecting a parser comprises:

[0109] A plurality of different parsers are selected based on a preferred selection criteria.

[0110] Example 16 is the computer-implemented method of any or all of the previous examples, and also includes:

[0111] The content items to be parsed are classified into predefined content type categories based on characteristics of the content items to be parsed, each predefined category having a corresponding parser identifier.

[0112] Example 17 is a computer-implemented method of any or all of the previous examples, wherein automatically selecting a parser comprises selecting a parser based on a parser identifier corresponding to a predefined category into which the content classifier logic unit classifies the content item to be parsed.

[0113] Example 18 is a computer-implemented method of any or all of the previous examples, wherein detecting data associated with the parser comprises:

[0114] Instruments parser performance data indicating parser latency and resource usage.

[0115] Example 19 is the computer-implemented method of any or all of the previous examples, and also includes:

[0116] generating a dynamically built model based on the data related to the parser and based on the parser performance data, the dynamically built model configured to output a parser deployment signal indicating the parser identified by the model to be deployed to the location; and

[0117] Deploy the resolver identified by the model to the location.

[0118] Example 20 is a computing system comprising:

[0119] a detection and monitoring logic unit that detects, from a content understanding system, data associated with a parser that parses content items to generate a textual summary of each of the content items, the parser-associated data indicating a type of content parsed by the content understanding system, and a parser identifier indicating a user-selected parser used to parse the content item;

[0120] a resolver tracking subsystem, which identifies available resolvers;

[0121] a content classifier logic unit configured to classify the content item to be parsed into predefined content type categories based on characteristics of the content item to be parsed, each predefined category having a corresponding parser identifier;

[0122] a user selection criteria detection logic unit configured to detect a priority input indicating a selection criteria to be prioritized;

[0123] a parser selection logic unit for: receiving a content item to be parsed; accessing a parser tracking system; and automatically selecting a parser from among available parsers for parsing the content item to be parsed based on a parser identifier, based on a predefined category and based on a prioritized selection criteria; and

[0124] A deployment subsystem that generates control signals to deploy the selected parser.

[0125] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. A computing system, include: at least one processor; as well as A memory storing instructions executable by the at least one processor, wherein when executed, the instructions cause the computing system to: receiving parser-related historical data representing previous parsing of a plurality of content items by a set of usable parsers, wherein a first content item of the plurality of content items has a text summary generated by a content understanding system based on the previous parsing of the first content item, wherein The historical data related to the parser includes: a content type identifier indicating a content type of the first content item parsed by the content understanding system; and a parser identifier indicative of a user-selected parser for parsing the first content item; Based on the historical data related to the parsers, identifying the set of parsers that can be used; for each parser in the set of usable parsers, detecting parser performance data indicative of parser latency and resource usage corresponding to previous parsing of the plurality of content items; receiving an indication of a second content item to be parsed by the content understanding system; automatically selecting a parser from the set of available parsers that is configured to parse the second content item based on the content type identifier, the parser identifier, and the detected parser performance data; and Control instructions are generated that control the content understanding system to parse the second content item using the selected parser.

2. The computing system according to claim 1, in, The instructions cause the computing system to: detecting a priority input indicating a selection criterion to be prioritized; and The parser is selected based on the prioritized selection criteria.

3. The computing system according to claim 2, in, The instructions cause the computing system to: An individual parser identifier is detected for each of a plurality of different corresponding content types, each individual parser identifier identifying a user-selected parser for parsing the corresponding content type.

4. The computing system according to claim 3, in, The second content item to be parsed comprises a complex content item having a plurality of different content types, and wherein the instructions cause the computing system to select a plurality of different parsers based on the parser identifiers corresponding to the plurality of different content types.

5. The computing system according to claim 4, in, The instructions cause the computing system to select the plurality of different parsers based on the prioritized selection criteria.

6. The computing system according to claim 1, in, The second content item includes content other than the first content item previously parsed by the content understanding system, and the second content item is received by the computing system after receiving the parser-related historical data from the content understanding system, and the instructions cause the computing system to: The second content item to be parsed is classified into a predefined content type category based on characteristics of the second content item to be parsed, the predefined category having a corresponding parser identifier.

7. The computing system according to claim 6, in, The instructions cause the computing system to select the parser based on the parser identifier corresponding to the predefined content type category.

8. The computing system according to claim 1, in, The instructions cause the computing system to: generating a dynamic build model based on the historical data related to the parser and based on the parser performance data, the dynamic build model configured to output a parser deployment signal indicating a selected parser and a selected location to deploy the selected parser; and Based on the resolver deployment signal, the selected resolver is deployed.

9. The computing system according to claim 8, in, The parser-related historical data includes position data indicating a given position at which the user-selected parser was deployed to parse the first content item, and the instructions cause the computing system to: based on the given location at which the user-selected parser is deployed to parse the first content item, selecting a location to deploy the selected parser; and Prior to receiving an indication of the second content item, the selected parser is pre-deployed to the selected location.

10. A computer-implemented method, include: receiving historical usage data associated with a parser, the historical usage data associated with the parser representing previous parsing of a plurality of content items by a set of usable parsers, wherein a first content item of the plurality of content items has a text summary generated by a content understanding system based on a previous parsing of the first content item, and the historical usage data associated with the parser comprises: a content type identifier indicating a content type of the first content item parsed by the content understanding system, and a parser identifier indicating a user-selected parser for parsing the first content item; Based on the historical usage data associated with the parsers, identifying the set of parsers that can be used; for each parser in the set of usable parsers, detecting parser performance data indicative of parser latency and resource usage corresponding to previous parsing of the plurality of content items; receiving an indication of a second content item to be parsed; automatically selecting a parser from the set of available parsers for parsing the second content item based on the content type identifier, the parser identifier, and the detected parser performance data; and Control instructions are generated that control the content understanding system to parse the second content item using the selected parser.

11. The computer-implemented method of claim 10, further comprising: include: A priority input is detected indicating a user-prioritized selection criterion, wherein automatically selecting the parser comprises selecting the parser based on the prioritized selection criterion.

12. The computer-implemented method of claim 11, in, Detection and parser-related data include: An individual parser identifier is detected for each of a plurality of different corresponding content types, each individual parser identifier identifying a user-selected parser for parsing the corresponding content type.

13. The computer-implemented method of claim 12, in, The second content item comprises a complex content item having a plurality of different content types, and wherein automatically selecting a parser comprises: A plurality of different parsers are selected based on the parser identifiers corresponding to the plurality of different content types.

14. A computing system, include: at least one processor; as well as A memory storing instructions executable by the at least one processor, wherein when executed, the instructions provide the following: A detection and monitoring logic unit is configured to detect parser-related historical data, the parser-related historical data representing previous parsing of a plurality of content items by a set of parsers, wherein the content understanding system is configured to generate a text summary for each of the plurality of content items based on the previous parsing, and the parser-related historical data includes: a content type identifier indicating a content type of the plurality of content items parsed by the content understanding system; and a parser identifier indicative of a user-selected parser for parsing the plurality of content items; and for: for each parser in the set of parsers, detecting parser performance data indicative of parser latency and resource usage corresponding to previous parsing of the plurality of content items; A parser tracking subsystem that identifies parsers that can be used; a content classifier logic unit configured to classify the content item to be parsed into a predefined content type category having a corresponding parser identifier based on characteristics of the content item to be parsed; a user selection criteria detection logic unit configured to detect a priority input from a user indicating a selection criteria that is prioritized; a parser selection logic unit configured to automatically select a parser from the set of parsers for parsing the content item to be parsed based on the parser identifier, based on the predefined content type category, based on the prioritized selection criteria, and based on the detected parser performance data; and A deployment subsystem generates a control signal to control the selected parser to parse the content item.

Citation Information

Patent Citations

  • System and method for the ingestion of industrial internet data

    US20180246944A1

Cited By

  • File content analysis and data management

    US20220309184A1