Information synthesis system for generating knowledge base, method for updating knowledge base using information synthesis system, and method for updating maintenance knowledge base
The tasks are distributed to mass workers and machines through information synthesis systems, and the use of templates and machine learning models to extract domain knowledge from unstructured data sources, solving the problem of relying on domain experts in the existing technology, and achieving efficient and accurate product diagnosis and maintenance.
Patent Information
- Application Number
- CN201911414154.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2018-12-28
- Filing Date
- 2019-12-27
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2039-12-27
AI Technical Summary
The prior art relies on field experts in product diagnosis and repair, which makes it costly and difficult to effectively utilize non-expert resources, especially in the absence of error codes to accurately diagnose and repair product problems.
Using an information synthesis system, the decomposition and efficient diagnostic repair of low-cognitive tasks are achieved by distributing tasks to mass workers and machine processing, and templates and machine learning models are used to extract domain knowledge from unstructured data sources, and the knowledge base is automatically verified and updated.
Reliance on field experts has been reduced, the efficiency and accuracy of diagnosis and maintenance have been improved, and non-expert resources can be used effectively to update the knowledge base dynamically to adapt to the diagnostic and maintenance needs of different fields.
Smart Images

Figure CN111400485B_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to a system for infusing domain knowledge and generating tasks for crowd-sourced and machine-sourced data excerpting. Background Art
[0002] Products may be made up of many components, some of which may be repairable and / or replaceable. Furthermore, products are becoming increasingly complex to operate. During operation, for various reasons, a product may not function as intended. For example, a component may become worn or damaged, resulting in improper or inoperative operation of the product. Some products may include self-diagnostic features. These features may cause the product to store and / or display an error code that may indicate a problem. In other cases, a product may exhibit a problem or symptom without storing an error code.
[0003] A typical process may include reading an error code and initiating repairs based on the error code. In some cases, an error code may indicate a number of issues. Further diagnosis and troubleshooting may be required to determine the source of the problem. In other cases, the problem / symptom may not have an associated error code. In some cases, the problem may not be understood through the manufacturer's diagnostic and repair procedures. Product diagnostic and repair documentation is typically generated by domain experts. For example, a manufacturer may task an expert with generating a product repair manual. The repair manual may include diagnostic and repair procedures. Summary of the Invention
[0004] An information synthesis system for generating a knowledge base includes: a computing system programmed to: distribute a template including a task for extracting information from an unstructured source to a task executor, receive a task result from the task executor as a response in the template, identify domain knowledge representations that appear in the task result but are missing from the knowledge base, and generate a template that defines the task and includes the domain knowledge representations for extracting additional information from the unstructured source.
[0005] Tasks can be defined as human-only tasks, machine tasks, and machine-guided tasks. The computing system can be further programmed to distribute templates based on the availability of task performers and the accuracy of the task performers. Tasks can include extracting information from unstructured data sources. Tasks can include extracting information contained in at least a portion of a video. The computing system can be further programmed to verify the task results received from each of the task performers and identify the task results as invalid in response to the task completion time being less than a predetermined percentage of the duration of the portion of the video. The computing system can be further programmed to verify the task results from each of the task performers and identify the task results as invalid in response to the following: (i) the task results are the same for a predetermined number of responses; (ii) the task results include terms identifying components missing from the original source corresponding to the task results; and (iii) the task results are unique compared to task results submitted by other task performers. The computing system can be further programmed to maintain a data chain for the tasks, the data chain including, for each of the tasks: data defining the original source: the relevant portion of the original source; an excerpt of the relevant portion; and a final summary derived from the excerpt. The computing system can be further programmed to facilitate training of the machine learning model by providing the data chain to one or more machine learning models as training input. The computing system can be further programmed to predict the accuracy of the task performer before distributing the template.
[0006] A method for updating a knowledge base by a computing system includes: maintaining a penalty score for task performers; and dispatching a task to the task performer in response to the penalty score being less than a predetermined threshold. The method further includes: increasing the penalty score for the task performer to a value greater than the predetermined threshold in response to the task performer providing more than a predetermined number of responses to the task containing a representation of a specific domain that does not appear in an original source associated with the task.
[0007] The method may further include: in response to receiving more than a predetermined number of identical responses from the task executor for different tasks, increasing the penalty score for the task executor to a value greater than a predetermined threshold. The method may further include: in response to the task executor providing more than a predetermined number of responses to the same task that are unique compared to responses submitted by other task executors, increasing the penalty score for the task executor to a value greater than a predetermined threshold. The method may further include: in response to the task executor completing the video excerpt task in a time that is less than a predetermined percentage of the running time of the allocated video segment, increasing the penalty score for the task executor. The method may further include invalidating responses that contribute to the penalty score exceeding the predetermined threshold.
[0008] A method for synthesizing information from unstructured data sources to update a maintenance knowledge base includes: identifying relevant portions of an original source having domain-specific knowledge related to the maintenance knowledge base; and creating a template including a task for extracting each of the relevant portions. The method further includes: distributing the template to task performers based on their availability and accuracy; and aggregating solutions from the template completed by the task performers to create a maintenance solution described as an action verb followed by a component name. The method further includes: updating the maintenance knowledge base with domain-specific representations from the maintenance solutions that are not present in the maintenance knowledge base; and creating and distributing new templates based on the domain-specific representations.
[0009] The method may further include creating a new machine learning model using the original source, the relevant portion, the summary, and the repair solution as training data for updating the machine learning model. The repair solution may be described as an action verb followed by a component name. The original source may be a document accessed on a website. The original source may be a video accessed on a website. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 Describes possible configurations for an information synthesis system for developing a knowledge base.
[0011] Figure 2 Depicts a possible block diagram for a process for an information synthesis system.
[0012] Figure 3A and Figure 3B Depicts possible display output for a first example template.
[0013] Figure 4A and Figure 4B Depicts possible display output for the second example template.
[0014] Figure 5B and Figure 5B Depicts possible display output for the third example template.
[0015] Figure 6 Depicts possible display output for the fourth example template.
[0016] Figure 7 Depicts possible display output for the fifth example template.
[0017] Figure 8 Depicts possible display output for a sixth example template.
[0018] Figure 9 Depicts possible display output for a seventh example template.
[0019] Figure 10Depicts possible display output for an eighth example template.
[0020] Figure 11 Describes a possible sequence of actions for an abstract workflow.
[0021] Figure 12 A block diagram depicting the types of tasks that can be managed by a task configurator.
[0022] Figure 13 Depicting a block diagram for maintaining trajectory data for use in updating a knowledge base.
[0023] Figure 14 Describes a possible sequence of operations for implementing an information synthesis system. DETAILED DESCRIPTION
[0024] Embodiments of the present disclosure are described herein. However, it is to be understood that the disclosed embodiments are merely examples, and that other embodiments may take a variety of forms and alternative forms. The figures are not necessarily to scale; some features may be magnified or minimized to show details of particular components. Therefore, the specific structural and functional details disclosed herein should not be interpreted as limiting, but merely as a representative basis for teaching those skilled in the art to employ the present invention in various ways. As will be understood by those of ordinary skill in the art, the various features illustrated and described with reference to any one of the figures may be combined with features illustrated in one or more other figures to produce embodiments that are not explicitly illustrated or described. The combination of illustrated features provides representative embodiments for typical applications. However, for specific applications or implementations, various combinations and modifications of features consistent with the teachings of the present disclosure may be desirable.
[0025] Existing methods for generating diagnostic and repair instructions have generally relied on experts in a specific field. For example, a knowledge base for automotive repairs may rely on the expertise of trained mechanics or others with similar domain knowledge. Domain expertise is advantageous because experts understand the standard terminology and procedures typical of the domain. However, such approaches have the disadvantage of requiring different experts for each domain and may not necessarily leverage the efforts of non-experts. Using non-experts in these types of activities can reduce costs. Several benefits can be realized by applying crowdsourcing models to generate diagnostic and repair knowledge bases.
[0026] Crowdsourcing systems can be used to complete high-cognitive tasks (e.g., understanding the meaning of a photo or written text) by breaking them down into relatively low-cognitive tasks that can be easily completed by ordinary human workers without specialized skills (e.g., without extensive domain knowledge). Numerous workers can be registered in crowdsourcing marketplaces (such as Amazon Mechanical Turk). Various design patterns (e.g., the find-extract-verify model) and quality control mechanisms (e.g., majority voting, gold standard injection) are available to maintain consistent quality across the system by minimizing the variation in the quality of task results from the crowd workers.
[0027] Disclosed herein are systems and methods for injecting domain knowledge into a semi-crowdsourced excerpting pipeline for diagnostic and repair knowledge from unstructured data sources (such as forums on the web) and / or multimedia information (such as videos). The excerpting pipeline can utilize a chain of tasks performed by crowd workers via a user interface and by machine processing via software operations. The system can automatically distribute tasks to task executors and aggregate processing results from the task executors. The task executors can be crowd workers and / or machines. Tasks can include: extracting symptoms and errors described by humans in an excerpt template or detected by a diagnostic tool, searching for relevant information sources, extracting the most relevant parts from the selected source and excerpting the extracted information in a template, and grouping similar information into the same or similar solutions.
[0028] Figure 1 An information synthesis system (ISS) 100 is depicted, which can be configured to generate a knowledge base for a given domain. ISS 100 can be configured to generate a knowledge base for use in diagnosing and repairing products or systems. ISS 100 may include at least one computing system 102. Computing system 102 may also be referred to as a knowledge base server. Computing system 102 may include at least one microprocessor unit 104 configured to execute instructions. Computing system 102 may include volatile memory 106 and non-volatile memory 108 for storing instructions and data. Computing system 102 may include a network interface 110 configured to provide communication with a network router 111. For example, network router 111 may be a wired or wireless Ethernet router. In some configurations, network router 111 establishes a local network to connect to one or more local servers 126. Network router 111 may further be configured to provide a communication interface to an external network 116. In some configurations, computing system 102 may exist as a remote server in a cloud computing architecture, such as Amazon Web Services (AWS).
[0029] The external network 116 may be referred to as the World Wide Web or the Internet. The external network 116 may establish a standard communication protocol between computing devices. The external network 116 may allow information and data to be easily exchanged between computing devices and the network. At least one server 120 may communicate with the external network 116. Each server 120 may host a website or web page from which information may be derived. For example, a server 120 may host a web page with information related to a field of interest. There may be many such servers 120 with information. The server 120 may host one or more unstructured data sources that provide data in varying formats, including blogs, forums, articles, images, audio, and / or video. The data may be considered unstructured because there may not be a common format between the sources. For example, each website may be arranged differently. Field-related information may be repeated on different web pages / websites.
[0030] One or more task processing machines 118 (e.g., 118A, 118B) can communicate with the external network 116. Task processing machines 118 (e.g., 118C) can also communicate with the local network established by the network router 111. Task processing machines 118 can be configured to execute received tasks or programs. Task processing machines 118 can be computing systems programmed to receive programs / instructions and data, process the data according to the programs / instructions, and output the processed data. Computing system 102 can also perform the functions of task processing machine 102. That is, computing system 102 can be configured to execute tasks and programs generated by the system.
[0031] Crowd workers 121 can utilize workstations 122 to access external networks 116. Crowd workers 121 may not be expected to have any domain expertise. Crowd workers 121 can be registered in one or more crowdsourcing marketplaces (such as Amazon Mechanical Turk). Crowdsourcing marketplaces can be implemented on one of servers 120. Crowdsourcing marketplaces can allow task requesters to upload tasks for completion by crowd workers 121. Crowd workers 121 can access crowdsourcing marketplaces using workstations 122. Workstations 122 can be personal computing devices that include a user interface for input and output. For example, workstations 122 can be computers with a display and a keyboard. Workstations 122 can include tablets and cell phones.
[0032] Computing system 102 can be directed and managed by one or more administrator users or system administrators 124 through a terminal / workstation / user interface 114 in communication with or coupled to computing system 102. User interface 114 can be configured to allow system administrator 124 to access and change information and programs in computing system 102. Computing system 102 can communicate with a knowledge base 112. Knowledge base 112 can represent domain-specific knowledge collected and organized by ISS 100. Knowledge base 112 can be configured to store domain knowledge and representations. Computing system 102 can be programmed with a reasoning engine that applies rules and logic to find information stored in knowledge base 112. For example, knowledge base 112 can be information related to diagnosing and repairing a particular product. Knowledge base 112 can reside on a physical storage device and / or can reside in memory within computing system 102. Computing system 102 can be programmed to access the information contained in knowledge base 112 and present it to users and administrators. For example, the information contained in the knowledge base 112 may be accessible via a web interface, such that external users may access the information via the external network 116. Access may be general (e.g., available to everyone) or may be restricted (e.g., limited to specific individuals such as registered product repair experts).
[0033] One or more domain experts 130 may utilize a workstation 132 to access the external network 116. A domain expert 130 may be a person with knowledge in a particular domain. The domain expert 130 may provide domain expertise to ensure that the knowledge base 112 includes an appropriate level of domain knowledge.
[0034] The computing system 102 can be configured to build or generate a knowledge base 112 using information found on the server 120. In some configurations, the computing system 102 can be programmed to perform searches for domain-related information via the external network 116. The searches can be directed by search terms entered by a system administrator 124 or retrieved from the knowledge base 112. The system administrator 124 can inject domain knowledge (such as terms of art in a related field) to guide the searches. The searches can result in one or more websites or uniform resource locators (URLs) / web addresses containing information related to the search terms. Once potential sources of domain-related data are identified, the sources can be examined to determine whether they provide relevant information that can improve the knowledge base 112.
[0035] The computing system 102 can be programmed to facilitate the generation of tasks that can be assigned to task performers (e.g., crowd workers 121 and task processing machines 118) for completion. Tasks can define workflows or pipelines for processing and synthesizing information. Tasks can be configured to generate summaries of information sources related to the domain for which knowledge is being collected. The summarization pipeline can be adapted to predict the accuracy of task performers (e.g., crowd workers 121 and / or machines 118) and can be configured to change over time to improve the quality of the solution. Tasks can be structured so that the crowd workers 121 do not require extensive domain-specific knowledge to complete the task.
[0036] New data discovered / learned throughout the pipeline can be used to update the knowledge base 112 and to support and design additional tasks. An example of information stored in the knowledge base 112 can be a dictionary consisting of equipment or vehicle components, which includes representative terms, synonyms, acronyms, multimedia content, along with different attributes / relationships between terms such as product information, symptom descriptions, and repair solutions. The knowledge base 112 can store information about error codes, symptom descriptions, and related information such as diagnostics and repair instructions.
[0037] The computing system 102 can implement various computational methods for information synthesis. Question answering (QA) research focuses on methods and systems for automatically answering queries posted by humans in natural language. In addition to factoid and list QA, complex, interactive QA (ciQA) can also be utilized. Semi-automated QA methods (and their crowd-based variants) can focus on answering short factual queries rather than completing complex meaning-understanding processes.
[0038] Multi-document excerpting aims to use computational techniques to extract information from multiple texts written on the same topic using feature-based, clustering-based, graph-based, and knowledge-based approaches. However, such approaches have limitations in coping with the complex, yet brief, and sparse data that may be encountered on the web, and do not engage in the complex synthesis that humans perform cognitively to achieve cohesive and coherent output.
[0039] While crowdsourcing has been shown to be effective, crowdsourcing systems have not yet been systematically adapted for different domains. Crowdsourcing can be more effective when domain knowledge is injected into the crowdsourcing system. By injecting domain knowledge into the structure when needed, the crowdsourcing system does not require domain experts to complete tasks. The systems and methods disclosed herein generally relate to the design of human cognitive tasks for crowdsourcing, machine processing and extraction of multi-source unstructured data in text or multimedia form, and the use of domain knowledge representation in system design to assign human cognitive tasks and machine learning tasks that result in high-quality diagnostic and repair knowledge in terms of efficiency and accuracy.
[0040] Server 120 may host content related to a related domain. For example, there may be websites and forums for appliance or vehicle repair that include detailed information about various repair and diagnostic techniques. The content may be unstructured, as there is no formal organization of the information. Knowledge base 112 may be enhanced by incorporating information related to products or types of products from server 120. Some techniques may be described by hosted content for similar products, and these techniques may be adaptable to the product for which knowledge base 112 is being synthesized.
[0041] The systems and methods disclosed herein can be applied to applications that provide domain-specific diagnostic and repair knowledge for products such as automobiles, heating systems, and household appliances. The diagnostic and repair knowledge can be made accessible on a website or via a diagnostic tool (e.g., a scan tool used by a technician in a motor vehicle workshop). Disclosed herein are systems and methods adapted for designing human cognitive tasks targeting a mass worker 121 that does not necessarily possess extensive domain knowledge, and designing machine processing / learning tasks with domain knowledge representations to synthesize complex information for diagnostic and repair knowledge for a specific domain of interest. The disclosed systems and methods can decompose information synthesis tasks into multiple microtasks for the mass worker 121 and the task processing machine 118. The decomposition can be configurable and / or dynamic. The systems and methods can be configured to update domain knowledge representations learned from earlier mass worker outputs and machine processing outputs.
[0042] Disclosed herein are systems and methods for incorporating domain knowledge into the design of a semi-crowd-sourced ISS 100. The ISS 100 can be configured to output a knowledge base 112 for use in diagnosing and repairing a product. The ISS 100 can be configured to create or generate tasks that can be performed by humans 121 or machines 118. The generated tasks can be low-cognitive tasks for a crowd worker 121 who may not have prior knowledge of repair and diagnosis related to a specific domain. The generation of low-cognitive tasks allows for execution by a wider pool of crowd workers 121, as specialized domain knowledge is not necessary to complete the tasks.
[0043] The ISS 100 may be integrated with automated information processing capabilities. That is, the ISS 100 may be implemented on a computing system 102 that may be programmed to automatically process and generate information.
[0044] Figure 2 A block diagram depicts features or processes that may be implemented as part of computing system 102. The processes described may interact with each other. Further, the processes may be updated by system administrator 124. The processes may be stored in memory of computing system 102 and executed periodically and / or when there is demand for execution (e.g., available input and / or desired output). Computing system 102 may implement an operating system to manage task execution and sequencing.
[0045] The computing system 102 may include a task generation process 206. The task generation process 206 may be configured to define and generate tasks to be completed to update the knowledge base 112. A task may define specific data or knowledge to be sought from the citizen workers 121 or the task execution machines 118. A task may be defined as a request to execute specific instructions to return information or data. A task may further define the format of the data to be returned. Tasks may be defined by the system administrator 124 based on the type of information being sought.
[0046] The computing system 102 may include a template definition / generation process 204. The ISS 100 may define one or more templates that include tasks for extracting information from text-based diagnostic and repair knowledge from humans 121 and / or machines 118. In some adaptations, the templates may be designed by a system designer or administrator 124. In some adaptations, the templates may be automatically generated by a machine (e.g., a local server 126). The template definition / generation process 204 may be programmed to facilitate the development of templates. The templates and information related to the templates may be stored in a template database 202. The template definition / generation process 204 may be programmed to automatically generate templates. For example, a template may include configurable fields related to a specific domain. The template definition / generation process 204 may be programmed to insert domain-specific values derived from the knowledge base 112 into the configurable fields. The template may include one or more tasks to be completed by a task executor. For example, a task may be a specific query posed by the template.
[0047] Templates can be designed to subdivide larger tasks into smaller, more manageable tasks that can be performed by citizen workers 121. Furthermore, tasks can be configured to generate low-cognitive tasks that can be completed by citizen workers 121 without detailed expert knowledge. For example, a task can be directed toward extracting or describing information found in an assigned source. Tasks can take the form of short queries or multiple-choice queries.
[0048] The template may include information and / or instructions for guiding the Volkswagen worker 121 in how to complete one or more tasks. The template may include features for collecting information (such as a symptom summary). The symptom summary may include data such as the occurrence, location, and / or related conditions associated with the problem. The template may be configured to extract the relevant product brand, model, model year, and other properties that identify the product (such as engine type or fuel type). The template may be configured to extract a repair solution. For example, a repair solution may be formulated as <action verb> followed by <component of the object being repaired>. The template may be configured to utilize domain-specific terminology and representations.
[0049] Figure 3A and Figure 3B Depicts a possible display output 300 for a first example template 302. The first example template 302 can define a text-based interface. In some configurations, the first example template 302 (and subsequent examples) can be generated or defined as a Hypertext Markup Language (HTML) document (e.g., a web page). The first example template 302 can be displayed on a user interface or display of a workstation 122 of a citizen worker 121 assigned to complete the first example template 302. The first example template 302 can include an instruction portion 304 for providing instructions to the citizen worker 121. In some applications, the instruction portion 304 can include one or more links to websites or documents to be browsed and processed by the citizen worker 121. The instruction portion 304 can also provide information to assist the citizen worker 121 in completing the task.
[0050] The first example template 302 may include one or more queries 306. A query 306 may be a specific inquiry for specific information, such as a symptom or condition. A query 306 may incorporate domain-specific representations derived from previous task responses. The query 306 posed may depend on responses from previous templates.
[0051] The first example template 302 may include one or more answer or input sections 308 for receiving input from the crowd worker 121. The first example template 302 may be configured to pose specific queries that the crowd worker 121 is expected to answer. The first example template 302 may be configured to receive textual input from the crowd worker 121 via the input section 308. For example, the input section 308 may include a field or box for the crowd worker 121 to type text into. The input section 308 may also include predefined selection boxes that allow the crowd worker 121 to browse a list of entries and select one or more of them for input. The first example template 302 may also include a multiple-choice section 310, which may pose specific queries with corresponding checkboxes that the crowd worker 121 may select in response. The query 306 may be a yes / no query and / or a multiple-choice query. Upon completion, the first example template 302 may be submitted to the computing system 102 for further review and processing. The computing system 102 may be configured to automatically process the response or to store the response for later review by the system administrator 124. The template may pose a query 306 to identify whether the information relates to a specific make, model, model year, and / or other properties of the product, such as engine type or fuel type. The template may pose a query 306 intended to identify a specific component.
[0052] Figure 4A and Figure 4B A second display output 400 for a second example template 402 is depicted. The second example template 402 can be displayed on a user interface or display of a workstation 122 of a mass worker 121 assigned to complete the second example template 402. The second example template 402 can define an instruction portion 403. The instruction portion 403 can define information related to completing the task associated with the second example template 402. For example, the second example template 402 can include a problem definition, specific product information, and specific instructions regarding the expected output or response. The second example template 402 can request that the task performer search for web pages related to the defined problem. The second example template 402 can further provide a list of web pages that have already been submitted to reduce the chance of duplicate search results being submitted.
[0053] The second example template 402 may further include a specific request 404. For example, the specific request 404 may be for a URL resulting from a requested search. The second example template 402 may further include a specific request response field 406 configured to allow the task performer to type or paste in a response to the specific request 404. The specific request response field 406 may be configured to accept text input from the workstation 122. Additional request / response fields may be defined. For example, the additional request / response fields may be configured to elicit search terms used by the task performer.
[0054] Figure 5A and Figure 5B Depicted is a third display output 500 for a third example template 502. The third example template 502 can be displayed on a user interface or display of a workstation 122 of a mass worker 121 assigned to complete the third example template 502. The third example template 502 can include one or more multiple-choice sections 504. The multiple-choice section 504 can recite a query or statement followed by a plurality of possible answers with corresponding check boxes or circles. The task performer can select one or more answers to apply to the query. For example, the query can be a specific query about a website (e.g., directed toward a specific product), or can be a specific query about a specific error code or problem. In some template definitions, the answer can be "yes" or "no."
[0055] The third example template 502 may include an instruction statement 505 that provides instructions to the task performer. The instruction statement 505 may include specific domain knowledge or representation. For example, a diagnostic and maintenance application may reference a specific error code to guide the task performer's response. The third example template 502 may include a response field 506 that is used to insert text or an image in response to the instruction statement 505. The response field 506 may be configured to accept pasted text or images. For example, the instruction statement 505 may request the task performer to copy information related to the solution to the specific error code suggested in the reference into the response field 506.
[0056] The third example template 502 may include a snippet request 508. The snippet request 508 may instruct the task performer to extract a previous response in a specific format. For example, the snippet request may instruct the task performer to state an action verb followed by a noun (e.g., a component name). The third example template 502 may define an action verb input field 510 and a component name input field 512, which allow text to be entered in response to the snippet request 508. In diagnostic and repair applications, the third example template 502 may further request information regarding whether the action identified by the response to the snippet request 508 has been confirmed to have resolved or repaired the associated problem.
[0057] Figure 6A fourth display output 600 for a fourth example template 602 is depicted. The fourth example template 602 can be displayed on a user interface or display of a workstation 122 of a mass worker 121 assigned to complete the fourth example template 602. The fourth example template 602 can be configured to present a plurality of solution statements 606. The fourth example template 602 can include an instruction field 604 for providing instructions to a task performer for completing the task. For example, the solution statement 606 can present a list of action verb / component combinations along with a checkbox to confirm or delete each combination in the list. The instruction field 604 can instruct the task performer to select similar combinations. In some configurations, the solution statement 606 can include at least one solution that is unrelated to the other presented solutions. The unrelated solution can be automatically inserted by the computing system 102 or manually inserted by the system administrator 124. Inserting the unrelated solution can help ensure that the task performer is performing the task accurately. For example, if the task performer selects an unrelated solution, the penalty score associated with the task performer can be increased.
[0058] Figure 7 A fifth display output 700 is depicted for a fifth example template 702. The fifth example template 702 can be displayed on a user interface or display of a workstation 122 of a crowd worker 121 assigned to complete the fifth example template 702. The fifth example template 702 can be a subsequent task associated with the fourth example template 602. For example, the fifth example template 702 can be displayed in response to selecting a "Next" button in the fourth example template 602.
[0059] The fifth example template 702 may include a selected solution summary 706 that lists the solution selected from the fourth example template 602. The fifth example template 702 may include an instruction section 704 to provide instructions to the task performer for completing the task. For example, the instruction section 704 may instruct the task performer to generate a title for the solution in the form of an action verb followed by a component name. The fifth example template 702 may include a suggestion field 710 that presents useful information to the task performer. For example, the suggestion field 710 may include a list of the most frequently used action verbs. Frequently used action verbs may be obtained from the knowledge base 112, and the suggestions may change over time as the knowledge base 112 is updated. The suggestion field 710 may help ensure consistency in the information presented by the knowledge base 112. The fifth example template 702 may include one or more input fields 708 configured to receive data inserted by the task performer. For example, the input field 708 may provide a field for entering an action verb and a component name.
[0060] Figure 8A sixth display output 800 for a sixth example template 802 is depicted. The sixth example template 802 can be displayed on a user interface or display of a workstation 122 of a mass worker 121 assigned to complete the sixth example template 802. The sixth example template 802 can include an instruction field 804 for providing instructions to a task performer for completing the task. The sixth example template 802 can include an image comparison field 806 that displays one or more images. The images can include different views of one or more components. The images can be derived from the knowledge base 112 as images associated with a particular component. The images can have been derived from the performance of a previous task. For example, the instruction field 804 can present instructions for comparing one or more sets of images. In this example, images of a low-pressure fuel level sensor and a fuel pump are displayed, and the task performer is asked whether the images represent the same component. The sixth example template 802 can include selection buttons 808 (e.g., "Yes" and "No") that can be used to indicate or record the task performer's response.
[0061] The template definition / generation process 204 can be configured to define one or more templates for extracting information from video-based diagnostic and repair knowledge from human workers 121 and / or machines 118. The video can be divided into multiple smaller segments (clips) for processing. The video-based templates can include fields for extracting the start time of the video segment, the end time of the video segment, and a brief (e.g., one-line) description of the video segment.
[0062] Some tasks / templates may be configured for processing by domain experts 130. Domain experts 130 may be able to inject domain knowledge into the task pipeline to ensure that the knowledge base 112 contains relevant information. In addition, domain experts 130 may filter and combine results from crowd workers 121. Figure 9A seventh display output 900 depicts a seventh example template 902. The seventh example template 902 may be displayed on a user interface or display of a workstation 122 of a domain expert 130 assigned to complete the seventh example template 902. The seventh example template 902 may include a link 904 to an information source (e.g., a website / webpage). For example, the information may be a forum related to the product. The domain expert 130 may be prompted to merge previously identified solutions from previous task executions. The seventh example template 902 may include an add solution interface, including an add solution button 912 and a solution entry field 908, which allows a reviewer to enter new instructions or descriptions. The seventh example template 902 may include a selection for a merge interface 910 (e.g., a virtual button), which allows the task performer to select items to be merged. For example, solutions may be identified as "Replace fuel pressure sensor," "Replace low pressure level fuel sensor," and "Replace sensor." The task performer may recognize that these solutions are identical and select them for merging into a single solution. The seventh example template 902 may include options to delete the solution (eg, move to trash) or show additional information related to the solution (eg, show a clip).
[0063] Figure 10 An eighth display output 1000 depicts an eighth example template 1002. The eighth example template 1002 can be displayed on a user interface or display of a workstation 122 of a human worker 121 assigned to complete the eighth example template 1002. The eighth example template 1002 can include an embedded video 1008 or a link to a video to be processed by the human worker 121 or machine 118. For example, the video 1008 can be an instructional video including step-by-step repair or diagnostic instructions for a particular product or system. The instruction section 1004 can prompt the human worker 121 to segment the video 1008 according to the instruction steps or logical sections defined by the video 1008. The eighth example template 1002 can include a new instruction interface 1006 or a button allowing the task performer to enter new instructions or descriptions into a newly created description field 1010. The eighth example template 1002 can be configured to record the elapsed time within the video 1008 when the new instruction interface 1006 is selected to associate the elapsed time with the new instruction. In this way, the video 1008 can be excerpted and the instruction steps can be documented.When completed, the template can be submitted to the computing system 102 for further review and processing.
[0064] Refer again Figure 2, the computing system 102 may include a task configurator process 208. The task configurator process 208 may be configured to distribute tasks and / or templates to one or more task executors. The task executors may include the crowd workers 121 and the task processing machines 118. The task configurator process 208 may be further configured to automate the distribution of tasks with respect to the accuracy and capabilities of the crowd workers 121 and the task processing machines 118. The task configurator process 208 may be configurable to distribute differently designed tasks that process the same input and produce the same output format. Task types may include human-only tasks, machine-guided human tasks, and machine-only tasks (see Figure 12 ). The system may be configurable to perform multiple tasks using a single type in parallel or sequentially, or to perform multiple tasks using a combination of different task types in parallel or sequentially.
[0065] Figure 12 Depicts the task configurator process 1202 and examples of different types of tasks that may appear. Human-only tasks 1204 may be defined that are intended to be assigned to and performed by one or more of the crowd workers 121 and / or domain experts 130. For example, a task in a workflow may be defined to group similar information from multiple sources into a single group, and provide a title for the group in the template (e.g., "Action Verb," "Component Name"). A task may be defined as a human-only task that provides a graphical user interface to a human (e.g., crowd worker 121) requesting the selection of similar information from a complete list and then requesting the insertion of a title for the grouped sentences. Some tasks may require processing by a domain expert 130. Different types of tasks may be distinguished by the level of domain knowledge required. The task configurator process 1202 may determine the level of knowledge and assign tasks accordingly.
[0066] Machine-only tasks 1208 may be defined that are intended to be assigned to and performed by one or more of the task processing machines 118. For example, a task may be defined as a machine-only task where the machine processing is designed to receive a number of sentences or phrases and group the sentences based on a prediction accuracy score.
[0067] Machine-directed human tasks 1206 may be defined that are intended to be assigned to and performed by a combination of a crowd worker 121 and a task processing machine 118. For example, a task may be defined as a machine-directed task in which the task processing machine 118 processes the initial grouping of sentences, and the crowd worker 121 then verifies the results of the machine processing and provides a caption by interacting with the system via a graphical user interface.
[0068] The system's tasks can be pluggable or interchangeable, allowing tasks of different types (e.g., human, machine, machine-guided) to be swapped for one another. Different types of tasks can be defined as receiving one or more inputs 1210 and generating one or more outputs 1212 using a common format. For interchangeable tasks, the inputs 1210 for each task can be identical, and the outputs 1212 can be identical. The task configurator process 1202 can then select the task executor for each task. For example, a new machine-guided task or a machine-only task can replace an existing human-only task without disrupting the workflow. Swapping between different types can be determined dynamically. The task configurator process 1202 can be configured to dynamically switch between task types depending on the availability and accuracy of information processed by machines or human workers. The task configurator process 1202 can assign tasks based on the availability of each task type. For example, some tasks may exist as human-only tasks 1204 without corresponding machine-only tasks 1208. Over time, additional machine-only tasks may be developed and easily inserted into the workflow. Tasks can be assigned to each type of task performer to compare outputs. This can be used to validate the outputs of different task performers. It can be used in workflows when validating tasks. Task configurator process 1202 can assign tasks to the most efficient task performer (e.g., the fastest task performer that provides a specific level of accuracy).
[0069] The task configurator process 1202 can implement a probability model to predict the accuracy of each task type before the task is dispatched and a probability model to verify the output from each task executed. The task configurator process 1202 can determine the workflow of the tasks (e.g., the dispatch of tasks) based on the predictions from the model to optimize quality and efficiency. The task configuration and selection can also be changed manually by the system administrator 124 or automatically based on rules / algorithms programmed into the task configurator process 1202.
[0070] Refer again Figure 2 , the computing system 102 may include a solution aggregation process 210. The solution aggregation process 210 may be configured to aggregate solutions received from task executors. A solution may be a response provided by a task executor in a template. For example, data inserted into a template by a task executor may be processed and compared with the output of similar tasks or templates. Some tasks may be sent to multiple task executors for completion. In some cases, similar tasks may be generated and assigned to ensure that solutions are consistent across task executors. The solution aggregation process 210 may compare solutions to determine whether additional tasks should be generated to combine or further validate the solution.
[0071] Computing system 102 may include a knowledge base update process 212 configured to update information in knowledge base 112 using solutions provided by task performers. For example, knowledge base 112 may be configured to include a library of domain-specific terminology or domain-specific representations. Components may be described by different labels or descriptions in different sources. It may be useful to capture various representations in the usage of each component. For example, an oxygen sensor may appear as "O2 sensor" or "Lambda sensor" in different sources. Knowledge base update process 212 may be configured to update knowledge base 112 using different domain knowledge representations. Awareness of different representations can result in more relevant search results and assist in discovering additional sources for data mining. Knowledge base update process 212 may implement a natural language processing algorithm to compare knowledge representations from knowledge base 112 with knowledge representations provided in solutions. Knowledge base update process 212 may be programmed to identify two compound nouns as potential synonyms of each other in response to nouns meeting predetermined criteria. The criteria may include an image search engine providing more than a predetermined number of public URLs in response to searches using different terms. Criteria may include that the nouns match each other exactly except for the ending word and the ending word that represents a set. Criteria may include that all unigrams of the two nouns are the same. Criteria may include that the unigram of one noun is a subset of the other noun and the ending words of the two nouns are the same and are common words. One or more tasks may be designed to compare terms with presented images (e.g. Figure 8 ) to facilitate the processing of identifying knowledge representations in different domains.
[0072] The computing system 102 may include a knowledge base source identification process 214. The knowledge base source identification process 214 may be configured to identify source material related to the knowledge base 112. The knowledge base source identification process 214 may be configured to monitor solutions provided by task performers for additional domain knowledge representations. For example, the domain knowledge representation may suggest additional search terms to identify web-based sources that include updated domain knowledge representations. The knowledge base source identification process 214 may be automated and / or directed by the system administrator 124.
[0073] The computing system 102 may include a quality control process 216 configured to evaluate and predict the quality of solutions and task performers. The quality control process 216 may implement methods for improving the quality control of tasks performed by the crowd workers 121 by injecting domain knowledge. Various strategies may be implemented, such as gold standards, majority voting, and prediction. The quality control process 216 may be configured to evaluate the accuracy of a given solution and the accuracy of the task performers. The quality control process 216 may also monitor the timeliness of the task performers in providing solutions. Timeliness may be determined as the time from task assignment to task completion. The quality control process 216 may be configured to verify task results and responses. The quality control process 216 may implement strategies for predicting the accuracy of task performers before distributing tasks and templates.
[0074] The quality control process 216 can be configured to manage the quality of the crowd workers 121 by implementing a worker profile management feature. The quality of each of the crowd workers 121 can be dynamically calculated during task estimation or review. The quality profile of the crowd worker 121 can include an overall penalty score for all task types and individual task penalty scores for each task type. The penalty score can be calculated based on the number of tasks with correct answers across the total number of task submissions, the correctness of the gold standard task, the consistency rate of answers for the same task among different crowd workers 121, and a sample estimate of the task. Based on the penalty score, the task configurator process 208 can be configured to temporarily or permanently prevent the assignment of a task to a worker.
[0075] The quality control process 216 can be configured to maintain a penalty score for each of the task performers. The penalty score can be used by the task configurator process 208 to assign tasks to the task performers. The task configurator process 208 can be programmed to distribute tasks to the task performers in response to the corresponding penalty score being less than a predetermined penalty threshold. The predetermined penalty threshold can be a value indicating that the task performer provided quality output. A penalty score exceeding the predetermined penalty threshold can indicate poor quality output by the task performer. The task configurator process 208 can be programmed to stop or prevent tasks from being distributed to task performers having a corresponding penalty score exceeding the predetermined penalty threshold. The quality control process 216 can increase or decrease the penalty score corresponding to each of the task performers.
[0076] Quality can be assessed through quality control mechanisms based on collaboration between people and machines. Quality can be assessed by validating human input using machine learning models with domain knowledge. In situations where quality prediction may be difficult, machine learning models can be defined as virtual crowd workers for consensus.
[0077] The computing system 102 may implement a method for presenting a ranked list of repair solutions based on source accuracy predictions and accuracy predictions of outputs from task performers (citizen workers, machines, domain experts). For example, solutions from task performers with higher predicted accuracy values may be placed more prominently in the list. The quality control process 216 may calculate source accuracy predictions based on: proposed solutions from laypeople, proposed solutions from experts, confirmation that the proposed solution fixes the problem, proposed solutions based on guesswork, proposed solutions based on solutions that fix earlier occurrences of the same or similar problem, and the number (incidence) of redundant solutions from multiple information sources.
[0078] Quality control process 216 can further implement a fraudster identification strategy. A fraudster may be a worker who is not performing properly. For example, a fraudster may be motivated to minimize the amount of work performed while maximizing the amount of compensation received. A fraudster may intentionally submit inaccurate information and not spend an appropriate amount of time on the assigned task. Quality control process 216 can be programmed to identify fraudsters and stop assigning tasks to identified fraudsters.
[0079] The quality control process 216 can implement a method for identifying fraudsters among Volkswagen workers for repair and diagnostic excerpts. For example, task solutions for text excerpts can be monitored to identify fraudsters. In response to receiving more than a predetermined number of responses from a task performer that are identical for different tasks (e.g., a Volkswagen worker provides the same solution to multiple repair inquiries), a penalty score for the task performer can be increased to a value greater than a predetermined penalty threshold. The corresponding task result or response can be invalidated or isolated. Invalidated responses can be not further processed. Isolated responses can be stored for possible later use, suspending further results from the same task performer.
[0080] The quality control process 216 can compare the representation of a specific domain in the provided solution and the original source. In response to a task performer providing more than a predetermined number of responses to a task that includes a representation of a specific domain that does not appear in the original source associated with the task (e.g., a crowd worker repeatedly provides a solution that includes a component that does not appear in the original document), a penalty score for the task performer can be increased to a value greater than a predetermined penalty threshold. If the criteria are met for more than a predetermined number of tasks, the penalty score can be increased above the predetermined penalty threshold. The corresponding task result or response can be invalidated or quarantined.
[0081] In response to a task performer providing more than a predetermined number of responses to the same task that are unique compared to responses submitted by other task performers (e.g., a crowd worker repeatedly provides unique solutions that are not provided by other crowd workers), the penalty score for the task performer may be increased to a value greater than a predetermined penalty threshold. If the criteria are met for more than a predetermined number of tasks, the penalty score may be increased above the predetermined penalty threshold. The corresponding task results or responses may be invalidated or quarantined.
[0082] Criteria for increasing the penalty score indicate that the citizen worker did not complete the task correctly. In a configuration with video excerpts, the penalty score can be increased in response to the task performer completing the video excerpt task in less than a predetermined percentage (e.g., half) of the runtime of the assigned video segment (e.g., the citizen worker completed the task in less than half the total length of the video). The video excerpt task can be invalidated in response to the task completion time being less than a predetermined percentage of the duration of the portion of the source video. This indicates that the citizen worker did not view the entire video segment associated with the task.
[0083] In some cases, the quality control process 216 can reduce the penalty score. For example, submitting a response that does not meet the penalty criteria can cause the penalty score to be reduced. For example, providing a solution verified by another source can cause the penalty score to be reduced. The rates of increase and decrease of the penalty score can be different.
[0084] Computing system 102 may include a system administrator user interface process 222. System administrator user interface process 222 may be configured to facilitate management of ISS 100 by system administrator 124. System administrator user interface process 222 may include a display interface for browsing information related to knowledge base 112. System administrator user interface process 222 may include an interface for creating and managing templates and tasks, and browsing solutions provided by task performers.
[0085] The computing system 102 may include a citizen worker user interface process 224 that may be configured to facilitate task completion by the citizen workers 121. The citizen worker user interface process 224 may include an interface that enables the citizen workers 121 to browse solutions and enter them into a template. For example, the citizen worker user interface process 224 may define a web-based interface for task completion.
[0086] The computing system may include a machine-to-machine interface process 226 configured to manage interactions with the task-processing machine 118. The machine-to-machine interface process 226 may implement a communication protocol between the computing system 102 and the task-processing machine 118. The machine-to-machine interface process 226 may distribute programs for execution on the task-processing machine 118.
[0087] ISS 100 may further include features for editing / reviewing diagnostic and repair solutions. Features may include user interfaces (UIs) and software components that represent data processing traces of original sources, extracted information, excerpts, and grouping similar solutions by template.
[0088] The computing system 102 may include a machine learning model update process 218 configured to manage the creation and training of machine learning models. The machine learning model update process 218 may be configured to increase efficiency through machine learning by repurposing knowledge gained from earlier executions through a multi-step validation process. The machine learning model update process 218 may be updateable to allow new machine tasks to be designed offline and injected into the workflow.
[0089] Domain knowledge representations can be updated using data chains (e.g., data tracking through backtracking). For example, the system may initially have a domain knowledge representation of automotive components with representative component names (e.g., oxygen sensor), acronyms (e.g., o2 sensor), and / or synonyms (e.g., Lambda sensor). By processing information through tasks, the system can learn and update informal usage of terms (e.g., o2s) or frequent misspellings (e.g., 02 sensor, where the letter o is replaced by the number zero) that refer to the same automotive component. The machine learning model update process 218 can be configured to learn new knowledge by keeping track / records of the original information, the excerpted information, and groups / clusters of multiple excerpts that resulted in a final solution / answer to the excerpted problem.
[0090] Figure 13A possible block diagram 1300 is depicted illustrating data that may be stored as part of a trajectory database 1320. The trajectory database 1320 may be a non-volatile memory storage device used to store data chains in a workflow. The trajectory database 1320 may include a first source 1304 and a second source 1310. The first source 1304 and the second source 1310 represent original source data or links to original source data. As previously described, the original sources are processed to obtain relevant portions. The trajectory database 1320 may include a first relevant portion 1306 derived from the first source 1304. The trajectory database 1320 may include a second relevant portion 1312 derived from the second source 1310. As previously described, the relevant portions may be excerpted. The trajectory database 1320 may include a first summary 1308 derived from the first relevant portion 1306. The trajectory database 1320 may include a second summary 1314 derived from the second relevant portion 1312. As previously described, the excerpted portions may be combined and grouped to form a final summary or title. Trajectory database 1320 may include final summary 1316 derived from first summary 1308 and second summary 1314 . Figure 13 Depicts the final summary as a combination of the two original sources. Note that trajectory database 1320 can include many such data structures. For example, multiple related parts can be derived from the original sources and result in additional final summaries. Trajectory database 1320 can represent a chain or trajectory of components that result in the final domain knowledge representation. Each task or template can be associated with a corresponding element within the chain or trajectory. Maintaining a track of components further allows for later analysis that can help redesign or automate the workflow.
[0091] The trajectory database 1320 can provide information to the KB update module 212. For example, the final summary 1316 can be provided to the KB update module 212. The KB update module 212 can search the final summary 1316 to determine whether it contains information that does not appear in the knowledge base 112. The KB update module 212 can update the knowledge base 112 with domain knowledge representations (such as component names, dictionary entries, symptoms, error codes, and repair information). The updating of domain knowledge can be performed automatically or semi-automatically by a machine (for example, the newly discovered information is reviewed / confirmed by a domain expert before being permanently used as updated domain knowledge).
[0092] The trajectory database 1320 can provide information to the machine learning model update process 218 for updating the machine learning model. The machine learning model update process 218 can be configured to create a new machine learning model by using the data chain tracked in different tasks of the workflow. For example, a collection of relevant parts (e.g., first relevant part 1306 and second relevant part 1312) selected from original information sources (e.g., first source 1304 and second source 1310), associated summaries (e.g., first summary 1308 and second summary 1314), and final summary 1316 can be input as training data for a new machine learning algorithm. The data chain includes inputs (e.g., original sources) and final outputs (e.g., final summary) that can be used to train the machine learning model. For example, the machine learning model can be repeatedly executed using the training inputs. The machine learning model output can be compared to the expected output to determine the error. The weighting factors or gains within the machine learning model can be adjusted in response to the error. The process can be repeated until the error falls below a predetermined value.
[0093] Information processing workflows may include Figure 11 Multiple tasks or processes depicted in . Figure 11 Depicted is a possible set of steps or tasks for extracting workflow 1100. Each step or task can be performed by a human (eg, human worker 121) or a machine (eg, task processing machine 118).
[0094] At task 1102, an operation may be performed to transform an original problem description from a diagnostic / scanning tool or information system into a template. The template may define one or more tasks to be completed. The operation may be performed by a system administrator 124 or by a software program or application. At a later stage of information synthesis, the operation may be performed automatically by a machine. For example, the original problem description may be associated with an error code associated with a product. The error code may be read from a diagnostic tool or displayed or indicated by the product itself. A template may be created that attempts to extract information related to a specific error code. Inquiries such as: "What caused error code X for product Y?", "How do I fix error code X for product Y?", or "How do I diagnose error code X for product Y?" may be posed.
[0095] At task 1104, an operation may be performed to search for relevant information sources. Search terms related to existing domain-specific representations may be used to guide the search. The initial search may reveal additional domain-specific representations that may be used to guide additional searches. Relevant information sources may include websites and / or web pages. The source may be an unstructured data source. An unstructured data source may not have a logical order or presentation of information. For example, a web-based discussion forum is sorted by replies or posts. In order to extract useful information, each post or reply must be parsed. Posts or replies may not always provide relevant information. The search process may generate one or more sources related to the original question. For example, one or more web addresses may be identified and stored. The web addresses may be incorporated into a template. Figure 4A and Figure 4B One possible example of a template that may be created to facilitate searching is provided.
[0096] At task 1106, operations can be performed to extract the most relevant data from the selected source.A template can be created to identify the selected source and request a review to determine the portion that is relevant to the original question. Figure 3A and Figure 3B Possible examples of templates that can facilitate information extraction from a source are provided. Relevant data may include those portions of the source that are most applicable to the original problem definition. For example, the information source may include additional information that is not relevant to the problem. The task output may be the identification of relevant portions for later processing. For example, a specific paragraph or section may be identified. The template may ask specific queries to facilitate the extraction of relevant information.
[0097] At task 1108, operations may be performed to extract the relevant data in a template. A template may be created to request a summary of the relevant portion of a source. For example, a template / task assigned to Crowdworker 121 may provide instructions for extracting the relevant data in a predetermined format, as previously described. Crowdworker 121 may process the task and provide the requested summary. Multiple tasks may be assigned, identifying different sources and / or different relevant portions. Figure 5A and Figure 5B Possible examples of templates that may be created for excerpted data are provided.
[0098] At task 1110, operations can be performed to group similar information. Task results can be analyzed to determine whether there is similar information that can be grouped together. A template can be created to generate tasks to facilitate identifying and grouping similar information. The template can generate a list of similar information and assign a task performer to identify the similar information. Figure 6 and Figure 9 Provides possible examples of templates that facilitate combining or grouping data.
[0099] At task 1112, operations can be performed to create a title / short description in the template. The task performer can review the grouping of information and provide a title or description in the format of <action verb><noun>. Figure 7 Provides possible examples of templates for facilitating title generation.
[0100] At task 1114, operations can be performed to search for images for components for each solution title.A template or task can be created to cause a search for images of components associated with a title / description.
[0101] At task 1116, operations can be performed to combine groups with common images into a single group. The images can be reviewed to determine whether any groups are associated with a common image. A task / template can be created to request analysis of a collection of images. For example, a task can request that the task performer identify images as being the same component. Figure 8 Provides possible examples of templates for image identification and composition.
[0102] At task 1118, operations may be performed to finalize and rank the solutions. Operations may be performed to rank the final solutions. Penalty scores may be used to rank the solutions based on the provided solutions and / or past performance of the task performer. Highly ranked solutions may be incorporated into the knowledge base 112. Lower ranked solutions may trigger additional tasks for further verification.
[0103] Figure 11 A workflow describes a way to process information, reducing unstructured data sets to target domain knowledge sets in a common and consistent format. Workflows can be facilitated by templates that identify specific, manageable tasks that can be completed by task performers. Workflows can be similar for many products. Templates can be tailored to use domain-specific representations for a given product. Workflows can be executed to generate a domain-specific knowledge base for any product or system.
[0104] Figure 14A possible sequence of operations that can be performed by ISS 100 is depicted. Depending on the level of automation implemented, the operations can be performed by computing system 102 and / or system administrator 124. At operation 1402, tasks and templates can be generated. For example, domain-specific representations can be inserted into blank or generic templates. In some configurations, the tasks and templates can be reviewed by system administrator 124. Templates can be generated using information currently in knowledge base 112. Computing system 102 can identify domain knowledge representations currently missing from knowledge base 112 from previous task results and generate a template that defines the task, which includes domain knowledge representations for extracting information from unstructured sources.
[0105] At operation 1404, the tasks / templates are distributed to the task performers. The computing system 102 may maintain data on the quality of the task performers interacting with the ISS 100. The computing system 102 may distribute the tasks / templates based on the availability of the task performers and the accuracy of the task performers. The computing system 102 may prioritize assigning tasks to the task performers with the highest quality designations. In addition, the computing system 102 may maintain scheduling information about the task performers. For example, the computing system 102 may maintain data on outstanding tasks assigned to the task performers. The computing system 102 may determine the workload of each task performer to determine whether the task performer has the ability to perform the task before the assigned deadline. The computing system 102 may also determine the type of task that should be assigned (e.g., human only or machine only). For example, the computing system 102 may check the availability of machine-only tasks that can complete the tasks to be assigned. In some cases, the computing system 102 may determine that the task should be distributed to multiple task performers to obtain more solutions that can be combined. When the task executor is determined, the computing system 102 may post the task to the task executor via a corresponding interface in an appropriate manner.
[0106] At operation 1406, computing system 102 may receive a solution / response to the task. For example, the task performer may have completed the appropriate input fields in the template. At operation 1408, computing system 102 may verify the response. For example, computing system 102 may compare the submitted response with other similar responses. Computing system 102 may perform a quality control process to determine whether the response appears to be valid. As part of the verification process, computing system 102 may update the quality control designation for the task performer. If the quality control check fails, the response may be rejected. The rejected task may be sent back to the task performer or may be sent to system administrator 124 for further review and action.
[0107] At operation 1410, computing system 102 may update the data chain for machine learning training. For example, computing system 102 may update trajectory database 1320 with relevant information associated with the task. Computing system 102 may identify the task associated with each piece of information. For example, multiple tasks may be used to create a chain from the original source to the final summary information. Computing system 102 may associate each task with a corresponding data segment.
[0108] At operation 1412, the computing system may aggregate and summarize the responses. Computing system 102 may be programmed to process the responses to group similar responses. Further, computing system 102 may implement a natural language processing routine to group similar response data. For example, for summary data expressed as an action verb plus a noun, computing system 102 may compare responses for words of similar meaning and use a preferred word to group responses. If a word appears in the knowledge base more than a predetermined number of times and / or more than other words of similar meaning, it may be a preferred word.
[0109] At operation 1414, computing system 102 may update knowledge base 112. For example, computing system 102 may identify a piece of information that is not currently in knowledge base 112. Computing system 102 may then update knowledge base 112 to include the new information. When new information is discovered, computing system 102 may determine whether additional tasks are required. For example, a response may have generated an alternate name for a component. Additional knowledge may be obtained by searching for information based on the alternate name. In this way, additional knowledge may be discovered and added to knowledge base 112.
[0110] The described information synthesis system facilitates the creation of a knowledge base pertaining to a product or system. Workers who are not domain experts can be used to generate the knowledge base based on existing information. This results in lower costs for generating the knowledge base. Furthermore, the system is easily adaptable to process new information. For example, emerging new domain-specific terms or expressions can be used to search for and create additional information fragments for the knowledge base. Furthermore, the templates defining tasks can be similar for different products. The templates can be easily adapted to other products by changing the domain-specific expressions and terminology.
[0111] The processes, methods or algorithms disclosed herein may be transferable to, or implemented by, a processing device, a controller or a computer that may include any existing programmable electronic control unit or a dedicated electronic control unit. Similarly, processes, methods or algorithms may be stored as data and instructions executable by a controller or computer by including but not limited to information permanently stored on a non-writable storage medium (such as a ROM device) and many forms of information modifiable on a writable storage medium (such as a floppy disk, a tape, a CD, a RAM device and other magnetic and optical media). Processes, methods or algorithms may also be implemented in software executable objects. Alternatively, suitable hardware components (such as application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), state machines, controllers or other hardware components or devices) or a combination of hardware, software and firmware components may be used to fully or partially embody processes, methods or algorithms.
[0112] Although exemplary embodiments have been described above, these embodiments are not intended to describe all possible forms covered by the claims. The words used in the specification are descriptive rather than restrictive, and it is understood that various changes can be made without departing from the spirit and scope of the present disclosure. As previously described, the features of the various embodiments can be combined to form further embodiments of the present invention that may not be explicitly described or illustrated. Although various embodiments may have been described as providing advantages or being preferred over other embodiments or prior art implementations with respect to one or more desirable characteristics, those skilled in the art recognize that one or more features or characteristics can be compromised to achieve desirable overall system properties, depending on the specific application and implementation. These properties may include, but are not limited to, cost, strength, durability, life cycle cost, marketability, appearance, packaging, size, serviceability, weight, manufacturability, ease of assembly, etc. Therefore, embodiments described as being less desirable than other embodiments or prior art implementations with respect to one or more characteristics are not outside the scope of the present disclosure and may be desirable for a particular application.
Claims
1. An information synthesis system for generating a knowledge base, comprising: A computing system comprising: a task configurator processing unit configured to distribute a template including a task for extracting information from an unstructured source to a task executor; a solution aggregation processing unit configured to receive a task result from a task executor as a response in a template; a knowledge base source identification processing unit configured to identify domain knowledge representations that appear in the task results but are missing from the knowledge base to enable updating of the domain knowledge representations; and a template definition and / or generation processing unit configured to generate a template defining a task and including an updated domain knowledge representation for extracting additional information from an unstructured source, The task executors include at least one of domain experts, mass workers and machines.
2. The system of claim 1, wherein: Tasks are defined as human-only tasks, machine tasks, and machine-guided tasks.
3. The system of claim 1, wherein: The computing system is further programmed to distribute the templates based on the availability of the task performers and the accuracy of the task performers.
4. The system of claim 1, wherein: The task involves extracting information from unstructured data sources.
5. The system of claim 1, wherein: The task consists of extracting information contained in at least a portion of a video.
6. The system of claim 5, wherein: The computing system is further programmed to validate the task results received from each of the task performers and identify the task results as invalid in response to the task completion time being less than a predetermined percentage of the duration of the portion of the video.
7. The system of claim 1, wherein: The computing system is further programmed to validate the task result from each of the task performers and identify the task result as invalid in response to: (i) the task result being the same for a predetermined number of responses; (ii) the task result includes a term identifying a component that is missing from an original source corresponding to the task result; as well as (iii) The task result is unique compared to the task results submitted by other task performers.
8. The system of claim 1, wherein: The computing system is further programmed to maintain a data chain for the tasks, the data chain comprising, for each of the tasks: data defining an original source: a relevant portion of the original source; an excerpt of the relevant portion; and a final summary derived from the extracts.
9. The system of claim 8, wherein: The computing system is further programmed to facilitate training of the machine learning models by providing the data chain to the one or more machine learning models as training input.
10. The system of claim 1, wherein: The computing system is further programmed to predict the accuracy of the task performer before distributing the template.
11. A method for updating a knowledge base using the information synthesis system for generating a knowledge base according to any one of claims 1 to 10, comprising: By computing system: Maintaining penalty scores for task performers; dispatching the task to a task executor in response to the penalty score being less than a predetermined threshold; as well as In response to a task performer providing more than a predetermined number of responses to a task containing a representation of a particular domain that does not appear in an original source associated with the task, increasing a penalty score for the task performer to a value greater than a predetermined threshold.
12. The method of claim 11, further comprising: In response to receiving more than a predetermined number of identical responses from a task performer for different tasks, a penalty score for the task performer is increased to a value greater than a predetermined threshold.
13. The method of claim 11, further comprising: In response to a task performer providing more than a predetermined number of responses to the same task that are unique compared to responses submitted by other task performers, a penalty score for the task performer is increased to a value greater than a predetermined threshold.
14. The method of claim 11, further comprising: In response to the task performer completing the video excerpt task in a time that is less than a predetermined percentage of the runtime of the allocated video segment, a penalty score for the task performer is increased.
15. The method of claim 11, further comprising: Responses that contribute to a penalty score exceeding a predetermined threshold are invalidated.
16. A method for synthesizing information from unstructured data sources to update a maintenance knowledge base, comprising: identifying, by means of the knowledge base source identification process, relevant portions of original sources having domain-specific knowledge relevant to the maintenance knowledge base; creating a template including tasks for extracting each of the relevant parts by means of a template definition and / or generation process; Distributing the templates to task performers based on their availability and accuracy via a task configurator process; aggregating solutions from the templates completed by the task performers via a solution aggregation process to create a repair solution described as an action verb followed by a component name; updating the maintenance knowledge base with domain-specific representations from the maintenance solution that are not present in the maintenance knowledge base by means of a knowledge base update process; as well as creating new templates based on the updated domain-specific representation by means of a template definition and / or generation process and distributing the new templates by means of a task configurator process, The task executors include at least one of domain experts, mass workers and machines.
17. The method of claim 16, further comprising: Create a new machine learning model using the original source, relevant parts, summaries, and repair solutions as training data for updating the machine learning model.
18. The method of claim 16, wherein: Repair solutions are described as action verbs followed by the component name.
19. The method of claim 16, wherein: The original source is the document accessed on the website.
20. The method of claim 16, wherein: The original source is the video accessed on the website.
Citation Information
Patent Citations
Mobile crowd sensing online incentive method combined with credit update
CN108364190A