Generating Similarity Scores Between Different Document Schemas
Patent Information
- Application Number
- JP2024513780
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-09-01
- Filing Date
- 2022-08-31
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-08-31
AI Technical Summary
Existing document repositories struggle to effectively search and compare documents with different schemas due to the difficulty in identifying semantic similarities, as existing methods primarily focus on syntactic comparisons, leading to missed connections in meaning across different fields.
A system that indexes documents with data cleanup to remove irrelevant metadata, uses inverted indexes for efficient searching, and generates intelligent queries across different schemas based on field mappings and document frequency scores to calculate weighted similarity scores.
Enhances the likelihood of finding semantically similar documents by accurately mapping fields across different schemas, improving the relevance of search results.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to U.S. patent application Ser. No. 17 / 464,534, entitled “GENERATING SIMILARITY SCORES BETWEEN DIFFERENT DOCUMENT SCHEMAS,” filed Sep. 1, 2021, which is incorporated by reference in its entirety. [Background technology]
[0002] background A document repository may contain a large number of documents in a persistent storage system. These documents may contain structured and unstructured data and may conform to many different schema types. For example, a document repository representing a knowledge base may contain FAQs, white papers, web pages, emails, and / or other information that can be used to address various problems in an operating environment. Although a document repository may store a large amount of information, effectively searching this information is also very difficult because the document repository may contain many different types of documents that are difficult to analyze uniformly.
[0003] An existing method of identifying documents that may be related to a source document is to generate a similarity score. A similarity score is a metric calculated by a repository's search interface that represents a measure of how syntactically similar two documents are. The source document can be compared to each of the individual documents in the document repository to generate a similarity score for each document in the repository. These scores can then be used to identify documents that are most likely to be similar to the source document. Summary of the Invention [Problem to be solved by the invention]
[0004] Quick Overview The embodiments described herein allow a document repository comprised of documents with many different schemas to be searched and compared to an input document to generate a similarity score. The similarity score can be used to identify the document in the document repository that is most similar to the input document. The schema of the input document can be identified and used to obtain a configuration specific to that schema. The configuration can include information defining how queries can be automatically generated and sent to the document repository so that searches can be performed across different fields in documents with different schemas. These queries can be concatenated and sent to the document repository. The weighted scores generated for the resulting documents can be aggregated to generate a final similarity score for each document.
[0005] Rather than simply searching documents for syntactic similarity, this configuration increases the likelihood that queries sent to the document repository will yield semantic similarities, and therefore similar meanings or concepts expressed in the documents. This configuration can be tailored to specifically map high frequency n-grams of particular source fields to particular target fields of documents with a specified schema in the document repository.
[0006] The document repository may be designed to include an interface that allows for document indexing. Existing document repositories can be crawled and / or documents can be submitted to be indexed in the inverted index. A data cleanup process may remove extraneous information and metadata that is not related to the semantic meaning of the documents before indexing occurs. The system may also include a search interface for the inverted index, as well as a document frequency API that can be used to obtain the document frequency of specific words. This document frequency can be used to generate a frequency score for individual words. This frequency score can be used to select which words in the target field to use in generating search queries.
[0007] The configuration itself may include a separate section for each schema that can be used as a source for a search query. The search fields of the target document may be used to generate a specified number of queries that provide individual n-grams or other field values from the source fields and can be concatenated to form a single query. The resulting similarity scores may be weighted according to values stored in the configuration. The various mappings between source and target fields, and between different schemas in the knowledge repository, may be aggregated to form a final similarity score for each document. These similarity scores may then be used to order and present the results to the requesting user or device.
[0008] A further understanding of the nature and advantages of the various embodiments can be realized by reference to the remaining portions of the specification and the drawings, in which like reference numerals are used to refer to like components throughout the several views. In some cases, a sub-label is associated with a reference numeral that denotes one of multiple similar components. When referring to a reference numeral without specifying an existing sub-label, the intention is to refer to all such multiple similar components. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 illustrates a system for submitting documents to a document repository for generating similarity scores, according to some embodiments. [Diagram 2] FIG. 2 illustrates a collection of documents with different schemas according to some embodiments. [Diagram 3] FIG. 1 illustrates a system for a document repository that may be used when generating queries for a similarity score process, according to some embodiments. [Figure 4] FIG. 1 is a diagram of a similarity scoring system, according to some embodiments. [Figure 5A] FIG. 2 illustrates an example of a configuration for a particular schema, according to some embodiments. [Figure 5B]FIG. 2 illustrates an example of a configuration for a particular schema, according to some embodiments. [Figure 6] FIG. 2 illustrates a flowchart of a method for calculating similarity scores for documents according to some embodiments. [Figure 7] FIG. 1 is a simplified block diagram of a distributed system for implementing some embodiments. [Figure 8] FIG. 1 is a simplified block diagram of components of a system environment in which services provided by components of the system of one embodiment may be offered as cloud services. [Figure 9] FIG. 1 illustrates an exemplary computer system in which various embodiments may be implemented. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Detailed Description The embodiments described herein allow a document repository consisting of documents with many different schemas to be searched and compared to an input document to generate a similarity score. The similarity score can be used to identify the document in the document repository that is most similar to the input document. The schema of the input document can be identified and used to obtain a configuration specific to that schema. The configuration can include information defining how queries can be automatically generated and sent to the document repository so that searches can be performed across different fields in documents with different schemas. These queries can be concatenated and sent to the document repository. The weighted scores generated for the resulting documents can be aggregated to generate a final similarity score for each document.
[0011] FIG. 1 illustrates a system 100 for submitting documents 104 to a document repository 106 for generating similarity scores, according to some embodiments. A client system 102 can submit documents 104 to a server, a web-based system, or a cloud-based system, which may be collectively referred to as a “server” or “server system.” The documents 104 can represent any type of document, including structured and / or unstructured data. As an example, the documents 104 can represent an incident report or trouble ticket received by an incident management system. The documents 104 can be generated by the client system 102. Alternatively, the documents 104 may be generated by a server that manages the document repository 106 and / or operates the incident management system in response to information transmitted by the client system 102. For example, the client system 102 can submit information from a web form that is used to populate fields in the documents 104 to generate an incident report by the client system 102 or the incident management system.
[0012] The document 104 may be received by the server system to locate documents in the document repository 106 that are responsive to the information in the document 104. Continuing with the example of an incident management system, the document 104 may represent a description of a problem or other incident related to a service provided by a service provider. The document repository 106 may include documents 108 such as white papers, solutions to common problems, knowledge base articles, and other information that may correspond to the problem described in the document 104 and / or other problems previously handled by the system. The similarity score represents a metric that indicates how closely the information in the document 104 is related to the information in each document 108 in the document repository 106. The higher the similarity score, the more likely one of the particular documents 108 provides information related to the topic of the document 104.
[0013] In such a system, the similarity score calculated by the existing system may simply run a comparison algorithm between the document 104 and the document 108 in the document repository 106 to compare individual words. This is very effective in finding documents in the document repository 106 that are syntactically similar to the document 104. However, a technical problem exists in that the existing methods do not find documents in the document repository 106 that are semantically similar to the meaning expressed by the document 104. For example, the existing techniques may identify a document 108 that uses similar terminology as the document 104, but is not related to the particular problem expressed in the semantics of the document 104. Another technical problem is that the existing techniques do not intelligently map language from a particular field in the document 104 to another particular field in the document 108. Because semantic ideas may be expressed differently in different fields, and ideas that match both a particular source field and a target field should be weighted more than others, the existing techniques often miss important connections between the ideas expressed in the document 104 and the document 108.
[0014] The embodiments described herein solve these and other technical problems by using a defined configuration and instruct the system how to generate intelligent queries that can connect the meaning of information expressed in the source of the document 104 to the meaning of identified documents 108 in the document repository 106 that are more likely to solve the problem expressed in the document 104. Furthermore, these embodiments solve the technical problem of generating accurate comparisons and similarity scores between documents with different schemas. In structured documents, comparisons between all fields in various documents can be inefficient and cumbersome. These embodiments provide targeted queries between specific fields between different schemas. Since information may be stored in different fields in different documents, the configuration defines the target fields and their corresponding source fields where the comparison of information is most efficient.
[0015] FIG. 2 illustrates a collection of documents with different schemas, according to some embodiments. A document 201 received by the document repository 106 may have a first schema. As defined herein, "schema" may refer to the structure of the document 200. For example, the schema may define the number of field-value pairs found in the document 201. Each of the fields 202 may be associated with a corresponding value 204 that is unique to each document instance. The fields 202 may define the data type of the value 204. For example, a first field 202-1 may include a label such as "username" and may define the type as "text string" such that a corresponding first value 204-1 may include a text string with a particular value of username. The first schema of the document 201 may define all the field-value pairs, but individual documents using that schema may define particular values for the corresponding field-value pairs. The schema may also define other structural elements of the document, including styles, images, backgrounds, divisions, static text, and other document elements.
[0016] As used herein, the terms "first" and "second" are used simply to distinguish various elements, such as different schemata. These terms do not imply any order, priority, importance, or other characteristics of these elements, but serve only to distinguish one element from another. For example, a first schema and a second schema may refer to two documents having separate schemata. The first schema and the second schema may be the same or different schemas, such that both documents may have the same schema or may have different schemas.
[0017] Traditionally, problems have arisen when a document 201 is submitted to a document repository 106. The document repository 106 may include multiple document collections 205. Each of the document collections 205 may be associated with an individual schema. For example, collection 205-1 may be associated with a first schema, and each document in collection 205-1 may share the same first schema. The other collections 205-n may each include a different schema. In traditional systems, document 201 could only be compared to other documents in the document repository 106 that share the same schema. This allows for similarity comparisons between corresponding values of field-value pairs. However, this significantly reduces the number of documents in the document repository 106 that can accurately respond to a request to generate a similarity score for document 201. The embodiments described herein can generate a semantically matching similarity score between document 201 having a first schema and any collection of documents in the document repository 106 having a second schema.
[0018] FIG. 3 illustrates a system 300 for a document repository that may be used in generating queries for a similarity score process, according to some embodiments. The system 300 may first include a document indexing interface 302. The document indexing interface 302 may receive a request 320 to index a new document to be added to the document repository. Additionally, the document indexing interface 302 may access existing documents in an existing document repository 318 to crawl and index the documents in the document repository 318. A data cleanup process 310 may be used to remove information from the documents that is not related to the semantic meaning of the documents before the indexing process takes place. For example, the data cleanup process 310 may perform various data cleanup steps, such as removing JavaScript, HTML code, CSS code, and other code or elements related to the display of the documents, the structure or format of the documents, or other metadata. After the data cleanup process 310, the documents may be provided to an indexing process 314 that generates a reverse or inverted index 316 of the document repository 318.
[0019] The inverted index 316 stores a list of each document in the document repository 318 that contains a particular word. The inverted index may include a database index that stores a mapping from content, such as individual words, to a location in a set of documents (as opposed to a forward index, which maps from documents to content). The purpose of the inverted index is to enable fast full-text searches at the expense of increased processing as documents are added to the document repository 318. The system may also include an inverted index search interface 304, which allows the system 300 to receive requests 322 to query the inverted index 316. The requests 322 may include words found in one or more documents in the document repository 318. The inverted index 316 may access a list of particular words and return a list of documents that contain the words. The embodiments described herein may also allow the requests 322 to specify a particular field in each document. For example, the requests 322 may include words to be searched for in the SUBJECT field of a particular document schema. The inverted index 316 may be generated to be associated with a particular collection of documents that all have the same schema. Alternatively, the inverted index 316 may be generated such that collections of documents with individual schemas can be searched and indexed as separate collections from one another. The inverted index 316 can be searched using queries, including Boolean queries, phrase queries, word queries, single value queries, and / or any other type of query.
[0020] The system 300 may also include a document frequency interface 306. This interface may be implemented using an application programming interface (API) that retrieves document frequencies. The document frequency interface 306 or API may search the document repository 318 for a given word to retrieve the number of documents in which the word is found. In some embodiments, the document frequency may be used to generate a document frequency score for a particular word. This score is generated by multiplying (1) the frequency with which a particular word is found in a source document by (2) the inverse measure of the number of documents that contain the word in a particular document collection. As described in more detail below, the frequency score for a particular word may be used to generate queries.
[0021] In some embodiments, system 300 having each of the above-described interfaces may be implemented using Apache® SOLR software or may be built on top of the Apache® Lucene search system. However, these particular software solutions are provided only as useful examples and are not intended to be limiting. Many other software systems capable of implementing similar functionality as described herein may be used.
[0022] Figure 4 shows a diagram of a similarity scoring system 400, according to some embodiments. The process performed by similarity scoring system 400 can assume that the document repository has been properly indexed and processed, as described above in connection with Figure 3. Thus, similarity scoring system 400 can receive document frequencies and submit requests to the various interfaces described above to perform searches of the inverted index.
[0023] A document 402 may be submitted to the similarity scoring system 400. The document 402 may be received from a client device and may represent any type of document, such as an incident report as described above in the example of FIG. 1. The similarity scoring system 400 may determine a particular configuration associated with the document 402 (404). For example, the configuration data store 406 may store configurations associated with each type of schema that may be received by the system or stored in a document repository. The schema of a particular document 402 may be determined by inspecting metadata or by identifying and matching field-value pairs in the document 402 with known schemas. Once a schema is identified, it may be submitted to the configuration data store 406 to obtain a configuration specific to that schema. Thus, the configuration data store 406 may store a configuration for each schema defined in the similarity scoring system 400.
[0024] The similarity scoring system 400 may then generate 408 a plurality of queries based on the configuration. Particular examples of configurations and how the configurations may be used to generate a plurality of queries are described below in conjunction with FIGS. 5A and 5B. In general, a configuration may include a set of fields that can be used as instructions for generating queries from a source schema of the document 402. The queries may target any of the schema types stored in the document repository. For example, if the document 402 has a schema A, the configuration may include fields that act as instructions for generating a set of queries between documents with schema A and documents with schema A, between documents with schema A and documents with schema B, and so on. Thus, the configuration may include instructions for mapping queries from the schema of the source document 402 to multiple other schemas that may exist in the document repository.
[0025] Generating the query may include receiving a document frequency score from the document frequency interface 306 described above. The document frequency score may be used to generate queries that are most likely to generate responsive answers. For example, the document frequency score may be used to generate queries for words in the source documents 402 that are most likely to be found in the document repository. Multiple queries may be generated for each field-to-field combination between the source documents 402 and fields in a particular schema indicated by the configuration.
[0026] The similarity scoring system 400 may then execute the queries (410). These queries may be submitted together as a combined (e.g., "OR") set of queries that are submitted to the reverse index search interface 304. For example, some embodiments may create a master query that combines all the queries together. This query is executed and the returned documents may receive a score. As described below, the configuration may include a weight that is applied to each score. The scores returned by the search may apply a weight to the scores from the index. For example, some search interfaces may receive a weight as a multiplier that boosts the return score. These scores may then be aggregated for each document to generate a final similarity score for each document. Note that some embodiments do not need to normalize the scores, but instead documents may be compared to each other using weighted scores, without requiring normalization. The scores are then displayed and / or used to order the document results presented to the requesting client system or user interface.
[0027] 5A-5B show an example of a configuration 500 for a particular schema, according to some embodiments. In this example, configuration 500 is selected for a source document having schema A. The schema itself may be an object having an object type that can be used to identify configuration 500 from multiple different configurations associated with different source document schemas. Configuration 500 may be part of a larger configuration file that defines many different configurations for different schemas. Configuration 500 may be stored as a structured document, such as XML.
[0028] The configuration 500 can be used to generate multiple search fields 502 that can be executed as queries on a document repository. The search fields 502 can use a source field 504 of schema A as a source for the query. In this example, a first source field 504-1 can identify a TITLE field of a source document, and words in the TITLE field can be used to search different fields of a specified schema type in the document repository. The TITLE field can have a type of "text," indicating that it stores a text string. Each of the queries generated against the first source field 504-1 can use words from the first source field 504-1 as a source when constructing the query. For each schema, one or more source fields 504 in the schema can be identified by the configuration and used to generate the query. For example, in addition to the TITLE field, a second search field 504-2 can identify a CONTENT field, and a third source field 504-3 can identify an AUTHOR field as a source in schema A.
[0029] For source fields, the configuration 500 may specify different query types 506. Each of the query types 506 may specify the number of words to be used in each query. For example, a first query type 506-1 may specify the type as 1-SHINGLE, indicating a query that matches n-grams (i.e., 1-grams) of order "1" from the source field to various target fields of the target schema. A second query type 506-2 may specify the type as 3-SHINGLE, to search for 3-grams from the source TITLE field of schema A. Another query type 516 may specify the type as a SINGLE_VALUE type, indicating that a single value from the source must match a single value in the target field. For example, the author's name may need to match exactly between the source and target fields.
[0030] The query type 506 may also identify the number of queries to be generated. For example, a second query type 506-2 may identify four queries to be generated as tri-grams from the source TITLE field of Schema A. To determine which of the tri-grams from the TITLE field to use, the system may query the document frequency interface described above to obtain the document frequency of each word in each tri-gram of the TITLE field. The query may then be generated using the tri-gram that produces the highest document frequency score, which is a combination of the individual frequency scores of the individual words. As explained above, the document frequency score may be the product of the frequency with which a particular tri-gram word appears in the source (TITLE) field and the inverse of the number of documents in the document repository that contain the tri-gram word.
[0031] Finally, each of the query types may identify one or more queries 508, 510, 514, 518 that may be generated for each query type. For example, the first query 508-1 may be composed of ten individual queries according to the first query type 506-1. Each of the individual queries may correspond to the one-gram (e.g., individual words in the TITLE field) with the highest document frequency score. The queries may target specific fields in a particular schema in the document repository. For example, the first query 508-1 may generate ten queries that each search for different words in the TITLE field of documents in the document repository having Schema A. Note that the example schemas of Figures 5A and 5B search fields from Schema A compared to fields from other documents in the document repository having the same Schema A. However, this is provided by way of example only and is not intended to be limiting. The configuration 500 may also include other queries that target objects having different schemas (e.g., Schema B) that are not specifically shown in these figures.
[0032] To generate query 508-1, the 10 single words with the highest frequency scores can be combined with an "OR" operator to form a single query that searches for any of the single words in the TITLE field of Schema A. The result score generated by this query for each document can be multiplied by a weight (e.g., 7). This weighting allows the configuration to specify matches between fields that are more indicative of similarity in semantic meaning. Finally, each of one or more queries 508, 510, 514, 518 can be concatenated, combined, and / or submitted to an index to generate a weighted similarity score. In some embodiments, the weights can be set by a user or automatically by a machine learning model.
[0033] 6 illustrates a flowchart 600 of a method for calculating a similarity score for a document, according to some embodiments. The method may include receiving 602 a first document having a first schema. The document and the first schema may be received as described above in FIGS. 1 and 2. For example, the first schema may define a format for a service request to be received by an incident management system, among other example operating environments.
[0034] The method may also include accessing a configuration of a first schema (604). The configuration may define how to generate a plurality of queries from a first document to a collection of documents having a second schema. The second schema may be the same as or different from the first schema. As described above in FIGS. 5A-5B, the schema may define field-value pairs in the document format, value types, field names, metadata, and / or other information related to the structure or format of the documents. In a particular example, the configuration may include a query type that defines an n-gram level (e.g., 1-gram, 3-gram, etc.) for some of the queries to be generated. The configuration may also include a number of queries to be generated for each query type. The words or n-grams selected for the queries are based on a frequency score. The frequency score may represent the product of the number of times a word appears in the source field and the inverse of the number of documents in which the word appears in the document repository.
[0035] The method may further include generating a plurality of queries based on the configuration (606). The multiple queries may be linked using a conjunction or "OR" operator to form a single query that may be executed against the document repository. The document repository may include a collection of documents with different schemas that are part of the knowledge base. The method may further include combining results of the multiple queries into a similarity score for the first document (608). The results of each query may be weighted according to the weights provided in the configuration. The individual scores for each target document may then be combined as a weighted combination into a single similarity score for each document, which may be used to sort or present the result documents to a user or client device.
[0036] It should be understood that the specific steps illustrated in FIG. 6 provide a particular method of generating a similarity score according to various embodiments. Other sequences of steps may be performed according to different embodiments. For example, alternative embodiments may perform the steps outlined above in a different order. Additionally, the individual steps illustrated in FIG. 6 may include multiple sub-steps that may be performed in various orders depending on the individual step. Additionally, additional steps may be added or removed depending on the particular application. Many variations, modifications, and alternatives are within the scope of the present disclosure.
[0037] Each of the methods described herein may be implemented by a computer system. Each of the steps of these methods may be performed automatically by the computer system and / or may be provided with input / output involving a user. For example, a user may provide inputs for each step of the method, and each of these inputs may respond with a particular output that requests such input, and the output is generated by the computer system. Each input may be received in response to a corresponding requested output. Additionally, inputs may be received as a data stream from a user, from another computer system, retrieved from a memory location, retrieved over a network, requested from a web service, etc. Similarly, outputs may be provided as a data stream to a user, provided to another computer system, stored in a memory location, transmitted over a network, provided to a web service, etc. In short, each of the steps of the methods described herein may be performed by a computer system and may include any number of inputs, outputs, and / or requests to and from the computer system, with or without the involvement of a user. Those steps that do not involve a user are said to be performed automatically by the computer system without human intervention. Thus, it will be appreciated that the steps of each method described herein in light of this disclosure may be modified to include inputs to and outputs from a user, or may be performed automatically by a computer system without human intervention where the decisions are made by a processor. Additionally, some embodiments of each method described herein may be implemented as a series of instructions stored on a tangible, non-transitory storage medium to form a tangible software product.
[0038] 7 shows a simplified diagram of a distributed system 700 for implementing one of the embodiments. In the illustrated embodiment, the distributed system 700 includes one or more client computing devices 702, 704, 706, and 708 that are configured to execute and operate client applications, such as web browsers, proprietary clients (e.g., Oracle Forms), etc., over one or more networks 710. A server 712 may be communicatively coupled to the remote client computing devices 702, 704, 706, and 708 over the network 710.
[0039] In various embodiments, server 712 may be adapted to execute one or more services or software applications provided by one or more components of the system. In some embodiments, these services may be provided to users of client computing devices 702, 704, 706, and / or 708 as web-based or cloud services, or under a software-as-a-service (SaaS) model. Users operating client computing devices 702, 704, 706, and / or 708 can in turn utilize one or more client applications to interact with server 712 and utilize the services provided by these components.
[0040] In the illustrated configuration, the software components 718, 720, and 722 of the system 700 are shown as being implemented on the server 712. In other embodiments, one or more of the components of the system 700 and / or the services provided by these components may be implemented by one or more of the client computing devices 702, 704, 706, and / or 708. A user operating a client computing device can then utilize one or more client applications to use the services provided by these components. These components may be implemented in hardware, firmware, software, or a combination thereof. It should be understood that a variety of different system configurations are possible that may differ from the distributed system 700. Thus, the embodiment illustrated in the figure is an example of a distributed system for implementing an embodiment system and is not intended to be limiting.
[0041] The client computing devices 702, 704, 706, and / or 708 may be portable handheld devices (e.g., iPhone®, mobile phones, iPad®, computing tablets, personal digital assistants (PDAs)), or wearable devices (e.g., Google Glass® head mounted displays), running software such as Microsoft Windows Mobile® and / or various mobile operating systems such as iOS, Windows Phone, Android, BlackBerry 10, Palm OS, and are Internet, email, short message service (SMS), Blackberry®, or other communication protocol enabled. The client computing devices may be general purpose personal computers including, for example, personal computers and / or laptop computers running various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux® operating systems. The client computing devices may be workstation computers running any of the various commercially available UNIX® or UNIX-like operating systems including, for example, but not limited to, various GNU / Linux operating systems such as Google Chrome OS. Alternatively, or in addition, client computing devices 702, 704, 706, and 708 may be any other electronic devices, such as thin-client computers, Internet-enabled gaming systems (e.g., Microsoft Xbox game consoles with or without Kinect® gesture input devices), and / or personal messaging devices, capable of communicating over network 710.
[0042] Although the exemplary distributed system 700 is shown with four client computing devices, any number of client computing devices may be supported. Other devices, such as devices with sensors, etc., may interact with the server 712.
[0043] The network 710 in the distributed system 700 can be any type of network and can support data communications using any of a variety of commercially available protocols, including, but not limited to, TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (Systems Network Architecture), IPX (Internet Packet Exchange), AppleTalk, and the like. By way of example only, the network 710 can be a local area network (LAN) based on Ethernet, Token Ring, and the like. The network 710 can be a wide area network and the Internet. It can include virtual networks, including, but not limited to, virtual private networks (VPNs), intranets, extranets, public switched telephone networks (PSTNs), infrared networks, wireless networks (e.g., networks operating under any of the Institute of Electrical and Electronics Engineers (IEEE) 802.11 protocol suite, Bluetooth, and / or other wireless protocols), and / or any combination of these and / or other networks.
[0044] Servers 712 may be comprised of one or more general purpose computers, dedicated server computers (including, by way of example, PC (personal computer) servers, UNIX servers, mid-range servers, mainframe computers, rack-mounted type servers, etc.), server farms, server clusters, or any other suitable arrangement and / or combination. In various embodiments, servers 712 may be adapted to execute one or more services or software applications described in the preceding disclosure. For example, servers 712 may correspond to servers performing the processes described above in accordance with one embodiment of the present disclosure.
[0045] Server 712 may run operating systems including any of those discussed above, as well as any commercially available server operating system. Server 712 may also run any of a variety of additional server applications and / or mid-tier applications, including a HyperText Transport Protocol (HTTP) server, a File Transfer Protocol (FTP) server, a Common Gateway Interface (CGI) server, a JAVA server, a database server, etc. Exemplary database servers include, but are not limited to, those commercially available from Oracle, Microsoft, Sybase, IBM (International Business Machines), etc.
[0046] In some implementations, the server 712 may include one or more applications for analyzing and consolidating data feeds and / or event updates received from users of the client computing devices 702, 704, 706, and 708. By way of example, the data feeds and / or event updates may include, but are not limited to, Twitter® feeds, Facebook® updates, or real-time updates received from one or more third party information sources and continuous data streams, which may include real-time events associated with sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, and the like. The server 712 may also include one or more applications for displaying the data feeds and / or real-time events via one or more display devices of the client computing devices 702, 704, 706, and 708.
[0047] Distributed system 700 may also include one or more databases 714 and 716. Databases 714 and 716 may reside in a variety of locations. As an example, one or more of databases 714 and 716 may reside on a non-transitory storage medium local to (and / or present on) server 712. Alternatively, databases 714 and 716 may be remote from server 712 and communicate with server 712 via a network-based or dedicated connection. In one set of embodiments, databases 714 and 716 may reside within a storage area network (SAN). Similarly, files required to perform functions attributed to server 712 may be stored locally on server 712 and / or remotely, as appropriate. In one set of embodiments, databases 714 and 716 may include relational databases, such as those provided by Oracle, adapted to store, update, and retrieve data in response to SQL-formatted commands.
[0048] 8 is a simplified block diagram of one or more components of a system environment 800 in which services provided by one or more components of an embodiment system may be provided as cloud services, according to one embodiment of the disclosure. In the illustrated embodiment, the system environment 800 includes one or more client computing devices 804, 806, and 808 that may be used by users to interact with a cloud infrastructure system 802 that provides cloud services. The client computing devices may be configured to run client applications, such as, for example, a web browser, a proprietary client application (e.g., Oracle Forms), or other applications that may be used by users of the client computing devices to interact with the cloud infrastructure system 802 to use services provided by the cloud infrastructure system 802.
[0049] It should be understood that the cloud infrastructure system 802 depicted in the figure may have other components than those shown. Additionally, the system depicted in the figure is only one example of a cloud infrastructure system that may incorporate some embodiments. In other embodiments, cloud infrastructure system 802 may have more or fewer components than depicted in the figure, may combine two or more components, or may have a different configuration or arrangement of components.
[0050] Client computing devices 804 , 806 , and 808 may be devices similar to those described above for 702 , 704 , 706 , and 708 .
[0051] Although the exemplary system environment 800 is shown with three client computing devices, any number of client computing devices may be supported. Other devices, such as devices with sensors, may interact with the cloud infrastructure system 802.
[0052] Network 810 can facilitate communication and data exchange between clients 804, 806, and 808 and cloud infrastructure system 802. Each network can be any type of network capable of supporting data communications using any of a variety of commercially available protocols, including those described above for network 710.
[0053] Cloud infrastructure system 802 may comprise one or more computers and / or servers, which may include those described above for server 712.
[0054] In an embodiment, the services provided by the cloud infrastructure system may include a number of services made available to users of the cloud infrastructure system on demand, such as online data storage and backup solutions, web-based email services, hosted office suites and document collaboration services, database processing, managed technical support services, etc. The services provided by the cloud infrastructure system may be dynamically scaled to meet the needs of the users. A particular instantiation of a service provided by the cloud infrastructure system is referred to herein as a "service instance." In general, a service available to a user from a cloud service provider's system over a communications network such as the Internet is referred to as a "cloud service." Typically, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's own on-premise servers and systems. For example, the cloud service provider's system may host applications, which users may order and use on demand over a communications network such as the Internet.
[0055] In some examples, services in a computer network cloud infrastructure may include protected computer network access to storage devices, hosted databases, hosted web servers, software applications, or other services provided by a cloud vendor to users. For example, the services may include password-protected access to remote storage devices on the cloud over the Internet. As another example, the services may include a web services-based hosted relational database and a scripting language middleware engine for personal use by a network developer. As another example, the services may include access to an email software application hosted on a cloud vendor's website.
[0056] In one embodiment, cloud infrastructure system 802 may include a suite of application, middleware, and database services products that are delivered to customers in a self-service, subscription-based, elastically scalable, reliable, highly available, and secure manner. One example of such a cloud infrastructure system is the Oracle Public Cloud, offered by the present assignee.
[0057] In various embodiments, cloud infrastructure system 802 may be adapted to automatically provision, manage, and track customer subscriptions to services provided by cloud infrastructure system 802. Cloud infrastructure system 802 may provide cloud services through different deployment models. For example, the services may be provided under a public cloud model, where cloud infrastructure system 802 is owned by an organization that sells cloud services (e.g., owned by Oracle) and the services are made available to the general public or to companies in different industries. As another example, the services may be provided under a private cloud model, where cloud infrastructure system 802 is operated only for a single organization and may provide services to one or more entities within the organization. Cloud services may also be provided under a community cloud model, where cloud infrastructure system 802 and the services provided by cloud infrastructure system 802 are shared by several organizations within an associated community. Cloud services may also be provided under a hybrid cloud model, which combines two or more different models.
[0058] In some embodiments, the services offered by cloud infrastructure system 802 may include one or more services offered in the Software as a Service (SaaS) category, Platform as a Service (PaaS) category, Infrastructure as a Service (IaaS) category, or other service category, including hybrid services. A customer may order one or more services offered by cloud infrastructure system 802 via a subscription order. Cloud infrastructure system 802 then performs processing to provide the services of the customer's subscription order.
[0059] In some embodiments, the services provided by the cloud infrastructure system 802 may include, but are not limited to, application services, platform services, and infrastructure services. In some examples, application services may be provided by the cloud infrastructure system through a SaaS platform. The SaaS platform may be configured to provide cloud services that fall into the SaaS category. For example, the SaaS platform may provide the ability to build and deliver a set of on-demand applications on an integrated development and deployment platform. The SaaS platform may manage and control the underlying software and infrastructure to provide the SaaS services. Customers may utilize the services provided by the SaaS platform to utilize applications that run on the cloud infrastructure system. Customers may obtain application services without purchasing separate licenses or support. A variety of different SaaS services may be provided. Examples include, but are not limited to, services that provide solutions for sales performance management, enterprise integration, and business flexibility for large organizations.
[0060] In some embodiments, platform services may be provided by the cloud infrastructure system through a PaaS platform. The PaaS platform may be configured to provide cloud services that fall under the PaaS category. Examples of platform services may include, but are not limited to, services that allow organizations (e.g., Oracle) to consolidate existing applications on a shared common architecture and build new applications that leverage shared services provided by the platform. The PaaS platform may manage and control the underlying software and infrastructure for providing the PaaS services. Customers may consume the PaaS services provided by the cloud infrastructure system without purchasing separate licenses or support. Examples of platform services include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), etc.
[0061] By utilizing the services provided by the PaaS platform, customers can use programming languages and tools supported by the cloud infrastructure system and also control the services deployed. In some embodiments, the platform services provided by the cloud infrastructure system may include database cloud services, middleware cloud services (e.g., Oracle Fusion Middleware services), and Java cloud services. In one embodiment, the database cloud services may support a shared services deployment model that allows organizations to pool database resources and provide databases as a service in the form of database clouds to customers. The middleware cloud services may provide a platform for customers to develop and deploy various business applications, and the Java cloud services may provide the cloud infrastructure system with a platform for customers to deploy Java applications.
[0062] A variety of different infrastructure services may be provided by IaaS platforms within a cloud infrastructure system that facilitate the management and control of underlying computing resources, such as storage, networking, and other fundamental computing resources, for customers that utilize the services provided by the SaaS and PaaS platforms.
[0063] In an embodiment, cloud infrastructure system 802 may also include infrastructure resources 830 for providing resources used to provide various services to customers of the cloud infrastructure system. In one embodiment, infrastructure resources 830 may include a pre-integrated and optimized combination of hardware, such as servers, storage, and networking resources, to run the services offered by the PaaS and SaaS platforms.
[0064] In some embodiments, resources in cloud infrastructure system 802 may be shared by multiple users and dynamically reallocated per request. Additionally, resources may be allocated to users in different time zones. For example, cloud infrastructure system 830 may make resources of the cloud infrastructure system available to a first set of users in a first time zone for a specified amount of time, and then allow the same resources to be reallocated to another set of users in a different time zone, thereby maximizing resource utilization.
[0065] In an embodiment, several internal shared services 832 may be provided that are shared by different components or modules of cloud infrastructure system 802 and by services provided by cloud infrastructure system 802. These internal shared services may include, but are not limited to, security and identity services, integration services, enterprise repository services, enterprise manager services, virus scanning and whitelist services, high availability, backup and recovery services, services enabling cloud support, email services, notification services, file transfer services, etc.
[0066] In an embodiment, cloud infrastructure system 802 may provide comprehensive management of cloud services (e.g., SaaS, PaaS, and IaaS services) in the cloud infrastructure system. In one embodiment, cloud management functions may include functions for provisioning, managing, and tracking customer subscriptions received by cloud infrastructure system 802, etc.
[0067] In one embodiment, as shown, cloud management functionality may be provided by one or more modules, such as an order management module 820, an order coordination module 822, an order provisioning module 824, an order management and monitoring module 826, and an identity management module 828. These modules may include or be provided using one or more computers and / or servers, which may be general purpose computers, dedicated server computers, server farms, server clusters, or any other suitable arrangement and / or combination.
[0068] In an example operation 834, a customer using a client device, such as client device 804, 806, or 808, may interact with cloud infrastructure system 802 by requesting one or more services offered by cloud infrastructure system 802 and ordering a subscription for one or more services offered by cloud infrastructure system 802. In an embodiment, the customer may access cloud user interfaces (UIs), cloud UI 812, cloud UI 814, and / or cloud UI 816, and place a subscription order via these UIs. Order information received by cloud infrastructure system 802 in response to the customer's order may include information identifying the customer and the one or more services offered by cloud infrastructure system 802 to which the customer intends to subscribe.
[0069] After an order is placed by a customer, the order information is received via cloud UI 812, 814, and / or 816.
[0070] At operation 836, the order is stored in order database 818. Order database 818 may be one of several databases operated by cloud infrastructure system 818 and in conjunction with other system elements.
[0071] At operation 838, the order information is forwarded to the order management module 820. In some cases, the order management module 820 may be configured to perform billing and accounting functions related to the order, such as validating the order and booking the order after validation.
[0072] At operation 840, information regarding the order is communicated to the order coordination module 822. The order coordination module 822 can utilize the order information to coordinate the provisioning of services and resources for the customer's order. In some examples, the order coordination module 822 can use the services of the order provisioning module 824 to coordinate the provisioning of resources to support the subscribed services.
[0073] In an embodiment, the order adjustment module 822 enables management of business processes associated with each order and applies business logic to determine whether the order should proceed to provisioning. At operation 842, upon receiving a new subscription order, the order adjustment module 822 sends a request to the order provisioning module 824 to allocate resources and configure the resources required to fulfill the subscription order. The order provisioning module 824 enables the allocation of resources for services ordered by a customer. The order provisioning module 824 provides a level of abstraction between the cloud services provided by the cloud infrastructure system 800 and the physical implementation layer used to provision resources to provide the requested services. Thus, the order adjustment module 822 can be decoupled from implementation details such as whether services and resources are actually provisioned on the fly or are pre-provisioned and allocated / assigned only upon request.
[0074] At operation 844 , once the services and resources have been provisioned, a notification of the services provided may be sent by the order provisioning module 824 of the cloud infrastructure system 802 to the customer on the client device 804 , 806 , and / or 808 .
[0075] At operation 846, the customer's subscription order may be managed and tracked by the order management and monitoring module 826. In some cases, the order management and monitoring module 826 may be configured to collect usage statistics for the services in the subscription order, such as the amount of storage used, the amount of data transferred, the number of users, system uptime and system downtime, etc.
[0076] In an embodiment, cloud infrastructure system 800 may include an identity management module 828. Identity management module 828 may be configured to provide identity services, such as access management and authorization services in cloud infrastructure system 800. In some embodiments, identity management module 828 may control information regarding customers who wish to utilize services provided by cloud infrastructure system 802. Such information may include information authenticating the identity of such customers and information describing actions those customers are authorized to perform on various system resources (e.g., files, directories, applications, communication ports, memory segments, etc.). Identity management module 828 may also include management of descriptive information regarding each customer and how that descriptive information can be accessed and modified by who.
[0077] 9 illustrates an exemplary computer system 900 in which various embodiments may be implemented. System 900 may be used to implement any of the computer systems described above. As shown, computer system 900 includes a processing unit 904 that communicates with a number of peripheral subsystems via a bus subsystem 902. These peripheral subsystems may include a processing acceleration unit 906, an I / O subsystem 908, a storage subsystem 918, and a communication subsystem 924. The storage subsystem 918 includes a tangible computer-readable storage medium 922 and a system memory 910.
[0078] Bus subsystem 902 provides a mechanism that allows the various components and subsystems of computer system 900 to communicate with each other as intended. Although bus subsystem 902 is shown diagrammatically as a single bus, alternative embodiments of the bus subsystem may utilize multiple buses. Bus subsystem 902 may be any of several types of bus structures including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. For example, such architectures may include an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MCA) bus, an Enhanced ISA (EISA) bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus. It may be implemented as a mezzanine bus manufactured in accordance with the IEEE P1386.1 standard.
[0079] The processing unit 904 may be implemented as one or more integrated circuits (e.g., conventional microprocessors or microcontrollers) and controls the operation of the computer system 900. One or more processors may be included in the processing unit 904. These processors may include single-core processors or multi-core processors. In an embodiment, the processing unit 904 may be implemented as one or more independent processing units 932 and / or 934 with a single-core processor or a multi-core processor included in each processing unit. In other embodiments, the processing unit 904 may be implemented as a quad-core processing unit formed by integrating two dual-core processors into a single chip.
[0080] In various embodiments, the processing unit 904 may execute various programs in response to program code and may maintain multiple simultaneously executing programs or processes. At any time, some or all of the program code being executed may reside in the processor 904 and / or in the memory subsystem 918. Through appropriate programming, the processor 904 may provide various functions as discussed above. The computer system 900 may further include a processing acceleration unit 906, which may include a digital signal processor (DSP), a special purpose processor, or the like.
[0081] The I / O subsystem 908 can include user interface input devices and user interface output devices. User interface input devices can include keyboards, pointing devices, such as mice or trackballs, touch pads or touch screens integrated into displays, scroll wheels, click wheels, dials, buttons, switches, keypads, audio input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices can include motion sensing and / or gesture recognizers, such as Microsoft Kinect® motion sensors, that allow users to control and interact with input devices, such as Microsoft Xbox® 360 game controllers, through a natural user interface using gestures and voice commands. User interface input devices can also include eye gesture recognizers, such as the Google Glass® blink detector, that detects eye activity from a user (e.g., “blinking” during picture taking and / or menu selection) and translates eye gestures as input to an input device (e.g., Google Glass®). Additionally, the user interface input devices may include voice recognition sensing devices that allow a user to interact with a voice recognition system (e.g., the Siri® navigator) through voice commands.
[0082] User interface input devices may include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, game pads and graphic tablets, and audio / visual devices, such as, but not limited to, speakers, digital cameras, digital video cameras, portable media players, webcams, image scanners, fingerprint scanners, barcode readers 3D scanners, 3D printers, laser range finders, eye-tracking devices, etc. Additionally, user interface input devices may include medical imaging input devices, such as, for example, computed tomography, magnetic resonance imaging, position emission tomography, medical ultrasound machines, etc. User interface input devices may also include audio input devices, such as, for example, MIDI keyboards, digital musical instruments, etc.
[0083] User interface output devices may include non-visual displays such as a display subsystem, indicator lights, or audio output devices. The display subsystem may be a flat panel device such as one using a cathode ray tube (CRT), liquid crystal display (LCD) or plasma display, a projection device, a touch screen, etc. In general, use of the term "output device" is intended to include all possible types of devices and mechanisms for outputting information from computer system 900 to a user or to another computer. For example, user interface output devices may include, but are not limited to, various display devices that visually convey text, graphics, and audio / video information, such as monitors, printers, speakers, headphones, automobile navigation systems, plotters, voice output devices, and modems.
[0084] Computer system 900 may include a storage subsystem 918 that comprises software elements shown as currently residing in system memory 910. System memory 910 may store program instructions that are loadable and executable on processing unit 904, as well as data generated during the execution of these programs.
[0085] Depending on the configuration and type of computer system 900, the system memory 910 may be volatile (such as random access memory (RAM)) and / or non-volatile (such as read only memory (ROM), flash memory, etc.). RAM typically contains data and / or program modules that are immediately accessible to and / or currently being operated on and executed by the processing unit 904. In some implementations, the system memory 910 may include a number of different types of memory, such as static random access memory (SRAM) or dynamic random access memory (DRAM). In some implementations, a basic input / output system (BIOS), containing the basic routines that help to transfer information between elements within the computer system 900, such as during start-up, may typically be stored in ROM. By way of example and not limitation, the system memory 910 also illustrates application programs 912, which may include client applications, web browsers, mid-tier applications, relational database management systems (RDBMS), and the like, as well as program data 914, and an operating system 916. As an example, operating system 916 may include various versions of Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various commercially available UNIX® or UNIX-like operating systems (including but not limited to various GNU / Linux operating systems, Google Chrome® OS, etc.), and / or mobile operating systems such as iOS, Windows® Phone, Android® OS, BlackBerry® 10 OS, and Palm® OS operating systems.
[0086] The storage subsystem 918 may also provide a tangible computer-readable storage medium for storing the basic programming and data constructs that provide the functionality of some embodiments. Software (programs, code modules, instructions) that, when executed by a processor, provide the functionality described above may be stored in the storage subsystem 918. These software modules or instructions may be executed by the processing unit 904. The storage subsystem 918 may also provide a repository for storing data used in accordance with some embodiments.
[0087] Storage subsystem 900 may also include a computer-readable storage medium reader 920 that may be further coupled to a computer-readable storage medium 922. Together, and optionally in combination with system memory 910, computer-readable storage medium 922 may comprehensively represent remote, local, fixed, and / or removable storage devices and storage media for containing, storing, transmitting, and retrieving computer-readable information on a temporary and / or more permanent basis.
[0088] The computer readable storage medium 922 containing the code or portions of code may include any suitable medium including storage media and communication media, such as, but not limited to, volatile and non-volatile, removable and non-removable media implemented in any manner or technology for storing and / or transmitting information. This may include tangible computer readable storage media, such as RAM, ROM, Electronically Erasable Programmable ROM (EEPROM), flash memory or other memory technology, CD-ROM, Digital Versatile Disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage, or other tangible computer readable media. This may also include intangible computer readable media, such as a data signal, data transmission, or any other medium that can be used to transmit the desired information and that can be accessed by the computing system 900.
[0089] As an example, the computer readable storage medium 922 may include a hard disk drive that reads or writes to a non-removable non-volatile magnetic medium, a magnetic disk drive that reads or writes to a removable non-volatile magnetic disk, an optical disk drive that reads or writes to a removable non-volatile optical disk such as a CD-ROM, DVD, Blu-Ray® disk, or other optical medium. The computer readable storage medium 922 may include, but is not limited to, a Zip® drive, a flash memory card, a Universal Serial Bus (USB) flash drive, a Secure Digital (SD) card, a DVD disk, a digital video tape, and the like. The computer readable storage medium 922 may also include a solid state drive (SSD) based on non-volatile memory such as a flash memory-based SSD, an enterprise flash drive, a solid state ROM, an SSD based on volatile memory such as solid state RAM, dynamic RAM, static RAM, DRAM-based SSD, a magnetoresistive RAM (MRAM) SSD, and a hybrid SSD that uses a combination of DRAM and a flash memory-based SSD. The disk drives and their associated computer-readable media may provide non-volatile storage of computer-readable instructions, data structures, program modules, and other data for computer system 900.
[0090] The communications subsystem 924 provides an interface to other computer systems and networks. The communications subsystem 924 serves as an interface for transmitting and receiving data from the computer system 900 to and from other systems. For example, the communications subsystem 924 may enable the computer system 900 to connect to one or more devices via the Internet. In some embodiments, the communications subsystem 924 may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using cellular technology, advanced data network technologies such as 3G, 4G, or EDGE (Enhanced Data Rates for Global Evolution)), WiFi (IEEE 802.11 family standard, or other mobile communications technologies, or any combination thereof), global positioning system (GPS) receiver components, and / or other components. In some embodiments, the communications subsystem 924 may provide a wired network connection (e.g., Ethernet) in addition to or instead of a wireless interface.
[0091] In some embodiments, the communications subsystem 924 may also receive incoming communications in the form of structured and / or unstructured data feeds 926, event streams 928, event updates 930, etc., on behalf of one or more users who may use the computer system 900.
[0092] By way of example, the communications subsystem 924 may be configured to receive data feeds 926 in real time from users of social networks and / or other communications services, such as web feeds, such as Twitter® feeds, Facebook® updates, rich site summary (RSS) feeds, and / or real-time updates from one or more third-party information sources.
[0093] Additionally, the communications subsystem 924 may be configured to receive data in the form of a continuous data stream, which may include an event stream 928 of real-time events and / or event updates 930, which may be continuous in nature or unlimited with no explicit end. Examples of applications that generate continuous data may include, for example, sensor data applications, financial tickers, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, automobile traffic monitoring, etc.
[0094] The communications subsystem 924 may also be configured to output structured and / or unstructured data feeds 926, event streams 928, event updates 930, etc. to one or more databases that may be in communication with one or more streaming data source computers coupled to the computer system 900.
[0095] The computer system 900 may be one of a variety of types, including a handheld portable device (e.g., an iPhone® mobile phone, an iPad® computing tablet, a PDA), a wearable device (e.g., a Google Glass® head mounted display), a PC, a workstation, a mainframe, a kiosk, a server rack, or other data processing system.
[0096] Due to the ever-changing nature of computers and networks, the description of the computer system 900 shown in the figure is intended as a specific example only. Many other configurations are possible with more or fewer components than the system shown in the figure. For example, customized hardware may be used and / or particular elements may be implemented in hardware, firmware, software (including applets), or a combination thereof. Additionally, connections to other computing devices, such as network input / output devices, may be used. Based on the disclosure and teachings provided herein, other methods and / or methods for implementing various embodiments should become apparent.
[0097] In the preceding description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of various embodiments. However, it will be apparent that some embodiments may be practiced without some of these specific details. In other instances, well-known structures and devices are shown in block diagram form.
[0098] The foregoing description provides exemplary embodiments only and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the foregoing description of various embodiments provides an enabling disclosure for implementing at least one embodiment. It should be understood that various changes can be made in the function and arrangement of elements without departing from the spirit and scope of the several embodiments as set forth in the appended claims.
[0099] Specific details are given in the foregoing description to provide a thorough understanding of the embodiments. However, it will be understood that the embodiments can be practiced without these specific details. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the embodiments.
[0100] It should also be noted that the particular embodiments may be described as a process, which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. Additionally, the order of operations may be rearranged. A process terminates when an operation is completed, but may include additional steps not included in the diagram. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, or the like. When a process corresponds to a function, its termination may correspond to a return of the function to the calling function or the main function.
[0101] The term "computer-readable medium" includes, but is not limited to, portable or fixed storage devices, optical storage devices, wireless channels, and various other media that can store, contain, or carry instructions and / or data. A code segment or machine-executable instructions may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing or receiving information, data, arguments, parameters, or memory contents. The information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means, such as memory sharing, message passing, token passing, network transmission, etc.
[0102] Furthermore, the embodiments may be implemented by hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof. When implemented by software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks may be stored in a machine-readable medium. A processor may perform the necessary tasks.
[0103] In the foregoing specification, features are described with reference to specific embodiments thereof, but it should be recognized that not all embodiments are limited thereto. Various features and aspects of the several embodiments may be used individually or in combination. Moreover, the embodiments may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. The specification and drawings are therefore to be regarded as illustrative rather than restrictive.
[0104] Furthermore, for illustrative purposes, the methods have been described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in an order different from that described. It should also be understood that the methods described above may be performed by hardware components or embodied in a sequence of machine-executable instructions, which may be used to cause a machine, such as a general-purpose or special-purpose processor or logic circuitry that is programmed with the instructions, to perform the methods. These machine-executable instructions may be stored on one or more machine-readable media, such as, for example, a CD-ROM or other type of optical disk, a floppy diskette, a ROM, a RAM, an EPROM, an EEPROM, a magnetic or optical card, a flash memory, or any other type of machine-readable medium suitable for storing electronic instructions. Alternatively, the methods may be performed by a combination of hardware and software.
Claims
1. When executed by one or more processors, the one or more processors receive a first document having a first schema, and access a configuration of the first schema, the configuration defining a method for generating a plurality of queries from the first document to a set of documents having a second schema, generate the plurality of queries based on the configuration, and combine results of the plurality of queries into a similarity score of the first document, the program including instructions for causing an operation to be performed.
2. The program according to claim 1, wherein the first schema is different from the second schema.
3. The program according to claim 1, wherein the first schema defines a service request, the set of documents is part of a knowledge base, and the second schema defines a text document including a solution to the service request.
4. The program according to claim 1, wherein the configuration defines a method for generating queries from the first document to a set of documents having a plurality of different schemas.
5. The program according to claim 1, wherein the plurality of queries are submitted to a search interface including an inverted index for receiving boolean queries and phrase queries, and an application programming interface (API) for receiving a word and returning a number of documents in the set of documents in which the word is used.
6. The program according to claim 1, wherein the first schema defines a plurality of field and value pairs.
7. The program according to claim 1, wherein the configuration includes a query type defining an n-gram level of a first subset of the plurality of queries for a first field of the first document.
8. The program according to claim 7, wherein the configuration further includes a number N of queries generated for the query type.
9. Generating the number N of queries for the query type includes determining a word frequency score of the first field from the set of documents, identifying the words in the first field having the highest N frequency scores, and generating N queries from the words in the first field having the highest N frequency scores, the program according to claim 8.
10. The program according to claim 9, wherein the frequency score is determined based on the number of times a word appears in the first document and the number of documents in the set of documents in which the word appears. **Claim 11** The program according to claim 7, wherein the configuration further includes one or more target fields in the second schema for the query type. **Claim 12** The program according to claim 11, wherein the configuration further includes a weight applied to a similarity score of a query generated from the first target field among the one or more target fields. **Claim 13** The program according to claim 12, wherein the weight is set for the configuration by a machine learning model. **Claim 14** The program according to claim 1, wherein the configuration is one of a plurality of configurations, and the plurality of configurations correspond to a plurality of different schemas. **Claim 15** The program according to claim 1, wherein the first document is received as part of a search request for identifying documents in the set of documents similar to the first document. **Claim 16** The program according to any one of claims 1 to 15, wherein the operation further includes executing the plurality of queries on the set of documents. **Claim 17** The program according to claim 16, wherein the results of the plurality of queries include a score of a second document in the set of documents, and the score of the second document is generated in response to the plurality of queries. **Claim 18** The program according to claim 17, wherein combining the results of the plurality of queries into a similarity score includes generating a weighted combination of scores for the second document. **Claim 19** A system comprising: one or more processors; when executed by the one or more processors, receiving a first document having a first schema; accessing a configuration of the first schema, the configuration defining a method for generating a plurality of queries from the first document to a set of documents having a second schema; generating the plurality of queries based on the configuration; combining the results of the plurality of queries into a similarity score of the first document; one or more memory devices including instructions for causing the one or more processors to perform operations including the above. Claim 20 A method for calculating a similarity score of a document, the method comprising: receiving a first document having a first schema; and accessing a configuration of the first schema, the configuration defining a method for generating a plurality of queries from the first document to a set of documents having a second schema; the method comprising: generating the plurality of queries based on the configuration; and combining the results of the plurality of queries into a similarity score of the first document.