A system and method for smart content categorization in a content management system.
An AI-driven tool for content categorization in content management systems uses feature vectors to automate the classification process, improving efficiency and accuracy by learning from past data and recommending categories, thus addressing the challenges of manual classification.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ORACLE INT CORP
- Filing Date
- 2026-02-09
- Publication Date
- 2026-06-02
AI Technical Summary
The task of correctly classifying and reclassifying content in content management systems is costly and error-prone, especially as the volume of content and number of taxonomies increase, requiring a complex and time-consuming process.
An AI-driven tool that uses feature vectors to automatically categorize content by learning from past data, generating clusters, and recommending categories based on similarity to existing content, assisted by a graphical user interface for content authors.
Facilitates efficient and accurate content categorization by leveraging AI to reduce manual effort and minimize errors, enhancing the content management process.
Smart Images

Figure 2026090373000001_ABST
Abstract
Description
[Technical Field]
[0001] Copyright Notice Parts of the disclosure in this patent document contain subject matter protected by copyright. The copyright holder has no objection to anyone reproducing the patent document or patent information disclosure so that it may be published in the file or records of the Patent and Trademark Office, but otherwise retains all copyrights.
[0002] Priority claims and cross-references of related applications: This application claims priority to U.S. Provisional Patent Application No. 63 / 084,174, filed on September 28, 2020, titled "SYSTEM AND METHOD FOR SMART CATEGORIZATION OF CONTENT IN A CONTENT MANAGEMENT SYSTEM," and U.S. Patent Application No. 17 / 486,524, filed on September 27, 2021, also titled "SYSTEM AND METHOD FOR SMART CATEGORIZATION OF CONTENT IN A CONTENT MANAGEMENT SYSTEM," and also claims priority to Indian Provisional Patent Application No. 17 / 486,524, filed on October 18, 2018, titled "SMART CONTENT RECOMMENDATIONS FOR AUTHORS." U.S. Patent Application No. 16 / 581,138, filed on September 24, 2019, titled "Smart Content Recommendations for Content Authors," claims priority to Patent Application No. 201841039495. A continuation-in-part application claiming priority, U.S. Patent Application No. 16 / 657,395, filed on 18 October 2019, titled "Techniques for Ranking Content Item Recommendations," has been filed. This relates to the above-mentioned applications, and each of them and their contents are incorporated herein by reference.
[0003] This application generally relates to online commerce environments and the management and distribution of content data, and in particular to smart categorization / classification of content in content management systems. [Background technology]
[0004] background: Creators and authors of original content intended for online publishing and / or transmission may use a wide variety of software-based tools and technologies to generate, edit, and store their newly created content.
[0005] In a content management system, various types of content (e.g., documents, structured content like blogs, articles, press releases, and media files like images and videos) often need to be evaluated / categorized based on their content. This categorization / classification occurs across a hierarchical set of categories or nodes. For example, a real estate lease agreement document might be evaluated / categorized under Legal Documents → Real Estate → Contracts. It is also possible for the same document (or content) to be categorized / classified more than once simultaneously. For example, the same contract document might exist under Valid Contracts → Signed.
[0006] Categories are grouped under an organizational concept called taxonomy. Content tends to have many taxonomies that reflect the business organization. When new documents or content items are added, or when new taxonomies arise, or when there are significant changes in the content organization, the task of correctly classifying or reclassifying the content is assigned to the end user (or content author). This can be a costly and error-prone task as the volume of content and the number of taxonomies increase. [Overview of the project] [Means for solving the problem]
[0007] overview: According to one embodiment, the systems and methods described herein can be used, for example, in conjunction with a content management system to provide recommendations for categorizing / classifying content into user-defined categories, thereby providing content managers with the opportunity to easily place new content into accurate categories based on pre-evaluated / categorized content.
[0008] Classifying vast amounts of content online is a complex task, fraught with challenges such as single-pass constraints on data and the requirement for fast response times. According to one embodiment, content users categorize similar content through logical clusters, such as hierarchical taxonomy trees, placing similar content in the same node / category within the taxonomy tree. Over time, as both the number of nodes and content entities in the taxonomy tree increases, similar content entities will coexist within nodes. Given this state of content organization, content already evaluated / categorized within taxonomies can be used by computer algorithms to determine where newly created / edited content may belong.
[0009] According to an embodiment, a recommendation system or tool uses artificial intelligence (AI) technology to continuously learn from past data and assist in placing newly created / edited content into relevant categories by automatically categorizing / classifying the content. The recommendation tool can be implemented and applied across various domains by generating feature vectors from the content, creating clusters in the feature space based on pre-categorized content, and recommending a category for new content by calculating the feature space distance from the clusters.
[0010] Aspects of the present disclosure relate to an AI-driven tool configured to function as a smart digital assistant for recommending images, text content, and other related media content from a content repository. Certain embodiments may include a front-end software tool having a graphical user interface (GUI) to complement a content authoring interface used to author original media content (e.g., blog posts, online articles, etc.). In some cases, additional GUI screens and features may be incorporated into an existing content authoring software tool, for example, as a software plugin. The smart digital content recommendation tool communicates with several backend services and a content repository to, for example, analyze text and / or visual inputs, extract keywords or topics from the inputs, classify and tag the input content, and store the classified / tagged content in one or more content repositories.
[0011] (e.g., directly by a software tool and / or indirectly by calling a backend service) various Additional techniques that may be performed in certain embodiments include converting the input text and / or image into vectors within a multi-dimensional vector space and comparing the input content with multiple repository contents to discover some relevant content options within the content repository. Such comparisons may include a complete and comprehensive deep search and / or a more efficient tag-based filtered search. Finally, relevant content items (e.g., images, audio and / or video clips, links to related articles, etc.) are retrieved and presented to the content author for review and may be embedded within the original authored content.
[0012] Although the description herein mainly shows application to text content, according to various embodiments, this approach can be extended to other types of content, such as multimedia or images / videos, by metadata extraction.
[0013] Brief description of the drawing: The nature and advantages of embodiments in accordance with the present disclosure can be further understood by reference to the remainder of this specification in conjunction with the accompanying drawings.
[0014] In the accompanying drawings, similar components and / or features may have the same reference levels. Further, different components of the same type may be distinguished by a dash and a second label following the reference label to identify similar components. When a first reference label is used in this specification, the description is applicable to any one of the similar components having the same first reference label regardless of the second reference label.
Brief Description of the Drawings
[0015] [Figure 1] A diagram showing an exemplary computer system architecture including a data integration cloud platform in which certain embodiments of the present disclosure may be implemented. [Figure 2]This figure shows an exemplary screen of a customized dashboard in a user interface used to configure, monitor, and control a service instance, according to a particular embodiment of this disclosure. [Figure 3] This is an architectural diagram illustrating a data integration cloud platform in accordance with a specific embodiment of this disclosure. [Figure 4] This figure shows an exemplary computing environment configured to perform content classification and recommendation in accordance with a particular embodiment of the present disclosure. [Figure 5] Another diagram shows an exemplary computing environment configured to perform content classification and recommendation according to a particular embodiment of the present disclosure. [Figure 6] This flowchart shows a process for generating feature vectors based on content resources in a content repository, according to a specific embodiment of the present disclosure. [Figure 7] This figure shows an exemplary image that identifies multiple image features according to a particular embodiment of the present disclosure. [Figure 8] This figure shows an example of a text document illustrating a keyword extraction process according to a specific embodiment of the present disclosure. [Figure 9] This figure shows a process for generating and storing image tags according to a specific embodiment of the present disclosure. [Figure 10] This figure shows a process for generating and storing image tags according to a specific embodiment of the present disclosure. [Figure 11] This figure shows a process for generating and storing image tags according to a specific embodiment of the present disclosure. [Figure 12] This flowchart shows another process for comparing feature vectors to identify related content in a content repository, according to a particular embodiment of the present disclosure. [Figure 13] This figure shows a technique for converting an image file into a feature vector, according to a specific embodiment of the present disclosure. [Figure 14]This figure shows an exemplary vector space in which feature vectors are populated according to a particular embodiment of the present disclosure. [Figure 15] This figure shows a deep feature space vector comparison according to a specific embodiment of the present disclosure. [Figure 16] This figure shows a filtered feature space vector comparison according to a specific embodiment of the present disclosure. [Figure 17] This figure shows a filtered feature space vector comparison according to a specific embodiment of the present disclosure. [Figure 18] This figure shows a process for receiving and processing text input to identify a relevant image or article, according to a specific embodiment of the present disclosure. [Figure 19] This is an illustrative diagram showing a comparison of extracted keywords and image tags according to a specific embodiment of the present disclosure. [Figure 20] This figure shows an example of keyword analysis in a 3D word vector space according to a specific embodiment of the present disclosure. [Figure 21] This figure shows a vector space analysis of keywords and tags according to a specific embodiment of the present disclosure. [Figure 22] This figure shows an example of a homophone image tag according to a specific embodiment of the present disclosure. [Figure 23] This figure shows an exemplary deambiguation process according to a particular embodiment of the present disclosure. [Figure 24] This figure shows an exemplary deambiguation process according to a particular embodiment of the present disclosure. [Figure 25] This figure shows a process for comparing feature vectors to identify relevant content in a content repository, according to a specific embodiment of this disclosure. [Figure 26] This figure shows a process for comparing feature vectors to identify relevant content in a content repository, according to a specific embodiment of this disclosure. [Figure 27]This figure shows a process for comparing feature vectors to identify relevant content in a content repository, according to a specific embodiment of this disclosure. [Figure 28] This figure shows a process for comparing feature vectors to identify relevant content in a content repository, according to a specific embodiment of this disclosure. [Figure 29] This is an illustrative diagram of a text document illustrating a topic extraction process according to a specific embodiment of the present disclosure. [Figure 30] This is an illustrative diagram of a text document illustrating a topic extraction process according to a specific embodiment of the present disclosure. [Figure 31] This figure shows a process for identifying relevant articles based on input text data, according to a specific embodiment of this disclosure. [Figure 32] This figure shows a process for identifying relevant articles based on input text data, according to a specific embodiment of this disclosure. [Figure 33] This figure shows a process for identifying relevant articles based on input text data, according to a specific embodiment of this disclosure. [Figure 34] This figure shows a process for identifying relevant articles based on input text data, according to a specific embodiment of this disclosure. [Figure 35] This figure shows a process for identifying relevant articles based on input text data, according to a specific embodiment of this disclosure. [Figure 36] This figure shows an exemplary semantic text analyzer system according to a particular embodiment of the present disclosure. [Figure 37] This figure shows an exemplary user interface screen illustrating image recommendations provided to the user during content creation, in accordance with a particular embodiment of this disclosure. [Figure 38] This figure shows an exemplary user interface screen illustrating image recommendations provided to the user during content creation, in accordance with a particular embodiment of this disclosure. [Figure 39]This is a simplified diagram showing a distributed system for implementing a specific embodiment in accordance with this disclosure. [Figure 40] This is a simplified block diagram showing one or more components in a system environment, where the services provided by one or more components of the system may be provided as cloud services, according to a particular embodiment of the present disclosure. [Figure 41] This figure shows an exemplary computer system in which various embodiments can be implemented. [Figure 42] This figure shows an exemplary computing environment configured to evaluate and rank content items from a content repository in response to input content received from a user or client system, according to a particular embodiment of the present disclosure. [Figure 43] This flowchart shows a process for identifying and ranking content items related to user content, according to a specific embodiment of this disclosure. [Figure 44] This figure shows an exemplary screen of a content authoring user interface according to a specific embodiment of the present disclosure. [Figure 45] This figure shows an exemplary table of sets of matching content items identified by a content recommendation system according to a particular embodiment of the present disclosure. [Figure 46] This figure shows another exemplary table of a set of matching content items, including ranking scores, according to a particular embodiment of the present disclosure. [Figure 47] This figure shows another exemplary screen of a content authoring user interface according to a particular embodiment of the present disclosure. [Figure 48] This figure shows an example of a content management system environment according to one embodiment. [Figure 49] This figure shows an exemplary use of a content management system for managing and distributing content data, according to one embodiment. [Figure 50]This is a smart content classification flow diagram according to a certain embodiment. [Figure 51] This is a flowchart illustrating the creation of a taxonomy according to a certain embodiment. [Figure 52] This is a taxonomy modification flowchart according to a certain embodiment. [Figure 53] This figure shows a sample taxonomy tree, along with a category diagram, according to one embodiment. [Figure 54] This figure further illustrates a sample taxonomy tree according to one embodiment, along with a category diagram. [Figure 55] This figure further illustrates a sample taxonomy tree according to one embodiment, along with a category diagram. [Figure 56] This figure shows the configuration of an automatic classification threshold diagram according to a certain embodiment. [Figure 57] This diagram shows a configuration that triggers (bulk) reclassification of content in a repository diagram, according to one embodiment. [Figure 58] This figure shows the problems associated with conventional clustering, which allows clusters to be arranged in any shape according to a certain embodiment. [Figure 59] This figure shows microclustering according to one embodiment. [Figure 60] This figure shows the sample topic distribution of clothing customers according to a certain embodiment. [Figure 61] This figure shows a visualization of cluster radius according to one embodiment. [Figure 62] This figure shows a cluster representation according to a certain embodiment. [Figure 63] This figure shows the macro-stages of categorization according to one embodiment. [Figure 64] This diagram shows the categorization process when a higher-level category is selected through macro steps. [Figure 65] This figure shows how cluster weights can decay over time according to one embodiment, and also illustrates an example of a decay window model. [Figure 66] This is a graph illustrating how shadow clusters can appear according to a certain embodiment. [Figure 67] This is a graph illustrating how shadow clusters can appear according to a certain embodiment. [Figure 68] This figure shows a suggestion to the user to create a new category from uncategorized content, according to one embodiment. [Figure 69] This is a flowchart illustrating a method for smart content categorization in a content management system according to one embodiment. [Modes for carrying out the invention]
[0016] Detailed explanation: In one embodiment, the present invention is shown as an example, not as an limitation, in the figures in the accompanying drawings, where similar reference numerals indicate similar elements. When referring to “an,” “one,” or “some” embodiments in this disclosure, it is not necessary to It should be noted that this does not refer to the same embodiment, and where it does, it means at least one. While specific implementation examples are discussed, it should be understood that these specific implementation examples are provided for illustrative purposes only. Those skilled in the art will recognize that other components and configurations may be used without departing from the scope and spirit of the invention.
[0017] The following description includes specific details for illustrative purposes to help you fully understand various implementation examples. However, it will become clear that various implementation examples can be carried out without these specific details. For example, circuits, systems, algorithms, structures, technologies, networks, processes, and other components may be shown as components in the form of block diagrams to avoid obscuring the implementation examples with unnecessary details. The diagrams and descriptions are not intended to be limiting.
[0018] Some examples disclosed in relation to the diagrams of this disclosure may be described as processes shown as flowcharts, flow diagrams, data flow diagrams, structure diagrams, sequence diagrams, or block diagrams. Sequence diagrams or flowcharts may describe operations as a continuous process, although many operations may be performed in parallel or simultaneously. In addition, the order of operations may be rearranged. A process terminates when its operations are complete, but may have additional steps not shown in the diagram. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. If a process corresponds to a function, the termination of that process may correspond to returning the corresponding function to a calling function or main function.
[0019] Processes described herein, such as those described with reference to the figures of this disclosure, may be implemented in software (e.g., code, instructions, programs), hardware, or a combination thereof, executed by one or more processing units (e.g., processor cores). The software may be stored in memory (e.g., on a memory device, on a non-temporary computer-readable storage medium). In some examples, processes shown in the sequence diagrams and flowcharts herein may be implemented by any of the systems disclosed herein. A particular set of processing steps in this disclosure is not intended to limit the scope. Other sequences of steps may be performed according to alternative examples. For example, alternative examples in this disclosure may perform the steps outlined above in a different order. Furthermore, individual steps shown in the figures may include multiple substeps that can be performed in various orders suitable for the individual step. Furthermore, additional steps may be added or removed depending on the particular application. Those skilled in the art will recognize many variations, modifications, and alternatives.
[0020] In some examples, each process in the diagrams of this disclosure may be executed by one or more processing units. A processing unit may include one or more processors, including a single-core or multi-core processor, one or more cores of a processor, or a combination thereof. In some examples, a processing unit may include one or more dedicated coprocessors, such as a graphics processor or a digital signal processor (DSP). In some examples, some or all of the processing units may be application-specific integrated circuits (ASICs) or field-programmable integrated circuits. Customized gate arrays (Field programmable gate arrays: FPGAs) It can be implemented using a circuit.
[0021] The specific embodiments described herein may be implemented as part of a Data Integration Platform Cloud (DIPC). Generally speaking, Integration involves combining data from different data sources and providing users with unified access and a unified view of that data. This process is frequent and is crucial in many situations, such as merging existing legacy databases with commercial entities. As the volume of data continues to grow to match the ability to analyze it to provide useful results ("big data"), data integration is becoming more frequent in enterprise software systems. For example, consider a web application where users can query various types of travel information (e.g., weather, hotels, airlines, demographics, crime statistics, etc.). The enterprise application can combine many heterogeneous data sources using unified views and virtual schemas within DIPC, without needing to store all of these different data types in a single schema in a single database, and present them to the user in a unified view.
[0022] DIPC is a cloud-based platform for data transformation, integration, replication, and management. It moves batch and real-time data between cloud and on-premises data sources while maintaining data consistency with default tolerances and resilience. DIPC can connect to various data sources and be used to prepare, transform, replicate, manage, and / or monitor data from these sources when they are combined into one or more data warehouses. DIPC can work with any type of data source and support any type of data in any format. DIPC is a Platform as a Service (PaaS), or Infrastructure as a Service. Using the Infrastructure as a Service (IaaS) architecture, We can provide cloud-based data integration for the prize.
[0023] DIPC can offer several different utilities, including transferring entire data sources to new cloud-based deployments and enabling easy access from cloud platforms to cloud databases. It can stream data in real time to keep it up-to-date with new data sources, while maintaining any number of distributed data sources in sync. The load may be split among the synchronized data sources to ensure extremely high availability to end users. The underlying data management system can be used to reduce the amount of data moved across the network for deployment to database clouds, big data clouds, third-party clouds, etc. Reusable Extract, Load, and Transform (ELT) functions and templates can be executed using a drag-and-drop user interface. A real-time test environment can be used to replicate data to ensure extremely high data availability to end users. It can be created to perform reporting and data analysis within the cloud on a data source. Data migration can be performed with zero downtime using replicated, synchronized data sources. Synchronized data sources can also be used for seamless disaster recovery while maintaining availability.
[0024] Figure 1 shows a computer system architecture that utilizes DIPC to integrate data from various existing platforms, according to several embodiments. The first data source 102 may include a cloud-based storage repository. The second data source 104 may include an on-premises data center. To provide uniform access and views to the first data source 102 and the second data source 104, DIPC 108 can copy data from the first data source 102 and the second data source 104 using an existing library of high-performance ELT functions 106. DIPC 108 can also extract, enrich, and transform the data once it is stored in the new cloud platform. DIPC 108 further enables access to any big data utilities that reside within or are accessible by the cloud platform. In some embodiments, the replicated data sources within the cloud platform can be used for testing, monitoring, management, and big data analysis, while the original data sources 102 and 104 may continue to provide access to customers. In some embodiments, data management may be provided for profiling, cleansing, and managing data sources within a set of customized dashboards existing within the user interface.
[0025] Figure 2 shows a customized dashboard in the user interface that can be used to configure, monitor, and control service instances in DIPC 108. The summary dashboard 202 may provide controls 204 that allow the user to create a service instance. The user can then be presented with a series of progressive web forms, sequentially presenting the types of information used to create the service instance. In the first step, the user will be asked to provide an email address and a service name and description with a service edition type. The user may also be asked about the cluster size, which specifies the number of virtual machines used in the service. Depending on the service edition type, it will be determined which applications are installed on the virtual machines. In the second step and in the corresponding web form, the user may provide a running cloud database deployment to store the schema of the DIPC server. Later, the same database can also be used to store data entities and perform integration tasks. In addition, the storage cloud may be designated and / or provisioned as a backup utility. The user may also provide credentials that can be used to access existing data sources used for data integration. In the third step, the provisioning information may be verified and the service instance may be created. Next, the new service instance may appear in the summary area 206 of the summary dashboard 202. From there, the user can access any information about any running data integration service instance.
[0026] Figure 3 shows an architectural diagram of DIPC according to several embodiments. Requests may be received via a browser client 302, which may be implemented using the component's Java Script® Extension Toolkit (JET) set. Alternatively, or additionally, the system may also receive requests via a DIPC agent 304 operating in the customer's on-premises data center 306. Agent 304 may include data integration agents 308 and 310 for replication services such as Oracle's GoldenGate® service. Each of these agents 308 and 310 may retrieve information from the on-premises data center 306 during normal operation and return the data to DIPC using the connectivity service 312.
[0027] Incoming requests can be passed through DIPC to a sign-in service 314, which may include load balancing or other utilities for routing requests. The sign-in service 314 may also use identity management services, such as an identity cloud service 316, to provide security and identity management for the cloud platform as part of an integrated enterprise security fabric. The identity cloud service 316 may manage user identities for both the cloud deployments and on-premises applications described in this embodiment. In addition to the identity cloud service 316, DIPC may also use a PaaS Service Manager (PSM) tool 318 to provide an interface for managing the lifecycle of platform services in cloud deployments. For example, the PSM tool... Using RU318, you can create and manage instances of data integration services on a cloud platform.
[0028] DIPC may be implemented on a web logical server 320 to build and deploy enterprise applications in a cloud environment. DIPC may include a local repository 322 that stores data policies, design information, metadata, and audit data related to information passing through DIPC. DIPC may also include a monitoring service 324 for populating the local repository 322. A catalog service 326 may include a set of machine-readable open APIs to provide access to many SaaS and PaaS applications in cloud deployments. The catalog service 326 may include distributed indexing such as Apache Solr®. It may also be available to search application 338 that uses a matching service. The connection service 328 and mediation service 330 can manage connections and can translate, verify, and route the logic of information passing through the DIPC. The information within the DIPC is in the Event Driven Architecture (EDA) and corresponding It may also be transmitted using message bus 332.
[0029] DIPC may also include an orchestration service 334. The orchestration service 334 may enable automated tasks by calling REST endpoints, scripts, third-party automation frameworks, etc. These tasks may then be executed by the orchestration service 334 to provide DIPC functionality. The orchestration service 334 may use runtime services to import, transform, and store data. For example, the ELT runtime service 334 may run the ELT functionality library described above, and the replication runtime service 342 may copy data from various data sources to the cloud-deployable DIPC repository 316. In addition, DIPC may include a code generation service 336 that automatically generates code for both the ELT functionality and the replication functionality.
[0030] Smart Content - Smart Content Recommendations As mentioned above, when users create / authorize original media content (e.g., articles, press releases, emails, blog posts, etc.), it is often useful to enhance the authored content with relevant additional content such as related images, audio / video clips, links to related articles, or other content. However, Searching for such supplementary content, and even embedding it within a user's original authored content, can be challenging in several ways. The first challenge may involve finding safe and reliable supplementary content from trusted sources and ensuring that the user / author is authorized to incorporate that content into their work. In addition, locating and embedding any relevant content from such a safe and authorized content repository into the user / author's original authored content can be an inefficient process requiring a great deal of manual work.
[0031] Therefore, the specific aspects described herein relate to smart digital content recommendation tools. In certain embodiments, a smart digital content recommendation tool may be an artificial intelligence (AI)-driven tool configured to process and analyze input content (e.g., text, images) from a content author in real time and recommend relevant images, additional text content, and / or other relevant media content (e.g., audio or video clips, graphics, social media posts, etc.) from one or more trusted content repositories. The smart digital content recommendation tool may communicate with several backend services and content repositories to, for example, analyze text and / or visual input, extract keywords or topics from the input, classify and tag the input content, and store the classified / tagged content in one or more content repositories.
[0032] The additional aspects described herein may be performed directly and / or indirectly by calling various backend services via smart digital content recommendation tools, each running on a client operated by a content author, and may include: (a) receiving original content as input in the form of text and / or images; (b) extracting keywords and / or topics from the original content; (c) determining and storing associated keyword and / or topic tags about the original content; (d) converting the original content (e.g., input text and / or images) into vectors in a multidimensional vector space; (e) comparing such vectors with multiple other content vectors representing each of the additional content in the content repository in order to discover and identify various potentially related additional content relating to the original content input authored by the user / author; and finally, (f) retrieving the identified additional content via smart digital content recommendation tools and presenting it to the author. In some embodiments, each additional content item (e.g., an image, a link to a related article or webpage, an audio or video file, graphics, a social media post, etc.) may be displayed and / or thumbnailed by a smart digital content recommendation tool in a GUI-based tool that allows the user to drag and drop or place the additional content within the user's original authored content, including positioning, formatting, and resizing the content.
[0033] Referring here to Figure 4, a block diagram is shown illustrating the various components of the system 400 for smart content classification and recommendation, including a client device 410, a content input processing and analysis service 420, a content recommendation engine 425, a content management system 435, and a content retrieval and embedding service 445. In addition, the system 400 includes one or more content repositories 440 for storing content files / resources and one or more vector spaces 430. As will be described in more detail below, a vector space may also refer to a multidimensional data structure configured to store one or more feature vectors. In some embodiments, the recommendation engine 42 5. The associated software components and services 420 and 445, the content management system 435, and the content repository 440 (which may store one or more data stores or other data structures) may be implemented and stored as a backend server system located away from the frontend client device 410. Thus, the interaction between the client device 410 and the content recommendation engine 425 may be an internet-based web browsing session or a client-server application session, during which the user may input original authored content via the client device 410 and receive content recommendations from the content recommendation engine 425. The receipt of such content recommendations may occur in the form of additional content retrieved from the content repository 440 and linked to or embedded in the content authoring user interface on the client device 410. Additionally or alternatively, the content recommendation engine 425 and / or the content repository 440 and related services may be implemented as dedicated software components running on the client device 410.
[0034] The various computing infrastructure elements shown in this example (e.g., the content recommendation engine 425, software components / services 420, 435, and 445, and the content repository 440) may correspond to a high level of computer architecture created and maintained by an enterprise or organization that provides internet-based services and / or content to various client devices 410. The content described herein (which may also be referred to as content resources and / or content files, content links, etc.) may be stored in one or more content repositories, retrieved and categorized by the content recommendation engine 425, and provided to a content author in a client device 410. In various embodiments, a wide variety of media types or file types of content may be input as original content by a content author in a client device 410, and similarly, a wide variety of media types or file types of content may be stored in the content repository 440 for recommendation / embedding of the front-end user interface in a client device 410. These diverse media types, which are authored by or recommended to content authors, may include text (e.g., writing articles, blogs), images (selected by or for the author), audio or video content resources, graphics, and social media content (e.g., posts, messages, or tweets).
[0035] In some embodiments, the system 400 shown in Figure 4 may be implemented as a cloud-based multi-tier system, where upper-tier user devices 410 can request and receive access to network-based resources and services via the content processing / analysis component 420. In this case, the application server may be deployed and run on an underlying set of resources (e.g., cloud-based, SaaS, IaaS, PaaS, etc.) including hardware and / or software resources. In addition, while a cloud-based system may be used in some embodiments, in other examples, the system 400 may use an on-premises data center, server farm, distributed computing system, and various other non-cloud computing architectures. Some or all of the functions described herein regarding the content processing / analysis component 420, the content recommendation engine 425, the content management system 435, the content retrieval and embedding component 445, and the generation and storage of the vector space 430 are based on the Simple Object Access Protocol. Representational State Transfer (REST) services and / or web services including SOAP or APIs It is publicly available via web services and / or via the Hypertext Transfer Protocol (HTTP) or HTTP Secure Protocol. This can be performed by the web content that is opened. Thus, although not shown in Figure 4 in order not to obscure the components that will be shown with additional details, the computing environment 400 may include additional client devices 410, one or more computer networks, one or more firewalls 435, a proxy server, and / or other intermediate network devices, thereby facilitating interaction between the client devices 410, the content recommendation engine 425, and the backend content repository 440. Another embodiment of a similar system 500 is shown in more detail in Figure 5.
[0036] Referring briefly to Figure 5, another exemplary diagram of the computing environment 500 shows a data flow / data transformation diagram for performing content classification and recommendation. Thus, the computing environment 500 shown in this example may correspond to one feasible implementation example of the computing environment 400 described above in Figure 4. In Figure 5, some of the blocks of the diagram shown represent specific data states or data transformations, rather than the structural hardware and / or software components described above in Figure 4. Thus, block 505 may represent input content data received via the user interface. Block 510 represents a set of keywords determined by the system 400 based on the input content 505. As described above, the keywords 510 may be determined by the input processing / analysis component 420 using one or more keyword extraction and / or topic modeling processes, and a text feature vector 515 may be generated based on the determined keywords 510.
[0037] Continuing with the example shown in Figure 5, several additional feature vectors 520 may be retrieved from the content repository 440. In this example, the additional feature vectors 520 may be selected from the content repository 440 by running one or more neural network-trained image models and providing the determined keywords 510 to the trained models. The resulting feature vectors 520 may be further narrowed down to exclude those with feature vector probabilities less than z% based on the output of the trained models, resulting in a subset of retrieved feature vectors 525. A feature space comparison 530 may then be performed between the test feature vector 515 and the subset 525 of retrieved feature vectors. In some embodiments, as shown in this example, the retrieved feature vector 525 closest to the test feature 515 may be identified using nearest Euclidean distance calculation. Based on the feature space comparison 530, one or more recommendations 530 may be determined. Each recommendation 530 is based on an associated feature vector 525 that has a similar threshold to the test feature vector 515, and each recommendation 530 corresponds to an image in the content repository 440.
[0038] The components shown in System 400 for providing AI-based and feature vector analysis-based content recommendation and services to client devices 410 may be implemented in hardware, software, or a combination of hardware and software. For example, web services may be generated, deployed, and run within a data center 440 using underlying system hardware or software components such as data storage devices, network resources, computing resources (e.g., servers), and various software components. In some embodiments, web services may correspond to various software components running on the same underlying computer server, network, data store, and / or within the same virtual machine. Web-based services are provided within the content recommendation engine 425. Some of the content, computing infrastructure instances, and / or web services may utilize dedicated hardware and / or software resources, while others may share underlying resources (e.g., a shared cloud). In either case, higher-level specific services (e.g., user applications), and even users on client devices, do not always need to be aware of the underlying resources used to support the services.
[0039] In such implementation examples, various application servers, database servers and / or cloud storage systems, as well as other infrastructure components such as web caches and network components (not shown in this example), provide and monitor the classification and vectorization of content resources, and further manage the underlying storage / server / network resources, using various hardware and / or software components (e.g., application programming interfaces (APIs), cloud storage systems, etc.). The underlying resources of the content repository 440 may include, for example, a set of non-volatile computer memory devices implemented as databases, file-based storage, etc., a set of network hardware and software components (e.g., routers, firewalls, gateways, load balancers, etc.), a set of host servers, and various software resources such as storage, installation, build, templates, and configuration files for software images corresponding to different versions of various platforms, servers, middleware, and application software, and may be stored within the content repository and / or cloud storage system. The data center housing the application servers of the recommended engine 425, the vector space 430, and the related services / components may also include hardware and software infrastructure to support various internet-based services, such as Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS), along with additional resources such as hypervisors, host operating systems, resource managers, and other cloud-based applications. In addition, the underlying hardware of the data center may be configured to support several internal shared services, such as security and identity services, integration services, repository services, enterprise management services, virus scanning services, backup and recovery services, notification services, and file transfer services.
[0040] As described above, web-based content recommendations can be provided to client devices 410 from a content recommendation engine 542 (which may be implemented via one or more content recommendation application servers) according to various embodiments described herein, using many different types of computer architectures (cloud-based, web-based, hosting, multi-tier computing environments, distributed computing environments, etc.). However, in certain implementation examples, a cloud computing platform may be used to provide certain advantageous features for generating and managing web-based content. For example, a cloud computing platform can provide adaptability and scalability for rapidly providing, configuring and deploying many different types of computing infrastructure instances, in contrast to non-cloud-based implementations with fixed architectures and limited hardware resources. Furthermore, public cloud platforms, dedicated cloud platforms, and hybrid cloud platforms of public and dedicated may be used in various embodiments to leverage the features and advantages of individual architectures.
[0041] In addition, as shown in this example, system 400 also includes a content management system 435. In some embodiments, the content management system 435 may include a distributed storage processing system, one or more machine learning-based classification algorithms (and / or non-machine learning-based algorithms), and / or a storage architecture. As will be described in more detail below, in some embodiments, the content management system 435 may access content resources (e.g., web-based articles, images, audio files, video files, graphics, social media content, etc.) via one or more content repositories 440 (e.g., network-based document stores, web-based content providers, etc.). For example, within system 400, dedicated JavaScript or other software components may be installed to run on one or more application servers, database servers, and / or cloud systems that store content objects or network-based content. These software components may be configured to retrieve content resources (e.g., articles, images, web pages, documents, etc.) and send them to the content management system 435 for analysis and classification. For example, whenever a user within the operating organization of system 400 imports or creates new content such as an image or article, the software component may return the content to the content management system 435 for various processing and analysis (e.g., image processing, keyword extraction, topic analysis, etc.) as described below. In addition, although this example shows the content management system 435 as being implemented separately from the content recommendation engine 425 and the content repository 440, in other examples the content management system 435 may be implemented locally on one of the storage devices that houses the content recommendation engine 425 and / or the content repository 440, and therefore does not need to receive content sent separately from those devices, but may still analyze and classify the content resources stored or provided by each system.
[0042] One or more vector spaces 430 may also be generated and used to store feature vectors corresponding to different content items in the content repository 440, and to compare feature vectors for the original authored content (e.g., received from a client device 410) with feature vectors for additional content items in the content repository 440. In some embodiments, multiple multidimensional feature spaces 430 may be implemented in the system 400, such as a first feature space 430a for text input / article topics and a second feature space 430b for images. In other embodiments, additional separate multidimensional feature spaces 430 may be generated for different types of content media (e.g., a feature space for audio data / files, a feature space for video data / files, a feature space for graphics, a feature space for social media content, etc.). As described below, comparison algorithms can be used to determine the distance between vectors in the feature spaces. Thus, in the feature space for image feature vectors, an algorithm can be used to identify the image closest to the received input image, and in the feature space for text feature vectors, an algorithm can be used to identify the text (e.g., an article) closest to the received input text block, and so on. Additionally or alternatively, the comparison algorithm may use keywords / tags in a vector space to determine similarity between different media types.
[0043] In various implementations, System 400 may be implemented using one or more computing systems and / or networks. These computing systems may include one or more computers and / or servers. These one or more computers and / or servers may be general-purpose computers, dedicated server computers (e.g., desktop servers, UNIX® servers, midrange servers, mainframe computers, rack-mount servers, etc.), server farms, server clusters, etc. The distributed servers may be any other suitable configuration and / or combination of computing hardware. The content recommendation engine 425 may run an operating system and / or a variety of additional server applications and / or middle-tier applications, including, for example, a Hypertext Transport Protocol (HTTP) server, a File Transport Service (FTP) server, a Common Gateway Interface (CGI) server, a Java® server, a database server, and other computing systems. The content repository 440 may include, for example, a commercially available database server from Oracle, Microsoft, etc. Each component within System 400 may be implemented using hardware, firmware, software, or a combination of hardware, firmware, and software.
[0044] In various implementations, each component within System 400 may include at least one memory, one or more processing units (e.g., processors), and / or storage. Processing units may be implemented as appropriate in hardware (e.g., integrated circuits), computer-executable instructions, firmware, or a combination of hardware and instructions. In some examples, various components of System 400 may include several subsystems and / or modules. Subsystems and / or modules within the Content Recommendation Engine 425 may be implemented in hardware, software running on hardware (e.g., program code or instructions executable by the processor), or a combination thereof. In some examples, the software may be stored in memory (e.g., non-temporary computer-readable media), memory devices, or any other physical memory, and may be executed by one or more processing units (e.g., one or more processors, one or more processor cores, one or more Graphics Process Units (GPUs), etc.). Examples of executable instructions or firmware implementations may include computer-executable or machine-executable instructions written in any suitable programming language capable of performing the various operations, functions, methods, and / or processes described herein. Memory may store program instructions that can be loaded and executed on the processing unit, and data generated during the execution of these programs. Memory may be volatile (e.g., random access memory (RAM)) and / or non-volatile (e.g., read-only memory (ROM), flash memory). Memory may be implemented using any type of persistent storage device, such as computer-readable storage media. In some examples, computer-readable storage media may be configured to protect the computer from electronic communications containing malicious code.
[0045] Referring now to Figure 6, a flowchart shows the process for generating feature vectors based on content resources in the content repository 440 and for storing these feature vectors in the feature space 430. As will be described below, the steps in this process may be performed by one or more components within the computing environment 400, such as the content management system 435, and the various subsystems and subcomponents implemented therein.
[0046] In step 602, content resources may be retrieved from the content repository 440 or other data store. As described above, individual content resources (which may also be referred to as content or content items) may correspond to data objects of any variety of content types, such as text items (e.g., text files, articles, emails, blog posts, etc.), images, audio files, video files, 2D or 3D graphic objects, and social media data items. In some embodiments, content items are owned and operated by a specific trusted organization. Content may be retrieved from a specific content repository 440, such as a proprietary data store. The content repository 440 may be an external data source, such as an internet web server or other remote data store, but a system 400 that retrieves and vectorizes content from a local and / or independently controlled content repository 440 may implement several technical advantages in the operation of system 400. Some technical advantages include ensuring that the content from the repository 440 is stored and accessible as needed, and that users / authors are authorized to use and reproduce the content from the repository 440. In some cases, the retrieval in step 602 (and subsequent steps 604-608) may be triggered in response to a new content item stored in the content repository 440, and / or in response to a modification of an item in the content repository 440.
[0047] In step 604, the content items extracted in step 602 may be parsed / analyzed / etc. to extract a set of item features or characteristics. The type of parsing, processing, feature extraction and / or analysis performed in step 604 may depend on the type of content item. For image content items, an artificial intelligence-based image classification tool may be used to identify specific image features and / or generate image tags. As shown in the image example in Figure 7, the image analysis may identify multiple image features (e.g., smile, waitress, counter, vending machine, coffee cup, cake, hand, food, person, cafe, etc.), and the image may be tagged with each of these identified features. For text-based content items such as blog posts, letters, emails, and articles, the analysis performed in step 604 may include keyword extraction and processing tools (e.g., stemming, synonym lookup, etc.), as shown in Figure 8. One or both types of analysis (i.e., feature extraction from images as shown in Figure 9 and keyword / topic extraction from text content as shown in Figure 8) may be performed via REST-based services or other web services using analysis, machine learning algorithms and / or artificial intelligence (AI), such as an AI-based cognitive image analysis service or a similar AI / REST cognitive text service used for the text content in Figure 8. Similar techniques may be used in step 604 for other types of content items such as video files, audio files, graphics, or social media posts. In this case, a dedicated web service of System 400 is used to extract and analyze specific features (e.g., words, objects in images / videos, facial expressions, etc.) depending on the media type of the content item.
[0048] In step 606, after extracting / determining specific content features (e.g., visual objects, keywords, topics, etc.) from a content item, a feature vector may be generated based on the extracted / determined features. Using various transformation techniques, each set of features associated with a content item may be transformed into a vector that can be input into a common vector space 430. The transformation algorithm may output a predetermined vector format (e.g., a 1 × 4096-dimensional vector). Then, in step 608, the feature vector may be stored in one or more of the vector spaces 430 (e.g., a topic vector space 430a for text content, an image vector space 430b for image content, and / or a combined vector space for multiple content types). Each vector space and each feature vector stored in the corresponding space may be generated and maintained by the content management system 435, the content recommendation engine 425, and / or other components of the system 400.
[0049] In some embodiments, a subset of extracted / determined content features may also be stored as tags associated with content items. An exemplary process for generating and storing image tags based on images, and conversely, for detecting images based on image tags. Exemplary processes for searching are shown in Figures 9 to 11. While these examples relate to image content items, similar tagging processes and / or keyword or topic extraction may be performed for text content items, audio / video content items, etc. As shown in Figure 9, in step 901, an image may be created and / or uploaded to, for example, a content repository 440. In step 902, the image may be sent from the content repository 440 to an artificial intelligence (AI) based REST service. This AI-based REST service is configured to analyze the image and extract topics, themes, specific visual features, etc. Based on the identified image features, the AI REST service may determine one or more specific image tags, and in step 903, the image tags may be sent back to the content repository and, in step 904, may be stored in the image or associated with the image. Figure 10 shows the same process as described in Figure 9 for generating and storing image tags based on an image. In addition, Figure 10 shows several exemplary features 1001 that may be implemented within an AI REST service in a particular embodiment, including an image tag determination / retrieval component, an Apache MxNet component, and a cognitive image service. After determining one or more tags for a content item, these tags may be stored again in the content repository 440 or a separate storage location. For example, referring to the exemplary image in Figure 7, more than a dozen potential image features may be extracted from this single image, and all of them may be incorporated into a feature vector. However, the AI REST service and / or content repository 440 may determine that tagging the image using only a few of the most prevalent themes of the image (e.g., coffee, retail) is optimal for content matching.
[0050] Referring briefly to Figure 11, another exemplary process for generating and storing image tags based on multiple images is shown in relation to the descriptions in Figures 9 and 10. Figure 11 shows the reverse process, in which multiple image tags are used to retrieve matching images from the content repository 440. In step 1101, the content authoring user interface 415 or other front-end interface may determine one or more content tags based on input received through the interface. In this example, a single content tag ("waitress") is determined from the received user input, and in step 1102, the content tag is sent to a search API associated with the content repository 440. The search API may be implemented within one or more separate layers of a computing system, including a content input processing / analysis component 420, a content recommendation engine 425, and / or a content management system 435. In step 1103, data identifying the matching image determined by the search API may be sent back to the content authoring user interface 415 so as to be integrated within the interface or presented to the user.
[0051] Therefore, once steps 602-608 for multiple content resources in the content repository 440 are completed, each vector corresponding to a content item in the repository 440 may be populated in one or more vector spaces 430. In addition, in some embodiments, a separate set of metadata tags may be generated for some or all of the content items and stored in the vector space 430 as separate objects from the vectors. Such tags may be stored in any data storage or component shown in Figure 4, or in a separate data store, and each tag may be associated with a content item in the repository 440, a vector in the vector space 430, or both.
[0052] Referring here to Figure 12, another flowchart is shown illustrating a second process for receiving original authored content from a user via a client device 410, extracting features and / or tags from the content in real time (or near real time) during the user's authoring session, vectorizing the authored content (also in real time or near real time), and comparing the vectors of the original authored content with one or more existing vector spaces 430 in order to identify and retrieve relevant / associated content from one or more available content repositories 440. The steps in this process may also be performed by one or more components within the computing environment 400, for example, a content recommendation engine 425 working with the client device 410, an input processing / analysis component 420 and an extraction / embedding component 445, as well as various subsystems and subcomponents implemented therein.
[0053] In step 1202, the original authored content may be received from the user via the client device 410. As described above, the original authored content may include text typed by the user, new images created or imported by the user, new audio or video input recorded or imported by the user, new graphics created by the user, etc. For this reason, step 1202 may be similar to step 602 described above. However, while the content in step 602 may be content retrieved from the repository 440 and pre-authored / stored, in step 1202, the content may be newly authored content received via a user interface such as a web-based text input control, image importer control, image creation control, or audio / video creation control.
[0054] In step 1204, the content received in step 1202 (e.g., original authored content) may be processed by, for example, the input processing / analysis component 420. Step 1204 may be the same as or identical to step 604 described above with respect to parsing steps, processing steps, keyword / data feature extraction steps, etc. For example, in the case of text input received in step 1202 (e.g., blog posts, characters, emails, articles, etc.), the processing in step 1204 may include parsing the text, identifying keywords, stemming, synonym analysis / extraction, etc. In another example, if images are received in step 1202 (may be uploaded by the user, imported from another system, and / or manually created or modified by the user via the content author user interface 415), step 1204 may include the step of identifying specific image features and / or generating image tags using an AI-based image classification tool as described above. These analyses in step 1204 may be performed via REST-based services or other web services using analysis, machine learning algorithms, and / or AI, such as AI-based cognitive image analysis services and / or AI / REST cognitive text services. Similar technologies / services may be used in step 1204 for other types of content items such as video files, audio files, graphics, or social media posts. In this case, a dedicated web service of System 400 is used to extract and analyze specific features (e.g., words, objects in images / videos, facial expressions, etc.) depending on the media type of the content item. Step 1204 may also include any of the tagging processes described herein for tagging text blocks, images, audio / video data, and / or other content with tags corresponding to any identified content topics, categories, or features.
[0055] In step 1206, based on the content received in step 1202, One or more vectors compatible with one or more of the vector spaces 430 may be generated. Step 1206 may be the same as or identical to step 606 described above. As described above, the vectors may be generated based on specific features in the content (and / or tags) identified in step 1204. The vector generation process in step 1206 may use one or more data transformation techniques that can transform a set of features associated with the originally authored content item into a vector compatible with one of the common vector spaces 430. For example, the technique shown in Figure 13 may transform the image input received in step 1202 into a feature vector in a predetermined vector format (e.g., a 1 × 4096-dimensional vector) in step 1206. As shown in Figure 13, the image may be provided as input to a model, which is configured to extract and learn features in the image (e.g., using convolution, pooling, and other functions) and output a feature vector representing the input image. For example, as described in the paper "Visualizing and Understanding Convolutional Networks" (2014) by Matthew D. Zeiler and Rob Fergus of New York University, which is incorporated herein by reference, within a convolutional neural network, the initial layers of the neural network may detect simple features from an image, such as linear edges, while later layers may detect more complex shapes and patterns. For example, the first and / or second layers in a convolutional neural network may detect simple edges or patterns, while later layers may detect actual, complex objects present in the image (such as cups, flowers, or dogs). As an example, when receiving and processing an image of a face using a convolutional neural network, the first layer may detect edges in various directions, the second layer may detect different parts of a given face (e.g., eyes, nose, etc.), and the third layer may obtain a feature map of the entire face.
[0056] In step 1208, the feature vectors generated in step 1206 can be compared to compatible feature vector spaces 430 (or spaces 430a-430n) populated during the processes 602-608 described above. For example, Figure 14 shows an exemplary vector space populated with multiple feature vectors corresponding to multiple images. In this example, each dot may represent a vectorized image, and the circles (and corresponding dot colors) in Figure 14 may represent one of three exemplary tags associated with the image. In this case, the image tags are "Coffee," "Mountain," and "Bird," and it should be understood that these tags are not mutually exclusive (i.e., an image may be tagged with one tag, two tags, or all three tags). In addition, it should be understood that the layout of these tags and the multidimensional vector space in Figure 14 is illustrative only. In various embodiments, the number or type of tags that may be used, or the number of dimensions of the vector space 430, are not limited.
[0057] To perform a vector space comparison in step 1208, the content recommendation engine 425 may calculate the Euclidean distance between each of the feature vectors generated in step 1206 and each of the other feature vectors stored in the vector space / space 430. Based on the calculated distances, the engine 425 may rank the feature vectors in ascending order of feature space distance, so that the smaller the distance between two feature vectors, the higher the rank. Such a technique may enable the content recommendation engine 425 to determine the highest-ranked set of feature vectors in the vector space 430. The highest-ranked feature vectors are those whose features / characteristics etc. are most similar to the feature vectors generated in step 1206 based on the input received in step 1202. In some cases, in step 1208, a predetermined number (N) of highest-ranked feature vectors (e.g., 5 most similar articles, 10 most similar images, etc.) may be selected, and in other cases, all feature vectors that satisfy a certain closeness threshold (e.g., distance between vectors < threshold (T)) may be selected. ) may be selected.
[0058] In some embodiments, the vector comparison in step 1208 may be a “deep feature space” comparison as shown in Figure 15. In these embodiments, the feature vectors generated in step 1206 may be compared without considering any tags or other metadata. In other words, the deep feature comparison may compare the feature vectors generated in step 1206 with all other features stored in the vector space 430. While the deep feature comparison can ensure that the closest vector is found in the vector space 430, this type of comparison may require additional processing resources and / or additional time to return the vector result. This is especially true for large vector spaces that may contain thousands or millions of feature vectors, each representing a separate content object / resource stored in the repository 440. For example, calculating the Euclidean distance between two image feature vectors of size 1 × 4096 requires the system 400 to execute approximately 10,000 addition and multiplication instructions. Therefore, if there are 10,000 images in the repository, 10,000,000 operations must be performed.
[0059] Therefore, in other embodiments, the vector comparison in step 1208 may be a “filtered feature space” comparison as shown in Figures 16 and 17. In a filtered feature space comparison, the vector space can first be filtered based on tags (and / or other properties such as resource media type, creation date, etc.) to identify a subset of feature vectors in the vector space 430 that have tags (and / or other properties) that match the tags of the feature vectors generated in step 1206. The feature vectors generated in step 1206 can then be compared only with the feature vectors in the subset that have matching tags / properties. Thus, a filtered feature space comparison can be performed more quickly and efficiently than a deep space comparison, although there may be no nearby feature vectors that are filtered out but not compared.
[0060] As described above, step 1208 may include comparing the feature vector generated in step 1206 with a single vector space or multiple vector spaces. In some embodiments, the feature vector generated in step 1206 may be compared with a vector space of the corresponding type. For example, if a text input is received in step 1202, the resulting feature vector may be compared with a vector in topic vector space 430a. Furthermore, if an image is received as input in step 1202, the resulting feature vector may be compared with a vector in image vector space 430b, and so on. In some embodiments, it may be possible to compare a feature vector corresponding to one type of input with a vector space containing different types of vectors (for example, to identify the image resource most closely related to the text-based input, or vice versa). For example, Figure 18 illustrates the process of receiving a text input in step 1402 and, in step 1408, retrieving both similar images (e.g., closest to image vector space 430b) and similar articles (e.g., closest to topic vector space 430b).
[0061] In embodiments that involve retrieving and / or comparing tags associated with content resources, problems can arise when tags from one resource relate to each other but do not exactly match the corresponding tags / keywords / characteristics of another resource. An example of this potential problem is shown in Figure 19. In this case, keywords extracted from the original authored text content resource are compared to a set of image tags stored for a set of image content resources. In this example, the extracted keywords are "Everest", "Base Camp", and "Summit". Neither "Mountain" nor "Himalaya" is a strict match to the image tags ("Mountaineer", "Cappuccino", or "Macaw"). In some embodiments, word / phrase parsing and processing techniques such as word stemming, word definition, and / or synonym lookup and analysis may be used to detect matches between related but unmatched terms. However, these techniques may also fail for related keywords / tags. Therefore, in some embodiments, the content processing / analysis component 420 and / or content recommendation engine 425 may perform word vector comparison to address this problem. As shown in the example in Figure 20, keywords extracted from the text document in Figure 19 may be analyzed in a three-dimensional word vector space, and the distance between those keywords and each of the image tags may be calculated. As shown in Figure 21, the keyword-versus-tag vector space analysis performed in Figure 20 shows that the image tag "Mountaineer" is an extracted keyword in the word vector space. Because it is sufficiently close to the tag, it can be determined that it should be considered an image tag match for filtered feature space comparison.
[0062] Another potential problem that may arise in embodiments for retrieving and / or comparing tags associated with content resources is caused by homophones of keywords and / or resource tags. Homophones are words or phrases (or homonyms) that have the same spelling but different meanings and are unrelated. An example of a homophone of an image tag is shown in Figure 22. In this case, the first image is "Crane," meaning a bird with long legs and a long neck. The first image is tagged with the word "crane," and the second image is tagged with the same word "crane," which refers to a machine with a protruding arm used to move heavy objects. It has been deleted. In this case, the content processing / analysis component 420 and / or the content recommendation engine 425 may perform a word semantic deambiguation process on the two image tags to determine which meaning of the word "crane" is being referred to. In this example, The word semantic ambiguity removal process for two different "Crane" tags is shown in Figure 22. To that end, Wordnet database entries (or other definition data) associated with each tag. You can search for "Ta)" first.
[0063] An exemplary word semantic ambiguity process is shown in Figures 23 and 24. This process involves de-ambiguating other keywords and / or the word "Crane" within the authored document. The content processing / analysis component 420 and / or content recommendation engine 425 use the specific context (e.g., explanation, part of speech, tense, etc.) to determine the most likely meaning of the word "crane" in the authored text document. This allows us to determine which of the "Crane" image tags has been authored. It is possible to determine whether it is related to a text document. For example, referring to Figure 23, several related keywords 2302 are extracted from the illustrated input text 2301. The first extracted keyword ("Crane") is related to an image tag in the content repository 440. In addition to being comparable, in this example, two matching tags 2303 correspond to two "Crane" tagged images 2304a and 2304b in the content repository 440. They are separated.
[0064] As shown in Figure 24, to address this potential problem of ambiguity in word meaning, the deambiguation process may continue to compare one or more additional keywords 2302 extracted from the input content 2301 with other content tags in the two matching images 2304a and 2304b. In this example, additional extracted keywords such as “mechanical,” “machine,” “lifting,” and “construction” may be compared with content tags and / or extracted features associated with each of the images 2304a and 2304b. As shown in Figure 24, these additional comparisons may clarify the initial keyword match of “Crane,” thus clarifying the content. The recommended system 425 returns image 2304a of the bird "crane". Instead, image 2304b of a construction crane is returned.
[0065] In other examples, a similar deambiguation process may be performed using image similarity. For example, the content processing / analysis component 420 and / or content recommendation engine 425 may determine which “crane” is the appropriate related image by using authored images. Common image features can be identified between an image associated with the content (e.g., a drawn or authored image) and two different "Crane" images. These de-ambiguation techniques... The process may also be combined in various ways, for example, by comparing keywords extracted from an authored text document with visual features extracted from an image. Therefore, in an authored text document that quotes "crane", "boom" And the related word "pulley" can visually correspond to the image of the "crane" below, provided that the boom and pulley are visually identifiable within that image. The authored text document quotes "crane," and also "beak" and "feather." When the related word "root" is included, the keyword "crane" refers to the beak and feathers. If the roots are visually identifiable within the image, then the image of "crane" in the upper section is visually identifiable. They can be perceived as matching.
[0066] Referring here to Figures 25-28, an end-to-end example of performing the process in Figure 12 is shown, specifically an example of an embodiment for retrieving a set of relevant images for an article authored by a user via the user interface 415. First, in 2501 (Figure 25), the user types text about the article into the Demo Editor (Alditor) user interface. In 2502, several keywords are extracted from the article text, and in 2503, the extracted keywords are compared by the AI Rest service with image tags stored for the image library in the image content repository 440. Figures 26 and 27 show the same exemplary process as in Figure 25, along with additional details regarding the operation of the AI Rest service. As shown in Figure 26, the AI Rest service uses the techniques described above to find one or more tags (e.g., "mountaineer") related to the authored article. It is identified as such. Then, as shown in Figure 27, in some embodiments, this step can be performed using a combination of different software services, for example, a first cognitive text REST service used to determine a set of keywords from text input and a second internal REST service used to map the keywords to image tags. Each of these services may be implemented within the content recommendation engine 425 and / or via an external service provider. Then, in step 2505 (shown in Figure 28), the content recommendation engine 425 may send the determined image tags to a search API associated with the image content repository 440. In some cases, the search API may be implemented within a cloud-based content hub, for example, Oracle Content Management (OCM). In step 2506, the search API may retrieve a set of relevant images based on tag matching, and in step 2507, the retrieved images (or thumbnail images) may be sent back and embedded in the user interface in 415 (in screen area 2810).
[0067] The examples shown in Figures 25-28 illustrate a specific embodiment of retrieving a set of associated images related to an article authored by a user, but it should be understood that the steps in Figure 12 can similarly be performed to retrieve other types of content. For example, similar steps may be performed to retrieve an article (or other text document) related to text entered by a user via the user interface 415. In other embodiments, other media types (e.g., audio files, video clips, graphics, social media) may be retrieved. Related content resources (such as remedia posts) may be retrieved. In addition, if the user imports / creates other types of input other than text into the user interface (e.g., drawn or uploaded images, spoken audio input, video input, etc.), similar steps may be performed to retrieve a wide variety of related content resources (e.g., related articles, images, videos, audio, social media, etc.) depending on the configuration of the content recommendation engine 425 and / or the user's preferences.
[0068] For example, another exemplary embodiment is shown here with reference to Figures 29–35. Here, the process steps of Figure 12 are performed to retrieve a set of relevant articles (or other text content resources) based on the original authored text input received via the user interface 415 (e.g., a user's blog post, email, article, etc.). As shown in Figure 29, the user authorizes a new article via the user interface 415, and the set of article topics is identified by an AI-based REST service invoked by the content recommendation engine 425. As shown in Figure 30, the identified article topics can be compared with previously identified topics with respect to a set of articles in the article content repository 440. In these examples, Figure 29 is shown according to one embodiment, and Figure 30 shows another embodiment. Figure 29 is only a subset of Figure 30, and it can be omitted without issue. Thus, article topics can be determined and stored as metadata or other associated data objects using techniques similar to the process of determining and associating image features / tags with images (described above in Figure 6). Similarly, the article content repository 440 may have metadata or other associated storage for each article stored in the repository 440, including article topic, date, keywords, author, publication, etc. In the example shown in Figure 30, articles related to deaths on Mount Everest are identified as potentially related to a newly created user article based on article topic matching. Figures 31-35 show an end-to-end process using the system 400 to find articles related to a user-input article, similar to the steps shown in Figures 25-28 for finding related images. In step 3101 (Figure 31), the user creates a new article via the user interface 415.In step 3102, the article text is sent to one or more software services by a content recommendation engine 425 (e.g., an AI-based REST service), and in step 3103 (Figure 32), the software services analyze the article text using cognitive text service functionality to determine one or more topics of the article. In step 3104, the determined article topics are sent back to the content recommendation engine 425, and in step 3105, the recommendation engine 425 sends both the article text and the identified topics to a separate API (e.g., in a cloud-based content hub), and in step 3106, the article is stored in a repository 440 for future reference and may be indexed based on the identified topics. Also in step 3106 (Figure 33), the article's existing repository 440 can be searched via a search API to identify potentially related topics based on a topic matching process (Figure 34). Finally, in step 3107 (Figure 35), articles identified as potentially related to the newly created article may be sent back (either entirely or simply via links) to be embedded within the user interface 415 (for example, in the user interface area 3510).
[0069] As shown in the examples above in Figures 29 to 33, certain embodiments described herein may include the identification of topics in newly created text documents and / or topics in text documents stored in the content repository 440, as well as the comparison and identification of topic similarity and match. In the various embodiments described herein, various techniques, including explicit semantic analysis, may be used for text topic evaluation and topic "similarity" techniques. As shown in Figures 29 and 30, in some cases, this Una technology uses large-scale data sources (for example, Wikipedia) This can provide a fine-grained semantic representation of an unlimited amount of natural language text, representing the meaning of natural concepts derived from data sources in a high-dimensional space. For example, text classification techniques may be used to explicitly represent the meaning of any text from the perspective of Wikipedia-based concepts. The semantic representation may also be a feature vector of a text snippet transformed by topic modeling. By using Wikipedia (or another large data source), a larger vocabulary (e.g., a "bag of words") may be included in the system to cover a large range of multi-word phrases. The Wikipedia-based concepts may also be the titles of Wikipedia pages used as classes / categories while classifying a given text snippet. In the case of a text snippet, the most approximate Wikipedia page title that can be used as a class / category of the text (e.g., "Mount Everest", "Stephen Hawking", "Car Accident", etc.) may be returned. The effectiveness of such techniques can be automatically evaluated by calculating the degree of semantic relationships between fragments of natural language text.
[0070] One advantage of these text classification / relationship evaluation techniques is the use of large, publicly available knowledge sources (e.g., Wikipedia or other encyclopedias), which provides access to a highly organized mass of human knowledge that is regularly modified and developed, and pre-encoded into publicly available sources. Semantic interpreters can also be constructed using machine learning techniques based on Wikipedia and / or other sources. These semantic interpreters map fragments of natural language text to a weighted set of Wikipedia concepts ordered by their relevance to the input. Thus, the input text can also be represented as weighted vectors of concepts, called interpretation vectors. The meaning of text fragments is therefore interpreted in terms of their similarity to a host of Wikipedia concepts. The semantic relationships of the text can then be computed, for example, by comparing those vectors in the space defined by the concepts, using a cosine metric. Such semantic analysis can be clear in the sense that clear concepts may be based on human cognition. Since user input can be received as plain text via the user interface 415, conventional text classification algorithms may be used to rank the concepts represented by these articles according to their relevance to a given text fragment. Therefore, online encyclopedias (e.g., Wikipedia) may be used directly without requiring an understanding of deep language or knowledge of pre-classified common semantics. In some embodiments, each Wikipedia concept may be represented as an attribute vector of words occurring in the corresponding article. Entries in these vectors may include, for example, a term frequency-inverse document frequency (TFIDF) skim. Weights may be assigned using 'm'. These weights can quantify the strength of the association between a word and a concept. To facilitate semantic interpretation, an inverted index may be used that maps each word to a list of concepts in which it appears. The inverted index may also be used to discard unimportant associations between words and concepts by removing concepts whose weight for a given word is below a certain threshold. The semantic interpreter may be implemented as a centroid-based classifier that can rank Wikipedia concepts by relevance based on the text fragment it receives. For example, a semantic interpreter in a content recommendation engine may receive an input text fragment T and represent the fragment as a vector (e.g., using a TFIDF scheme). The semantic interpreter may iterate through the text words, extract the corresponding entries from the inverted index, and merge them into a weighted vector. The entries in the weighted vector may reflect the relevance of the corresponding concepts to the text T. To compute the semantic relationships of pairs of text fragments, their vectors may be compared, for example, using a cosine metric.
[0071] In other examples, similar methods for generating features for categorizing texts may involve supervised learning tasks. In this case, words appearing in training documents can be the features being used. Thus, in some examples, Wikipedia concepts can be used to augment a "bag of words." On the other hand, computing the semantic relationships of text pairs is essentially a "one-off" task, and therefore the "bag of words" representation may be replaced with a concept-based representation. These techniques and other The relevant technologies are incorporated herein by reference in their entirety for all purposes by Evgeniy Gabrilovich and Shaul Markovitch (Israel Institute of Technology). For more details, please refer to the paper "Computing Semantic Relatedness using Wikipedia-based Explicit Semantic Analysis" by the Department of Computer Science Technology, and other related literature listed herein. This is described in detail. Using the techniques described in this paper and other papers, a filtered subset of Wikipedia may have one concept for each article, which is the title of the article. When the content recommendation engine 425 receives a text document via the user interface 415, the text may first be summarized. In these cases, each unique word in the text may be weighted based on the frequency and inverse frequency of the word in the article after stop words have been removed and the words have been stemmed. Each word may be compared to determine which Wikipedia article (concept) it appears in, so that the content recommendation engine 425 may generate a concept vector for that word. By combining the concept vectors for all words in the text document, a weighted concept vector for the text document can be formed. The content recommendation engine 425 can then measure the similarity between each word concept vector and the text concept vector. Furthermore, all words that exceed a certain threshold may be selected as "keywords" for the document.
[0072] Referring here to Figure 36, an exemplary semantic text analyzer system 3600 is shown, illustrating the techniques used by the analyzer system 3600 to perform semantic text summarization in the particular embodiment described above. Such a system 3600 may be incorporated and / or separate by being accessed by the content recommendation engine 425 in various implementation examples.
[0073] In some implementations, explicit semantic analysis for calculating semantic relationships is used for a different purpose: to calculate a text summary of a given text document. More specifically, the text summary is derived based on word embeddings. In other words, the context of n-grams (e.g., words) is taken in for the purpose of determining semantic similarity, as opposed to typical similarity measures such as cosine or edit distance on a string for a "bag of words".
[0074] A given text document may be an article, web page, or other piece of text for which a text summary is desired. Similar to the classification approach described herein, the text is not limited to the language in which it is written and may include symbols, numbers, charts, tables, equations, formulas, etc., that are readable by others.
[0075] A text summary approach using explicit semantic analysis generally follows these steps: (1) Grammatical units (e.g., sentences or words) are extracted from a given text document using any known technique for identifying and extracting such units; (2) Each extracted grammatical unit and text document is represented as a weighted vector of knowledge base concepts; and (3) Semantic relationships between the entire text document and each grammatical unit are represented using the weighted vectors. (4) The grammatical units that are most semantically relevant to the entire text document are selected to be included in the text summary of the text document. In some cases, a weighted vector representation of knowledge base concepts may correspond to topic modeling. In this case, each sentence or word may first be converted into a feature vector, and then the difference / similarity of the features may be calculated in a high-dimensional vector space. Various methods can be used to convert words into vectors, for example, WORD2VEC or Latent Dirichlet Allocation. There are possible methods.
[0076] Figure 36 shows text summarization using explicit semantic analysis. First, a text summarizer is constructed based on knowledge base 3602. Knowledge base 3602 can be general or domain-specific. An example of a general knowledge base is a collection of encyclopedia articles, such as a collection of Wikipedia articles or a collection of other encyclopedias about text articles. However, knowledge base 3602 may also be domain-specific, such as a collection of text articles specific to a particular technical field, such as a collection of articles related to medicine, science, engineering, or finance.
[0077] Each entry in knowledge base 3602 is represented as an attribute vector of n-grams (e.g., words) that appear within the entry. Each entry in the attribute vector is assigned a weight. For example, the weights may be used in a scoring scheme that uses inverse document frequency against word frequency. The weights in the attribute vector for a given entry conceptually quantify the strength of the association between the n-grams (e.g., words) of that entry and the entry itself.
[0078] In some implementations, a scoring scheme using inverse document frequency against word frequency calculates the weights for a given n-gram t of a given article document d, as expressed by the following formula:
[0079]
number
[0080] Here, tf t,d This represents the frequency of n-gram t in document d. t represents the document frequency of n-gram t in knowledge base 3602. M represents the total number of documents in the training set, and L represents the document frequency of n-gram t in knowledge base 3602. d The length of document d is expressed as a number, L avg k represents the average length in the training corpus, and K and b are free parameters. In some implementations, k is approximately 1.5 and b is approximately 0.75.
[0081] The above is an example of an inverse document frequency scoring scheme against word frequency that can be used to weight attributes in an attribute vector. Other statistical measures that reflect how important attributes (e.g., n-grams) are to articles in knowledge base 3602 may be used. Other TF / IDF variations, such as BM25F which takes anchor text into account, may be used with specific types of knowledge bases, such as a knowledge base of web pages or other sets of hyperlinked documents.
[0082] The weighted inverted index builder computer 3604 constructs a weighted inverted index 3606 from attribute vectors representing articles in the knowledge base 3602. The weighted inverted index 3606 maps each distinct n-gram represented in the set of attribute vectors to the conceptual vector of the concept (article) in which the n-gram appears. Each concept in the conceptual vector can be weighted according to the strength of the association between the concept and the n-gram to which the conceptual vector is mapped by the weighted inverted index 3606. In this example, the indexer computer 3604 discards less important associations between n-grams and concepts by using the inverted index 3606 to remove concepts from the concept vector whose weight for a given n-gram is below a threshold.
[0083] To generate a text summary of a given text document 3610, grammatical units 3608 are extracted from the given text document 3610, and the semantic relationship between each grammatical unit and the given text document 3610 is calculated. Several grammatical units that have a high degree of semantic relationship with the given text document 3610 are selected to be included in the text summary.
[0084] The number of grammatical units selected for inclusion in a text summary can vary based on a wide range of factors. One approach is to select a predetermined number of grammatical units. For example, this predetermined number may be configured by the system user or learned by a machine learning process. Another approach is to select all grammatical units that have a degree of semantic relation to a given text document 3610 that exceeds a predetermined threshold. This threshold may be configured by the system user or learned by a machine learning process. Yet another feasible approach is to determine the grammatical unit with the highest degree of semantic relation to the given text document 3610, and then select all other grammatical units whose difference between their degree of semantic relation to the given text document 3610 and the highest degree is below a predetermined threshold. The grammatical unit with the highest degree and any other grammatical units below the predetermined threshold are selected for inclusion in the text summary. Again, this threshold may be configured by the system user or learned by a machine learning process.
[0085] In some implementations, the grammatical unit with the highest or relatively high degree of semantic relevance to a given text document 3610 is not necessarily selected for inclusion in the text summary. For example, a first grammatical unit with a lower degree of semantic relevance to the given text document 3610 than a second grammatical unit may be selected for inclusion in the text summary, and the second grammatical unit may not be selected for inclusion in the text summary if the first grammatical unit is not sufficiently different from the grammatical units already selected for inclusion in the text summary. The degree of difference of grammatical units to existing text summaries can be measured in a wide variety of ways, such as using a lexical approach, a probabilistic approach, or a hybrid of a lexical and probabilistic approach. Using a measure of difference to select grammatical units for inclusion in the text summary can prevent multiple similar grammatical units from being included in the same text summary.
[0086] In some implementations, other techniques may be used, and are not limited to any particular technique, to select several grammatical units for inclusion in a text summary as a function of their semantic relation to a given text document 3610 and their differences from one or more other units. For example, if the number of grammatical units having semantic relation to a given text document 3610 exceeds a threshold, the differences of compound grammatical units for combinations of grammatical unit numbers may be measured, and the number of grammatical units that are most different from each other may be selected for inclusion in the text summary. As a result, the grammatical units selected for inclusion in the text summary are, as a whole, highly semantically related to the text document, but different from each other. This is a more useful text summary than one that includes grammatical units that are highly semantically related but similar, because similar grammatical units are more likely to be redundant to each other in terms of the information they convey than different grammatical units.
[0087] Another possibility is to calculate a composite similarity / difference scale for grammatical units and then select the grammatical units to include in the text summary based on those composite scores. For example, the composite scale could be a weighted average of a semantic relation scale and a difference scale. Possible composite scales calculated as a weighted average are as follows:
[0088] (a*similarity)+(b*difference) Here, the parameter "similarity" represents the semantic relationship of grammatical units to the entire input text 3610. For example, the parameter "similarity" could be a similarity estimate 3620 calculated for the grammatical units. The parameter "difference" represents a difference scale of the differences between grammatical units for a set of one or more grammatical units. For example, a set of one or more grammatical units could be a set of one or more grammatical units already selected to be included in the text summary. Parameter a represents the weighted average of the weights applied to the similarity scale. Parameter b represents the weighted average of the weights applied to the difference scale. A composite scale is one that effectively balances the similarity and difference scales against each other. They can be balanced equally against each other (e.g., a=0.5 and b=0.5). Alternatively, the similarity scale may be given more weight (e.g., a=0.8 and b=0.2).
[0089] The grammatical units extracted from a given text document may be sentences, phrases, paragraphs, words, n-grams, or other grammatical units. In this case, if the grammatical unit 3608 extracted from the given text document 3610 is a word or an n-gram, the process may be considered keyword generation, as opposed to text summarization.
[0090] The text summarizer 3612 accepts a text fragment, which is a given text document 3610 or its grammatical unit. This text fragment is represented as an "input" vector of its weighted attributes (e.g., words or n-grams). Each weight in the input vector is for a corresponding attribute (e.g., a word or n-gram) identified in the text fragment and represents the strength of the association between the text fragment and the corresponding attribute. For example, these weights may be calculated according to a TF-IDF scheme, etc.
[0091] In some implementation examples, the attribute weights in the input vector are calculated as follows:
[0092]
number
[0093] Here, tf t,d is the frequency of n-gram t in text fragment d. Parameters k, b, L d , and L avg This is the same as before, except that it concerns knowledge base 3602 rather than a classification training set. In some implementations, k is approximately 1.5 and b is approximately 0.75.
[0094] Other weighting schemes are also possible, and it should be noted that the embodiments are not limited to any particular weighting scheme when forming the input vectors. Forming the input vectors may also involve unit-length normalization with respect to the training data item vectors as described above.
[0095] The text summarizer 3612 non-zero input vector formed based on text fragments The weighted attributes are iterated over, and attribute vectors corresponding to the attributes are extracted from the weighted transposed index 3606. These extracted attribute vectors are then merged into a weighted vector representing the concept of the text fragment. This weighted vector of the concept will hereafter be referred to as the "concept" vector.
[0096] The attribute vectors extracted from the weighted transposed index 3606, which corresponds to the attributes of the input vector, are also each weight vectors. However, the weights in the attribute vectors quantify the strength of the association between each concept in the knowledge base 3602 and the attribute mapped to the attribute vector by the transposed index 3606.
[0097] The text summarizer 3612 creates a conceptual vector of text fragments. The conceptual vector is a vector of weights. Each weight in the conceptual vector represents the strength of the association between each concept in the knowledge base 3602 and the text fragment. The conceptual weights in the conceptual vector are calculated by the text summarizer 3612 as the sum of the values for each attribute that have been weighted non-zero in the input vector. Each value for an attribute in this sum is calculated as the product of (a) the attribute weight in the input vector and (b) the concept weight in the attribute vector for that attribute. Each conceptual weight in the conceptual vector reflects the association of the concept with respect to the text fragment. In some implementation examples, the conceptual vector is normalized. For example, the conceptual vector may be normalized with respect to unit length or conceptual length (e.g., class length above).
[0098] The text summarizer 3612 can generate a conceptual vector 3616 for the input text 3610 and a conceptual vector 3614 for each of the grammatical units 3608. The vector comparator 3618 compares the conceptual vectors 3614 generated for the grammatical units with the conceptual vectors 3616 generated for the input text 3610 using a similarity scale to generate a similarity estimate 3620. In some implementations, a cosine similarity scale is used. Implementations are not limited to any particular similarity scale, and any similarity scale that allows measuring the similarity between two non-zero vectors may be used.
[0099] The similarity estimate 3620 quantifies the degree of semantic relationship between a given grammatical unit and the input text 3610 from which that grammatical unit was extracted. For example, the similarity estimate 3620 may be a value between 1 and 0, including values closer to 1 for a higher degree of semantic relationship and values closer to 0 for a lower degree of semantic relationship.
[0100] A similarity estimate 3620 can be calculated for each of the grammatical units 3608. The similarity estimates 3620 generated for each of the grammatical units 3608 can be used to select one or more of the grammatical units 3608 to include in the text summary of the input text 3610 (or to select one or more keywords for keyword generation about the input text 3610).
[0101] For example, there are various applications of the aforementioned techniques in text summarization, which provides accurate text summaries of longer texts such as news articles, blog posts, journal articles, and web pages.
[0102] In any or all of the above embodiments, after one or more content resources (e.g., images, articles, etc.) are identified as potentially relevant to the content currently being created by the user via the user interface 415, the relevant content resources are sent back to the content recommendation engine 425, where they may be retrieved, modified, and embedded in the user interface 415, for example, by the content retrieval / embedding component 445. The retrieval / embedding component 445 may optionally be used to include the potentially relevant content resources in the content currently being created. Image recommendations can be provided to the user via a user interface 415 so that they can be selected. Two exemplary user interfaces are shown in Figures 37 and 38, where image recommendations are provided to the user while content is being created. Figure 37 shows a media recommendation pane containing images selected by the user via the user interface based on the text of the content currently being authored. In Figure 38, visual feature analysis has been used to select a set of images potentially related to a first image selected by the user ("filename.JPG"). Using similar techniques and user interface screens, it may be possible to enable the user to select and drag and drop images, links to articles and other text documents, audio / video files, etc., and embed them in the content currently being created by the user.
[0103] Figure 39 shows a simplified diagram of a distributed system 3900 in which the various examples described above may be implemented. In the illustrated example, the distributed system 3900 includes one or more client computing devices 3902, 3904, 3906, and 3908 connected to a server 3912 via one or more communication networks 3910. The client computing devices 3902, 3904, 3906, and 3908 may be configured to run one or more applications.
[0104] In various embodiments, the server 3912 may be adapted to run one or more services or software applications that enable one or more operations associated with the content recommendation system 400. For example, a user may use client computing devices 3902, 3904, 3906, and 3908 (corresponding, for example, to the content author device 410) to access one or more cloud-based services provided via the content recommendation engine 425.
[0105] In various examples, server 3912 may also provide other services or software applications, which may include non-virtual and virtual environments. In some examples, these services may be provided to users of client computing devices 3902, 3904, 3906, and 3908 as web-based services or cloud services, or under a Software as a Service (SaaS) model. Users operating client computing devices 3902, 3904, 3906, and 3908 may then interact with server 3912 using one or more client applications to utilize the services provided by these components.
[0106] In the configuration shown in Figure 39, server 3912 may include one or more components 3918, 3920, 3922 that implement the functions performed by server 3912. These components may include one or more processors, hardware components, or software components that can be executed by a combination thereof. It should be recognized that a wide variety of system configurations different from the exemplary distributed system 3900 are possible.
[0107] Client computing devices 3902, 3904, 3906, and 3908 may include various types of computing systems, such as portable handheld devices like smartphones and tablets, general-purpose computers like personal computers and laptops, workstation computers, wearable devices like head-mounted displays, gaming systems such as portable game consoles and internet-enabled game consoles, thin clients, various messaging devices, sensors, and other sensing devices. These computing devices may run on various mobile operating systems (e.g., Microsoft Windows® Mobile®, iOS®, Wind). The client device may run various types and versions of software applications and operating systems (e.g., Microsoft Windows®, Apple Macintosh®, UNIX® or UNIX-like operating systems, Linux® or Linux-like operating systems), including ows Phone®, Android®, BlackBerry®, and Palm OS®. It can run a wide variety of applications, such as S) applications, and can use various communication protocols. The client device may provide an interface that allows the user of the client device to interact with the client device. The client device may also output information to the user through this interface. Although Figure 39 shows only four client computing devices, any number of client computing devices may be supported.
[0108] Network 3910 within the distributed system 3900 may be any type of network known to those skilled in the art, capable of supporting data communication using any of the various available protocols, including but not limited to TCP / IP (Transmission Control Protocol / Internet Protocol), SNA (System Network Architecture), IPX (Internet Packet Switching), AppleTalk®, etc. For example, network 3910 may be a local area network (LAN), an Ethernet®-based network, Token Ring, a wide area network, the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., a network operating under any of the IEEE 802.11 protocol suites, Bluetooth®, and / or any other wireless protocol), and / or any combination of these and / or other networks.
[0109] Server 3912 may consist of one or more general-purpose computers, dedicated server computers (including, for example, PC (personal computer) servers, UNIX® servers, midrange servers, mainframe computers, rack-mount servers, etc.), server farms, server clusters, or any other suitable configuration and / or combination. Server 3912 may include one or more virtual machines running a virtual operating system, or other computing architectures with virtualization, such as one or more flexible pools of logical storage that can be virtualized to maintain virtual storage for the server. In various examples, Server 3912 may be adapted to run one or more services or software applications that perform the operations described above.
[0110] Server 3912 may run any of the above-mentioned operating systems, as well as any server operating system available on the market. Server 3912 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP (Hypertext Transfer Protocol) servers, FTP (File Transfer Protocol) servers, CGI (Common Gateway Interface) servers, JAVA® servers, and database servers. Examples of database servers include those from Oracle®, Microsoft®, Sybase®, IBM® (International Business Machines), and others available on the market. This includes, but is not limited to, items that are available.
[0111] In some implementations, server 3912 may include one or more applications for analyzing and organizing data feeds and / or event updates received from users of client computing devices 3902, 3904, 3906, and 3908. For example, data feeds and / or event updates may include, but are not limited to, feeds or real-time updates received from one or more third-party sources and continuous data streams, including real-time events related to sensor data applications, financial stock market boards, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, and automotive traffic monitoring. Server 3912 may also include one or more applications for displaying data feeds and / or real-time events via one or more display devices on client computing devices 3902, 3904, 3906, and 3908.
[0112] The distributed system 3900 may also include one or more data repositories 3914, 3916. These data repositories may provide mechanisms for storing various types of information, such as the information described by the various examples above. The data repositories 3914, 3916 may reside in various locations. For example, the data repository used by server 3912 may be located locally with server 3912 or in a remote location from server 3912, communicating with server 3912 via a network-based connection or a dedicated connection. The data repositories 3914, 3916 may be of different types. In some examples, the data repository used by server 3912 may be a database, such as a relational database, such as a database provided by Oracle Corporation® and other manufacturers. One or more of these databases may be adapted to allow the storage, updating, and retrieval of data to and from the database in response to commands in SQL format.
[0113] In some examples, one or more of the data repositories 3914, 3916 may also be used by the application to store application data. The data repositories used by the application may be of various types, such as a key-value store repository, an object store repository, or a general-purpose storage repository supported by the file system.
[0114] In some examples, a cloud environment may offer one or more services as described above. Figure 40 is a simplified block diagram of one or more components of a system environment 4000 that can provide these and other services as cloud services. In the example shown in Figure 40, the cloud infrastructure system 4002 may provide one or more cloud services that a user may request using one or more client computing devices 4004, 4006, and 4008. The cloud infrastructure system 4002 may include one or more computers and / or servers, which may include those described above for server 3912 in Figure 39. The computers in the cloud infrastructure system 4002 in Figure 40 may be organized as general-purpose computers, dedicated server computers, server farms, server clusters, or any other suitable configuration and / or combination.
[0115] Network 4010 can facilitate data communication and exchange between clients 4004, 4006, and 4008 and the cloud infrastructure system 4002. Network 4010 may include one or more networks. The networks may be of the same type or different types. Network 4010 facilitates communication. To that end, it may support one or more communication protocols, including wired and / or wireless protocols.
[0116] The example shown in Figure 40 is merely one example of a cloud infrastructure system and is not intended to be limiting. In other examples, the cloud infrastructure system 4002 may have more or fewer components than those shown in Figure 40, may combine two or more components, or may have components in different configurations or arrangements. For example, Figure 40 shows three client computing devices, but in other examples, the number of client computing devices that may be supported is arbitrary.
[0117] The term "cloud service" is generally used to refer to services made available to users on demand via communication networks such as the Internet, through a service provider's system (e.g., cloud infrastructure system 4002). Typically, in a public cloud environment, the servers and systems that make up the cloud service provider's system are different from the customer's own on-premises servers and systems. The cloud service provider's system is managed by the cloud service provider. Therefore, customers can use the cloud services provided by the cloud service provider without purchasing separate licenses, support, or hardware and software resources for the service. For example, the cloud service provider's system can host applications, and users can order and use applications on demand and self-service via the Internet without purchasing infrastructure resources to run the applications. Cloud services are designed to provide easy and scalable access to applications, resources, and services. Several providers offer cloud services. For example, several cloud services, such as middleware services, database services, Java cloud services, and others, are offered by Oracle Corporation®.
[0118] In various examples, the cloud infrastructure system 4002 may provide one or more cloud services using various models, including a hybrid service model, a Software as a Service (SaaS) model, a Platform as a Service (PaaS) model, and an Infrastructure as a Service (IaaS) model. The cloud infrastructure system 4002 may include a set of applications, middleware, databases, and other resources that enable the provisioning of various cloud services.
[0119] The SaaS model enables the delivery of applications or software as a service to customers over a communication network such as the internet, without requiring customers to purchase the underlying hardware or software for the application. For example, the SaaS model may be used to allow customers to access on-demand applications hosted on a cloud infrastructure system 4002. Examples of SaaS services offered by Oracle Corporation® include, but are not limited to, various services for human resource / capital management, customer relationship management (CRM), enterprise resource planning (ERP), supply chain management (SCM), enterprise performance management (EPM), analytics services, social applications, and others.
[0120] The IaaS model generally involves providing infrastructure resources (such as servers, storage, hardware, and networking resources) to customers as a cloud service. It is used to provide flexible computing and storage capabilities. Various IaaS services are offered by Oracle Corporation (registered trademark).
[0121] The PaaS model is generally used to provide a platform and environment resource as a service, enabling customers to develop, run, and manage applications and services without having to procure, build, or manage such resources themselves. Examples of PaaS services offered by Oracle Corporation® include, but are not limited to, Oracle Java Cloud Service (JCS), Oracle Database Cloud Service (DBCS), data management cloud services, various application development solution services, and others.
[0122] In some examples, resources within the cloud infrastructure system 4002 may be shared by multiple users and dynamically reallocated as needed. In addition, resources may be allocated to users in different time zones. For example, the cloud infrastructure system 4002 may allow a first set of users in a first time zone to utilize resources in the cloud infrastructure system for a specified number of hours, and then reallocate the same resources to another set of users in a different time zone, thereby maximizing resource utilization.
[0123] The cloud infrastructure system 4002 may provide cloud services through different deployment models. In a public cloud model, the cloud infrastructure system 4002 may be owned by a third-party cloud service provider, and the cloud services are provided to any customer in the general public. In this case, the customer may be an individual or a company. In some other embodiments, in a private cloud model, the cloud infrastructure system 4002 may function within an organization (for example, within a corporate organization), and the services are provided to customers within this organization. For example, these customers may be various departments within a company, such as the human resources department or the payroll department, or individuals within the company. In other specific embodiments, in a community cloud model, the cloud infrastructure system 4002 and the services provided may be shared among several organizations within the community concerned. Various other models, such as hybrid models of the above models, may be used.
[0124] Client computing devices 4004, 4006, and 4008 may be similar to those described above for client computing devices 3902, 3904, 3906, and 3908 in Figure 39. Client computing devices 4004, 4006, and 4008 in Figure 40 may be configured to operate client applications such as a web browser, a proprietary client application (e.g., Oracle Forms), or any other application that can be used by users of the client computing devices to interact with the cloud infrastructure system 4002 in order to use services provided by the cloud infrastructure system 4002.
[0125] In various examples, the cloud infrastructure system 4002 can also provide “big data” and related computing and analytical services. The term “big data” is generally used to refer to extremely large datasets that can be stored and manipulated by analysts and researchers to visualize, detect trends, and / or otherwise interact with the data. The analysis that the cloud infrastructure system 4002 can perform involves using, analyzing, and manipulating large datasets to detect various trends, behaviors, relationships, etc., within this data. This may include visualization. This analysis may be performed by one or more processors, which may, in some cases, process the data in parallel and run simulations using the data. The data used in this analysis may be structured data (e.g., data stored in a database or data structured according to a structured model) and / or unstructured data (e.g., data blobs (binary large objects)). It may include the object.
[0126] As shown in the embodiment of Figure 40, the cloud infrastructure system 4002 may include infrastructure resources 4030 used to facilitate the provision of various cloud services provided by the cloud infrastructure system 4002. The infrastructure resources 4030 may include, for example, processing resources, storage or memory resources, networking resources, etc.
[0127] In some cases, these resources may be grouped together into resource sets or resource modules (also called "pods") to facilitate efficient provisioning of these resources to support various cloud services provided by the cloud infrastructure system 4002 to different customers. Each resource module or pod may contain a pre-integrated and optimized combination of one or more types of resources. In some cases, different pods may be pre-provisioned for different types of cloud services. For example, a first set of pods may be provisioned for a database service, and a second set of pods, which may contain different resource combinations than those in the first set of pods, may be provisioned for a Java service, etc. In some cases, the resources allocated to provision these services may be shared among those services.
[0128] The cloud infrastructure system 4002 itself may internally use services 4032 that are shared by various components of the cloud infrastructure system 4002 and facilitate the provisioning of services by the cloud infrastructure system 4002. These internally shared services may include, but are not limited to, security and identity services, integration services, enterprise repository services, enterprise manager services, virus scanning and whitelisting services, high-availability backup and recovery services, services that enable cloud support, email services, notification services, and file transfer services.
[0129] In various examples, the cloud infrastructure system 4002 may include multiple subsystems. These subsystems may be implemented in software, hardware, or a combination thereof. As shown in Figure 40, the subsystem may include a user interface subsystem w that enables users or customers of the cloud infrastructure system 4002 to interact with the cloud infrastructure system 4002. The user interface subsystem 4012 may include various different interfaces, such as a web interface 4014, an online store interface 4016 where cloud services offered by the cloud infrastructure system 4002 are advertised and available for purchase by consumers, and other interfaces 4018. For example, a customer may use a client device to request (service request 4034) one or more services offered by the cloud infrastructure system 4002 using one or more of interfaces 4014, 4016, and 4018. For example, a customer may access the online store, browse the cloud services offered by the cloud infrastructure system 4002, and select one of the services offered by the cloud infrastructure system 4002 that the customer wishes to subscribe to. A customer may place a subscription order for one or more services. A service request may include information identifying the customer and one or more services that the customer wishes to subscribe to. For example, a customer may place a subscription order for the services described above. As part of the order, the customer may provide information identifying, among other things, the amount of resources the customer needs and / or the time frame in which they apply.
[0130] In some examples, such as the one shown in Figure 40, the cloud infrastructure system 4002 may include an Order Management Subsystem (OMS) 4002 configured to process new orders. As part of this process, the OMS 4020 may be configured to prepare the order for provisioning by, among several operations, creating a customer account if one does not already exist, receiving billing and / or account information from the customer to be used to charge the customer in order to provide the requested services to the customer, verifying the customer information, reserving the order for the customer after verification, and coordinating various workflows.
[0131] If properly validated, OMS4020 may invoke Order Provisioning Subsystem (OPS)4024, which is configured to provision resources for this order, including processing resources, memory resources, and networking resources. Provisioning may include allocating resources for the order and configuring those resources to facilitate the services requested by the customer order. The method of provisioning resources for an order and the types of resources provisioned may depend on the type of cloud service ordered by the customer. For example, following a certain workflow, OPS4024 may be configured to determine the specific cloud service being requested and to identify the number of pods that would have been pre-configured for that particular cloud service. The number of pods to be allocated for an order may depend on the size / volume / level / scope of the requested service. For example, the number of pods to allocate may be determined based on the number of users the service should support, the duration for which the service is requested, etc. The allocated pods may then be customized to suit the specific customer making the request in order to provide the requested service.
[0132] The cloud infrastructure system 4002 may send a response or notification 4044 to the requesting customer to indicate when the requested service will be available. In some examples, the customer may be sent information (e.g., a link) that will enable the customer to begin using and utilizing the benefits of the requested service.
[0133] The cloud infrastructure system 4002 may provide services to multiple customers. For each customer, the cloud infrastructure system 4002 manages information related to one or more subscription orders received from the customer, maintains customer data related to the orders, and provides the requested services to the customer. The cloud infrastructure system 4002 may also collect usage statistics about the customer's use of the subscribed services. For example, statistics may be collected on the amount of storage used, the amount of data transferred, the number of users, and the amount of system uptime and system downtime. This usage information may be used to charge the customer. Billing may be done, for example, on a monthly basis.
[0134] The cloud infrastructure system 4002 may provide services to multiple customers in parallel. The cloud infrastructure system 4002 may store information about these customers, which may include copyright information. In some examples, the cloud infrastructure system 4002 manages customer information and isolates the managed information so that information about one customer cannot be accessed by another customer. The system includes an Identity Management Subsystem (IMS) 4028 configured to provide various security-related services, such as identity services, including information access management, authentication and authorization services, and services for managing customer identities and roles and related capabilities.
[0135] Figure 41 shows an example of a computer system 4100 that can be used to implement the various examples described above. In some examples, any of the various servers and computer systems described above can be implemented by using computer system 4100. As shown in Figure 41, computer system 4100 includes various subsystems, including a processing subsystem 4104 that communicates with several other subsystems via a bus subsystem 4102. These other subsystems may include a processing acceleration unit 4106, an I / O subsystem 4108, a storage subsystem 4118, and a communication subsystem 4124. The storage subsystem 4118 may include a non-temporary computer-readable storage medium 4122 and system memory 4110.
[0136] The bus subsystem 4102 provides a mechanism for various components and subsystems of the computer system 4100 to communicate with each other as intended. Although the bus subsystem 4102 is schematically shown as a single bus, alternative bus subsystems may utilize multiple buses. The bus subsystem 4102 may be one of several types of bus structures, including a memory bus or memory controller, peripheral bus, and local bus, using one of various bus architectures. For example, such architectures include the Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and the IEEE P1386.1 standard. Therefore, the manufactured mezzanine bus may include a Peripheral Component Interconnect (PCI) bus, etc.
[0137] The processing subsystem 4104 controls the operation of the computer system 4100 and may include one or more processors, application-specific integrated circuits (ASICs), or field-programmable gate arrays (FPGAs). The processors may include single-core or multi-core processors. The processing resources of the computer system 4100 can be organized into one or more processing units 4132, 4134, etc. A processing unit may include one or more processors, including single-core or multi-core processors, one or more cores from the same or different processors, a combination of cores and processors, or other combinations of cores and processors. In some examples, the processing subsystem 4104 may include one or more dedicated coprocessors, such as graphics processors or digital signal processors (DSPs). In some examples, some or all of the processing units of the processing subsystem 4104 can be implemented using customized circuits such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs).
[0138] In some examples, a processing unit within the processing subsystem 4104 can execute instructions stored in system memory 4110 or a computer-readable storage medium 4122. In various examples, a processing unit can execute various program or code instructions and maintain multiple programs or processes running simultaneously. At any given time, some or all of the program code to be executed is stored in system memory 4110 and / or potentially one or more computer-readable storage devices. It may reside permanently on the read / store medium 4110. With appropriate programming, the processing subsystem 4104 can provide the various functions described above. In an example where the computer system 4100 is running one or more virtual machines, each virtual machine may be assigned to one or more processing units.
[0139] In some examples, a processing acceleration unit 4106 may be optionally provided to perform customized processing to accelerate the overall processing performed by the computer system 4100, or to offload a portion of the processing performed by the processing subsystem 4104.
[0140] The I / O subsystem 4108 may include devices and mechanisms for inputting information into and / or outputting information from or via computer system 4100. Generally, the use of the term “input device” is intended to include all conceivable types of devices and mechanisms for inputting information into computer system 4100. User interface input devices may include, for example, pointing devices such as keyboards, mice or trackballs, touchpads or touchscreens integrated into displays, scroll wheels, click wheels, dials, buttons, switches, keypads, voice input devices with voice command recognition systems, microphones, and other types of input devices. User interface input devices may also include motion sensing and / or gesture recognition devices, such as Microsoft Kinect® motion sensors, Microsoft Xbox® 360 game controllers, and devices that provide interfaces for receiving input using gestures and voice commands, enabling users to control and interact with input devices. The user interface input device may also include an eye gesture recognition device that detects eye movements from the user (e.g., blinking while taking a picture and / or making a menu selection) and translates the eye gestures into input to the input device. In addition, the user interface input device may include a voice recognition sensing device that enables the user to interact with a voice recognition system via voice commands.
[0141] Other examples of user interface input devices include, but are not limited to, three-dimensional (3D) mice, joysticks or pointing sticks, gamepads and graphic tablets, as well as auditory / visual devices such as speakers, digital cameras, digital camcorders, portable media players, webcams, image scanners, fingerprint scanners, barcode readers, 3D scanners, 3D printers, laser rangefinders, and eye-tracking devices. In addition, user interface input devices may include medical imaging input devices such as computed tomography, magnetic resonance imaging, positional emission tomography, and medical ultrasound devices. User interface input devices may also include audio input devices such as MIDI keyboards and digital musical instruments.
[0142] In general, the use of the term "output device" is intended to include all conceivable types of devices and mechanisms for outputting information from the computer system 4100 to a user or another computer. User interface output devices may include non-visual displays such as display subsystems, indicator lights, or audio output devices. Display subsystems may include flat panel devices such as those using cathode ray tubes (CRTs), liquid crystal displays (LCDs), or plasma displays, projection devices, touchscreens, etc. For example, user interface output devices include monitors, printers, speakers, headphones, and car navigation systems. This may include, but is not limited to, various display devices that visually convey text, graphics, and audio / video information, such as display systems, plotters, audio output devices, and modems.
[0143] The storage subsystem 4118 includes a repository or datastore for storing information used by the computer system 4100. The storage subsystem 4118 provides a tangible, non-temporary, computer-readable storage medium for storing basic programming and data configurations that provide some example functionality. Software (e.g., programs, code modules, instructions) that, when executed by the processing subsystem 4104, provides the above-described functionality may be stored in the storage subsystem 4118. The software may be executed by one or more processing units of the processing subsystem 4104. The storage subsystem 4118 may also include a repository for storing data used in accordance with this disclosure.
[0144] The storage subsystem 4118 may include one or more non-temporary memory devices, including volatile memory devices and non-volatile memory devices. As shown in Figure 41, the storage subsystem 4118 includes system memory 4110 and computer-readable storage medium 4122. The system memory 4110 may include several memories, including volatile primary random access memory (RAM) for storing instructions and data during program execution, and non-volatile read-only memory (ROM) or flash memory for storing fixed instructions. In some implementation examples, a basic input / output system (BIOS) containing basic routines to assist in the transfer of information between elements within the computer system 4100, such as during startup, is typically located in ROM. It may be stored. Typically, RAM contains data and / or program modules currently being operated and executed by the processing subsystem 4104. In some implementations, system memory 4110 may include several different types of memory, such as static random access memory (SRAM) and dynamic random access memory (DRAM).
[0145] As an example, without limitation, as shown in Figure 41, the system memory 4110 may load running application programs 4112, program data 4114, and operating systems 4116, which may include client applications, web browsers, middle-tier applications, relational database management systems (RDBMS), etc. As an example, the operating system 4116 may include Microsoft Windows®, Apple Macintosh®, and / or Linux operating systems, various UNIX® or UNIX-like operating systems available on the market (including, but not limited to, various GNU / Linux operating systems, Chrome® OS, etc.), and / or mobile operating systems, such as iOS®, Windows® Phone, Android® OS, BlackBerry® OS, and Palm® OS operating systems.
[0146] The computer-readable storage medium 4122 can store programming and data structures that provide several example functions. The computer-readable medium 4122 may store computer-readable instructions, data structures, program modules, and other data for the computer system 4100. Software (programs, code modules, instructions) that, when executed by the processing subsystem 4104, provides the above functions may be stored in the storage subsystem 4118. As an example, the computer-readable storage medium 4122 may be a hard disk drive, magnetic disk drive, CD-ROM, DVD, B The computer-readable storage medium 4122 may include, but is not limited to, Zip® drives, flash memory cards, Universal Serial Bus (USB) flash drives, Secure Digital (SD) cards, DVD discs, digital videotapes, and the like. The computer-readable storage medium 4122 may also include flash memory-based SSDs, enterprise flash drives, solid-state drives (SSDs) based on non-volatile memory such as solid-state ROM, SSDs based on volatile memory such as solid-state RAM, dynamic RAM, and static RAM, DRAM-based SSDs, magnetoresistive RAM (MRAM) SSDs, and hybrid SSDs that use a combination of DRAM and flash memory-based SSDs. The computer-readable storage medium 4122 may store computer-readable instructions, data structures, program modules, and other data for the computer system 4100.
[0147] In some examples, the storage subsystem 4118 may also include a computer-readable storage medium reader 4120 that can be further connected to a computer-readable storage medium 4122. The reader 4120 may be configured to receive and read data from memory devices such as disks and flash drives.
[0148] In some cases, the computer system 4100 may support virtualization technologies, including but not limited to the virtualization of processing and memory resources. For example, the computer system 4100 may provide support for running one or more virtual machines. The computer system 4100 may run programs such as hypervisors that facilitate the configuration and management of virtual machines. Each virtual machine generally runs independently of other virtual machines. Virtual machines may be allocated memory, computing (e.g., processors, cores), I / O, and networking resources. Each virtual machine runs its own operating system, which may typically be the same as or different from the operating systems run by other virtual machines run by the computer system 4100. Thus, multiple operating systems may potentially run simultaneously by the computer system 4100.
[0149] The communication subsystem 4124 provides interfaces to other computer systems and networks. It functions as an interface for sending and receiving data between other systems and the computer system 4100. For example, the communication subsystem 4124 may enable the computer system 4100 to establish communication channels to one or more client computing devices via the internet in order to send and receive information with one or more client computing devices.
[0150] The communication subsystem 4124 may support both wired and / or wireless communication protocols. For example, in some cases, the communication subsystem 4124 may include radio frequency (RF) transceiver components for accessing wireless voice and / or data networks (e.g., using cellular telephone technology, advanced data network technologies such as 3G, 4G, or EDGE (High Speed Data Rate for Global Evolution), WiFi (IEEE 802.11 family of standards), or other mobile communication technologies, or any combination thereof), a global positioning system (GPS) receiver component, and / or other components. In some cases, the communication subsystem 4124 may provide a wired network connection (e.g., Ethernet®) in addition to or instead of a wireless interface.
[0151] The communication subsystem 4124 can receive and transmit data in various formats. For example, in some cases, the communication subsystem 4124 may receive input communications in the form of structured data feeds and / or unstructured data feeds 4126, event streams 4128, event updates 4130, etc. For example, the communication subsystem 4124 may be configured to receive (or transmit) data feeds 4126 in real time from users of social media networks, and / or from users of other communication services such as feeds, updates, web feeds such as Rich Site Summary (RSS) feeds, and / or real-time updates from one or more third-party sources.
[0152] In some examples, the communication subsystem 4124 may be configured to receive data in the form of a continuous data stream, which may include an event stream 4128 and / or event update 4130 of real-time events that are inherently continuous or infinite and do not have a clear end. Examples of applications that generate continuous data include, for example, sensor data applications, financial stock market boards, network performance measurement tools (e.g., network monitoring and traffic management applications), clickstream analysis tools, and automotive traffic monitoring.
[0153] The communication subsystem 4124 may be configured to output structured and / or unstructured data feeds 4126, event streams 4128, event updates 4130, etc., to one or more databases that can communicate with one or more streaming data source computers coupled to the computer system 4100.
[0154] The computer system 4100 may be one of a variety of types, including handheld portable devices (e.g., iPhone® cellular phone, iPad® computing tablet, PDA), wearable devices (e.g., head-mounted displays), personal computers, workstations, mainframes, kiosks, server racks, or other data processing systems.
[0155] Because the nature of computers and networks is constantly changing, the computer system 4100 shown in Figure 41 is described only as an example. Many other configurations with more or fewer components than the system shown in Figure 41 are possible. Those skilled in the art will understand other embodiments and / or methods for realizing various examples based on the disclosure and teachings herein.
[0156] Smart Content - Ranking Results In a content recommendation system configured to recommend specific content (content items) in response to input from a client or user, the process of evaluating and ranking content items against each other plays a crucial role in the overall recommendation process. For example, a user may enter one or more search terms into a search engine query interface, and the recommendation system may recommend one or more relevant content items (e.g., web pages, documents, images, etc.) that most closely match the entered search terms. In another example, a user may use a content authoring system to author original content such as emails, documents, online articles, or blog posts. The content authoring system may be configured to recommend related content items that may be relevant to the content authored by the user. For example, as mentioned above, the content recommendation system may, if the author wishes, recommend one or more of the recommended content items to the content authored by the user. The system may recommend to the user relevant images or links to web pages, etc., so that they can be incorporated into the content. Several examples of such techniques are described in the paragraph above. As these examples show, ranking and recommending specific content items can be applied to many different use cases in response to receiving manually entered user input and / or automated input from other processes. Content items to be ranked and / or recommended may correspond to images, web pages, other media files, documents, digital objects, etc., that the recommendation system can use during a search. These content items may be stored in one or more repositories that are accessible to the recommendation system, which may be dedicated repositories or public repositories (e.g., the internet).
[0157] However, ranking content items is a crucial task. For example, consider a recommendation system that uses tag matching technology to perform recommendations. In such a system, content items available for search by the recommendation system are tagged, and the content items may be stored in one or more repositories along with their tags. Tagging may be performed by a content item tagging service / application. For content items, one or more tags associated with a content item indicate the content contained within that item. A value (sometimes called a tag probability) may also be associated with each tag, and this value provides a measure (e.g., probability) of the content indicated by the tags appearing in the content item. Upon receiving user input (e.g., search terms / phrases, user-authored content) that will result in recommendations for content items, the user input may be analyzed to identify one or more tags that should be associated with the user input. The recommendation system may then use tag matching technology to identify a set of content items from the content items available for search that have associated tags that match the tags associated with the user input. The recommendation system may then use some ranking algorithm to rank the content items within the identified set and display the results to the user.
[0158] However, the effectiveness of the ranking algorithms used by recommendation systems may be limited in certain use cases and may not produce optimal results. For example, consider the case where multiple tags are associated with user input. For instance, a user searches for the words "coffee" and "human" in an image search engine. When typing, the tags "coffee" and "person" may be associated with the user input. In certain embodiments, the search term itself may be treated as a tag associated with the user input. For brevity, assume that the collection of content items available for search includes tagged images. The recommendation system may use these two tags to retrieve a set of matching content items from the collection of content items (e.g., images). In this case, a content item is considered a match if at least one tag associated with it matches a tag associated with the user input. Consider a scenario (Example 1) where multiple matching content items are retrieved by the recommendation system, and each content item is associated with both the "coffee" and "person" tags. One possible way to rank these retrieved content items is to (a) add the values associated with the "coffee" and "person" tags for each matching content item, and then (b) rank the content items based on their associated added sums. However, this presents a problem because multiple values associated with matching tags for multiple matching content items may sum up to the same value. This situation is highly likely to occur because most tagging services normalize multiple probabilities in only one step. For example, three matching images could have associated tag values as follows: Specifically, Image A (("coffee", 0.5), ("person", 0.5)); Image B (("coffee", 0.2) The images are ,("person", 0.8)); and image C(("coffee", 0.7),("person", 0.3)). As described above, the sum of the matching tag values for each of these matching images is "1", and therefore there is no way to rank one of these images relative to the other using conventional addition techniques. Consequently, simply ranking these images based on the sum of their matching tag values cannot be used for image ranking.
[0159] Extending the above example, a set of matching images could also contain multiple images that share the same tag value and are associated with only one tag (for example, only "coffee" or only "person"). This also presents problems when ranking images. For example, consider a case where three matching images may have the following associated tag values: Image A (("coffee", 0.5)), Image B (("person", 0.5)). Again, there is no way to rank one of these images relative to the others using conventional techniques.
[0160] The above situation worsens if there are three or more tags associated with user input. For example, if a user types the words "coffee," "person," and "cafe" into an image search engine, three search tags are associated with the user input: "coffee," "person," and "cafe." The recommendation system may pick a set of matching content items from a collection of content items (e.g., images). In this case, a content item is considered a match if at least one tag associated with that content item matches a tag associated with the user input. The number of matching tags for a matching content item can vary from just one matching tag to several matching tags (in the example of "coffee," "person," and "cafe," there are up to three matching tags). In this scenario as well, if there are multiple matching content items, the values associated with the matching tags can sum up to the same value. For example, three matching images may have the following associated tag values: Specifically, these are image A (("person", 0.8), ("coffee", 0.2)); image B (("person", 0.2), ("coffee", 0.2), ("cafe", 0.6)); and image C (("person", 0.5), ("cafe", 0.5)). As described above, the sum of the matching tag values for each of these matching images is "1", and therefore there is no way to rank one of these images relative to the other using conventional additive techniques.
[0161] Therefore, in many cases, simple tag matching techniques may not return optimal content recommendations for a variety of reasons. For example, a particular image (or other content item) in a repository may be tagged with only one or two content tags, while other images / content items may be tagged with numerous tags, potentially containing tens or hundreds of tags for a single content item. In such cases, traditional tag matching techniques may over-recommend highly tagged content items (for example, because they are more likely to contain at least one tag that matches the entered word) and / or under-recommend such items (for example, because even when matching one or more content tags, the majority of those tags still do not match the entered word). Similarly, input content provided by a user or client system may contain only a few input words (for example, clearly entered search terms or topics extracted from a larger input text) or a relatively large number of input words, depending on the input data received. In such cases, traditional tag matching technologies may fail to identify specific related content items within the repository (for example, because there are too few input words that match the tags of the relevant content items), or they may incorrectly recommend less relevant content items (for example, because these less relevant content items contain one or more matching tags).
[0162] In certain embodiments, improved techniques for evaluating, ranking, and recommending tagged content items are described herein. In some embodiments, a content recommendation system may receive input content from a client device, such as a search query or newly authored text input. One or more tags may be included in the input content received from the client device, or may be associated with such input content, and / or may be determined and extracted from such input content based on preprocessing and analysis techniques performed on such input content. In addition, the content recommendation system may have access to a content repository that stores multiple tagged content items, such as images, media content files, links to web pages, and / or other documents. In some cases, the content repository may store data that identifies tagged content items, and for each tagged content item, it may further store associated tag information for each item. In this case, the tag information for a content item includes information that identifies one or more tags associated with the content item, and tag values for each associated tag.
[0163] In response to receiving input data for which recommendations should be made, the content recommendation system may retrieve a set of matching tagged content items from the content repository from a searchable collection of content items. In this case, a content item is considered a matching content item if at least one content tag associated with that content item matches a tag associated with the input content. For each matching tagged content item retrieved from the content repository, the content recommendation system may then calculate two scores: (1) a first score (also called the tag count score) based on the number of tags associated with the content item that match the tags associated with the input content, and (2) a second score (also called the tag value-based score or TVBS) based on the tag value for each of the matching tags for the content item. The content recommendation system then calculates a final ranking score for each of the matching content items based on the first and second scores for the matching content item. The final ranking scores calculated for the set of matching content items are then used to generate a ranking list of the matching content items. This ranking list is used to identify a recommended subset of matching content items that should be output to the user or client system.
[0164] Referring now to Figure 42, a block diagram of a computing environment 4200 is shown, which includes a content recommendation system 4220 implemented to evaluate and rank content items from a content repository 4230 in response to input content received from a user or client system 4210 according to a particular embodiment. This example also shows various components and subsystems within the content recommendation system 4220, including a graphical user interface (GUI) 4215. The client system 4210 is this GUI. The user can interact with the content recommendation system 4220 via 4215 to provide input content and receive data identifying a subset of recommended content items. In certain embodiments, the GUI 4215 may be the GUI of a separate client application 4215 (e.g., a web browser application) used by the user to author content. In this embodiment, the content recommendation system 4220 may receive content provided or authored by the user from the client application. The content is an application programming interface (application progra) that enables the client application and the content recommendation system 4220 to interact with each other and exchange information. The content recommendation system 4220 receives data via the mming interface (API). It's okay.
[0165] The embodiments shown in Figure 42 are merely examples and are not intended to unduly limit the scope of the claimed embodiments. Those skilled in the art will recognize many possible modifications, alternatives, and variations. For example, in some implementations, the content recommendation system 4220 may have more or fewer systems or subsystems than those shown in Figure 42, or it may be a combination of two or more systems, or it may have systems in different configurations or arrangements. In some embodiments, the content recommendation system 4220 may be implemented as one or more computing systems, including separate systems using independent computing and network infrastructure with dedicated, specialized hardware and software. Alternatively or additionally, one or more of these components and subsystems may be integrated into a single system performing a distinct function. The various systems, subsystems, and components shown in Figure 42 may be implemented as software (e.g., code, instructions, programs), hardware, or a combination thereof, executed by one or more processing units (e.g., processors, cores) of each system. The software may be stored on non-temporary storage media (e.g., on memory devices).
[0166] At a high level, the content recommendation system 4220 is configured to receive user input content and, in response to that user input content, to recommend content items based on that user input content. These recommendations are made from a collection of content items that are available and accessible to the content recommendation system 4220 for recommendation purposes. The collection of content items may include images, various types of documents, media content, digital objects, etc. Based on tag information associated with the user input content and tag information associated with the collection of content items, the content recommendation system 4220 is configured to use tag matching technology to identify a set of matching content items for the user input content. The content recommendation system 4220 is then configured to rank the content items in the set of matching content items using the innovative ranking technology described herein. Based on the ranking, the content recommendation system 4220 is configured to identify and recommend a subset of matching content items that should be output to the user or client system.
[0167] The content recommendation system 4220 includes a content tagging subsystem 4222 configured to receive or retrieve content items available for recommendation by the content recommendation system 4220. Content items may include, but are not limited to, images, web pages, documents, media files, etc. Content items may be received or retrieved from one or more content repositories 4230. The content repositories 4230 may include a variety of public or private content repositories, such as libraries or databases, including image libraries, document stores, and local or wide area networks (e.g., the Internet) of web-based resources. One or more content repositories 4230 may be stored locally in the content recommendation system 4220, while other content repositories may be separate and remote from the content recommendation system 4220 and accessible to the content recommendation system 4220 via one or more computer networks.
[0168] In a particular embodiment, for each content item, the content tagging subsystem 4222 retrieves and analyzes the content of the content item, and also performs content analysis. The system is configured to identify one or more content tags (tags) that should be associated with an item. For each tag associated with a content item, the content tagging unit 4222 may also determine the tag value associated with the tag. In this case, this value provides a measure (e.g., probability) of the content indicated by the tag appearing in the content item. The tag value for a tag may correspond to a numerical scale that represents how applicable a particular content tag is to that content item. One or more tags and their corresponding tag values may be associated with a content item. For a content item with multiple associated tags and their corresponding tag values, these tag values may represent the relative prominence of the image topic or theme indicated by the tags in the image. For example, a first tag associated with a content item with a relatively high tag value may indicate that the content or feature indicated by the first tag is particularly relevant and prominent in the content item. In contrast, a second tag associated with the same content item with a lower tag value may indicate that the content or feature indicated by the second tag is less prominent or widespread in the content item compared to the content indicated by the first content tag. For example, an image content item may have two associated tags and values, such as ("person", 0.8) and ("coffee", 0.2). This indicates that the image contains content related to coffee (e.g., a coffee cup) and a person (e.g., a person drinking coffee), and further, that the person is depicted more prominently in the image than the depiction of the coffee (e.g., the majority of the image may depict the person, and the coffee cup may occupy only a small area of the image). Tag values may be expressed using different formats. For example, in some implementations, these tag values may be represented as floating-point numbers between 0.1 and 1.0. In some implementations, the sum of all tag values for tags associated with a particular content item may be a fixed constant value (e.g., summing to 1).
[0169] In some embodiments, the content tagging unit 4222 may use the services of a content tagging service to perform a tagging task on a content item, which includes identifying one or more tags to be associated with the content item and a tag value for each tag. In certain embodiments, the content tagging unit 4222 is implemented using one or more predictive machine learning models that take content items as input and are trained to predict tags for the content item and its associated tag value. In some embodiments, these tags may be selected from a set of pre-configured tags used to train the models. Various machine learning techniques using pre-trained machine learning models, and / or other artificial intelligence-based tools including AI-based text or image classification systems, topic or feature extraction, and / or any other combination of the techniques described above may be used to determine which tags should be associated with the content item and its corresponding tag value.
[0170] In some embodiments, content items retrieved from the content repository 4230 may already contain associated content tags and tag values. If the retrieved content item does not contain tag information, and / or if the content recommendation system 4220 is configured to determine additional tags for the content item, the content tagging unit 4222 may be used to update or generate new tags for the retrieved content item. The content tagging unit 4222 may generate tag information (e.g., one or more tags and associated tag values) for the content item using a wide variety of techniques. For example, the content tagging unit 4222 may use one or all of the aforementioned techniques, such as parsing, processing, feature extraction, and / or other analysis techniques, to analyze the retrieved content item and determine the content tags. It may be. The type of parsing, processing, feature extraction, and / or analysis may depend on the type of content item. For example, for text-based content items such as blog posts, letters, emails, articles, and documents, analysis may include keyword extraction and processing tools (e.g., stemming, synonym lookup, etc.), topic analysis tools, etc. If the content item is an image, artificial intelligence-based image classification tools may be used to identify specific image features and / or generate image tags. For example, the analysis of an image may identify multiple image features, and the image may be tagged with each of these identified features. One or both types of analysis (i.e., tag extraction from images and keyword / topic extraction from text content) may be performed via REST-based services or other web services using analysis, machine learning algorithms, and / or artificial intelligence (AI)-based technologies, such as AI-based cognitive image analysis services, or similar AI / REST cognitive text services used for text content. Similar technologies may be used for other types of content items such as video files, audio files, graphics, or social media posts. In this case, a dedicated web service may be used to extract and analyze specific features (e.g., words, objects in images / videos, facial expressions, etc.) depending on the media type of the content item.
[0171] In some embodiments, the content tagging unit 4222 may use one or more machine learning and / or artificial intelligence-based pre-trained models trained on training data to identify and extract content features used to determine tags and tag values for content items. For example, a model training system may generate one or more models that can be pre-trained using machine learning algorithms based on a training dataset containing a training dataset of previous input data (e.g., text input, images, etc.) and corresponding tags for the previous input data. In various embodiments, one or more different types of pre-trained models may be used, including classification systems that perform supervised or semi-supervised learning techniques, such as naive Bayes models, decision tree models, logistic regression models, or deep learning models, or any other machine learning or artificial intelligence-based prediction systems that can perform supervised or unsupervised learning techniques. For each machine learning model or model type, the pre-trained model may be run by one or more computing systems. During this period, a content item may be provided as input to one or more models, and the output from the models may identify one or more tags that should be associated with the content item, or the output of the models may be used to identify one or more tags that should be associated with the content item. Therefore, the content tagging unit 4222 may use, but is not limited to, a wide variety of tools or techniques, such as keyword extraction and processing (e.g., stemming, synonym search, etc.), topic analysis, feature extraction from images, machine learning and AI-based modeling tools and text or image classification systems, and / or any other combination of the above techniques for determining or generating tag information for each content item that can be used for recommendations (e.g., one or more tags and associated tag values).
[0172] In certain embodiments, content items available for recommendation and their associated tag information (for example, for each content item, one or more tags associated with the content item and corresponding tag values) may be stored in the data store 4223. In some embodiments, the content / tag information data store 4223 may store data identifying content items retrieved from the content repository 4230. This may include the items themselves (e.g., images, web pages, documents, media files, etc.) or, additionally / alternatively, references to the items (e.g., item identifiers). This may include the child, the network address from which the content item can be retrieved, the item description, the item thumbnail, etc. Figure 45 shows an example of the types of data that may be stored in the content / tag information data store 4223, and will be explained in more detail below.
[0173] The content recommendation system 4220 includes a tag identifier subsystem 4221. The tag identifier subsystem 4221 is configured to receive user input content from device 4210 and to determine one or more tags that should be associated with the user content. In some embodiments, the user content received from device 4210 may include associated tags. In some other embodiments, the tag identifier 4221 may be configured to process the input data to determine one or more tags that should be associated with the input data. For example, the tag identifier 4221 may use a data tagging service to identify a set of one or more tags that should be associated with the input data. The tag identifier 4221 may then provide the tags associated with the input data (and in some implementations, also with the user content) to the recommended content item identifier and ranking subsystem 4224 (which may also be referred to as the content item ranking unit 4224 for brevity) for further processing.
[0174] In some embodiments, the tag identifier 4221 may determine one or more tags to be associated with the received user content, as described above, using the various techniques described above employed by the content tagging unit 4222. In some embodiments, both the tag identifier 4221 and the content tagging unit 4222 may use the same superset of tags from which the tags to be associated with user input and content items are determined. In certain embodiments, the tag identifier 4221 and the content tagging unit 4222 may use the same data tagging service to identify the tags to be associated with user content and content items, respectively. In yet another embodiment, the subsystems of the tag identifier 4221 and the content tagging unit 4222 may be implemented as a single subsystem configured to perform similar (or identical) processing on content items received from the repository 4230 and input content received from the client system 4210.
[0175] As described above, when a content item is tagged by the content tagging unit 4222, for each content item, one or more tags to be associated with the content item are identified, along with a tag value for each tag. With regard to tagging for user content, in some embodiments, the tag identifier 4221 is configured to determine only the tags to be associated with user input, without associated tag values. In such embodiments, each tag is given equal weight for ranking performed by the content item ranking unit 4224 based on the tags associated with user input. In some other embodiments, both the tags and associated tag values may be determined for user content and may be used by the content item ranking unit 4224 to rank content item recommendations.
[0176] As described above, user content received and processed by tag identifier 4221 can take various forms. For example, user content may include the content of documents authored by the user (e.g., emails, articles, blog posts, documents, social media posts, images, etc.), or content created or selected by the user (e.g., multimedia files). Another example is that user input may be a document accessed by the user (e.g., a web page). Yet another example is that user content may be search terms entered by the user to perform a search (e.g., a browser-based search engine). In certain embodiments, for example, for search terms, these terms themselves may be used as tags.
[0177] As illustrated in Figure 42 and described above, the content item ranking unit 4224 receives as input information from the tag identifier 4221 that identifies one or more sets of tags associated with user content. Based on this tag information about the user content and on the content items available for recommendation, the content item ranking unit 4224 is configured to identify one or more content items that are most relevant and / or related to the input content using tag matching techniques. If multiple content items are identified as relevant or related to the user input, the content item ranking unit 4224 is further configured to rank the content items using the innovative ranking techniques described herein. Further details relating to the various techniques used by the content item ranking unit 4224 to score and rank content items are described in more detail below. The content item ranking unit 4224 is configured to generate a ranked list of content items that should be recommended to a user in response to user input received on behalf of the user. The ranked list of content items is then provided to the recommendation selector subsystem 4225 for further processing.
[0178] The recommendation selector 4225 is configured to select one or more specific content items to be recommended to the user in response to input content received from the client system 4210, using a ranked list of content items received from the content item ranking unit 4224. In certain scenarios, all content items in the ranked list may be selected for recommendation. In several other scenarios, a subset of ranked content items may be selected for recommendation, in which case the subset contains fewer content items than all content items in the ranked list, and one or more content items included in the subset are selected based on the ranking of the content items in the ranked list. For example, the recommendation selector 4225 may select up to X ranked content items (e.g., the top 5, top 10, etc.) from the ranked list for recommendation, where X is some integer less than or equal to the number of ranked items. In certain embodiments, the recommendation selector 4225 may select content items to be included in a subset to be recommended to the user based on the score associated with the content items in the ranked list. For example, only content items with scores exceeding a user-configurable threshold score may be selected to be recommended to the user.
[0179] Next, information identifying the content items selected for recommendation by the recommendation selector 4225 may be communicated from the content recommendation system 4220 to the user's user client device 4210. Then, information about the recommended content items may be output to the user via the user client device. For example, recommendation information may be output via a GUI 4215 displayed on the user client device or via an application 4215 executed by the user client device. For example, if user input corresponds to a search query entered by the user via a web page displayed by a browser executed by the user device, recommendation information may be output to the user via that web page or via additional web pages displayed by the browser. In certain embodiments, for each recommended content item, the information output to the user may include information identifying the content item (e.g., text information, image thumbnail, etc.) and information for accessing the content item. For example, information for accessing the content item may be in the form of a link (e.g., a URL) which, when selected by the user (e.g., by a mouse click), accesses the corresponding content item. This is displayed to the user via a user client device. In some embodiments, information identifying a content item and information for accessing the content item may be combined (for example, a recommended image thumbnail representation that can be selected by the user to identify an image content item and access the image itself).
[0180] In various embodiments, the content recommendation system 4220 includes its associated hardware / software components 4221-4225 and services, and may be implemented as a backend service located away from the frontend client device 4210. Interaction between the client device 4210 and the content recommendation system 4220 may be an internet-based web browsing session or a client-server application session, during which the user may input user content (e.g., search terms, original authored content, etc.) via the client device 4210 and receive content item recommendations from the content recommendation system 4220. Additionally or alternatively, the content recommendation system 4220 and / or the content repository 4230 and related services may be implemented as dedicated software components running directly on the client device.
[0181] In some embodiments, the system 4200 shown in Figure 42 may be implemented as a cloud-based tier-layer system, where upper-tier user devices 4210 can request and receive access to network-based resources and services via a content recommendation system 4220 residing on a backend application server deployed and running on an underlying set of resources (e.g., cloud-based, SaaS, IaaS, PaaS, etc.). Some or all of the functions of the content recommendation system 4220 described herein may be performed by or accessed through web services including representational state transfer (REST) services and / or simple object access protocol (SOAP) web services or APIs, and / or web content exposed via hypertext transfer protocol (HTTP) or HTTP secure protocol. Therefore, although not shown in Figure 42 to avoid obscuring components that will be shown with additional details, the computing environment 4200 may include additional client devices, one or more computer networks, one or more firewalls, proxy servers, routers, gateways, load balancers, and / or other intermediate network devices to facilitate interaction between the client device 4210 and the content recommendation system 4220 and the content repository 4230.
[0182] In various implementations, the system shown in computing environment 4200 may be implemented using one or more computing systems and / or networks, including dedicated server computers (desktop servers, UNIX servers, midrange servers, mainframe computers, rack-mount servers, etc.), server farms, server clusters, distributed servers, or any other appropriate configuration and / or combination of computing hardware. For example, content recommendation system 4220 may run various additional server applications and / or middle-tier applications, including an operating system and / or Hypertext Transport Protocol (HTTP) server, File Transport Service (FTP) server, Common Gateway Interface (CGI) server, Java® server, database server, and other computing systems. Any or all of the components or subsystems within content recommendation system 4220 may have at least one memory, one or more processing units Subsystems and / or modules within the Content Recommendation System 4220 may include (for example, a processor) and / or storage. Subsystems and / or modules within the Content Recommendation System 4220 may be implemented in hardware, software running on the hardware (for example, program code or instructions executable by a processor), or a combination thereof. In some examples, the software may be stored in memory (for example, non-temporary computer-readable media), memory devices, or any other physical memory, and may be executed by one or more processing units (for example, one or more processors, one or more processor cores, one or more graphics process units (GPUs), etc.). Examples of computer-executable instruction or firmware implementations described herein may include computer-executable instructions or machine-executable instructions written in any suitable programming language capable of performing the various operations, functions, methods, and / or processes described herein. Memory may store program instructions that can be loaded and executed on the processing unit, and data generated during the execution of these programs. Memory may be volatile (e.g., random-access memory (RAM)) and / or non-volatile (e.g., read-only memory (ROM), flash memory, etc.). Memory may be implemented using any type of persistent storage device, such as computer-readable storage media. In some examples, computer-readable storage media may be configured to protect the computer from electronic communications containing malicious code.
[0183] Figure 43 shows a simplified flowchart 4300 illustrating the processes performed by a content recommendation system for identifying and ranking content items related to user content, according to a particular embodiment. The processes shown in Figure 43 may be implemented by software (e.g., code, instructions, programs), hardware, or a combination thereof, executed by one or more processing units (e.g., processors, cores) of each system. The software may be stored on a non-temporary storage medium (e.g., on a memory device). The methods presented below in Figure 43 are illustrative and non-limiting. Figure 43 shows various processing steps performed in a particular sequence or order, but this is not intended to be limiting. In some alternative embodiments, the processes may be performed in some different order, or some steps may be performed in parallel. The processes shown in Figure 43 may be performed by one or more systems shown in Figure 42, such as a content recommendation system 4220. For example, in the embodiment shown in Figure 42, the tag identifier 4221 may perform processing in 4302 and 4304, the content item ranking unit 4224 may perform processing in 4306 to 4316, and the recommendation selector 4225 may perform processing in 4318 and 4320. However, it should be understood that the techniques and functions described in relation to Figure 43 are not necessarily limited to the implementation example within the specific computing infrastructure shown in Figure 42, but can be implemented using other compatible computing infrastructures described herein.
[0184] In 4302, the content recommendation system 4220 may receive input content from one or more users or client systems 4210. As described above with reference to Figure 42, the input content may be received from the client device 4210 via a graphical user interface 4215 (e.g., a web-based GUI) provided by the content recommendation system 4220. In another example, the input content may be received by a web server or backend service running within the content recommendation system 4220 based on data transmitted by a frontend application (e.g., a mobile application) installed on the client device 4210.
[0185] In some embodiments, the input content received in step 4302 is used by the user This can correspond to a set of search terms or phrases entered into the search engine user interface. In other embodiments, the input content can correspond to original content authored by a user and entered into a dedicated user interface. For example, new original content may include online articles, press releases, emails, blog posts, etc., and such content may be entered by the user via a software-based word processing tool, an email client application, a web development tool, etc. In yet another example, the input content received in step 4302 may be images, graphics, voice input, or any other text and / or multimedia content generated or selected by the user via the client device 4310.
[0186] Briefly referring to Figure 44, an exemplary user interface 4400 is shown, which includes a user interface screen 4410 that allows the user to input original authored content. In this example, the user interface screen 4410 is collectively referred to as the “Content Authoring User Interface,” but in various embodiments, the user interface 4410 may correspond to an interface for a word processor, article designer or blog post writer, email client application, etc. In this example, the user interface 4410 includes a first text box 4411 in which the user can input the title or subject of the authored content, and a second text box 4412 in which the user can input all the text about the input content (e.g., article, email body, document, etc.). In addition, the user interface 4410 includes selectable buttons 4413 that allow the user to start searching for related content items (e.g., images, related articles, etc.) that can be incorporated into the newly authored content. In some cases, selecting button 4413 or a similar user interface component may initiate the process shown in Figure 43 by first analyzing user content received via user interface 4410 (e.g., user content entered in 4411 and / or 4412) and sending it to the content recommendation system 4220. In other embodiments, content item recommendations can be updated in real time, continuously or periodically, by running a background process continuously (or periodically) within the front-end user interface to continuously (or periodically) analyze new text input received from the user (e.g., content entered by the user in 4411 and / or 4412), and restarting the process in Figure 43 in response to text updates.
[0187] Referring again to Figure 43, in 4304, the content recommendation system 4220 determines one or more tags for the input content received in step 4302. In some embodiments, the input content received in 4302 may already have tags associated with it or may have tags embedded in it, and the tag identifier 4221 may identify and extract a predetermined set of tags associated with the input content. If the input content received in step 4302 corresponds to a search term, the tag identifier 4221 may simply use the search term entered by the user as the tag (excluding certain words such as articles, conjunctions, prepositions, modifiers, etc.). In the case of pre-authored text content, or any other untagned input content received by the content recommendation system 4220, the tag identifier 4221 may be configured to analyze various characteristics of the input content received in 4302 and, based on this analysis, determine one or more tags that should be associated with the received input content. As mentioned above, the tag identifier 4221 can use a variety of technologies to determine one or more tags that should be associated with the input content received in 4302.
[0188] Referring again to the exemplary user interface 4400 shown in Figure 44, in this example, the input content received in 4302 is the text entered by the user in the subject / topic box 4411, namely, "Is Coffee Healthier For You Than This may correspond to the text entered in box 4412, "Tea? (Is coffee healthier to you than tea?)". Based on the subject content analysis in 4411 and the article text authored in 4411, along with any additional input content that may be provided, the content recommendation system 4220 may determine in 4304 that the tags "coffee", "tea", and "human" should be associated with the user content.
[0189] In 4306, based on the tags determined for the input content in 4302, the content recommendation system 4220 uses tag matching techniques to identify a set of matching content items from a collection of content items available for recommendation. In this case, a content item is considered a match and identified in 4306 if at least one tag associated with the content item matches a tag determined for the input content in 4304. The set of content items identified in 4306 may also be referred to as a matching set of content items and includes content items that are candidates recommended to the user in response to the input content received in 4302. In various embodiments, data processing techniques such as stemming and synonym lookup / comparison may be used as part of the matching process in 4306 to identify matching tags.
[0190] For example, in the embodiment shown in Figure 42, a collection of content items available for recommendation, along with their associated tag information (e.g., tags associated with the content items and associated tag values), may be stored in the content / tag information data store 4223. As part of the processing in 4306, the content recommendation system 4220 may compare one or more tags determined in 4304 with tags associated with content items available for recommendation, and identify those content items from a collection having at least one associated tag that matches the tags identified in 4304.
[0191] Continuing with the example in Figure 44, assuming that the user input content received in 4302 was identified in 4304 as having the tags "coffee," "person," and "tea," the exemplary Table 4500 in Figure 45 shows the set of matching content items identified by the content recommendation system 4220 (for example, by the content item ranking unit 4224) from a collection of content items available for recommendation. As can be seen from Table 4500, eight different image content items were identified as having at least one associated tag that matches at least one of the tags "coffee," "person," or "tea." As shown in this example, each matched image content item is identified using an image identifier 4501. The image description 4502 provided in Figure 45 is a description of the content of the matched image and is provided in Figure 45 so that it is not necessary to show the actual image. Each matched image has one or more associated tags 4503, and a tag value 4504 is associated with each tag. In the example in Table 4500, the tag values are floating-point values within a predetermined range (e.g., 0.0 to 1.0), and the sum of the tag values for each content item is the same total (e.g., 1.0). Such embodiments offer additional technical advantages, for example, ensuring that content items with more associated tags are not intentionally ranked higher or over-recommended based on the number of associated tags, by having a uniform sum of tag count score and TVBS. Furthermore, as described below... Thus, having tag values in the range of 0.0 to 1.0 ensures that when they are multiplied (for example, in the use cases described below), the resulting value allows content items to be ranked within a particular group or bucket, while ensuring that the highest-ranked content item within one group does not rank higher than the lowest-ranked content item in the next highest group.
[0192] Figure 45 shows that a content item (in this example, an image) is considered a match if at least one tag associated with it matches a tag associated with the input content. A matched content item may have tags associated with it that match one or more of the tags associated with the input content. A matched content item may also have other tags associated with it that are different from the tags for the input content (e.g., Image_1, Image_3, etc.).
[0193] In 4308, a tag count score (first score) is calculated for each matching content item identified in 4306, based on the number of tags associated with the content item that matches the tags determined for the input content. In some scenarios, each matching tag of a content item is given a value of 1, so the tag count score calculated in 4308 for a given content item is equal to the number of tags of the content item that match the tags associated with the input content. For example, in the case of the matching images identified in Figure 45, Table 4600 in Figure 46 identifies the tag count score 4602 for each matching image identified in Figure 45 (and also in Figure 46). For example, the tag count score for Image_1 is calculated based on the single tag associated with that image ("person"). The tag count score for Image_2 is "1" because the tag ("Co") matches the tag associated with the input content. As another example, the tag count score for Image_2 is "1" because the two tags associated with the image ("Co") The tag count score for Image_8 is "2" because the tags "hee" and "person" matched the tags associated with the input content. The one tag that was selected ("coffee") matched the tag associated with the input content, so the result is "1".
[0194] At 4310, matching content items are grouped (bucketed) into groups or buckets based on the tag count scores calculated for the content items at 4308. In certain embodiments, a group (or bucket) includes all content items having the same tag count score. In a scenario where each matching tag of a content item is given a value of 1, the content items included in each group or each bucket have the same number of matching tags. The processing at 4310 may be optional and may not be performed in certain embodiments.
[0195] Continuing to refer to the example shown in FIG. 46, the eight matching images can be grouped into two groups or buckets, including a first group or bucket that includes content items with a tag count score of 1 and a second group or bucket that includes content items with a tag count score of 2. The first group would include the images {Image_1, Image_3, Image_6, Image_8}. The second group would include the images {Image_2, Image_4, Image_5, Image_7}. Note that in this example, none of the matching images match all three tags ("coffee", "person", "tea") associated with the input content.
[0196] At 4312, for each group identified at 4310, a tag-value-based score (the second score) is calculated for each candidate content item within that group. In some embodiments, the tag-value-based score (TVBS) for a particular content item is based on and calculated using the tag values associated with the matching tags for that content item. Based on and calculated using the tag values associated with the matching tags for that content item.
[0197] In certain embodiments, the TVBS for a content item is calculated by multiplying the tag values associated with the tags of the content item that match the tags associated with the input content. For example, as follows.
[0198] For Image_1 in FIG. 46, TVBS = 0.93, For Image_2 in FIG. 46, TVBS = 0.5 * 0.5 = 0.25, For Image_3 in FIG. 46, TVBS = 0.65, For Image_4 in FIG. 46, TVBS = 0.35 * 0.60 = 0.21, and so on.
[0199] Table 4600 shows the TVBS4603 calculated for various matching images using the above technology. In a specific embodiment, the naive Bayes method is used to calculate the tag - value - based score for content items. For example, assuming that two tags tag1 and tag2 are determined for the input content, the tag - value - based score (TVBS) for a certain image content item can be expressed as follows.
[0200] Image i TVBS for = P(Image i |tag 1, tag2) = the probability of Image i when tags are tag1 and tag2 Expanding this for "n" tags gives the following.
[0201] P(Image i |tag 1, tag 2 …, tag n ) = the probability of Image n when tags are tag1 and tag2 and... tag i Assume that tags tag1, tag2,..., tag n are independent of each other.
[0202] TVBS for Image i = TVBS i = P(Image i |tag1, tag2,..., tagn ) =P(Image i |tag1)*P(Image i |tag2)*…*P(Image n |tag n ) According to simple Naive Bayes, the following holds true:
[0203]
number
[0204] in this case, P(Image i |tag t )=Image i tag about t probability P(Image i ) = All images are considered unique, and this term can be discarded or ignored. P(tag t ) = the frequency of tags in a collection of content items (i.e., tag t (The number of content items (e.g., images) in a collection of content items that are tagged and available for recommendation.) Because this term is in the denominator, the less frequently a tag exists in a collection of content items (i.e., the fewer content items that have this associated tag), the higher the TVBS score for images with that tag.
[0205] The above formula can be extended to the following:
[0206]
number
[0207] In equation 3 above, the numerator is the product of the probabilities that would result in a higher score given the same probability (if the same tag matches multiple images, the denominator will remain the same).
[0208] The following example shows an application example of Equation 3 for calculating TVBS for content items. Assume that the tags determined for the input content are "person" and "coffee". Further, assume that a collection of content items (images) available for recommendation includes three images with the following tags and tag values.
[0209] Image A: ("person", 0.5), ("coffee", 0.5) Image B: ("person", 0.1), ("coffee", 0.9) Image C: ("person", 0.8), ("coffee", 0.2) For Image A: P(human|image) = 0.5 and P(coffee|image) = 0.5 Applying Equation 3 gives the following.
[0210]
Number
[0211] In this case, "Frequency(human)" is the number of content items available for recommendation in the collection of content items and associated with the "person" tag. And "Frequency(coffee)" is the number of content items available for recommendation in the collection of content items and associated with the "coffee" tag.
[0212]
Number
[0213] TVBS for Image A = 0.028 Using a similar technique, the following is obtained.
[0214] For image B, TVBS = (0.1 / 3) * (0.9 / 3) = 0.01 TVBS for image C = (0.8 / 3) * (0.2 / 3) = 0.018 As this example shows, if the frequencies are the same, TVBS will be higher if the probabilities are similar (the denominator will remain the same if the same tag matches multiple images).
[0215] According to an extended example of Equation 3, the frequency of content items with a specific associated tag within a collection of content items is taken into consideration when calculating TVBS (and thus also affects the overall ranking score, as described below). The fewer the number of content items with a specific associated tag (i.e., the lower the frequency), the higher the TVBS score for images with that specific tag will be. In some embodiments, this is desirable because such content items with a "rare" or lower frequency of the tag are more likely to rank higher than content items with a higher frequency of the higher-ranked tag, thereby increasing the likelihood of these content items being included in the list of content items recommended to the user. Therefore, the TVBS value for a given content item is inversely proportional to the frequency of occurrence of a particular tag in a collection of content items (i.e., the number of content items in that collection that have the associated specific tag).
[0216] In 4314, the total ranking score is calculated for each matching content item identified in 4306, based on the tag count score calculated for the content item in 4308 and the TVBS calculated for the content item in 4312. In some embodiments, the total ranking score for a candidate content item may be calculated as the sum of the tag count score (calculated in 4308) and the TVBS (calculated in 4312) calculated for the content item. That is, for an image content item, iIn this case, it is as follows:
[0217] Ranking score (Image i )= TagsCountScore i + TVBS i In the example shown in Figure 46, column 4604 shows the total ranking score calculated for each matching image by adding the tag count score for that image (shown in column 4602) and the TVBS for that image (shown in column 4603). For example, the calculated total ranking score is as follows:
[0218] Image_1: 1 + 0.93 = 1.93 Image_2: 2 + 0.25 = 2.25 Image_3: 1 + 0.65 = 1.65 And so on.
[0219] In step 4316, the content recommendation system 4220 (for example, the content item ranking unit 4224 in the content recommendation system 4220) generates a ranking list of matching content items based on the total ranking score calculated for the content items in 4314. Therefore, referring to the examples in Figures 44-46, based on the total ranking score calculated for the matching images (in column 4604), the images are ranked as follows: (1) Image_2, (2) Image_5, (3) Image_4, (4)Image_7, (5)Image_6, (6)Image_1, (7)Image_8, (8)Image_3, Sea urchin can be ranked from the highest to the lowest rank.
[0220] In certain embodiments where the tag value for a tag is in the range of 0.0 to 1.0, the calculation of the total ranking score using the (TagCountScore+TVBS) approach is performed on the input content. Ensure that content items with more associated tags that match the tags associated with them are ranked higher than content items with fewer matching tags. For example, in the example in Figure 46, (the tags associated with the input content An image with a tag count score of 2 (corresponding to the two tags associated with the image that matched) will always have a total rank score and will therefore be ranked higher than an image with a tag count score of 1 (corresponding to the two tags associated with the image that matched) (corresponding to the tags associated with the input content). This is because, assuming that tag values are in the range of 0 to 1, the TVBS for an image calculated by multiplying the tag values associated with the matching tags cannot exceed 1. This also suggests that, with respect to a first group or bucket of content items corresponding to a first tag count score and a second group or bucket of content items corresponding to a second tag count score, if the first tag count score is higher than the second tag count score, each content item in the first group will be ranked higher than the content items in the second group (due to a higher total rank score). Therefore, in an example where three tags ("coffee," "person," and "tea") are determined for the input content, a content item with three content tag matches for the tags of the input content will always be ranked higher than a content item with two matching content tags, and each of those two matching content tag items will always be ranked higher than a content item with one matching content tag, and so on. Within each group or bucket, content items may also be ranked based on their TVBS, which is advantageous for both higher and more equal parameters for matching content tags. However, it should be understood that in other embodiments, various formulas or logic may be used to calculate the tag count score, TVBS, and total ranking score in order to realize different content item ranking priorities and policies.
[0221] In 4318, the content recommendation system 4220 (for example, recommendation selector 4225) may use the ranked list generated in 4316 to select one or more content items to be recommended to the user. In certain scenarios, all content items in the ranked list may be selected for recommendation. In some other scenarios, a subset of ranked content items may be selected for recommendation, in which case the subset does not include all content items in the ranked list, but one or more content items included in the subset are selected based on the ranking of the content items in the ranked list. For example, recommendation selector 4225 may select content items for recommendation that are ranked Xth (for example, the top 5, top 10, etc.) in the ranked list, where X is some integer less than or equal to the number of ranked items in the list. In certain embodiments, recommendation selector 4225 may select content items to be included in a subset to be recommended to the user based on the total ranking score associated with the content items in the ranked list. For example, only content items with associated scores above a user-configurable threshold score may be selected to be recommended to the user.
[0222] In 4320, the content recommendation system 4220 may communicate information about the content items selected in 4318 to the user device. This information may be called recommendation information as it includes information about content items that should be recommended to the user. The recommendation information communicated in 4320 may also include ranking information (e.g., the total ranking score associated with the selected content items). This information may be used on the user device to determine how information about the recommended content items (e.g., orders) will be displayed to the user via the user device. In some embodiments, the recommendation selector 4225 may include the content item itself or specific information that identifies the content item (e.g., content) as part of the recommendation information. The client device 4310 may receive any of the following: an item identifier and description, a thumbnail image, a network path or link for download, etc. In step 4302, input content is received from this client device 4310.
[0223] Next, information regarding the selected recommendation may be output to the user via the user device. For example, the recommendation information may be output via a GUI 4215 displayed on the user client device, or via an application 4215 executed by the user client device. For example, if user input corresponds to a search query entered by the user via a web page displayed by a browser executed by the user device, the recommendation information may be output to the user via a web page showing the search results or an additional web page displayed by the browser. In certain embodiments, for each recommended content item, the information output to the user may include information identifying the content item (e.g., text information, an image thumbnail, etc.) and information for accessing the content item. For example, the information for accessing the content item may be in the form of a link (e.g., a URL), and when the link is selected by the user (e.g., by a mouse click), the corresponding content item is accessed and displayed to the user via the user client device. In some embodiments, the information identifying the content item and the information for accessing the content item may be combined (e.g., a thumbnail representation of a recommended image that can be selected by the user to identify an image content item and access the image itself).
[0224] For example, referring to Figure 47, an exemplary user interface 4700 is shown, corresponding to an update of the user interface screen 4400 in Figure 44, which displays information related to recommended images. In this example, based on the title / subject 4711, body text 4712, and / or other optional input content, the content recommendation system 4220 selects four content item images from a ranking list that are ranked highly to recommend to the user. Information related to these top four ranked images is displayed in rank within a dedicated section 4714 of the user interface 4700 to indicate content item recommendations. In certain embodiments, thumbnail representations of the recommended images may be displayed in 4714. The user interface 4700 supports drag-and-drop functionality or other techniques to allow the user to incorporate one or more of the suggested images displayed in 4714 into the body text 4712 of authored content.
[0225] The processes described above shown in Figure 43 are not intended to be limiting. Various modifications may be provided in different embodiments. For example, in the embodiment described above shown in Figure 43, the processes in 4312, 4314, and 4316 are performed on all matching content items identified in 4306. In certain modifications, the tag count score calculated in 4308 may be used to filter out specific content items from subsequent processing. For example, if content items have different tag count scores, the content item with the lowest tag count score (or some other threshold) may be filtered out from subsequent processing in the flowchart. In some other embodiments, only the content item with the highest tag count score may be selected for subsequent processing, and other content items may be filtered out from subsequent processing. For example, the content item ranking unit 4224 may calculate TVBS only for the highest tag score group (as determined in step 4304), or it may calculate TVBS sequentially from the highest tag count score group to the lowest tag count score group, during which the calculation process reaches a threshold number or threshold tag count score for candidate content items. This could potentially cause the process to stop. Such filtering reduces the number of content items that need to be processed and allows the overall recommended behavior to run faster and more efficiently using fewer processing resources (e.g., processor, memory, and network resources).
[0226] In the method described above, shown in Figure 43, each tag associated with or determined for input content in 4304 is given equal weight for ranking performed by the content recommendation system 4220. Based on this assumption, the tag score for each matching content item was determined as the number of content tags that match the tags of the input content. Thus, each matching content tag associated with a given matching content item was given equal value / weight for the determination of the tag count score. However, in other embodiments, tags associated with input content may be given different weights. For example, if two tags are determined for input content, one tag may be given a higher weight than the other to indicate that it is "more important" for the input content. For example, in the above example shown in Figures 44-47, the tags ("person", "coffee", "tea") are determined for input content, and instead of giving these three tags equal importance, they are weighted as follows: person=1, coffee=2, and tea=4. This weighting may indicate the relative importance of the tags to the input content. For example, "tea" is weighted more heavily than "coffee," and "coffee" is weighted more heavily than "person." In certain embodiments, the logic used to calculate the tag count score for each content item may be modified to take into account the different weights assigned to the tags with respect to the input content. Following such a modified logic, each contribution of the matching tags in the content item is multiplied by the weight associated with the same tag with respect to the input content. For example, using the weighting (person=1, coffee=2, and tea=4) for the input content, the tag count score for the matching image in Figure 45 would be as follows:
[0227] Image_1: TagCountScore (unweighted) = "person" tag match = 1 TagCountScore (weighted) = "person" tag match = 1 * 1 = 1 Image_2: TagCountScore (unweighted) = "Coffee" and "People" tags match = 1 + 1 = 2 TagCountScore (weighted) = Matches with "coffee" and "person" tags = 2(1)+1 (1) = 3 Image_3: TagCountScore (unweighted) = "tea" tag match = 1 TagCountScore (weighted) = "tea" tag match = 4 (1) = 4 Image_4: TagCountScore (unweighted) = "Coffee" and "People" tags match = 1 + 1 = 2 TagCountScore (weighted) = Matches with "coffee" and "person" tags = 2(1)+1 (1) = 3 Image_5: TagCountScore (unweighted) = Matches with "Tea" and "Person" tags = 1 + 1 = 2 TagCountScore (weighted) = Matches with "Tea" and "Person" tags = 4(1) + 1(1) =5 Image_6: TagCountScore (unweighted) = "Coffee" tag match = 1 TagCountScore (weighted) = "Coffee" tag match = 2 (1) = 2 Image_7: TagCountScore (unweighted) = Matches with "Tea" and "Person" tags = 1 + 1 = 2 TagCountScore (weighted) = Matches with "Tea" and "Person" tags = 4(1) + 1(1) =5 Image_8: TagCountScore (unweighted) = "Coffee" tag match = 1 TagCountScore (weighted) = "Coffee" tag match = 2 (1) = 2 As a result of different tag count scores, the grouping or bucketing of content items performed in 4310 will vary considerably. For example, content items with content tags that match both the "tea" and "person" tags will be grouped together and assigned a tag score of 5, content items with only one content tag matching "tea" will be grouped together and assigned a tag score of 4, and content items with content tags that match both the "coffee" and "person" tags will be grouped together and assigned a tag score of 3, and so on. Thus, in such embodiments, the overall ranking of candidate content items is influenced not only by how many tags of a content item match the tags of the input content, but also by which particular tags of the content item match the input content tags and the relative importance weights assigned to those tags. In this example, images Image_5 and Image_7 would be the highest-ranked overall content items based on the highest tag count score of 5 (tea=4 + person=1).
[0228] Smart Content - Smart Categorization / Classification Classifying large amounts of content online is a complex task with challenges such as the constraint of a single pass to the data and the requirement for fast response times. According to one embodiment, content users categorize similar content through logical clusters, such as a hierarchical taxonomy tree, and place similar content in the same node / category of the taxonomy tree. Over time, as the number of content entities and nodes in the taxonomy tree increases, similar content entities will be found to exist side by side within the nodes. Given this state of content organization, content already present in evaluated / categorized taxonomies can be used by computer algorithms, such as categorization engines, to determine where newly created / edited content may belong.
[0229] According to one embodiment, the systems and methods described herein can be used, for example, in conjunction with a content management system to provide recommendations for categorizing / classifying content into user-defined categories, thereby providing content managers with the opportunity to effortlessly place new content into the correct categories with less effort, based on pre-evaluated / categorized content.
[0230] Recommendation systems or tools can use artificial intelligence (AI) technology to continuously learn from historical data and / or from newly input results from generated recommendations, and help place content into relevant categories through automatic categorization / classification of newly created / edited content.
[0231] The recommendation tool can be implemented and applied across various domains by generating feature vectors from content, creating clusters in the feature space based on pre-categorized content, and recommending categories for new content by calculating feature space distances from the clusters.
[0232] Figure 48 shows an exemplary use of a content management system environment according to one embodiment. .
[0233] More specifically, Figure 48 shows an exemplary content management system that may include a categorization engine for smart content categorization within the content management system.
[0234] As shown in Figure 48, according to one embodiment, for each of the multiple client devices 4800, 4802, and 4804 having user interfaces 4801, 4803, 4805 and physical device hardware 4806, 4807, 4808 (e.g., CPU, memory), a content access application 4810, 4811, 4812 to be executed thereon can be provided to the client device.
[0235] According to one embodiment, a client device can communicate with an application server 4830 which includes physical computer hardware 4831 (e.g., CPU, memory) and a content management system 4832 (4862).
[0236] According to one embodiment, a content access application on a client device can communicate with a content management system via a network 4860 (e.g., the internet or a cloud environment). The content access application can be configured to allow users 4850, 4852, and 4854 to view, upload, modify, or delete content such as content items 4820, 4822, and 4824 on each client device, or to access such content. For example, a user can add or upload new content to the content management system by interacting with the content access application on the relevant client device. The content can be sent to the content management system for tagging and storage, for example.
[0237] According to one embodiment, a content management system may be or include a platform for organizing and consolidating content that can be managed by several users or clients. According to one embodiment, a content management system may be configured to communicate with a content repository 4836 for storing content (or content items) 4840 and to deliver content to users via those client devices. According to one embodiment, the content repository may be a relational database management system (RDBMS), a file system, or a content management system. This could be any other data store that the system can access. The content could include, for example, documents, files, emails, notes, images, videos, slide presentations, conversations, and user profiles.
[0238] In one embodiment, a content management system can be configured to associate metadata with content. Metadata may include information about the content item, such as its title, author, publication date, historical data such as who accessed the item and when, and the location where the content is stored.
[0239] According to one embodiment, metadata can be stored in the metadata database 4838. According to one embodiment, the content management system can be configured to communicate with the metadata database to access the metadata stored therein and to store metadata generated by the system in the metadata database.
[0240] According to one embodiment, the content management system also uses search index 4839 It can also be configured to communicate with a search index. The search index may be configured to provide indexing and searching of content and data stored in a content repository and metadata database. According to one embodiment, the search index may be a relational database management system (RDBMS) or a search tool, such as Oracle Secure Enterprise Search (Oracle SES).
[0241] According to one embodiment, the content management system may further include a content management application 4833 and a categorization engine 4834. The categorization engine may include an artificial intelligence / machine learning engine and library 4835 and a user interface 4836 that can be used, for example, to display output indicating recommendations for content categorization and / or to receive input indicating selections for content categorization.
[0242] In one embodiment, a categorization engine can be used, for example, in conjunction with a content management system, to provide recommendations for categorizing / classifying content into defined (e.g., user-defined) categories, thereby providing content managers with the opportunity to effortlessly place new content into the correct categories based on pre-evaluated / categorized content. The categorization engine can use artificial intelligence (AI) technology and, for example, machine learning libraries to continuously learn from historical data and assist in placing content into relevant categories through automatic categorization / classification of newly created / edited content. The recommendation tool can be implemented and applied across various domains by generating feature vectors from content, creating clusters in the feature space based on pre-categorized content, and recommending categories for new content by calculating feature space distances from the clusters. The engine can, additionally or alternatively, automatically categorize content by its calculations.
[0243] As outlined, according to one embodiment, a categorization engine can be used to create new taxonomies, modify existing taxonomy structures, and / or categorize / classify existing and / or new content in bulk according to various recommendations or suggestions associated with various confidence scores or confidence levels (such as high confidence, medium confidence, or low confidence).
[0244] For example, according to one embodiment, low-confidence recommendations for categories / classifications associated with a particular set of content can be ignored by the system or the user, while high / medium-confidence recommendations for categories are acceptable.
[0245] According to one embodiment, the categorization engine may display such recommendations or suggestions in the user interface for categories / classifications associated with a particular set of content for content administrators to review and / or perform actions.
[0246] For example, while working with a new set of content, the system may suggest or recommend one or more categories / classifications for the new content based on the previous classification of various content, and the content administrator can then choose to accept or reject the assignment of those categories / classifications to the content.
[0247] According to one embodiment, such acceptance or rejection of categorization suggestions by users or content administrators can be stored in a database, such as the database of the content categorization engine. By utilizing such historical records, future recommendations for content categorization can be improved.
[0248] According to one embodiment, classification can be performed as a two-step process, comprising a macro-level classification step that can match the topic distribution and named entity distribution of documents with higher-level cluster nodes, and a micro-level classification step that can expand and compare categories within the feature space based on evaluations of microclusters.
[0249] Figure 49 shows an exemplary use of a content management system for managing and distributing content data, according to one embodiment.
[0250] According to one embodiment, a content management system (CM) is used. S) enables users to collaboratively manage digital content, including web content. An example of a content management system is Oracle® Content Management (OCM), which is provided by and can be accessed through a cloud service. A content management system provides various features, such as storing, managing, and publishing relevant digital content. Systems like OCM, for example, enable the rapid deployment and publication of content across multiple distribution channels.
[0251] In one embodiment, the distribution channel may be in any form as long as the content is delivered to the consumers of that content. Exemplary distribution channels include websites, blogs, HTML emails, storyboards, and mobile applications. To quickly deploy documents used by the relevant distribution channels, the system provides, for example, out-of-the-box templates, drag-and-drop components, sample page layouts, and site themes. These features, among others, enable users to assemble content from predefined building blocks into publishable documents. The content management system can use these components to generate the document's markup language and code (collectively referred to herein as "code").
[0252] According to one embodiment, a server application program interface (API) contract (management API contract) can be provided to establish interaction between a content management system and a commercial provider system (commercial provider). A server API contract functions as a contract between a content management system and a commercial provider (e.g., an online retailer), enabling communication and data exchange between the content management system and the commercial provider. According to one embodiment, based on interaction with interactive features, it may be possible to retrieve product-related data and other metrics from the commercial provider, or to send data associated with a request to the commercial provider to perform an action associated with that product.
[0253] According to one embodiment, a commercial provider system 4910 can be provided within the system to manage and distribute content data. The commercial provider system may include a product catalog 4911, APIs such as a server API 4912, and physical computing resources such as a CPU and memory 4913.
[0254] According to one embodiment, the content management system 4920 may include a content metadata database 4921 (e.g., OCM content and metadata database), APIs such as a server API 4922, and physical computer resources such as a CPU and memory 4923.
[0255] According to one embodiment, the content administrator 4935 interacts with the management system 4930 in such a system to create / manage content data 4903. It can be called and distributed. The management system may have or may have a user interface 4931 which may have a content mapping configuration 4932. The management system may also include physical computer resources such as a CPU and memory 4933.
[0256] In one embodiment, a Server API Agreement (Management API Agreement) 49049 can function as an agreement between a content management system and a commercial provider. For example, a Server API Agreement can be configured to receive account information and credentials associated with user accounts at the commercial provider.
[0257] In one embodiment, a content management system can use received account credentials to authenticate an administrator user (e.g., a content administrator) of the content management system in a commercial provider system via a server API (management API), thereby enabling the content management system to access data stored in the commercial provider system. For example, the content management system may request via the server API that the commercial provider send data describing / defining a set of products offered for sale by the commercial provider. The commercial provider system can then send such data to the content management system via the server API. This data may include, for example, a list of products, along with data describing a list of products.
[0258] According to one embodiment, the content management system can publish the content 4904 through several different channels, for example, via a generated mobile application page 4905 or via a generated web page 4906.
[0259] According to one embodiment, an end user 4945 can access such a generated page via a client device 4940, for example, via a mobile application 4941 that displays the generated mobile application page 4905, or via a web application such as a browser 4942 that accesses and displays the generated web page 4906.
[0260] According to one embodiment, for example, publishable content (e.g., a web page or a mobile application (app) page) can be built using a code library such as Software Development Kit (SDK) 4907, which defines functional and visual content components such as a "buy" button.
[0261] Macro / micro classification process According to one embodiment, classification can be performed as a two-step process, comprising a macro-level classification step that can match the topic distribution and named entity distribution of documents with higher-level cluster nodes, and a micro-level classification step that can expand and compare categories within the feature space based on evaluations of microclusters.
[0262] Figure 50 shows a smart content classification flow diagram according to one embodiment. As shown in Figure 50, the automatic classification of content into known taxonomies (such as domain-specific ontologities), also referred to herein as "automatic classification," can be achieved by AI / ML systems.
[0263] New content (i.e., new content in a content management system, such as content generated in a content management system, or content management system When new content is created (5000) and uploaded to the system, during the prediction phase (5015), an engine such as the categorization engine described above can automatically suggest (5010) which category the newly created content should be categorized into.
[0264] According to one embodiment, after receiving selections (one or more) regarding the categorization of content, the content can be placed into such selected categories (5020).
[0265] Based on these selections, during the learning phase 5035, the cluster (for example, the cluster storing the memory or library of the categorization engine) can be updated accordingly (5030). Thus, the prediction phase 5015 can be updated according to the newly created other content 5000.
[0266] In other systems, such categorization is typically achieved by analyzing large amounts of similar data in the public domain. However, according to one embodiment, the aforementioned automated classification of content in a user-defined category hierarchy requires a novel approach, such as continuously observing pre-evaluated / categorized content, category metadata, and user behavior in response to categorization / classification proposals.
[0267] In one embodiment, unrated / uncategorized content, or content that is only slightly rated / categorized, presents a particular challenge in that it has little information for the system to rely on to make suggestions during a cold start. As businesses using content management systems evolve over time, older category hierarchies become less relevant and even outdated.
[0268] According to one embodiment, a system and tool for automatically classifying various types of content into categories based on clusters of related concepts using a microclustering approach are described herein. This approach performs better than conventional clustering approaches. Furthermore, the tool can learn from user behavior and adapt itself over time, providing several productivity improvements for content authors. According to one embodiment, the system incorporates the following set of use cases.
[0269] Creating a new taxonomy According to one embodiment, a new taxonomy can be created using the disclosed system and method, for example, as described below.
[0270] a. The system receives an instruction from the user indicating that they want to create a taxonomy. In certain embodiments, the taxonomy may include a hierarchical classification structure. Each node in the structure is a category.
[0271] b. The user receives an instruction indicating that they will classify / categorize at least some of their existing content items (e.g., content items already in a content management system) into the appropriate categories (e.g., one at a time or in bulk) to classify these taxonomies.
[0272] c. Receive new content (for example, uploaded or created within a content management system).
[0273] d. The system automatically receives a notification and begins learning from the actions described above. e. The system will begin recommending categories to which newly created content and existing uncategorized content may belong (either individually or in bulk).
[0274] f. The user accepts or rejects the suggestions (one at a time or all at once), and the user action notifies the system to learn from this and make better recommendations.
[0275] Figure 51 shows a taxonomy creation flowchart according to one embodiment. According to one embodiment, in step 5100, as described above, the system may receive an instruction indicating that the user creates one or more taxonomies. When creating a taxonomy, the instruction may further include categories under the taxonomy. In a particular embodiment, the taxonomy may include a hierarchical classification structure, where each node in the structure is a category.
[0276] In one embodiment, in step 5105, the system may receive an instruction indicating that the user classifies / categorizes at least some of the existing content items (e.g., content items already in the content management system) into the appropriate categories (e.g., one at a time or in batches) to classify / categorize these taxonomies. This may also include classifying / categorizing new content (e.g., new content uploaded to the content management system). New content does not necessarily receive classification or categorization.
[0277] According to one embodiment, in step 5110, the system automatically receives notification and begins learning from the actions described above. The system begins recommending categories to which newly created content and existing uncategorized content may belong. This can be done for individual content / on a per-content basis, or such recommendations can be done to categorize content in bulk.
[0278] According to one embodiment, in step 5115, the system may receive an instruction from the user indicating that they accept or reject the generated proposals (for example, one at a time or in batches).
[0279] Thus, according to one embodiment, in step 5120, a command indicating a user action notifies the system to learn from that user action and make better recommendations. The system can then improve its recommendations based on ML and AI, taking into account recorded acceptances or rejections of categorization suggestions.
[0280] Modification within the existing taxonomy structure According to one embodiment, modifications within an existing taxonomy structure can be made using the following process.
[0281] a. Users can create new categories, modify the structure of taxonomies, or categorize content.
[0282] b. The system generates suggestions (one at a time or in bulk) for existing unrated / uncategorized content or newly added content, including any newly added categories.
[0283] c. The user accepts or rejects the proposals (one at a time or in bulk), and the user The action will notify the system to make better recommendations.
[0284] Figure 52 shows a taxonomy modification flowchart according to one embodiment. According to one embodiment, in step 5200, as described above, the system may receive an instruction indicating that the user is adding a new category or modifying one or more existing taxonomies. The system may then receive an instruction indicating that the user is classifying existing or new content (for example, individually or in bulk) into the newly created category.
[0285] According to one embodiment, in step 5205, the system can automatically start generating suggestions and recommendations for categorizing existing content, including new content and newly created categories.
[0286] According to one embodiment, in step 5210, the system may receive an instruction indicating that the user accepts or rejects the generated proposals (for example, one at a time or in batches).
[0287] Thus, according to one embodiment, in step 5215, a command indicating a user action notifies the system to learn from that user action and make better recommendations. The system can then improve its recommendations based on ML and AI, taking into account recorded acceptances or rejections of categorization suggestions.
[0288] Content classification in bulk According to one embodiment, content can be categorized in bulk using the following process.
[0289] a. The system allows users to categorize / classify sets of content into suggested categories based on their confidence scores.
[0290] b. For a given category, the system can place recommended content items into three dissimilar buckets (high, medium, and low) based on their prediction confidence score.
[0291] c. Confidence score buckets provide a mechanism for grouping sets of content items so that bulk categorization is more easily achieved. For example, a user can accept all content categorization suggestions for content in a high confidence score bucket and reject all content categorization suggestions for other content in a low confidence score bucket.
[0292] d. In addition, users can skip accepting / rejecting suggestions through the overall classification threshold configuration. If selected, the system can automatically assign content to categories with a recommendation confidence score higher than the threshold set by the user.
[0293] e. Typically, suggestions are generated by backend jobs that run automatically and periodically. The system also allows repository administrators to manually trigger jobs, resulting in the recommendation system generating suggestions or immediately classifying content into the taxonomy assigned to the repository.
[0294] Figure 53 shows a sample taxonomy tree according to one embodiment, along with a category diagram.
[0295] As shown in Figure 53, according to one embodiment, as shown in the screenshot of the user interface 5300, (i) a sample taxonomy tree having the category "View Category Suggestions" 5305 is displayed in the smart suggestion picture. The user can be guided to a surface. This user interface 5300 is designed to help categorize content individually or in batches.
[0296] According to one embodiment, each category on the screen displays the number of suggestions in addition to the category itself. For example, content 5315 may be displayed with several suggested categories 5320. The content suggestions are divided into three buckets (high, medium, and low) based on the confidence score of the content that should be placed within the category.
[0297] Figure 54 further shows a sample taxonomy tree according to one embodiment, along with a category diagram.
[0298] As shown in Figure 54, according to one embodiment, as shown in the screenshot of the user interface 5400, (i) a sample taxonomy tree having the category “Browse Category Suggestions” can lead the user to a smart suggestion screen. This user interface 5400 is designed to help categorize content individually or in batches.
[0299] As shown in Figure 54, according to one embodiment, in this example, highly reliable content 5405 and moderately reliable content 5410 are selected. The "Assign to Category" link 5415 allows the user to place all selected content into their respective categories.
[0300] According to one embodiment, the categorization engine can further refine subsequent categorization proposals based on accepting proposals that fall within the range of high-confidence and medium-confidence proposals, as described above.
[0301] Figure 55 further shows a sample taxonomy tree according to one embodiment, along with a category diagram.
[0302] As shown in Figure 55, according to one embodiment, as shown in the screenshot of the user interface 5500, (i) a sample taxonomy tree having the category “Browse Category Suggestions” can lead the user to a smart suggestion screen. This user interface 5500 is designed to help categorize content individually or in batches.
[0303] As shown in Figure 55, according to one embodiment, the user can similarly select content (in this example, a score item with a low confidence score of 5505 is selected). The user can then select option 5510 to reject all low confidence score categorization suggestions.
[0304] According to one embodiment, as described above, the categorization engine can further refine subsequent categorization proposals based on the rejection of proposals that fall within the range of low-confidence proposals.
[0305] Figure 56 shows the configuration of an automatic classification threshold diagram according to one embodiment. As shown in Figure 56, according to one embodiment, the screen of the user interface 5600 As shown in the lean shot, this user interface 5600 is designed to help / allow the user to set an automatic classification threshold 5605 for confidence scores.
[0306] In other words, according to one embodiment, if the system and method are configured to enable the automatic classification or categorization of content items, the threshold may be set such that any content items with a confidence score above the set threshold can be automatically classified / categorized by the system without further user interaction. As shown in the figure, the taxonomy-level configuration for controlling the automatic classification setting is set between a “medium” confidence score and a “high” confidence score. This suggests that content with a confidence score above this threshold will be categorized / classified in the corresponding category. In addition, content items with a confidence score below the set threshold may be automatically discarded, or additionally or alternatively, may be presented to the user for input regarding the decision.
[0307] Figure 57 shows a configuration for triggering (bulk) reclassification of content in a repository diagram, according to one embodiment.
[0308] As shown in Figure 57, according to one embodiment, as shown in the screenshot of the user interface 5700, this user interface 5700 is designed to assist / allow the user to trigger, for example, a bulk classification of documents in a target repository 5705.
[0309] According to one embodiment, for example, the screenshot shows that content items in an existing repository (e.g., in a content management system) can be reclassified by a background process. The background process can run to propose new categorizations / classifications that should be submitted to the user for decision-making. Alternatively, for example, the background process can run for automated classification of content items, for example, based on a calculated confidence score.
[0310] Concepts and approaches In various situations, traditional clustering approaches do not work well from a content management system perspective. This is because categorizing / classifying content into the correct categories is difficult for several reasons, including but not limited to the following:
[0311] a. The amount of content within a content management system can increase dramatically over time.
[0312] b. Content can be obtained from various fields, and domain-specific solutions may not work.
[0313] c. Categories / nodes in the taxonomy tree are not restricted, and new categories are frequently introduced.
[0314] d. The number of content items within a category can vary from single digits to several thousand digits, and the same content may be assigned to multiple categories.
[0315] e. Content within a certain category may not be semantically similar to each other, and the same There may be multiple dissimilar subsets of similar content that share a category.
[0316] f. While the nearest neighbor approach can be accurate, it cannot be scalable when dealing with large amounts of data.
[0317] g. Traditional classification algorithms may not work in such scenarios because the distribution of content to categories is uneven and the labeled data is insufficient.
[0318] h. Traditional clustering approaches may not be suitable because, as more and more content is added, clusters can grow to arbitrary sizes and shapes, and rebuilding a cluster requires a path across all data points (content) present in the cluster, which is unacceptable in terms of scalability.
[0319] Figure 58 illustrates a problem associated with conventional clustering, which allows clusters to be of any shape according to one embodiment.
[0320] In one embodiment, data point N5805 is shown in a cluster map containing cluster 1 5800 and cluster 2 5810. From the described embodiment, it is clearly visible that data point N5805 should be part of cluster 1. However, in substance, the distance from N to the center of cluster 2 is actually shorter than the distance from N to the center of cluster 1. This suggests that cluster 2 appears to be more accurate.
[0321] According to one embodiment, with regard to both accuracy and scalability, the described approach provides general-purpose density-based single-pass microclustering that can discover clusters of arbitrary shapes (in terms of accuracy), reshape clusters without examining historical data (in terms of scalability), continuously learn and evolve from recent data, and accurately predict new content categories in bulk, thereby reducing the time and effort required to categorize content into the correct categories.
[0322] Adaptive microclustering According to one embodiment, the described approach comprises two parts or phases: (1) a learning phase in which features are extracted from documents and clusters are created, and (2) a predictive phase in which new content is automatically classified.
[0323] 1. Extract features from the document. In one embodiment, extracting features from a document involves creating and updating clusters (microclusters). In this step, the better the clusters can be represented, the greater the accuracy of the system. As illustrated in Figure 59, a single cluster center may not represent the entire cluster well on its own, and microclusters may be formed additionally inside the main cluster.
[0324] Figure 59 shows microclustering according to one embodiment. In the same example as in Figure 58 described above, cluster 1 in Figure 58, which is a larger cluster, is represented here by three points (e.g., microclusters), namely cluster 1 5900, cluster 1′ 5901, and cluster 1′′ 5902. In this case, the new data point N5905 is N Since the closest point is cluster 1'5901 and not cluster 2'5910, it will be categorized as belonging to cluster 1.
[0325] According to one embodiment, there are several steps performed to update the cluster and partition the microclusters within the cluster, as described below.
[0326] A. Generating feature vectors from text According to one embodiment, the first step is to transform the raw text document into feature vectors for further processing. The feature space of the text snippet is generated through a topic modeling system that outputs sparse text features (e.g., 200k dimensions). The text features are then projected into a lower-dimensional (2048) high-density feature space using random projection techniques for faster processing. If the content is sparse, the approach works on the content model itself, such as names and descriptions as well as other metadata used to define the taxonomy itself.
[0327] B. Extracting named entities from text According to one embodiment, a named entity recognizer (NER) system The system performs the task of automatically extracting entities such as the names of people, organizations, customer products and countries, and the titles of books or music albums.
[0328] In one embodiment, named entity recognition allows a system to extract key information from unstructured text and classify it into user-defined categories. For example, consider a case where documents from the pharmaceutical industry often refer to drug and chemical names, while documents from publishers include names of books, authors, and fictional characters. Such documents can be classified into their respective domains simply by looking at the entities. By considering the named entity score of a document in relation to the distribution of named entities in user-defined categories, the classification task can be made more effective, scalable, and accurate.
[0329] C. Extracting topics from text In one embodiment, topic extraction automatically discovers keywords and key phrases from a document by identifying term frequencies and grouping similar word patterns. User-defined taxonomy trees can grow significantly both horizontally and vertically, and statistical topic extraction models can be used for macro-level classification to select higher-order nodes in the tree, followed by feature-based microclassifiers, while proposing categories for the documents.
[0330] For example, customers of clothing may have a taxonomy tree hierarchy such as "Women > Clothing > Ethnic > Saree > Cotton > Chanderi" and "Household Linens > Bedding > Bed Covers > Cotton > Haiba". In this example, at the macro level, It is clear that comparing each document at the leaf level is not necessary to determine the category (all documents fall under either women's clothing or bed linens), but rather, a higher-level category can be selected by examining the topic distribution of items in the top 3-4 nodes, and furthermore, a micro-level (feature comparison) classifier can help suggest types of saris or bed covers (like chanderi or banarasi).
[0331] Figure 60 shows the sample topic distribution for clothing customers according to one embodiment.
[0332] According to one embodiment, as shown in the figure, the topic distribution includes, for example, a clothing cluster 6001, a household linen cluster 6020, and a household decor cluster 6030. Exemplary data points are shown around each cluster.
[0333] D. Create a cluster In one embodiment, all categories (nodes in a taxonomy tree) are assigned clusters. A cluster may include a summarized feature-space representation of feature vectors associated with the content present within that cluster, a distribution of named entities present in the documents within that cluster, and a distribution of topics extracted from the documents in that cluster.
[0334] According to one embodiment, a cluster is created with categories (initialized with extracted feature vectors, named entities from category metadata, and topics).
[0335] E. Definition of a cluster According to one embodiment, a cluster can be defined using the following process: Cluster features: According to one embodiment, a cluster is defined in terms of its centroid and radius. Centroid = the average of the feature vectors of the content present in the cluster. Radius = the average distance of the member feature vectors to the cluster centroid. A smaller radius indicates that the content present in the cluster is more closely related to each other.
[0336] In one embodiment, a list of named entities is attached to a cluster along with its frequency. The distribution is created by selecting the top n named entities from each document belonging to the cluster.
[0337] According to one embodiment, a topic distribution can be defined that identifies the most frequent topics attached to a cluster, similar to named entities. A topic distribution can also be created by extracting the most relevant topics from each document.
[0338] Figure 61 shows a visualization of cluster radius according to one embodiment. According to one embodiment, Figure 61 shows two exemplary clusters, cluster 1 6100 and cluster 2 6110. As shown in the figure, cluster 1 has a smaller radius than cluster 2. As mentioned above, a smaller radius means that the content present within the cluster is more closely related.
[0339] Update cluster features. According to one embodiment, as shown in Figure 61, a high-density, small-radius cluster (left) can achieve the best possible accuracy. To achieve this goal, the cluster is divided into microclusters as more and more content (e.g., new content) is added. If the new content consequently expands the microcluster radius, this new content does not need to be immediately added to the microcluster, but can instead simply be identified as a potential microcluster.
[0340] According to one embodiment, as content grows, potential microclusters may either grow into core microclusters as more and more similar content is added, or they may be ignored as outliers due to a decay factor over time. This microclustering approach helps discover disparate subsets of related content within a category.
[0341] The centroid and radius of a cluster can be updated online (without needing to consult historical data).
[0342] Figure 62 shows a cluster representation according to one embodiment. According to one embodiment, Figure 62 shows three typical clusters: cluster 1 6200, cluster 2 6210, and cluster 3 6230. As shown in the figure, each cluster has a different radius, and each radius reflects how closely related the content items within each cluster are. As mentioned above, a smaller radius means that the content present in the cluster is more closely related.
[0343] According to one embodiment, as content grows, potential microclusters may grow into core microclusters as more and more similar content is added, or they may be ignored as outliers such as O1 6240 and O2 6241 due to a decay factor over time. This microclustering approach helps to discover disparate subsets of related content within a category.
[0344] According to one embodiment, when a new document D is placed in a cluster, the associated named entity distribution and topic distribution for that cluster are updated by extracting topics and named entities from D.
[0345] 2. Classify new content into relevant categories. According to one embodiment, classification is a two-stage process, as described below, including a macro-level classification stage and a micro-level classification stage.
[0346] A. Macro-level classification According to one embodiment, as described above, each category node has a topic distribution and a named entity distribution associated with it. When a document is uploaded, topics and named entities are first extracted from the document, and a conditional probability score is calculated for each high-level category in the taxonomy tree.
[0347] Figure 63 shows a macro-level categorization according to one embodiment. More specifically, Figure 63 shows an embodiment in which the topic distribution and named entity distribution of documents coincide with the high-level cluster nodes (i.e., nodes 6300-6309).
[0348] According to one embodiment, given a list of topics such as "shirts" and "cotton" (and the weight of each topic for a document), the scoring function can be considered as the chance that a document may belong to the category "women's clothing". Here, n topics (e.g., topic-1, topic-2, ..., topic-n) and the weight of each topic (weight) topic-1 weight topic-2 , ..., weight topic-n Assuming document D has the following characteristics, the probability of that document joining category Catg-C can be calculated as follows:
[0349]
number
[0350] B. Microlevel classification According to one embodiment, during this stage, feature comparison is performed to classify all documents present in the child nodes of the macro-selected category into a finer-grained category.
[0351] Figure 64 shows categorization when higher-level categories are selected through macrosteps. In microsteps, according to one embodiment, the categories under the selected tree, the selected tree which is clothing 6401 --> women 6405, and the categories under the selected tree which are categories 6410-6415 are expanded and compared in feature space.
[0352] According to one embodiment, in order to categorize / classify new content into categories, the cosine similarity between the content and available microclusters can be calculated in the feature space, and the category to which the most similar microcluster belongs can be recommended as a new item.
[0353] In one embodiment, for example, as content users introduce new categories and more and more content is placed into those categories, the number of microclusters can reach very high values. Comparing new content to all existing microclusters in the high-dimensional feature space can be resource-intensive and degrade system performance. This problem can be solved with the decay window model approach used in data stream clustering.
[0354] Attenuation window model (weighted decay over time): In one embodiment, when using a decaying window model (weight decay over time), each piece of content is associated with a weight corresponding to its arrival time. When new content arrives, it is assigned the highest possible weight, and this weight decreases over time according to a time function (e.g., exponentially). A time function typically used for decaying window models is an exponential fading function. The weight of a microcluster can be calculated as the sum of the weights of the content present in it. While comparing new content with existing microclusters, the microclusters can first be sorted (in descending order) based on their weights, and then only the top n microclusters can be selected for feature similarity calculations. In this way, recently used, high-density microclusters can receive higher priority than less frequently used microclusters.
[0355] Figure 65 shows how cluster weights can decay over time according to one embodiment, and also provides an example of a decay window model.
[0356] According to one embodiment, as shown in Figure 65, there are three categories: Category 1 6510, Category 2 6520, and Category 3 6530. According to one embodiment, the damping functions applied to these categories are as follows:
[0357]
number
[0358] In one embodiment, considering the above equation, if category 1 6500 has 100 items added at t1, as shown in the figure as the thick score box 6501, then category 1 will have a score of 100 at t1, which will be attenuated as shown in the figure. If no further items are added to category 1 by t5, the score for the 100 items calculated by the attenuation function will be 6.25.
[0359] In one embodiment, considering the above equation, if category 2 6510 has 30 items added at t1, 20 items added at t2, and 10 items added at t4, as shown in the bold score boxes 6511, 6512, and 6513, then category 2 will have a score of 30 at t1, a score of 35 at t2, a score of 17.5 at t3, and a score of 17.75 at t4. Furthermore, if no more items are added to category 2 by t5, the score for the 60 items will be 9.375.
[0360] In one embodiment, considering the above equation, if category 3 6510 has 10 items added at t2, 10 items added at t3, 10 items added at t4, and 10 items added at t5, as shown by the thick-lined score boxes 6521, 6522, 6523, and 6524, then category 3 will have a score of 0 at t1, a score of 10 at t2, a score of 15 at t3, a score of 17.5 at t4, and a score of 18.75 at t5 with respect to 40 items.
[0361] Learning from user behavior In one embodiment, the user has the option to accept or reject the system's suggestions regarding content categorization / classification. Depending on the user action, the system may add the data to a machine learning database, for example, in the following ways:
[0362] If the user accepts the proposal In one embodiment, the most recent acceptance enhances the existing signal. The content is now part of a category cluster, and the content features are included in the cluster features. This is not a direct average update of the cluster features, and the attenuation coefficient plays a major role in this update. The previous cluster features are partially attenuated and then added to the most recently placed content features to compute the updated cluster features. This allows the system to give more weight to the most recently added content as it classifies the next batch. The frequencies of named entities and topics are also updated with respect to the categories.
[0363] If the user rejects the proposal According to one embodiment, if a user refuses to categorize / classify certain content, such refusal can ultimately lead to two scenarios.
[0364] i. The user rejects the item and places it in a different category. This action has the same effect as acceptance, and the category to which the user has placed the content at this point will consume weighted content features, and the updated cluster features will begin to recommend promoting similar content with a higher confidence level.
[0365] ii. The user rejects the content, and the content remains "uncategorized." In this situation, the system cannot know what to do with the content (or similar content) because it does not contribute to any cluster. To track and record such content, a category called "shadow cluster" is introduced. The shadow cluster includes all rejected and uncategorized content from all previously rejected proposals. When new content arrives, if this new content has the highest similarity score in the shadow cluster, the system does not propose any category for that content.
[0366] Figure 66 shows how shadow clusters can appear according to one embodiment. A rough sketch.
[0367] According to one embodiment, assigned / categorized / classified content items can be populated in cluster 1 6600, cluster 2 6610, and cluster 3 6620, as described above.
[0368] However, according to one embodiment, as described above, the shadow cluster 6640 may grow over time. These shadow clusters may grow indefinitely in shape and size, and as a result, the system may no longer be able to suggest categories for most of the newly created content. To avoid such a scenario, the same micro-clustering approach can be applied to the shadow clusters. Micro-clustering helps to group similar uncategorized content within the shadow clusters.
[0369] Figure 67 is a graph illustrating how shadow clusters may appear according to one embodiment.
[0370] More specifically, Figure 67 shows that, according to one embodiment, shadow cluster 1 6700, shadow cluster 2 6710, and shadow cluster 3 6720 can be generated as microclusters within a shadow cluster.
[0371] According to one embodiment, as microclusters within a shadow cluster grow, the recommendation system (e.g., a categorization engine) can begin generating suggestions for such new microclusters and provide these suggestions to the user. Such uncategorized content can form categories based on the most frequently used topics in the system, and the system can suggest names for those categories. The user can then form new categories from this content using the option to create new categories.
[0372] Figure 68 shows an example of a mechanism that suggests to the user that they create a new category from uncategorized content.
[0373] According to one embodiment, Figure 68 shows an exemplary screenshot of the user interface 6800. In the illustrated embodiment, uncategorized content 6805 is displayed. The user may be presented with a suggestion to create a new category 6810 for such uncategorized content (for example, within a taxonomy).
[0374] Figure 69 is a flowchart of a method for smart content categorization in a content management system according to one embodiment.
[0375] According to one embodiment, in step 6900, the method may include one or more computers that include a processor and provide access to a content management system.
[0376] According to one embodiment, in step 6910, the method may provide a content categorization engine on one or more computers, which can access taxonomies.
[0377] According to one embodiment, in step 6920, the method features content in a content management system by the recommendation system of the content categorization engine. The system can generate vectors, and the recommendation system can access the database of the content categorization engine and the AI / ML engine, and the generation of feature vectors is based on an evaluation of content that has been pre-categorized within the taxonomy at least.
[0378] According to one embodiment, in step 6930, the method can utilize the generated feature vector when categorizing new content into a taxonomy.
[0379] While specific embodiments have been described, various modifications, changes, alternative configurations, and equivalents are possible. The implementation examples described herein are not limited to operation within a few specific data processing environments, but can be freely implemented in multiple data processing environments. In addition, while the implementation examples have been described using a specific set of transactions and steps, it will be apparent to those skilled in the art that this is not intended to be limiting. Some flowcharts describe operations as sequential processes, but many of these operations can be performed in parallel or simultaneously. Furthermore, the order of operations may be rearranged. The processes may also have additional steps not shown in the diagrams. The various features and aspects of the implementation examples described above may be used individually or together.
[0380] Furthermore, while the implementation examples described herein have been explained using specific combinations of hardware and software, it should be understood that other combinations of hardware and software are also possible. Some of the implementation examples described herein may be implemented using hardware alone, software alone, or a combination thereof. The various processes described herein can be implemented on the same processor or on different processors in any combination.
[0381] Where a device, system, component, or module is described as being configured to perform a particular operation or function, such configuration can be achieved, for example, by designing electronic circuits to perform the operation; by programming programmable electronic circuits (such as a microprocessor) to perform the operation; by executing computer instructions or code, for example; or by designing a processor or core programmed to execute code or instructions or any combination thereof stored in a non-temporary memory medium. Processes can communicate using a variety of techniques, including but not limited to conventional techniques for inter-process communication, and different pairs of processes may use different techniques, or the same pair of processes may use different techniques at different times.
[0382] This disclosure provides certain details to ensure that the embodiments are fully understood. However, the embodiments may be carried out without these specific details. For example, well-known circuits, processes, algorithms, structures, and techniques are shown without unnecessary details to avoid ambiguity in the embodiments. This description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of other embodiments. Rather, the above description of embodiments will provide a description that enables the implementation of various embodiments for those skilled in the art. Various modifications may be made within the scope of the function and configuration of the elements.
[0383] Therefore, the specification and accompanying drawings should be considered illustrative rather than restrictive. However, it will be clear that additions, reductions, deletions, and other modifications and changes may be made to them without deviating from the broader spirit and scope of the disclosure. Thus, while specific implementation examples have been described, these are not intended to be limiting, and various modifications and equivalents are included within the scope of the disclosure.
[0384] The embodiments described herein may be implemented using one or more general-purpose or dedicated digital computers, computing devices, machines or microprocessors, or other types of computers including one or more processors, memory and / or computer-readable storage media programmed in accordance with the teachings of this disclosure. As will be apparent to those skilled in the art of software technology, suitable software coding can be readily prepared by a skilled programmer based on the teachings of this disclosure.
[0385] In some embodiments, the features described herein may be implemented entirely or partially in a cloud environment as part of a cloud computing system or as a service thereof. Such a cloud computing system may enable on-demand network access to a shared pool of configurable computing resources (e.g., networks, servers, storage, applications, and services) and may include, for example, characteristics as defined by the National Institute of Standards and Technology, such as on-demand self-service, broad network access, resource pooling, rapid scalability, and instrumentation services. Exemplary cloud deployment models may include public clouds, private clouds, and hybrid clouds, while exemplary cloud service models may include Software as a Service (SaaS), Platform as a Service (PaaS), and Data as a Service. Database as a Service (DBaaS), and infrastructure as a service. This may include Infrastructure as a Service (IaaS). In some embodiments, unless otherwise specified, the cloud may encompass, as used herein, embodiments of public cloud, private cloud, and hybrid cloud, as well as all cloud deployment models, including but not limited to cloud SaaS, cloud DBaaS, cloud PaaS, and cloud IaaS.
[0386] In some embodiments, a computer program product may be provided which is a non-temporary computer-readable storage medium containing instructions that can be used to program a computer to perform any of the processes described herein. Examples of such storage media may include, but are not limited to, hard disk drives, hard disks, fixed disks, or other electromechanical data storage devices, floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic or optical cards, nanosystems, or other types of storage media or devices suitable for the non-temporary storage of instructions and / or data.
[0387] The above description is provided for illustrative and explanatory purposes only and is not intended to be exhaustive or to limit the invention to the exact form disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments described above have been selected and described to best illustrate the principles of this teaching and their practical applications, thereby enabling those skilled in the art to understand various embodiments and variations suitable for specific intended uses. The scope is intended to be defined by the appended claims and their equivalents.
Claims
1. A system for smart content categorization in a content management system, One or more computers, including a processor, that provide access to the content management system, A content categorization engine provided on one or more computers and capable of accessing taxonomies, The system comprises a recommendation system including the content categorization engine, the recommendation system generates feature vectors from content in the content management system, and the recommendation system can access the database of the content categorization engine. The generation of the feature vectors is based at least on the evaluation of pre-categorized content within the taxonomy, The generated feature vectors are used in the system to categorize new content into the taxonomy.
2. The recommendation system according to claim 1, wherein the recommendation system creates clusters in the feature space based on pre-categorized content.
3. The recommendation system according to claim 2, wherein the recommendation system generates one or more recommendations for the new content to the taxonomy by calculating the feature spatial distance from the cluster.
4. The system according to claim 1, wherein the recommended system is used to create a new taxonomy or to modify the taxonomy.
5. The database of the content categorization engine includes a history of user acceptance of previous categorization recommendations. The system according to claim 1, wherein the database of the content categorization engine includes a history record of user rejections of previous categorization recommendations.
6. The recommendation system according to claim 5, wherein the recommendation system generates one or more recommendations for the new content to the taxonomy based on the historical records of user acceptance of previous categorization records and the historical records of user rejection of previous categorization recommendations.
7. The recommendation system according to claim 1, wherein the recommendation system generates recommendations for creating new categories within the taxonomy for multiple uncategorized content items.
8. A method for smart content categorization in a content management system, The steps include providing one or more computers, including a processor, that provide access to a content management system, The steps include providing one or more computers with a content categorization engine that can access the taxonomy, The recommendation system includes the step of generating feature vectors from content in the content management system using the content categorization engine, wherein the recommendation system can access the database of the content categorization engine, and the step of generating feature vectors includes at least the pre-categorization within the taxonomy. The method is based on an evaluation of the re-edited content, and furthermore, A method comprising the step of using the generated feature vectors to categorize new content into the aforementioned taxonomy.
9. The method according to claim 8, further comprising the step of creating clusters in the feature space based on the pre-categorized content within the taxonomy using the recommendation system.
10. The method according to claim 9, further comprising the step of generating one or more recommendations for the new content to the taxonomy by calculating the feature spatial distance from the cluster using the recommendation system.
11. The method according to claim 8, wherein the recommendation system is used to create a new taxonomy or to modify the taxonomy.
12. The database of the content categorization engine includes a history of user acceptance of previous categorization recommendations. The method according to claim 8, wherein the database of the content categorization engine includes a history record of user rejections of previous categorization recommendations.
13. The method according to claim 12, further comprising the step of generating one or more recommendations for the new content to the taxonomy based on the recommendation system, using the historical records of user acceptance of previous categorization records and the historical records of user rejection of previous categorization recommendations.
14. The method according to claim 8, further comprising the step of generating recommendations for creating new categories within the taxonomy for multiple uncategorized content items using the recommendation system.
15. A non-temporary computer-readable storage medium in which instructions are stored, wherein when an instruction is read and executed by a computer, the computer causes the computer to perform the following steps, and these steps are: The steps include providing one or more computers, including a processor, that provide access to a content management system, The steps include providing one or more computers with a content categorization engine that can access the taxonomy, The recommendation system includes the step of generating feature vectors from content in the content management system using the content categorization engine, the recommendation system can access the database of the content categorization engine, the step of generating feature vectors is based at least on an evaluation of pre-categorized content within the taxonomy, and the following steps further include: A non-temporary, computer-readable storage medium, comprising the step of using the generated feature vectors to categorize new content into the aforementioned taxonomy.
16. The following steps further: The non-temporary computer-readable storage medium according to claim 15, further comprising the step of creating clusters in the feature space based on the pre-categorized content within the taxonomy by the recommendation system.
17. The recommended system calculates the feature spatial distance from the cluster, thereby determining the taxon A non-temporary computer-readable storage medium according to claim 16, further comprising the step of generating one or more recommendations to Me about the new content.
18. The recommended system is used to create a new taxonomy or to modify the taxonomy, the non-temporary computer-readable storage medium according to claim 15.
19. The database of the content categorization engine includes a history of user acceptance of previous categorization recommendations. The non-temporary computer-readable storage medium according to claim 15, wherein the database of the content categorization engine includes a history record of user rejections of previous categorization recommendations.
20. The following steps further: A non-temporary computer-readable storage medium according to claim 19, comprising the step of generating one or more recommendations for the new content to the taxonomy based on the recommendation system, using the historical records of user acceptance of previous categorization records and the historical records of user rejection of previous categorization recommendations.