Classification of Data Objects Using Proximal Representation

The data object classification system enhances label prediction accuracy and speed by using a neighborhood representation with machine learning models, efficiently processing dynamic data objects with missing labels, and improving interpretability.

JP7711325B2Active Publication Date: 2025-07-22GOOGLE LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024532520
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-07-22
Estimated Expiration
2042-12-30

AI Technical Summary

Technical Problem

Existing machine learning models struggle to accurately and efficiently classify data objects, particularly in real-time, especially for dynamic data objects with missing labels, due to the need for processing high-dimensional categorical features and lack of interpretability.

Method used

A data object classification system that utilizes a neighborhood representation, generating a classification based on similarity scores and features of neighboring data objects within a dataset, employing machine learning models like neural networks to predict missing labels, and storing this information in a dense vector format for efficient processing.

Benefits of technology

Improves label prediction accuracy and inference speed for dynamic data objects by using a neighborhood representation, reducing memory usage and enhancing interpretability, allowing real-time classification and addressing reliability issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007711325000001
    Figure 0007711325000001
  • Figure 0007711325000002
    Figure 0007711325000002
  • Figure 0007711325000003
    Figure 0007711325000003
Patent Text Reader

Abstract

Methods, systems, and apparatus for classifying data objects, including a computer program encoded on a computer storage medium. One method includes maintaining a dataset including reference data objects, each having one or more labels, one or more features, or both, receiving a request to add a new data object to the dataset, the new data object having the one or more features but missing one or more labels, selecting N neighboring data objects based on similarity scores of the neighboring data objects with respect to the new data object, generating a neighborhood feature vector for the new data object, processing the neighborhood feature vector using a machine learning model to predict one or more labels for the new data object, and updating the dataset to include the new data object and to associate the one or more predicted labels with the new data object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the classification of data objects using machine learning and neural networks.

Background Art

[0002] A machine learning model receives an input and generates an output based on the received input and the values of the parameters of the model. For example, a machine learning model receives an image and generates a score for each set of classes. The score for a given class represents the probability that the image contains an object belonging to that class.

[0003] A machine learning model may, for example, be composed of a single level of linear or non-linear operations, or may be a machine learning model composed of multiple levels, such as a deep network, where one or more of them may be layers of non-linear operations. An example of a deep network is a neural network having one or more hidden layers.

Summary of the Invention

[0004] This specification describes a data object classification system that receives a data object and generates a classification of the data object. Specifically, the system accesses a dataset that stores reference data objects associated with known categories, and uses the known categories of similar reference data objects in the vicinity of the data object to generate a classification of the data object.

[0005] Generally, one innovative aspect of the subject matter described in this specification is to maintain a dataset that includes one or more labels, one or more features, or both, each having a reference data object, where each label of the reference data object defines each category of the reference data object, and each feature of the reference data object describes and maintains the characteristics of the reference data object, and to receive a request for adding a new data object to the dataset that has (i) one or more features but (ii) lacks one or more labels, and to select N neighboring data objects from the reference data objects based on the similarity scores of the neighboring data objects with respect to the new data object, where the similarity score of each neighboring data object is determined based on one or more features of the new data object and one or more features of the neighboring data object, and N is a natural number greater than or equal to 1, and for each neighboring data object within the N neighboring data objects, to generate a neighboring feature vector for the new data object using (i) one or more labels of the neighboring data object and (ii) the similarity score of the neighboring data object with respect to the new data object, and to process the neighboring feature vector using a machine learning model to predict one or more labels missing from the new data object, and to update the dataset to include the new data object and to associate one or more predicted labels with the new data object. It can be embodied in a method including these operations.

[0006] Other embodiments of this aspect include corresponding apparatuses, systems, and computer programs configured to execute the method aspects, encoded on a computer storage device. Each of these and other embodiments may optionally include one or more of the following features.

[0007] In some embodiments, maintaining a data set that includes reference data objects includes maintaining data that describes a heterogeneous graph that includes nodes connected by edges, each node corresponding to a different reference data object within the reference data objects, and each edge representing a relationship between two nodes connected by the edge.

[0008] In some embodiments, maintaining a data set that includes reference data objects includes maintaining the reference data objects as well as the associated labels and features of the reference data objects in a relational data set.

[0009] In some embodiments, the operation further includes determining a similarity score for each neighboring data object with respect to a new data object based on one of Euclidean distance or cosine similarity in an embedding space.

[0010] In some embodiments, the operation further includes determining a similarity score for each neighboring data object with respect to a new data object based on one of a pointwise mutual information (PMI) score or a bipartite score.

[0011] In some embodiments, the machine learning model includes one of a neural network, a logistic regression model, a support vector machine (SVM), or a decision tree or random forest model.

[0012] In some embodiments, generating a neighborhood feature vector for a new data object includes determining a concatenation of the respective similarity scores of a subset of N neighboring data objects having a particular category within each category defined by one or more labels.

[0013] In some embodiments, a subset of the N neighboring data objects includes neighboring data objects each having a first label that defines a trusted category, e.g., a positive label.

[0014] In some embodiments, a subset of the N neighboring data objects includes neighboring data objects each having a second label that defines a non-trusted category, e.g., a negative label.

[0015] As used herein, a data object having a positive label that defines a trusted category refers to a data object corresponding to a category that meets a trustable condition (e.g., meets a reliability threshold). A data object having a positive label may be "whitelisted," whereby use of the data object by a computer system, e.g., an operation or other processing, is permitted. A data object having a negative label that defines a non-trusted category refers to a data object corresponding to a category that does not meet a trustable condition (e.g., does not meet a reliability threshold). A non-trusted data object may be "blacklisted," whereby use of the data object by a computer system may not be permitted.

[0016] In some embodiments, each data object represents an image, video, audio, text, or web page.

[0017] In some embodiments, the value of N depends on the total number of labels that each reference data object has.

[0018] Certain embodiments of the subject matter described herein may be implemented to realize one or more of the following advantages. By configuring a data object classification system to predict the label of an input data object based on a neighborhood representation of the input data object determined from a dataset of reference data objects maintained by the system, both the accuracy of label prediction and the inference speed of the prediction process can be improved. In particular, label prediction can be performed in real time for dynamic data objects that are frequently modified over time and at very early stages in the entire life cycle of the data object (e.g., within seconds or milliseconds after the data object becomes available). Label prediction can also be performed using fewer memory resources than other approaches because it can avoid processing high-dimensional categorical features that are typically stored as sparse high-dimensional vectors using one-hot encoding. Instead, the label can be predicted based on a neighborhood representation that is efficiently stored and retrieved from memory in the form of a dense vector.

[0019] The neighborhood representation of an input data object may be generated by encoding, into a structured dense vector format, a subset of neighborhood reference data objects of the input data object selected from a dataset, the similarity scores (for the input data object), known labels, feature information, etc., which is a data - efficient and information - rich representation that facilitates not only the classification of the input data object using any of various known classifiers but also other downstream tasks involving processing of the representation. Further, the process of generating the neighborhood representation can be traced, thereby making it easier to examine the context of the classification, which improves the interpretability of the data object classification system. In many real - world production environments, this enables understanding how a particular label was assigned to an input data object, e.g., how the system predicted that a web page is a phishing or malicious web page, or how the system predicted that an image contains prohibited content. Thereby, it becomes possible to address potential reliability issues that arise as concerns regarding the decisions leading to a particular result.

[0020] Details of one or more embodiments of the subject matter described in this specification are set forth in the accompanying drawings and the following description. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

Brief Description of the Drawings

[0021]

Figure 1

Figure 2

Figure 3

Figure 4

Best Mode for Carrying Out the Invention

[0022] Like reference symbols and designations in the various drawings refer to like elements.

[0023] This specification describes a data object classification system that receives data objects and generates classifications of the data objects.

[0024] FIG. 1 shows an example of a data object classification system 100. The data object classification system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations, and the systems, components, and techniques described below may be implemented.

[0025] The data object classification system 100 generates label data for a given data object, e.g., label data 142 for a new data object 102. Each data object has one or more features that describe respective characteristics, attributes, or properties of the data object. The label data for a given data object identifies one or more category labels of the data object. Each of the category labels is a label of a category within an individual (e.g., finite) set of categories that the data object classification system 100 has determined the data object may belong to. Generally, the category label for a given category is a term that identifies or describes the category. For example, the category of an image may include the topic of the content depicted in the image, or the objects detected within the image.

[0026] The data object classification system 100 may be configured to generate label data for various data objects, for example, any type of data object that can be classified as belonging to one or more categories, with each category being associated with respective labels. When the label data is generated, the data object classification system 100 can use the label data in any of various ways.

[0027] For example, when the data object is an image, the data object classification system 100 may be a visual recognition system that determines whether the input image includes an image of an object belonging to an object category from a predetermined set of object categories. In this example, the label data for a given input image identifies one or more labels of the input image, and each label labels the respective object category to which the object depicted in the image belongs. In this example, the features of the data object may include pixel intensity values, edges, corners, ridges, interest points, and color histograms of the image, metadata that may be associated with the image, and the like.

[0028] As another example, when the data object is a video or a part of a video, the data object classification system 100 may be a video classification system that determines to which topic or topics the input video or part of the video is related. In this example, the label data for a given video or video part identifies one or more labels of the video or video part, and each label identifies the respective topic to which the video or video part is related. In this example, the features of the data object may include pixel intensity values of each frame of the video, visual features of the frames, and metadata that may be associated with the video.

[0029] As another example, when the data object is audio data, the data object classification system 100 may be a speech recognition system that determines the terms represented by a given utterance. In this example, the label data for a given utterance identifies one or more labels, each label being a term represented by the given utterance. In this example, the features of the data object may include time-domain and frequency-domain features, such as energy envelope and distribution, frequency content, harmonicity, pitch, and the like.

[0030] As another example, when the data object is text data, the data object classification system 100 may be a text classification system that determines to which topic or topics an input text segment is related. In this example, the label data for a given text segment identifies one or more labels for the text segment, each label identifying a respective topic to which the text or text segment is related. In this example, the features of the data object may include text features obtained through any suitable natural language processing techniques, such as lexical analysis, pattern recognition, sentiment analysis, and the like.

[0031] As another example, when the data object is a web page, the data object classification system 100 may be a web page classification system that determines which category a given web page (or the content of a given web page) belongs to. In this example, the label data of a given web page identifies one or more labels, and each label identifies a respective category to which the web page belongs. In this example, the features of the data object include URL features (e.g., host name, usage rate, subdomain, top-level domain, URL length, and presence of HTTPS), domain features (e.g., domain registration period, age of the domain, owner of the domain name, associated domain registrar, traffic of the website, and DNS records), website code features (e.g., presence or absence of website forwarding, and use of pop-up windows), website content features (e.g., hyperlinks to the target website, hyperlinks to non-target websites, media content, and text input fields), and website visitor features (e.g., IP address, visit time, and session duration).

[0032] In these examples, appropriate actions such as, for example, providing labels to the users of the system (e.g., as a response to a request / search query submitted by the user), transferring the data object, the label, or both to another system for further processing, flagging for deeper review in a human investigation process, and using the labeled data object in a map / navigation context can be performed by the system on the data object.

[0033] The data object classification system 100 includes a neighboring data object selection engine 110, a dataset 120 of reference data objects, a neighboring feature vector generation engine 130, and a classification subsystem 140.

[0034] The dataset 120 of reference data objects can be a labeled data object database (or other suitable data structure) that stores reference data objects in association with known label data of the reference data objects, for example, in a data storage device and / or cloud-based storage. The dataset 120 of reference data objects can be any data store suitable for storing a plurality of reference data objects and the labels and features associated with the plurality of reference data objects. In some embodiments, the dataset 120 can be a relational database having a multi-dimensional structure, where each feature of a data object is stored in a cell of the relational database defined by rows and columns. In an example where each data object represents a web page, the relational database can store relationships between features of web pages having different labels. For example, the relational database can store relationships between the geographical location of a web page having a particular label, the content included in a web page having a particular label, the layout of a web page having a particular label, and the functions available on a web page having a particular label, and / or combinations thereof.

[0035] In some embodiments, the dataset 120 can be a heterogeneous graph including a plurality of nodes and edges connecting the nodes, where both the nodes and the edges are of heterogeneous types. The heterogeneous graph can represent relationships between reference data objects and can be stored in one or more data stores. Each node in the heterogeneous graph represents a different reference data object (nodes of the same type represent reference data objects of the same type), and a pair of reference data objects in the heterogeneous graph are connected by an edge indicating the relationship between the two reference data objects represented by the pair of nodes.

[0036] FIG. 2 shows an example of a heterogeneous graph 200 representing reference data objects. The heterogeneous graph 200 identifies relationships between reference data objects.

[0037] For example, the heterogeneous graph 200 includes nodes 202 and 204 representing a first reference data object, nodes 206 and 208 representing a second reference data object, and a node 210 representing a third reference data object. The first, second, and third reference data objects can be of different types. For example, nodes 202 and 204 can each represent a web domain, and nodes 206 and 208 can each represent an IP address. An edge connecting a pair of nodes indicates that there is some measure of relevance between the reference data objects represented by the connected pair of nodes. For example, node 202 and node 206 are connected by edge 212, which indicates that the web domain represented by node 202 is hosted by the IP address represented by node 206.

[0038] Returning to FIG. 1, to classify the new data object 102, the neighborhood data object selection engine 110 selects an appropriate subset of the neighborhood data objects 122 from the dataset 120 of reference data objects based on the similarity score of each of the neighborhood data objects 122 to the new data object 102. The appropriate subset includes at least one data object in the dataset, but does not include all of the data objects in the dataset. That is, the neighborhood data object selection engine 110 selects a relatively small number of reference data objects that are most similar to the new data object 102 as the neighborhood data objects 122 from the dataset. A large number, for example, one million, ten million, one billion or more, of reference data objects can be maintained in the dataset 120 by the system 100, but only a small number of them may actually need to be selected, for example, only the top 50, 100, or 200 neighborhood data objects having the highest similarity score or meeting a similarity threshold may need to be selected.

[0039] For example, the neighborhood data object selection engine 110 can calculate a similarity score by evaluating a vector space-based similarity measure, such as pairwise cosine similarity or Euclidean distance, of vectors representing the features of the neighborhood data object 122 and the features of the new data object 102, respectively, or some other embedding (i.e., a structured representation within an embedding space). As another example, the neighborhood data object selection engine 110 can calculate a similarity score by evaluating a graph-based distance measure between a pair of nodes within a heterogeneous graph or another information graph representing the neighborhood data object 122 and the new data object 102, respectively. Other similarity scores can be used.

[0040] The neighborhood feature vector generation engine 130 uses a selected subset of the neighborhood data objects 122 to generate one or more neighborhood feature vectors 132 for the new data object 102. In particular, the neighborhood feature vector generation engine 130 can generate each neighborhood feature vector 132 using (i) one or more labels of each neighborhood data object 122 within the selected subset, (ii) a similarity score (with respect to the new data object), or both (i) and (ii).

[0041] The classification subsystem 140 receives one or more neighborhood feature vectors 132 and processes the neighborhood feature vectors 132 to generate label data 142 that identifies one or more category labels for the new data object 102. The classification subsystem 140 can include one or more machine learning models. Each classification machine learning model can be configured as, by way of just a few examples, a neural network model (e.g., a multi-layer perceptron model), a naive Bayes model, a support vector machine model, a linear regression model, a logistic regression model, or a k-nearest neighbor model. Each model can be configured to process a model input that includes the neighborhood feature vectors 132 according to model parameters to generate a model output that includes respective classification scores for each of a predetermined set of possible categories of the new data object 102.

[0042] In some embodiments that include a plurality of classification machine learning models, each model may correspond to a respective partition (or subset) of a predetermined set of possible categories of the new data object 102. In some of these embodiments, each model may process a different neighborhood feature vector 132 from each other model. Each model may be configured to generate a respective classification score for each category within each partition. For example, assume that the new data object 102 represents an image having a total of four possible categories, namely, (i) horse, (ii) dog, (iii) golden retriever, and (iv) German shepherd (each representing an object that may be present in the image). The first partition may include the more general categories, i.e., categories (i) and (ii), and the second partition may include the more specific categories, i.e., categories (iii) and (iv). In this example, the first model included in the classification subsystem 140 can generate a respective classification score for each of categories (i) and (ii), and the second model included in the classification subsystem 140 can generate a respective classification score for each of categories (iii) and (iv), and the classification score of a given category represents the likelihood that the new data object 102 belongs to the given category. Within each partition, the classification subsystem 140 may then select the category having the higher (or highest) classification score as the category determined by the classification subsystem of the new data object 102.

[0043] When the label data 142 of the new data object 102 is generated, the data object classification system 100 may store the new data object 102 in the dataset 120 of the reference data objects. For example, the system may store the new data object in association with the label data of the new data object, for example, in association with data identifying the label generated for the new data object. In some embodiments, instead of or in addition to storing the new data object in the dataset 120 of the reference data objects, the data object classification system 100 may associate a label or labels with the new data object and provide the labeled data object for use for some immediate purpose. For example, the system may provide the label to another system, for example, a client device whose application has requested a connection to a web page, and then that client device may use the label to approve or reject the application's connection request.

[0044] FIG. 3 is a flowchart of an exemplary process 300 for generating one or more labels for an input data object. For convenience, process 300 is described as being executed by one or more computer systems located in one or more locations. For example, a suitably programmed data object classification system, such as data object classification system 100 of FIG. 1, may execute process 300.

[0045] The system maintains a dataset that includes a plurality of reference data objects (step 302). The dataset is a labeled dataset that stores the reference data objects in association with data specifying one or more labels and one or more features of the data objects. Each label of the reference data objects in the dataset defines a respective known category of the reference data objects. Each feature of the reference data objects describes a characteristic of the reference data objects.

[0046] The label data, feature data, or both of the reference data objects within the dataset can be made available to the system in any of various ways. For example, the label data can be provided by a user of the system, who may be the same as or different from another user who provided the feature data of the reference data object. As another example, the feature data can be provided by a data collection system, and the label data may be determined by another data classification system from the feature data. As another example, the label data can be pre-determined by the system itself from the available feature data of the reference data object.

[0047] As a specific example for illustration, each data object may represent a web page, and one or more labels of each data object may be selected from one or more of (i) activity state labels that define the activity state of the web page (including, for example, active state labels, dormant state labels, expired state labels, etc.), (ii) legitimacy state labels that define the legitimacy state of the web page (including, for example, phishing labels, trusted labels, etc.), and one or more features of each data object may describe one or more of (i) the geographical location of the web page (defined, for example, by an IP address or area code), (ii) the content included in the web page, (iii) the layout of the web page, (iv) the functionality of the web page, (v) the age of the web page or the number of visits the web page has received, etc.

[0048] The system receives a request to add a new data object to the dataset (step 304). The new data object may have one or more features, but one or more labels may be missing. The request may be submitted by a user of the system, for example, via a wired or wireless network. The request may refer to the new data object, for example, by identifying the new data object and providing data that defines one or more features of the new data object through a user interface made available by the system. Unlike the feature data made available to the system when a request to add a new data object is received, one or more labels are typically not specified by the request and are therefore not immediately available to the system. To determine one or more missing labels of the new data object, the system accesses reference data objects stored in the dataset maintained by the system.

[0049] Specifically, the system selects N neighboring data objects from among the plurality of reference data objects based on the similarity scores of the neighboring data objects to the new data object (step 306). In particular, the system determines the similarity score of each reference data object in the dataset based on (i) one or more features of the new data object and (ii) one or more features of the reference data object. Once determined, the system sorts these similarity scores in descending order of score value and selects the N reference data objects having the highest score values accordingly (as the N neighboring data objects). Here, N is a predetermined natural number of 1 or more, for example, 10, 50, 100, etc. In some cases, N may have a predetermined fixed value. In other cases, the exact value of N may depend on the total number of features each reference data object has. Thus, when each reference data object has more labels that define more categories of the reference data object, the system generally selects more neighboring data objects.

[0050] Depending on how the reference data object is stored in the dataset, the system can calculate the similarity score in any of various ways. For example, the similarity score can be calculated by evaluating a vector space-based similarity measure, such as pairwise cosine similarity or Euclidean distance, of vectors or some other embedding that represent the features of the reference data object and the features of the new data object, respectively. As another example, the similarity score can be calculated by evaluating a graph-based distance measure between a pair of nodes in an information graph representing the reference data object and the new data object, respectively.

[0051] As a specific example, in an embodiment where the dataset maintains multiple reference data objects in a heterogeneous graph format, the system can calculate the similarity score as a pointwise mutual information (PMI) similarity score by using the technique described in commonly owned PCT patent application No. WO2021173114A1, which is incorporated herein by reference. As another specific example, in these embodiments, the system can alternatively calculate the similarity score as a bipartite similarity score by using the technique described in commonly owned U.S. Patent No. US10152557B2, which is incorporated herein by reference.

[0052] The system generates one or more neighborhood feature vectors for the new data object (step 308). In particular, the system generates each neighborhood feature vector using (i) one or more labels of the neighborhood data object, (ii) the similarity score of the neighborhood data object to the new data object for each neighborhood data object within the N neighborhood data objects, or both (i) and (ii).

[0053] For example, the system may generate a neighborhood feature vector by concatenating some or all of the similarity scores of N neighboring data objects, concatenating some or all of the labels of N neighboring data objects, or both. As another example, the system may generate a neighborhood feature vector by determining a concatenation of similarity scores, labels, or other values that may be derived from both, such as binary values or other numerical values.

[0054] In a particular example, the system may generate a neighborhood feature vector by determining a concatenation of the similarity scores of each subset of N neighboring data objects that have one or more particular categories out of all of the categories defined by one or more labels. In this particular example, the neighborhood feature vector may be calculated as follows. x1x2···x L x L+1 ···x KL x KL+1 ···x KL+N Here, K << N, and L can be any natural number less than or equal to the total number of labels of each data object, and is as follows. ● x1 = the highest similarity score among all the similarity scores of N neighboring data objects having the first particular label. ... ● x L = the highest similarity score among all the similarity scores of N neighboring data objects having the Lth particular label. ... ● x KL = the Kth highest similarity score among all the similarity scores of N neighboring data objects having the Lth particular label. ● x KL+1 = the highest similarity score among all the similarity scores of N neighboring data objects. ... ● x KL+N = the Nth highest similarity score among all the similarity scores of N neighboring data objects.

[0055] The system processes one or more neighborhood feature vectors using one or more machine learning models to predict one or more missing labels for a new data object (step 310). Since each of the neighborhood feature vectors generated in step 308 contains information about the similarity scores and labels of N neighboring data objects of the new data object, or is a vector representation derived therefrom, using the neighborhood feature vectors can facilitate faster and more accurate identification of one or more missing labels of the new data object.

[0056] Generally, each of the one or more machine learning models is a model configured to generate a classification score for a new data object. In some embodiments that include multiple machine learning models, different models may process different model inputs to generate different classification scores. For example, a first model may process a first model input including a first neighborhood feature vector according to first model parameters to generate respective classification scores for each category within a first partition of a predetermined set of categories, and a second model may process a second model input including a second neighborhood feature vector of the new data object according to second model parameters to generate respective classification scores for each category within a second partition of the predetermined set of categories. The system may select each of the one or more labels based on the classification scores, for example, by selecting the label corresponding to the highest classification score (either within a predetermined set of categories or within a partition of a predetermined set of categories), or by selecting the label corresponding to a classification score that exceeds a predetermined threshold.

[0057] When label data for a new data object is generated, the system updates the data set to include the new data object and associate one or more prediction labels with the new data object (step 312). That is, the system adds to the data set data that specifies (i) one or more features defined by the received request and (ii) one or more labels predicted by the system from the neighboring feature vectors, thereby adding the new data object (as a reference data object) to the data set.

[0058] Process 300 can be executed for each input data object to generate one or more labels for the input data object. The input data object can be a data object for which the desired label data, i.e., the label to be generated by the system for the data object, is not known. The system can also execute Process 300 for input data objects within a set of training data, i.e., input data objects for which the label to be predicted by the system is known, in order to train the system, i.e., to determine the trained values of the parameters of a machine learning model (s) configured to process neighboring feature vectors. In some embodiments, Process 300 can be repeatedly executed for input data objects selected from a set of training data as part of a gradient descent method using conventional machine learning training techniques, e.g., error backpropagation training techniques, to train the model.

[0059] FIG. 4 is a block diagram of an example of a computer system 400 that can be used to execute the operations described above. System 400 includes a processor 410, a memory 420, a storage device 430, and an input / output device 440. Each of the components 410, 420, 430, and 440 can be interconnected using, for example, a system bus 450. The processor 410 is capable of processing instructions for execution within the system 400. In one embodiment, the processor 410 is a single-threaded processor. In another embodiment, the processor 410 is a multi-threaded processor. The processor 410 is capable of processing instructions stored in the memory 420 or the storage device 430.

[0060] The memory 420 stores information within the system 400. In one embodiment, the memory 420 is a computer-readable medium. In one embodiment, the memory 420 is a volatile memory unit. In another embodiment, the memory 420 is a non-volatile memory unit.

[0061] The storage device 430 can provide mass storage for the system 400. In one embodiment, the storage device 430 is a computer-readable medium. In various different embodiments, the storage device 430 can include, for example, a hard disk device, an optical disk device, a storage device shared on a network by a plurality of computing devices (e.g., a cloud storage device), or some other large-capacity storage device.

[0062] The input / output device 440 provides input / output operations for the system 400. In one embodiment, the input / output device 440 may include one or more of, for example, a network interface device such as an Ethernet card, a serial communication device such as an RS-232 port, and / or a wireless interface device such as an 802.11 card. In another embodiment, the input / output device may include a driver device configured to receive input data and transmit output data to other input / output devices such as a keyboard, a printer, and the display device 370. However, other embodiments such as mobile computing devices, mobile communication devices, set-top box television client devices, etc. may also be used.

[0063] Although FIG. 4 illustrates an example of a processing system, embodiments of the subject matter and the functional operations described herein may be implemented in other types of digital electronic circuits, or in computer software, firmware, or hardware that include the structures disclosed herein and their structural equivalents, or one or more combinations thereof.

[0064] An electronic document (referred to simply as a document for brevity) does not necessarily correspond to a file. A document may be stored in a part of a file that holds other documents, a single file dedicated to the document, or multiple coordinated files.

[0065] Embodiments of the subject matter and the operations described in this specification can be implemented in digital electronic circuitry, or in computer software, firmware, or hardware that includes the structures disclosed in this specification and their structural equivalents, or combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions, encoded on a computer storage medium (or multiple media) for execution by, or to control the operation of, a data processing apparatus. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a mechanically generated electrical, optical, or electromagnetic signal, generated to encode information for transmission to a suitable receiver apparatus for execution by a data processing apparatus. The computer storage medium can be, or can include, a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them. Further, the computer storage medium is not a propagated signal, but the computer storage medium can be the source or destination of computer program instructions encoded on an artificially generated propagated signal. The computer storage medium can also be, or can include, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).

[0066] The operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on, or received from, one or more computer-readable storage devices.

[0067] The term "data processing apparatus" encompasses, by way of example, all kinds of devices, apparatuses, and machines for processing data, including programmable processors, computers, system-on-chips, or a plurality or combination of the foregoing. The apparatus may include special-purpose logic circuits, such as FPGAs (field-programmable gate arrays) or ASICs (application-specific integrated circuits). The apparatus may also include, in addition to hardware, code for creating an execution environment for the computer program, such as processor firmware, protocol stacks, database management systems, operating systems, cross-platform runtime environments, virtual machines, or code constituting one or a combination of these. The apparatus and the execution environment may implement various different computing model infrastructures, such as web services, distributed computing, and grid computing infrastructures.

[0068] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may or may not correspond to a file in a file system. The program can be stored as part of a file that holds other programs or data (such as one or more scripts stored in a markup language document), in a single file dedicated to the program, or in multiple cooperating files (such as files that store one or more modules, subprograms, or portions of code). A computer program can be deployed to be executed on one computer or on multiple computers located in one place or distributed across multiple places and interconnected by a communication network.

[0069] The processes and logic flows described herein can be executed by performing actions of operating on input data and generating output by one or more programmable processors executing one or more computer programs. The processes and logic flows can also be executed by special-purpose logic circuitry, such as an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit), and the apparatus can also be implemented as special-purpose logic circuitry.

[0070] Processors suitable for the execution of a computer program include, by way of example, both general-purpose microprocessors and special-purpose microprocessors. Generally, a processor receives instructions and data from a read-only memory or a random access memory or both. Essential elements of a computer are a processor that performs actions in accordance with instructions, and one or more memory devices that store instructions and data. Generally, a computer also includes or is operatively coupled to receive data from, transfer data to, or both, one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks. However, a computer need not have such devices. Further, a computer may be incorporated in another device, such as, by way of example only, a cellular phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive). Devices suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks. The processor and the memory may be supplemented by, or incorporated in, dedicated logic circuitry.

[0071] To provide interaction with a user, embodiments of the subject matter described herein may be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, and a keyboard and a pointing device, such as a mouse or trackball, by which the user may provide input to the computer. Other types of devices may also be used to provide interaction with the user. For example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback, and input received from the user may be in any form including acoustic, speech language, or tactile input. Additionally, the computer may interact with the user by sending documents to and receiving documents from the devices the user uses, such as by sending a web page to a web browser on a user's client device in response to a request received from the web browser.

[0072] Embodiments of the subject matter described herein may be implemented in a computing system that includes back-end components, such as a data server, or that includes middleware components, such as an application server, or that includes front-end components, such as a client computer having a graphical user interface or a web browser through which a user may interact with an implementation of the subject matter described herein, or in any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, such as by a communication network. Examples of communication networks include local area networks ("LANs") and wide area networks ("WANs"), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).

[0073] A computing system may include a client and a server. The client and the server are generally separated from each other and typically interact through a communication network. The relationship between the client and the server results from computer programs operating on respective computers and having a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a client device (e.g., for the purpose of displaying the data to a user interacting with the client device and receiving user input from the user). Data generated at the client device (e.g., as a result of user interaction) may be received at the server from the client device.

[0074] Although this specification contains many details of specific embodiments, these should not be construed as limitations on the scope of any invention or of what may be claimed, but rather as descriptions of features specific to particular embodiments of a particular invention. Features described in the context of individual embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately, or in any suitable sub-combination, in a plurality of embodiments. Further, features may be described above as acting in certain combinations and even initially claimed as such, but one or more features from a claimed combination may in some cases be excised from the combination, and the claimed combination may be directed to a sub-combination or variation of a sub-combination.

[0075] Similarly, in the drawings, operations are shown in a particular order, but this should not be understood to mean that such operations must be performed in the particular order or sequence shown, or that all of the shown operations must be performed, to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Further, the separation of various system components in the embodiments described above should not be understood to mean that such separation is required in all embodiments, and that the described program components and systems may generally be integrated into a single software product or packaged into multiple software products.

[0076] Thus, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. In addition, in the processes shown in the accompanying figures, the particular order or sequence shown is not necessarily required to achieve the desired result. In certain embodiments, multitasking and parallel processing may be advantageous.

Claims

**Claim 1** A method executed by one or more computers, comprising: maintaining a dataset including a plurality of reference data objects each having one or more labels, one or more features, or both, wherein each label of the reference data object defines a respective category of the reference data object, and each feature of the reference data object describes a characteristic of the reference data object; receiving a request to add a new data object to the dataset, the new data object having (i) one or more features but (ii) lacking one or more labels; selecting N of the neighboring data objects from the plurality of reference data objects based on a similarity score of the neighboring data objects with respect to the new data object, wherein the similarity score of each neighboring data object is determined based on the one or more features of the new data object and the one or more features of the neighboring data object, and N is a natural number greater than or equal to 1; for each neighboring data object within the N neighboring data objects, generating a neighboring feature vector for the new data object using (i) the one or more labels of the neighboring data object and (ii) the similarity score of the neighboring data object with respect to the new data object; processing the neighboring feature vector using a machine learning model to predict the one or more labels missing from the new data object; updating the dataset to include the new data object and associate the one or more predicted labels with the new data object; A method comprising the above steps. **Claim 2** The method of claim 1, wherein maintaining the dataset including the plurality of reference data objects includes maintaining data describing a heterogeneous graph including a plurality of nodes connected by edges, each node corresponding to a different one of the reference data objects within the plurality of reference data objects, and each edge representing a relationship between two nodes connected by the edge. **Claim 3** Maintaining the dataset including a plurality of reference data objects includes maintaining the plurality of reference data objects as well as the associated labels and features of the plurality of reference data objects in a relational dataset, the method according to claim 1.

4. The method according to claim 1, further comprising determining the similarity score of each neighboring data object with respect to the new data object based on one of Euclidean distance or cosine similarity in the embedding space.

5. The method according to claim 1, further comprising determining the similarity score of each neighboring data object with respect to the new data object based on one of the pointwise mutual information (PMI) score or the bipartite score.

6. The method according to claim 1, wherein the machine learning model includes one of a neural network, a logistic regression model, a support vector machine (SVM), or a decision tree or random forest model.

7. Generating the neighborhood feature vector of the new data object includes determining a concatenation of the similarity scores of each of a subset of the N neighboring data objects having a specific category among the respective categories defined by one or more labels, the method according to claim 1.

8. The method according to claim 7, wherein the subset of the N neighboring data objects includes neighboring data objects each having a positive label defining a trusted category.

9. The method according to claim 7, wherein the subset of the N neighboring data objects includes neighboring data objects each having a negative label defining a non-trusted category.

10. The method according to claim 1, wherein each data object represents an image, a video, an audio, a text, or a web page.

11. The method according to any one of claims 1 to 10, wherein the value of N depends on the total number of the labels that each reference data object has.

12. A system comprising one or more computers and one or more storage devices that store instructions which, when executed by the one or more computers, cause the one or more computers to perform operations, wherein the operations are maintaining a dataset that includes one or more labels, one or more features, or both, for each of a plurality of reference data objects, where each label of a reference data object defines the respective category of the reference data object, and each feature of the reference data object describes a characteristic of the reference data object; receiving a request to add to the dataset a new data object that has (i) one or more features but (ii) is missing one or more labels; selecting N of the neighboring data objects from the plurality of reference data objects based on a similarity score of the neighboring data objects with respect to the new data object, where the similarity score of each neighboring data object is determined based on the one or more features of the new data object and the one or more features of the neighboring data object, and N is a natural number greater than or equal to 1; for each neighboring data object within the N neighboring data objects, generating a neighboring feature vector for the new data object using (i) the one or more labels of the neighboring data object and (ii) the similarity score of the neighboring data object with respect to the new data object; processing the neighboring feature vector using a machine learning model to predict the one or more labels missing from the new data object; updating the dataset to include the new data object and associate the one or more predicted labels with the new data object; A system comprising the above. Claim 13 Maintaining the dataset including a plurality of reference data objects includes maintaining data that describes a heterogeneous graph including a plurality of nodes connected by edges, where each node corresponds to a different one of the reference data objects within the plurality of reference data objects, and each edge represents a relationship between two nodes connected by the edge. The system according to claim 12.

14. Maintaining the dataset including a plurality of reference data objects includes maintaining the plurality of reference data objects as well as associated labels and features of the plurality of reference data objects in a relational dataset. The system according to claim 12.

15. The system according to claim 12, further including determining the similarity score of each neighboring data object with respect to the new data object based on one of Euclidean distance or cosine similarity in the embedding space.

16. The system according to claim 12, further including determining the similarity score of each neighboring data object with respect to the new data object based on one of a pointwise mutual information (PMI) score or a bipartite score.

17. The system according to claim 12, wherein the machine learning model includes one of a neural network, a logistic regression model, a support vector machine (SVM), or a decision tree or random forest model.

18. Generating the neighborhood feature vector of the new data object includes determining a concatenation of the similarity scores of each of a subset of the N neighboring data objects having a particular category among the respective categories defined by one or more labels. The system according to claim 12.

19. The system according to any one of claims 12 to 18, wherein each data object represents an image, a video, an audio, text, or a web page.

20. A computer storage medium storing a computer program, where the computer program includes instructions that, when executed by one or more computers, cause the one or more computers to perform operations, and the operations are Maintaining a dataset that includes one or more labels, one or more features, or both, and a plurality of reference data objects each having them, wherein each label of the reference data object defines the respective category of the reference data object, and each feature of the reference data object describes the characteristics of the reference data object, and said maintaining; Receiving a request to add a new data object to the dataset that (i) has one or more features but (ii) lacks one or more labels; Selecting N of the neighboring data objects from the plurality of reference data objects based on a similarity score of the neighboring data objects with respect to the new data object, wherein the similarity score of each neighboring data object is determined based on the one or more features of the new data object and the one or more features of the neighboring data object, and N is a natural number greater than or equal to 1, and said selecting; For each neighboring data object within the N neighboring data objects, generating a neighboring feature vector of the new data object using (i) the one or more labels of the neighboring data object and (ii) the similarity score of the neighboring data object with respect to the new data object; Processing the neighboring feature vector using a machine learning model to predict the one or more labels missing from the new data object; Updating the dataset to include the new data object and associating the one or more predicted labels with the new data object; A computer storage medium comprising.

Citation Information

Patent Citations

  • Dataset adaptation for high-performance in specific natural language processing tasks

    US20190197128A1