Personalized data models leveraging closed data
By utilizing part of the enclosed data set in machine learning model training and freezing parameters associated with partial entity subsets, the contradiction between data privacy and model performance in the prior art is solved, and prediction effects with high accuracy and low complexity are achieved.
Patent Information
- Application Number
- CN201980017210.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2039-11-27
AI Technical Summary
The prior art is difficult to improve the prediction accuracy and performance of machine learning models while maintaining data privacy.
By using a portion of the enclosed data set to train a machine learning model, the specific steps include determining a subset of data associated with multiple entities, freezing parameters associated with a subset of entities, and training the model with the remaining data subset.
It realizes the prediction accuracy and performance of machine learning models while maintaining data privacy, and avoids the risks of model leakage and data sharing.
Smart Images

Figure CN113179659B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to personalized data models utilizing closed data. Background Art
[0002] The machine learning architecture may employ one or more models to predict outputs for received inputs. Some machine learning models may be trained using a data set. The data set may have subsets of data associated with public data and private data. During training of the machine learning model, the model may determine one or more parameters. Thus, each trained machine learning model may provide output predictions based on received inputs according to the current values of the one or more parameters. Summary of the invention
[0003] Some embodiments relate to a method for training a machine learning architecture, the method being implemented by one or more processing circuits. The method includes receiving a data set by the one or more processing circuits. In addition, the method includes determining, by the one or more processing circuits, a first portion of the data set associated with a plurality of entities. In addition, the method includes training, by the one or more processing circuits, an entity model using the first portion of the data set, wherein the entity model is trained to recognize one or more patterns in subsequently received data. In addition, the method includes determining, by the one or more processing circuits, a second portion of the data set associated with a first entity subset of the plurality of entities. In addition, the method includes determining, by the one or more processing circuits, a second entity subset, wherein the second entity subset does not include any entity in the first entity subset. In addition, the method includes: freezing, by the one or more processing circuits, one or more parameters associated with the second entity subset, such that the one or more parameters remain fixed during subsequent training of the entity model, and such that one or more non-frozen parameters of the trained entity model are not associated with the second entity subset, and training, by the one or more processing circuits, the entity model using the second portion of the data set.
[0004] In some embodiments, the method further includes determining, by one or more processing circuits, a third portion of the data set associated with a plurality of users, wherein the plurality of users includes a set of user identifiers and a plurality of user information, and training, by one or more processing circuits, a user model using the third portion of the data set, wherein the user model is trained to recognize one or more patterns in subsequently received data. In addition, the method includes receiving, by one or more processing circuits, an input data set. In addition, the method includes inputting, by one or more processing circuits, the input data set into a user model and an entity model, and generating, by one or more processing circuits, an output prediction based on the trained user model and the trained entity model, wherein the output prediction is specific to the first entity subset, and wherein the output prediction is an accuracy measure, the accuracy measure including a value. In addition, the output prediction is also generated based on using a user embedding vector generated by the user model and an entity embedding vector generated by the entity model. In addition, training each of the user model and the entity model also includes configuring at least one neural network.
[0005] In some embodiments, the method further includes determining, by the one or more processing circuits, a fourth portion of the data set associated with a third entity subset of the plurality of entities, wherein the third entity subset does not include any entity in the first entity subset or the second entity subset, and training, by the one or more processing circuits, a second entity model using the fourth portion of the data set, the second entity model being based on an entity model trained and frozen using the first portion of the data set. In addition, the fourth portion of the data set utilized in training the second entity model does not include any data from the second portion of the data set utilized in training the entity model.
[0006] In some implementations, the first portion of the data set does not include any data from the second portion of the data set, such that each subset of the plurality of entities includes a particular data set.
[0007] In some implementations, the method further includes configuring, by one or more processing circuits, a first neural network associated with the user model, and configuring, by one or more processing circuits, a second neural network associated with the entity model.
[0008] In some embodiments, a method for training a machine learning framework utilizes a two-stage technique, a first stage associated with training a physical model using a first portion of a data set, and a second stage associated with training the physical model using a second portion of the data set.
[0009] Some embodiments relate to a system having at least one processing circuit. The at least one processing circuit may be configured to receive a data set. In addition, the at least one processing circuit may be configured to determine a first portion of the data set associated with a plurality of entities. In addition, the at least one processing circuit may be configured to train an entity model using the first portion of the data set, wherein the entity model is trained to recognize one or more patterns in subsequently received data. In addition, the at least one processing circuit may be configured to determine a second portion of the data set associated with a first entity subset of the plurality of entities. In addition, the at least one processing circuit may be configured to determine a second entity subset, wherein the second entity subset does not include any entity in the first entity subset. In addition, the at least one processing circuit may be configured to freeze one or more parameters associated with the second entity subset, such that the one or more parameters remain fixed during subsequent training of the entity model, and such that one or more non-frozen parameters of the trained entity model are not associated with the second entity subset, and train the entity model using the second portion of the data set.
[0010] In some embodiments, the at least one processing circuit is further configured to determine a third portion of the data set associated with a plurality of users, wherein the plurality of users includes a set of user identifiers and a plurality of user information, and train a user model using the third portion of the data set, wherein the user model is trained to recognize one or more patterns in subsequently received data. In addition, the at least one processing circuit is configured to receive an input data set. In addition, the at least one processing circuit is configured to input the input data set into the user model and the entity model, and generate an output prediction based on the trained user model and the trained entity model, wherein the output prediction is specific to the first entity subset, and wherein the output prediction is an accuracy measure, the accuracy measure including a value.
[0011] In some embodiments, the at least one processing circuit is further configured to determine a fourth portion of the data set associated with a third entity subset of the plurality of entities, wherein the third entity subset does not include any entity in the first entity subset or the second entity subset, and to train a second entity model using the fourth portion of the data set, the second entity model being based on the entity model trained and frozen using the first portion of the data set. In addition, the at least one processing circuit is configured to determine a fifth portion of the data set associated with a fourth entity subset of the plurality of entities, wherein the fourth entity subset does not include any entity in the first entity subset, the second entity subset, or the third entity subset, and to train a third entity model using the fifth portion of the data set, the third entity model being based on the entity model trained and frozen using the first portion of the data set.
[0012] In some implementations, the first portion of the data set does not include any data from the second portion of the data set, such that each subset of the plurality of entities includes a particular data set.
[0013] Some embodiments relate to one or more computer-readable storage media having instructions stored thereon, which, when executed by at least one processing circuit, cause the at least one processing circuit to perform operations. The operations include receiving a data set. Furthermore, the operations include determining a first portion of the data set associated with a plurality of entities. Furthermore, the operations include training an entity model using the first portion of the data set, wherein the entity model is trained to recognize one or more patterns in subsequently received data. Furthermore, the operations include determining a second portion of the data set associated with a first entity subset of the plurality of entities. Furthermore, the operations include determining a second entity subset, wherein the second entity subset does not include any entity in the first entity subset. Furthermore, the operations include freezing one or more parameters associated with the second entity subset, such that the one or more parameters remain fixed during subsequent training of the entity model, and such that one or more non-frozen parameters of the trained entity model are not associated with the second entity subset, and training the entity model using the second portion of the data set.
[0014] In some embodiments, the operations further include determining a third portion of the data set associated with a plurality of users, wherein the plurality of users includes a set of user identifiers and a plurality of user information, and training a user model using the third portion of the data set, wherein the user model is trained to recognize one or more patterns in subsequently received data. Furthermore, the operations include receiving an input data set. Furthermore, the operations include inputting the input data set into the user model and the entity model, and generating output predictions based on the trained user model and the trained entity model, wherein the output predictions are specific to the first entity subset, and wherein the output predictions are accuracy measures, the accuracy measures including values.
[0015] In some implementations, the operations further include configuring a first neural network associated with the user model and configuring a second neural network associated with the entity model. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings are not intended to be drawn to scale. The same reference numbers and names in different drawings represent the same elements. For clarity, not every component may be labeled in every drawing. In the drawings:
[0017] Figure 1A is a block diagram depicting an implementation of a machine learning architecture according to an illustrative implementation;
[0018] Figure 1B is a block diagram depicting an embodiment of a data collection architecture according to an illustrative embodiment;
[0019] Figure 2 is a block diagram of an analysis system and associated environment according to an illustrative embodiment;
[0020] Figure 3 is a flow chart of a method for training a machine learning architecture according to an illustrative embodiment;
[0021] Figure 4 is a flow chart of a method for training a machine learning architecture according to an illustrative embodiment;
[0022] Figures 5A-5D According to various illustrative embodiments, Figure 1A Example learning curve plots associated with the machine learning architecture shown;
[0023] Figure 6 is a block diagram depicting an implementation of a machine learning architecture according to an illustrative implementation;
[0024] Figure 7 is a block diagram depicting an implementation of a machine learning architecture according to an illustrative implementation;
[0025] Figure 8 According to an illustrative embodiment, Figure 1A Example hidden layer representations of neural networks associated with the machine learning architecture shown;
[0026] Fig. 9 According to an illustrative embodiment, Figure 1A Example hidden layer representations of neural networks associated with the machine learning architecture shown;
[0027] Fig.10 According to an illustrative embodiment, Figure 1A Example hidden layer representations of neural networks associated with the machine learning architecture shown; and
[0028] Fig.11 is a block diagram of a computing system in accordance with an illustrative embodiment. DETAILED DESCRIPTION
[0029] The present disclosure relates to systems and methods generally related to training of machine learning architectures. In some embodiments, training of machine learning architectures may include training models using data sets collected by one or more processing circuits. In some embodiments, models are trained so that they can recognize one or more patterns in subsequently received data. The data sets may include many data subsets that can be used in the training of the machine learning architecture. In some embodiments, a subset of the data set may include an open data set. The open data set may consist of data collected by one or more processing circuits. For example, the collected data may include a business type (e.g., non-profit, government agency, healthcare provider). In various embodiments, a subset of the data set may also include a closed data set. A closed data set may be associated with multiple identifiers and may consist of data received by one or more processing circuits. For example, an identifier may include an entity identification number, and the collected data may include entity-specific data (e.g., customers, patients, mailing lists, purchase history). In some embodiments, an entity model may be trained using an open data set (e.g., a first portion of a data set). In some embodiments, the entity model may be trained again using a portion of a closed data set (e.g., a second portion of a data set) associated with a first entity subset that may include a single entity. Prior to training the entity model using the closed data set, a second entity subset may be determined that does not include any entity in the first entity subset. Additionally, prior to training the entity model using the closed data set, one or more parameters associated with the second entity subset may be frozen such that one or more non-frozen parameters of the trained entity model are not associated with the second entity subset. In various embodiments, the user model may also be trained using an open data set (e.g., a third portion of the data set).
[0030] In some systems, an open dataset is the only dataset used to train a machine learning architecture and ultimately generate output predictions for the machine learning architecture. However, the ability to incorporate a closed dataset in the training of a machine learning architecture allows output predictions to be generated based on training an open dataset associated with multiple entities and a portion of a closed dataset associated with a specific subset of entities, providing entity-specific enhanced output predictions for the entity. This approach allows the machine learning architecture to maintain the privacy of closed datasets specific to a subset of entities while providing significant improvements to their output predictions, thereby improving the accuracy of the predictions and the performance of the machine learning architecture. Therefore, aspects of the present disclosure address issues in data modeling privacy by maintaining the privacy of closed datasets used to generate output predictions specific to a subset of entities (i.e., data that should not be used to train a baseline model but can be used to train a portion of a model that is only related to a subset of entities).
[0031] In some systems, in order to maintain the privacy of a closed dataset so that it is not shared between entities, a separate closed dataset model is created for each entity so that a new personalized model specific to each entity can be trained. The system can then maintain the new personalized model as well as the entity model and the user model. However, the ability to incorporate a portion of a closed dataset into the training of a machine learning architecture enables output predictions to be generated based on a user model and an entity model (e.g., a baseline entity model), where the entity model only utilizes a portion of the closed dataset associated with a specific subset of entities to perform additional training, which provides enhanced performance and efficiency for the training of the machine learning architecture while reducing duplication in the entire model. This approach allows the machine learning architecture to be trained to maintain the privacy of a portion of a closed dataset specific to a subset of entities while providing an effective model that minimizes duplication, thereby improving the overall design of the machine learning architecture. Therefore, aspects of the present disclosure solve problems in data modeling architectures by designing data models that utilize a baseline training model (e.g., an entity model) to generate embedding vectors specific to a subset of entities.
[0032] Therefore, the present disclosure is directed to systems and methods for training machine learning architectures so that output predictions can be entity-specific. In some embodiments, the described systems and methods include utilizing one or more processing circuits. The one or more processing circuits allow for receiving a data set and subsequently training a model based on the received data set. The trained model can then be utilized to generate output predictions so that the output predictions can be an accuracy measure of the correlation between a specific entity and a specific user. In the present disclosure, the trained model includes a user model and an entity model (i.e., a two tower model). In some embodiments, the two tower model can provide output predictions of "matches" between matching pairs (i.e., user / entity pairs).
[0033] In some embodiments, a user model is trained using an open dataset associated with a user (e.g., a first phase). In parallel, an entity model is trained using an open dataset associated with multiple entities (e.g., a first phase). The entity model is further trained using a portion of a closed dataset that is specific to a subset of entities (e.g., a second phase). During the training process using a portion of the closed dataset, one or more parameters associated with the open dataset can be frozen (i.e., fixed so that they remain constant). This enables a second training of the entity model based on the closed dataset specific to the subset of entities (i.e., during the second phase). In addition, the entity model can be saved after training using the open dataset so that multiple entities can utilize a baseline training of the entity model before training the entity model using a closed dataset associated with a specific subset of entities.
[0034] In some embodiments, a training model may be generated using a neural network such that a first neural network is configured and associated with a user model and a second neural network is configured and associated with an entity model. In some embodiments, a second portion of a closed data set is received such that the second portion of the closed data set is associated with a second entity. The second entity model is then trained using a fourth portion of the data set associated with a third entity subset of the plurality of entities, wherein the third entity subset does not include any entity in the first entity subset or the second entity subset. The second entity model may be trained using the fourth portion of the data by one or more processing circuits, and wherein the second entity model may be based on an entity model trained and frozen using the first portion of the data set. This may then be performed using a third and fourth entity model, each trained entity model using a portion of the closed data set associated with a particular entity subset of the plurality of entities. In addition, each entity model may generate an embedding vector specific to the entity subset.
[0035] In cases where the systems discussed herein collect personal information about users and / or entities or can utilize personal information, the users and / or entities may be provided with an opportunity to control whether a program or feature collects user information and / or entity information (e.g., information about a user's social network, social actions or activities, occupation, user preferences, or a user's current location), or to control whether and / or how to receive content from a content server that may be more relevant to the user and / or entity. In addition, before certain data is stored or used, it may be processed in one or more ways to remove personally identifiable information. For example, the identity of the user may be processed so that personally identifiable information cannot be determined for the user, or the user's geographic location (such as city, ZIP code, or state level) may be summarized when obtaining location information so that the user's specific location cannot be determined. Therefore, the user and / or entity may control how information is collected about the user and / or entity and how the content server uses the information.
[0036] Reference now Figure 1A, a block diagram of a machine learning architecture 100 is shown, according to an illustrative embodiment. The machine learning architecture 100 is shown to include a user model 102, a user data set 104, a user identifier 106, an entity model 108, an open entity data set 110, an entity identifier set 112, a closed data set 114, an output prediction generator 116, and one or more entity-specific parameters 118. In some embodiments, the machine learning architecture 100 can be implemented using a machine learning algorithm (e.g., a neural network, a convolutional neural network, a recurrent neural network, a linear regression model, a sparse vector machine, or any other algorithm known to a person of ordinary skill in the art). In some embodiments, the learning algorithm can adopt a dual-tower approach. The machine learning architecture 100 can be communicatively coupled to other machine learning architectures (e.g., such as via a network 230, as shown in reference Figure 2 The machine learning architecture 100 may have an internal logging system that may be used to collect and / or store data (e.g., in the analysis database 220, as described in detail in reference to FIG. Figure 2 In various embodiments, the model may be trained on a recent N-day dataset, where the recent N-day dataset may include logs collected by an internal log system.
[0037] In some embodiments, the machine learning architecture 100 may be executed on one or more processing circuits, such as the following reference Figure 2 Those described in detail. Refer to Figures 1 and Figure 2 , one or more processing circuits may include a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), etc., or a combination thereof. The memory may include, but is not limited to, an electronic, optical, magnetic, or any other storage or transmission device capable of providing program instructions to the processor. The instructions may include code from any suitable computer programming language. In some embodiments, a portion of a data set associated with multiple entities may be used to train the entity model 108 (i.e., tower 1) so that it can recognize one or more patterns in subsequently received data. The subsequently received data may in turn be used to generate an embedding vector based on the multiple entities used to train the model.
[0038] In some embodiments, the embedding vector may be a series of floating point values representing the model's predictions. Specifically, the embedding vector may allow the model to represent the category generally in a transformed space (i.e., by taking a dot product). In some embodiments, the open dataset 110 may be a first portion of a dataset (e.g., stored in the analysis database 220) associated with previous data collected by one or more processing circuits. In general, training the entity model 108 using the open dataset 110 may be defined as stage 1.
[0039] In some embodiments, one or more processing circuits may receive an entity identifier set 112 (i.e., E2) and an open data set 110 (i.e., E1-O). The entity identifier set 112 may include a plurality of entity identifiers, which may include strings, numbers, and / or symbols that identify a particular entity. The open data set 110 may include features associated with a plurality of entities of the entity identifier set 112. For example, the open data set 110 may include industry types (e.g., manufacturing, retail, finance), enterprise sizes (e.g., small, medium, large), and / or geographic regions, such as countries (e.g., the United States, Canada, Germany). In a specific example, a plurality of entities may be classified as large manufacturing companies headquartered in the United States, or a plurality of entities may be classified as medium-sized retail companies headquartered in Germany. In another example, the open data set 110 may include medical centers / hospitals with trauma center levels (e.g., level 1, level 2, level 3) and bed capacities (e.g., less than 100 beds, 100 to 499 beds, 500 or more beds). In a particular example, multiple entities may be classified as Level 1 Trauma Centers with 500 or more beds.
[0040] In some embodiments, upon receiving the open dataset 110, the one or more processing circuits may use the open dataset 110 to train the entity model 108. Specifically, the open dataset 110 may be fed as an input to the entity model 108. During the training process, one or more parameters of the trained entity model 108 may be adjusted. In some embodiments, once the entity model 108 is trained, the one or more processing circuits may receive an input including an entity identifier and entity data associated with the entity identifier. After receiving the input, the one or more processing circuits may generate an entity embedding vector using the entity model 108.
[0041] In some embodiments, once phase 1 of training the entity model 108 is completed (i.e., benchmark training), the one or more processing circuits may determine a second portion of the data set associated with a first entity subset of the plurality of entities. In addition, the one or more processing circuits may also determine a second entity subset such that the second entity subset does not include any entity in the first entity subset. In some embodiments, each subset of entities may contain a single entity. After determining the first entity subset and the second entity subset, the one or more processing circuits may freeze the entity model 108. Freezing the entity model 108 may include fixing one or more parameters associated with the second entity subset such that the one or more parameters remain fixed during subsequent training of the entity model, and such that one or more non-frozen parameters of the trained entity model are not associated with the second entity subset. In some embodiments, the frozen entity model may be stored in a storage device (e.g., data store 209) such that the frozen entity model may be reused to reduce duplication across models and to increase storage capacity of the one or more processing circuits.
[0042] In some embodiments, after the training and freezing of the entity model 108 is completed, the one or more processing circuits may retrain the entity model 108 using a portion of the data set associated with a particular entity subset (i.e., a second portion of the data set associated with a first entity subset of multiple entities) so that it can recognize one or more patterns in subsequently received data. In various embodiments, the particular entity subset may include a single entity. After receiving the subsequent data, the one or more processing circuits may generate a user embedding vector using the entity model 108. In general, training the entity model 108 using a portion of the closed data set 114 (i.e., E1-C) can be defined as stage 2. In some embodiments, the portion of the closed data set 114 is the second portion of the data set associated with previous data collected and / or received by the one or more processing circuits.
[0043] In some embodiments, one or more processing circuits may receive a portion of a closed data set 114 associated with a particular entity subset. In particular, a portion of a closed data set 114 may include features associated with a particular entity subset. For example, a portion of a closed data set 114 may include sales history by customer, a list of recurring customers, and / or a mailing list. In a specific example, a particular entity identifier of the set of entity identifiers 112 may be Merchant 1, and a portion of the closed data set 114 may include a mailing list of all sales and campaigns over the past year. In another example, a portion of a closed data set 114 may include a list of customers and their locations. In a specific example, an entity identifier may be Company 1, and a portion of the closed data set 114 may include all of their customers over the past 5 years.
[0044] Upon receiving a portion of the closed data set 114, one or more processing circuits may train the entity model 108 using the portion of the closed data set 114. Specifically, a portion of the closed data set 114 may be fed as input to the entity model 108. During the training process, one or more entity-specific parameters of the entity model 108 may be adjusted and stored in a portion of the one or more entity-specific parameters 118. The one or more entity-specific parameters may be stored in a matrix so that they can be retrieved based on a specific entity subset. In some embodiments, the matrix E2 is an entity-specific N×M matrix of previous data collected and / or received by one or more processing circuits. In some embodiments, the one or more processing circuits may convert the entity-specific data of the matrix E2 into N embedding vectors of a fixed size, so that each specific entity may be associated with the N embedding vectors of a fixed size. In some embodiments, the matrix E2 may contain a size of M specific entities, so that each column of the matrix E2 may be entity-specific data (i.e., a portion of the one or more entity-specific parameters 118).
[0045] In some implementations, once the entity model 108 is trained, one or more processing circuits may receive an input including an entity identifier and entity data associated with the entity identifier. The one or more processing circuits may then utilize the entity model 108 to generate an embedding vector.
[0046] In some embodiments, the entity model 108 can be implemented using a machine learning algorithm, such as a neural network (i.e., NN), a deep neural network (i.e., DNN), a convolutional neural network (i.e., CNN), a recurrent neural network (i.e., RNN), a linear regression model, a sparse vector machine, or any other algorithm known to those of ordinary skill in the art. In some embodiments, once stage 1 of training the entity model 108 is completed, one or more processing circuits can then train the entity model 108 using a portion of the closed data set 114 without freezing one or more parameters.
[0047] In some embodiments, one or more processing circuits may train a user model 102 (i.e., Tower 2) using a portion of a data set associated with multiple users (i.e., U2) so that it can recognize one or more patterns in subsequently received data. After receiving the subsequent data, the one or more processing circuits may generate a user embedding vector using the user model 102. In some embodiments, the user data set 104 (i.e., U1) is a third portion of a data set associated with previous data collected by one or more processing circuits. In some embodiments, the one or more processing circuits may receive a user identifier set 106 and a user data set 104. The user identifier set 106 may include multiple user identifiers, which may include strings, numbers, and / or symbols that identify specific entities. The user data set 104 may include features associated with multiple user identifiers of the user identifier set 106. For example, the user data set 104 may include purchase history information, an email address, an Internet history, and / or a previous device history. In a specific example, one user identifier may be Person 1, and the user data set 104 may contain all purchases made by the user in the past 30 days, and the user's email address. In another example, user data set 104 may include previously watched movies, frequently visited stores, and previous food expenditures. In a specific example, one user identifier may be Person 2, and user data set 104 may contain all movies watched by the user in the past week, the five most common stores where the user made purchases, and the last 50 grocery stores where the user made purchases. In some implementations, user models 102 may be user-specific, such that each user may have their own user model 102. In some implementations, a particular user may have more than one user model 102 associated with more than one user data set (e.g., User Data Set 1 and User Data Set 2).
[0048] Once the set of user identifiers 106 and the user data set 104 are received, the one or more processing circuits may train the user model 102 using the set of user identifiers 106 and the user data set 104. Specifically, the set of user identifiers 106 and the user data set 104 may be fed as input to the user model 102. During the training process, one or more parameters of the user model 102 may be adjusted. In some embodiments, once the user model 102 is trained, the one or more processing circuits may receive input including a user identifier and data associated with the user. Using the received input, the one or more processing circuits may generate a user embedding vector using the user model 102.
[0049] In some embodiments, the entity model 108 can be implemented using a machine learning algorithm, such as a neural network (i.e., NN), a deep neural network (i.e., DNN), a convolutional neural network (i.e., CNN), a recurrent neural network (i.e., RNN), a linear regression model, a sparse vector machine, or any other algorithm known to those of ordinary skill in the art. In some embodiments, during subsequent training, one or more parameters of the user model 102 can be frozen.
[0050] Once the entity model 108 and the user model 102 are trained, one or more processing circuits of the machine learning architecture 100 can use the entity model 108 and the user model 102 to generate output predictions based on the received data. In some embodiments, the output prediction can predict the likelihood that the user will interact with a particular entity (i.e., perform a conversion). In some embodiments, the generated output prediction can be generated using an output prediction generator 116. Specifically, given the input to the user model 102 and the entity model 108, the output prediction generator 116 can use the dot product of the entity embedding vector and the user embedding vector to generate the output prediction. In some embodiments, the entity embedding vector and the user embedding vector can be represented as a series of floating point values. In some embodiments, the entity embedding vector and the user embedding vector are normalized so that the floating point values are scaled to variables between 0 and 1. At this point, the dot product of the output prediction generator 116 can produce a cosine distance between the vectors, which ranges from -1 (i.e., the least likely interaction) to +1 (i.e., the most likely interaction). In another embodiment, the Euclidean distance can be calculated to determine the likelihood of interaction.
[0051] The output prediction generator 116 may be defined as:
[0052] Phase 1:
[0053] Phase 2:
[0054] The notation symbols are as follows:
[0055] uid: user identifier
[0056] U: User model function
[0057] ·θ U : Parameters associated with the user model 102
[0058] · (uid): User embedding vector
[0059] f(uid): user data set 104
[0060] eid: entity identifier
[0061] E: Entity model function
[0062] ·θ E : Parameters associated with the entity model 108
[0063] · (eid): multiple entity embedding vectors
[0064] f(eid): Open Dataset 104
[0065] opData (uid, eid): the possibility of interaction between the user and a specific entity
[0066] Loss: loss function (e.g., sigmoid cross entropy)
[0067] · (eid): entity-specific embedding vector
[0068] clsdData(uid, eid): the possibility of interaction between the user and a specific entity
[0069] As shown, phase 1 is completed after training the user model 102 and training the entity model 108 using only the open data set 110. Also as shown, phase 2 is completed after training the user model 102 and the entity model 108 using the open data set 110 and a portion of the closed data set 114 associated with identifiers of a particular entity subset of the entity identifier set 112.
[0070] In some embodiments, the output prediction may predict how likely user 1 is to perform a conversion (e.g., purchase an item, provide requested information, etc.) relative to merchant X. For example, one or more processing circuits may receive a portion of an input data set and an input identifier associated with a user. In another example, one or more processing circuits may also receive a portion of an input data set and an identifier of a specific entity subset associated with the entity model 108. At this point, the one or more processing circuits may then generate an embedding vector using the user model 102 and the entity model 108, wherein the output generator 116 will generate an output prediction based on the likelihood of user 1 performing a conversion relative to merchant X. Thus, merchant X may target specific users using output predictions based on the likelihood of multiple users performing a conversion relative to their store. In another example, the output prediction may predict the likelihood that patient 1 will visit hospital Y. In another example, the output prediction may predict the likelihood that user 2 will click on content item Z.
[0071] In various embodiments, the one or more processing circuits may train a second entity model using the subsequently frozen baseline model (i.e., the model trained using the open dataset 110). In some embodiments, the second entity model may be an exact duplicate of the entity model 108 (after it completes stage 1). In some embodiments, the different portion of the closed dataset 114 is a fourth portion of the dataset associated with a third entity subset of the plurality of entities. In this regard, the second entity model may be a different entity model associated with the third entity subset.
[0072] In some embodiments, one or more processing circuits may receive the different portion of the closed data set 114 associated with the third entity subset. In various embodiments, the third entity subset may include a single entity. Upon receiving the different portion of the closed data set 114, the one or more processing circuits may train the second entity model using the different portion of the closed data set 114. Specifically, the different portion of the closed data set 114 may be fed as an input to the second entity model. During the training process, one or more parameters of the second entity model specific to the third subset (e.g., a portion of one or more entity-specific parameters 118, Company 2) may be adjusted. In some embodiments, once the second entity model is trained, the one or more processing circuits may receive an input including an entity identifier and entity data associated with the entity identifier. The received data may then be used in the second entity model to generate an entity embedding vector. For example, as shown in a portion of one or more entity-specific parameters 118, Company 1 may be one or more parameters utilized in the entity model 108, and Company 2 may be one or more parameters utilized in the second entity model. In this regard, one or more processing circuits may generate an embedding vector for a first subset of entities using the entity model 108, and one or more processing circuits may also generate an embedding vector for a third subset of entities using the second entity model. Furthermore, the first subset of entities may be associated with a particular entity (e.g., Company 1), and the third subset of entities may be associated with a different particular entity (e.g., Company 2). In some implementations, retraining the entity model 108 may occur for multiple entities, such that each subset of entities may generate an embedding vector for each particular subset of entities.
[0073] Reference now Figure 1B, a block diagram of a data collection architecture 150 according to an illustrative embodiment is shown. The data collection architecture 150 is shown to include a log layer 152, a user event layer 154, an interaction layer 156, and a database 158. In some embodiments, public data can be defined as data available to the public without certain restrictions (e.g., an open data set 110). In some embodiments, private data can be defined as restricted data (e.g., a closed data set 114). As shown, public data flows through each layer of the model in parallel, where the public data is ultimately stored in a central location so that the data can be shared without certain restrictions. Also as shown, private data flows through each layer of the model independently of public data, so that private data is ultimately stored in a private location, where the data is restricted to certain entities and / or users to prevent the data from leaking. In some embodiments, each layer of the data collection architecture 150 can be configured to run on one or more processing circuits (e.g., a computing system 201).
[0074] In some embodiments, the logging layer 152 can receive data from one or more processing circuits. For example, a client device can interact with an entity website. Each time the client device interacts with the website, a log can be saved and transmitted to the logging layer 152. In some embodiments, the logging layer 152 transmits the logs it receives to the user event layer 154. In some embodiments, all communications between layers or between one or more processing circuits can be completed through a network (e.g., network 230).
[0075] In some embodiments, the user event layer 154 can determine whether an event has occurred based on the received data. For example, a client device purchases an item from a physical website. In another example, the client device may then dial a phone number. In some embodiments, the received data can be stored and classified into a database 158. The database 158 can be used to provide closed data and open data to the interaction layer 156. In various embodiments, the data collection architecture 150 can be communicatively and operably coupled to the database 158. The database 158 can be composed of multiple databases so that the data can be separated based on certain attributes. For example, open data can be stored in a portion of the database 158, and closed data can be stored in another portion of the database 158.
[0076] In some embodiments, the interaction layer 156 can associate data based on user events and received logs. For example, associating logs that cause a client device to purchase an item. In some embodiments, the associated data can then be saved based on whether it is considered private or public. At this point, if the associated data is considered public, the data can be stored in a public database (e.g., a portion of database 158) so that it can be used to train models based on multiple entities. In another example, if the associated data is considered private, the data can be stored in a private database (e.g., a portion of database 158) so that it can be used to train models based on a specific subset of entities. In some embodiments, other architectures can be used to collect data.
[0077] Reference now Figure 2 , a block diagram of an analysis system 210 and an associated environment 200 according to an illustrative embodiment is shown. A user may use one or more user computing devices 240 to perform various actions and / or access various types of content, some of which may be provided through a network 230 (e.g., the Internet, a LAN, a WAN, etc.). As used herein, a "user" may refer to an individual who operates a user computing device 240 and interacts with resources or content items via the user computing device 240, etc. The user computing device 240 may be used to send data to the analysis system 210, or to access a website (e.g., using an Internet browser), a media file, and / or any other type of content. One or more entity computing devices 250 may be used by an entity to perform various actions and / or access various types of content, some of which may be provided through a network 230. The entity computing device 250 may be used to send data to the analysis system 210, or to access a website, a media file, and / or any other type of content.
[0078] The analysis system 210 may include one or more processors (e.g., any general or special purpose processors), and may include and / or be operably coupled to one or more temporary and / or non-temporary storage media and / or memory devices (e.g., any computer readable storage medium, such as magnetic storage, optical storage, flash memory, RAM, etc.). In various embodiments, the analysis system 210 may be communicatively and operably coupled to an analysis database 220. The analysis system 210 may be configured to query information from the analysis database 220 and store information in the analysis database 220. In various embodiments, the analysis database 220 includes various temporary and / or non-temporary storage media. The storage media may include, but are not limited to, magnetic storage, optical storage, flash memory, RAM, etc. The database 220 and / or the analysis system 210 may use various APIs to perform database functions (i.e., manage data stored in the database 220). These APIs may be, but are not limited to, SQL, ODBC, JDBC, etc.
[0079] Analysis system 210 may be configured to communicate with any device or system shown in environment 200 via network 230. Analysis system 210 may be configured to receive information from network 230. The information may include browsing history, cookie logs, television data, printed publication data, radio data, and / or online interaction data. Analysis system 210 may be configured to receive and / or collect interactions of user computing devices 240 on network 230. The information may be stored in data set 222.
[0080] The data set 222 may include data collected by the analysis system 210 by receiving interaction data from the entity computing device 250 or the user computing device 240. The data may be data input from a specific entity or user at one or more points in time (e.g., a patient, a customer purchase, an Internet content item). The data input may include data associated with multiple entities, multiple users, a specific entity, a specific user, etc. The data set 222 may also include data collected by various data aggregation systems and / or entities that collect data. In some embodiments, the data set 222 may include a closed data set 224 and an open data set 226. The closed data set 224 may be associated with multiple entities or multiple users, and the closed data set may be associated with entity-specific data or user-specific data.
[0081] The analysis system 210 may include one or more processing circuits (i.e., computer readable instructions executable by a processor) and / or circuits (i.e., ASICs, processor memory combinations, logic circuits, etc.) configured to perform various functions of the analysis system 210. In some embodiments, the processing circuit may be or include the model generation system 212. The model generation system 212 may be configured to generate various models and data structures from data stored in the analysis database.
[0082] Reference now Figure 3 , a flow chart of a training method 300 of a machine learning framework according to an illustrative embodiment is shown. The machine learning framework 100 can be configured to perform the method 300. In addition, any computing device described herein can be configured to perform the method 300.
[0083] In a general overview of method 300, at stage 310, one or more processing circuits receive a data set. At stage 320, the one or more processing circuits determine a first portion of the data set associated with a plurality of entities. At stage 330, the one or more processing circuits train an entity model using the first portion of the data set, wherein the entity model is trained to recognize one or more patterns in subsequently received data. At stage 340, the one or more processing circuits determine a second portion of the data set associated with a first entity subset of the plurality of entities. At stage 345, the one or more processing circuits determine a second entity subset, wherein the second entity subset does not include any entity in the first entity subset. At stage 350, the one or more processing circuits freeze one or more parameters associated with the second entity subset, such that the one or more parameters remain fixed during subsequent training of the entity model, and such that one or more non-frozen parameters of the trained entity model are not associated with the second entity subset. At stage 360, the one or more processing circuits train the entity model using the second portion of the data set.
[0084] Referring to method 300 in more detail, at stage 310, one or more processing circuits receive a data set. In some embodiments, the data set may include data collected from multiple sources. For example, the data set may be collected by one or more processing circuits. In some embodiments, the data set may be composed of many data subsets. In some embodiments, the data subsets may have different characteristics. For example, the data subset may be public data available to all the public. In another example, the data subset may be restricted private data and only available to certain entities and / or users. In some embodiments, the data subset may include certain restrictions.
[0085] At stage 320, one or more processing circuits determine a first portion of a data set (e.g., open data set 110) associated with multiple entities. In some embodiments, the first portion of the data set can be determined based on which data can be shared between entities. For example, the first portion of the data set can be associated with a business type (e.g., non-profit, government agency, healthcare provider). In some embodiments, the first portion of the data set does not include any identifying information about a particular subset of entities. For example, Company 1 can be a non-profit business, but can also have a list of all of its donors. Thus, Company 1's non-profit business status can be included in the first portion of the data set, where a list of all of its donors would not be included in the first portion of the data set.
[0086] At stage 330, one or more processing circuits train an entity model (e.g., entity model 108) using the first portion of the data set, wherein the entity model is trained to recognize one or more patterns in subsequently received data. In some embodiments, the one or more patterns may be any features or characteristics associated with the first portion of the data set.
[0087] In one example, an entity model may be trained based on one or more patterns associated with medical information, where the medical information may be defined as information associated with multiple medical institutions. In this regard, the medical information may include information about medical care received at multiple medical institutions, with all personally identifiable information removed.
[0088] In another example, an entity model can be trained based on one or more patterns associated with suspicious transaction and fraudulent purchase data, where a transaction and / or fraudulent purchase can be defined as any report made by an individual. In this regard, a suspicious transaction and / or fraudulent purchase can include any time an individual submits a report regarding a transaction or fraudulent purchase.
[0089] In some embodiments, one or more patterns can be modified so that the training of the entity model can be trained based on different patterns. In one example, the entity model can be trained based on one or more patterns associated with one or more attributed conversions, where a conversion can be defined as a meaningful user action completed by a content producer, and utilizing attributed conversions can include associating a conversion event with a previous content event. In this regard, an attributed conversion can include a conversion event that is preceded by a click from a client device on a content item from a content producer, which can be associated with a conversion through a click, where the conversion is then attributed to the click. On the other hand, if the conversion event is preceded by an impression from a client device on a content item from a content producer, it can be associated with a conversion through a view, where the conversion is then attributed to the impression.
[0090] At stage 340, one or more processing circuits determine a second portion of the data set associated with a first entity subset of the plurality of entities. In some embodiments, the second portion of the data set does not include any data associated with the first portion of the data set. Specifically, the second portion of the data set is constructed based on data specific to the first entity subset, so that the data specific to the first entity subset is not shared with the data specific to the second entity subset. In some embodiments, the first entity subset may be associated with a single entity. For example, Company 1 has data associated with all products sold last year (e.g., 1 billion products were sold). However, competitor Company 1 also has data associated with all products sold last year (e.g., 1 million products were sold). In this example, Company 1 contains most of the data specific to the company and should not be used in the benchmark model (i.e., the model trained using the open data set 110) to generate output predictions for competitor Company 1, and vice versa. Therefore, in this example, competitor Company 1 should not use a large amount of Company 1's data to generate output predictions that may ultimately affect Company 1's profitability and / or competitive advantage. In addition, Company 1 can expand its customer base based on the generated output predictions.
[0091] At stage 345, the one or more processing circuits determine a second entity subset, wherein the second entity subset does not include any entity in the first entity subset. In some embodiments, the second entity subset can be associated with a single entity. For example, the first entity subset can include Company 1, and the second entity subset can include Company 2. At stage 350, the one or more processing circuits freeze one or more parameters associated with the second entity subset such that the one or more parameters remain fixed during subsequent training of the entity model, and such that one or more non-frozen parameters of the trained entity model are not associated with the second entity subset. In some embodiments, freezing the one or more parameters can include fixing all parameters associated with the model trained using the first portion of the dataset (i.e., the open dataset 110).
[0092] At stage 360, one or more processing circuits train the entity model using the second portion of the data set. In some embodiments, training the entity model using the second portion of the data set is performed so that the entity can obtain an entity model that generates output predictions based on data specific to the entity (i.e., the second portion of the data set). In this regard, training the entity model using the second portion of the data set prevents model leakage of the entity-specific data and thereby maintains the privacy of each data set associated with each entity while producing more accurate output predictions. In some embodiments, training the entity model a second time using the second portion of the data set ensures that certain data remains private and is not shared or used to make output predictions for another entity.
[0093] In the same example above, the entity model can be trained again based on one or more patterns associated with medical information associated with a particular medical institution, where the medical information can include personally identifiable information about the patient. In this regard, medical information associated with a particular patient should not be disclosed to any entity (e.g., another medical institution) without the patient's permission. For example, a patient in a medical institution may have a disease that requires the care of multiple doctors. Doctors within a medical institution can work together to help resolve the medical condition, but the medical institution should not share the medical information with any other entity and / or individual without the patient's permission.
[0094] In the same example above, the entity model can be trained again based on one or more patterns associated with a particular financial company, where transactions and / or fraudulent purchases can be defined as any reports made by an individual directly to the particular financial company. In this regard, transactions and / or fraudulent purchases are specific to the financial company such that they should not be shared with other financial companies or any other entity, and examples of suspicious transactions and / or fraudulent purchases can include any time an individual submits a report of a transaction or fraudulent purchase to the particular financial company.
[0095] In some embodiments, one or more patterns can be modified so that the training of the entity model can be retrained based on different patterns. In the same example above, the entity model can be retrained based on one or more patterns associated with non-attributed conversions, where a conversion can be defined as a meaningful user action completed by a content producer, and the non-attributed conversion is not associated with a content item impression or a content item click event. In this regard, non-attributed conversions are specific to the content producer such that they are not shared with other content producers, and examples of non-attributed conversions can include purchases from client devices that are not associated with content item impressions or content item click events.
[0096] Reference now Figure 4 , a flow chart of a training method 400 of a machine learning architecture according to an illustrative embodiment is shown. The machine learning architecture 100 can be configured to perform the method 400. In addition, any computing device described herein can be configured to perform the method 400. The method 400 is similar to the method described above with reference to Figure 3 Similar features and functions as described in detail.
[0097] In the general overview of method 400, stages 410-460 are described above with reference to Figure 3Detailed description of stages 310-360 of the present invention is given in detail. However, at stage 470, one or more processing circuits determine a third portion of the data set associated with a plurality of users. At stage 480, the one or more processing circuits train a user model using the third portion of the data set, wherein the user model is trained to recognize one or more patterns in subsequently received data. At stage 490, the one or more processing circuits receive an input data set. At stage 492, the one or more processing circuits input the input data set into the user model and the entity model. At stage 494, the one or more processing circuits generate output predictions.
[0098] Referring to method 400 in more detail, at stage 470, one or more processing circuits determine a third portion of the data set associated with a plurality of users, wherein the plurality of users includes a set of user identifiers and a plurality of user information. Stage 470 is similar to the method described above with reference to Figure 3 The data may be similar to the features and functions described in detail at stage 420. However, at stage 470, the portion of the data is associated with multiple users.
[0099] At stage 480, one or more processing circuits train a user model using the third portion of the data set, wherein the user model is trained to recognize one or more patterns associated with the user. Stage 480 is similar to the above reference Figure 3 However, at stage 480, the user model is trained so that it can generate user embedding vectors associated with prediction scores based on multiple users used to train the user model (e.g., user model 102).
[0100] At stage 490, one or more processing circuits receive an input data set. In some embodiments, the input data set may have multiple portions associated with a user identifier. In some embodiments, the input data may have multiple portions associated with an entity identifier. In some embodiments, the entity identifier may determine which entity model to utilize. For example, if the entity identifier is Company 1, then the entity model that will be subsequently trained to be utilized will be the entity model that was trained using the portion of the closed data set associated with the Company 1 entity identifier. In another example, if the entity identifier is Company 2, then the entity model that will be subsequently trained to be utilized will be the entity model that was trained using the portion of the closed data set associated with the Company 2 entity identifier.
[0101] At stage 492, one or more processing circuits input the input data set into the user model and the entity model. In some embodiments, the entity model utilized can be an entity-specific entity model associated with an entity identifier. In some embodiments, the entity model utilized can be a baseline entity model associated with multiple entity identifiers (e.g., entity identifier set 106). For example, the entity model utilized can be associated with company M so that the output of the entity model is specific to company M. In another example, the entity model utilized can be associated with entity identifier set 112 so that the output of the entity model is not specific to the entity. At this point, company M can utilize the entity model associated with entity identifier set 112, but any other company other than company M cannot utilize the entity model associated with company M. In these examples, the specific data of company M is kept privately so that there is no model leakage (i.e., utilization by any other company).
[0102] At stage 494, one or more processing circuits generate an output prediction based on the trained user model and the trained entity model, wherein the output prediction is specific to the first entity subset, and wherein the output prediction is an accuracy measure, the accuracy measure comprising a value. Stage 494 is similar to the above reference Figure 1A Similar features and functionality as described in detail for the output prediction generator 116 of . In addition, the generated output predictions can provide targeted insights and detailed demographics of entities and users.
[0103] Reference now Figures 5A-5D , shows an example learning curve graph associated with the machine learning architecture 100 according to several illustrative embodiments. Figure 5A-5B Each of the plots is based on a true positive / false positive ratio of 1:5. Figure 5C-5D Each of the graphs is drawn based on a true positive / false positive rate of 1:100. In the illustrated embodiment, a true positive is defined as an input to the user model 102 and the entity model 108 where the model correctly predicts that the user has interacted with a particular entity when the user has interacted with the entity. In some embodiments, a false positive may be defined as an input to the user model 102 and the entity model 108 where the model predicts that the user has interacted with the entity when the user has not interacted with the particular entity. In some embodiments, a false negative may be defined as an input to the user model 102 and the entity model 108 where the model predicts that the user has not interacted with the entity when the user has interacted with the particular entity.
[0104] Described below Figure 5Ais an illustrative example of a precision-recall (PR) curve. The precision-recall curve can be used to evaluate the skill of a trained model (e.g., user model 102, entity model 108). The precision-recall curve can be used to evaluate recall (i.e., x-axis) versus precision (y-axis). The precision-recall curve can be defined as:
[0105]
[0106]
[0107] Where P PP is the positive example prediction ability, R is the recall rate, T P is the number of true positives, F P is the number of false positives, and F N is the number of false negatives. Therefore, a large area under the curve is the result of high recall and high precision (i.e., recall is a performance measure of the proportion of actual positive examples that are correctly identified, while precision is a performance measure of the proportion of positive examples that are correctly identified).
[0108] Thus, as shown, the precision-recall curve is plotted using three separate lines (e.g., line 502A, line 504A, and line 506A). Line 502A is plotted using the trained model after stage 1 of the machine learning architecture 100. Line 504A is plotted using the trained model after stages 1 and 2 of the machine learning architecture 100. Figure 8 As described in detail, line 506A is drawn using 3 trained model architectures. As shown, the trained model of line 504A produces significantly improved results than the trained model of line 502A because the model of line 504A can more accurately predict true positives. Also as shown, the trained model of line 504A cannot predict true positives as accurately as the model of line 506A. However, the trained model of line 506A has more disadvantages because each entity will require a completely different personalized model. In contrast, the trained model of 504A utilizes the two-stage technique of the machine learning architecture 100, in which completely different personalized models are not required, and therefore, allows the training of the machine learning architecture 100 to maintain the privacy of various parts of the closed data set 114 specific to a subset of entities, while providing an effective model that minimizes replication, so that the overall design of the machine learning architecture is improved.
[0109] Described below Figure 5Bis an illustrative example of a receiver operating characteristic (ROC) curve. The receiver operating characteristic curve can be used to evaluate the false positive rate (i.e., x-axis) versus the true positive rate (y-axis). The receiver operating characteristic curve can be defined as:
[0110]
[0111] Among them, T PR is the true positive rate, T P is the number of true positives, and F N is the number of false negatives. Thus, smaller values on the x-axis of the plot indicate lower false positives and higher true negatives, while larger values on the y-axis of the plot indicate higher true positives and lower false negatives (that is, a larger area under the curve (AUC) indicates how accurately the model predicts that interactions actually occurred and did not actually occur).
[0112] As shown, the precision-recall curve is plotted using three separate lines (e.g., line 502B, line 504B, and line 506B). Line 502B is plotted using the trained model after stage 1 of the machine learning architecture 100. Line 504B is plotted using the trained model after stages 1 and 2 of the machine learning architecture 100. As shown in FIG. Figure 8 As described in detail, line 506B is plotted using three trained model architectures. As shown, the trained model of line 504A produces significantly improved results than the trained model of line 502A because the model of line 504A can more accurately predict true positives. Also as shown, when the actual result is positive, the trained model of line 504A cannot predict true positives as accurately as the model of line 506A. However, as shown in reference Figure 5A As described in detail, the trained model of line 504B has more advantages than 506B. In addition, as shown, Figure 5C-5D Similar to reference Figure 5A-5B Similar features and functionality are described, but instead a true positive / false positive ratio of 1:100 is utilized.
[0113] Reference now Figure 6 , a block diagram of a machine learning architecture 600 according to an illustrative embodiment is shown. The machine learning architecture 600 is similar to the reference Figure 1ADetailed description of the invention is provided with similar features and functionality, specifically Phase 1 of training models (i.e., user model 602 and entity model 608). However, in some embodiments, once Phase 1 of training entity model 608 (i.e., baseline training) is completed, machine learning architecture 600 is configured to include personalized embedding vectors. The personalized embedding vectors are part of a subset of entity-specific parameters 618 that are similar to the reference embedding vectors. Figure 1A 118. However, instead of retraining the entity model in stage 2 using a portion of the closed dataset (e.g., closed dataset 114), the machine learning architecture 600 can directly generate output predictions using the personalized embedding vectors. In some embodiments, the machine learning architecture 600 outputs predictions similar to those described in detail in reference Figure 1A (specifically, output prediction generator 116) with similar features and functionality as described in detail.
[0114] However, the equation defining the output prediction generator of the machine learning architecture 600 may be defined as:
[0115] Phase 1:
[0116] Phase 2:
[0117] Wherein, the notation is similar to that of the output prediction generator 116, and the new notation is as follows:
[0118] P: Personalized entity model function
[0119] In some embodiments, the machine learning architecture 600 provides the Figure 1A Similar improvements described above. In some embodiments, the machine learning architecture 600 can provide improvements to the generated output predictions because the personalized embedding vectors can provide a direct impact on the generated output predictions. In addition, the personalized embedding vectors can provide a method for evaluating the output predictions without retraining and / or adding any models.
[0120] Reference now Figure 7 , according to an illustrative embodiment, a block diagram of a machine learning architecture 700 is shown. The machine learning architecture 700 is similar to the reference Figure 1AThe machine learning architecture 700 trains a personalized entity model (i.e., model P) with similar features and functionality as described in detail in the previous embodiment (specifically, stage 1 of training the model). However, instead of retraining the entity model in stage 2 using a different portion of the dataset (e.g., a portion of the closed dataset 114 that is specific to the entity), the machine learning architecture 700 trains another personalized entity model (i.e., model P). In some embodiments, the machine learning architecture 700 trains personalized entity models on a per-entity basis, such that each entity has its own separate model (e.g., personalized architecture 702, personalized architecture 704, and personalized architecture 706).
[0121] In some embodiments, once each model (i.e., user model, entity model, and personalized model) is trained, one or more processing circuits of the machine learning architecture 700 can employ each model to generate output predictions based on received data. The machine learning architecture 700 outputs predictions similar to the reference Figure 1A (specifically, output prediction generator 116) with similar features and functionality as described in detail.
[0122] However, the equation defining the output prediction generator of the machine learning architecture 700 may be defined as:
[0123] Phase 1:
[0124] Phase 2:
[0125] Wherein, the notation is similar to that of the output prediction generator 116, and the new notation is as follows:
[0126] · Personalized entity model for any given entity eid
[0127] ·θ U : Parameters associated with the personalized entity model
[0128] clsdData eid : The possibility of interaction between a user and a specific entity
[0129] In some embodiments, the machine learning architecture 700 requires a greater amount of resources than the machine learning architecture 100, which only requires a personalized dataset (i.e., the closed dataset 114) that includes all of the "private data" for each entity. Therefore, the need to train 3 models instead of 2 models may not be as advantageous from an efficiency and retention perspective.
[0130] Reference now Figure 8, shows an example hidden layer representation 800 of a neural network associated with a machine learning architecture 100 according to an illustrative embodiment. In some embodiments, the models can be visualized so that they can provide confidence and exploration opportunities for data pairs (e.g., user / entity). In some embodiments, the entity embedding vector can contain structure. For example, entities with similar vertical directions can be grouped together. In another example, unlike a method based purely on vertical similarity (where it can only be determined based on entities with exactly the same vertical direction), the position of the entity embedding vector on the example hidden layer representation 800 can be based on country and / or other factors. As shown, the example hidden layer representation 800 includes 10,000 embedding vectors. As shown, a small cluster [1] and another columnar structure [2] are also included. Therefore, clusters similar to [1] and [2] can be formed throughout the example hidden layer representation 1100, so that predictions can be made based on the position of each entity embedding vector.
[0131] Reference now Fig. 9 , shows an example hidden layer representation 900 of a neural network associated with the machine learning architecture 100 according to an illustrative embodiment. In some embodiments, the embedding vector may include a structure such that points grouped together may provide an indication of the characteristics of each point. In the example shown, open points that are located together may be entities, and filled points that are located together may be users. Thus, in this particular example, the location of the points provides insight into the user / entity relationship. One insight may include a particular angle depicted by the entity / user relationship. As shown, most entity / user relationships depict obtuse angles (i.e., negative inner products), which indicates that the entity / user relationships are unlikely to interact.
[0132] Reference now Fig.10 , shows an example hidden layer representation 1000 of a neural network associated with the machine learning architecture 100 according to an illustrative embodiment. The example hidden layer representation 1000 is similar to the reference Fig. 9 Similar features and functionality as described in detail. However, in this example shown, the filled points that are located together can be users of a given entity, the large open points can be a given entity, and the open points that are located together can be randomly sampled users. Therefore, in this specific example, the location of the points provides insight into the user / entity relationship. One insight can include a particular angle that the entity / user relationship is depicting. As shown, most entities / users of a given entity relationship depict an acute angle (i.e., a positive inner product), which indicates that the entities / users of the given entity relationship are likely to interact.
[0133] Fig.11A depiction of a computer system 1100 is shown, which may be used, for example, to implement an illustrative user device 240, an illustrative entity device 250, an illustrative analysis system 210, and / or various other illustrative systems described in the present disclosure. The computing system 1100 includes a bus 1105 or other communication component for communicating information, and a processor 1110 coupled to the bus 1105 for processing information. The computing system 1100 also includes a main memory 1115 (such as a random access memory (RAM) or other dynamic storage device) coupled to the bus 1105 for storing information and instructions to be executed by the processor 1110. The main memory 1115 may also be used to store location information, temporary variables, or other intermediate information during the execution of instructions by the processor 1110. The computing system 1100 may also include a read only memory (ROM) 1120 or other static storage device coupled to the bus 1105 for storing static information and instructions for the processor 1110. A storage device 1125 , such as a solid-state device, magnetic disk, or optical disk, is coupled to bus 1105 for persistent storage of information and instructions.
[0134] The computing system 1100 may be coupled to a display 1135, such as a liquid crystal display or an active matrix display, via the bus 1105 for displaying information to a user. An input device 1130, such as a keyboard including alphanumeric and other keys, may be coupled to the bus 1105 for communicating information and command selections to the processor 1110. In another embodiment, the input device 1130 has a touch screen display 1135. The input device 1130 may include a cursor control, such as a mouse, a trackball, or cursor direction keys, for communicating direction information and command selections to the processor 1110, and for controlling cursor movement on the display 1135.
[0135] In some implementations, computing system 1100 may include a communications adapter 1145, such as a network adapter. Communications adapter 1145 may be coupled to bus 1105 and may be configured to enable communications with a computing or communications network 1145 and / or other computing systems. In various illustrative implementations, communications adapter 1145 may be used to implement any type of network configuration, such as wired (e.g., via Ethernet), wireless (e.g., via WiFi, Bluetooth, etc.), pre-configured, ad hoc, LAN, WAN, etc.
[0136] According to various embodiments, in response to processor 1110 executing an instruction arrangement contained in main memory 1115, computing system 1100 may implement processes that implement the illustrative embodiments described herein. Such instructions may be read into main memory 1115 from another computer-readable medium, such as storage device 1125. Execution of the instruction arrangement contained in main memory 1115 causes computing system 1100 to perform the illustrative processes described herein. One or more processors in a multi-processing arrangement may also be used to execute the instructions contained in main memory 1115. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions to implement the illustrative embodiments. Thus, embodiments are not limited to any specific combination of hardware circuitry and software.
[0137] Despite Fig.11 An example processing system has been described in the specification, but the subject matter and implementation of the functional operations described in this specification may be performed using other types of digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them.
[0138] The embodiments of the subject matter and operations described in this specification may be performed using digital electronic circuits, or in computer software contained in tangible media, firmware or hardware (including the structures disclosed in this specification and their structural equivalents), or in a combination of one or more of them. The embodiments of the subject matter described in this specification may be implemented as one or more computer programs, that is, one or more modules of computer program instructions encoded on one or more computer storage media, for execution by a data processing device or for controlling the operation of a data processing device. Alternatively or additionally, the program instructions may be encoded on an artificially generated propagation signal, such as a machine-generated electrical, optical or electromagnetic signal, wherein the propagation signal is generated to encode information for transmission to a suitable receiver device for execution by a data processing device. The computer-readable storage medium may be a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination of one or more of them, or included therein. In addition, although the computer storage medium is not a propagation signal, the computer storage medium may be the source or destination of the computer program instructions encoded in the artificially generated propagation signal. The computer storage medium may also be one or more separate components or media (e.g., multiple CDs, disks or other storage devices), or included therein. Therefore, computer storage media is both tangible and non-transitory.
[0139] The operations described in this specification may be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.
[0140] The term "data processing device" or "computing device" includes all kinds of devices, equipment and machines for processing data, including, for example, a programmable processor, a computer, a system on a chip, or multiple systems on a chip, or a combination of the foregoing. The device may include special-purpose logic circuits, for example, an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). In addition to hardware, the device may also include code that creates an execution environment for the computer program in question, such as code that constitutes processor firmware, a protocol stack, a database management system, an operating system, a cross-platform runtime environment, a virtual machine, or a combination of one or more of them. The device and execution environment can implement a variety of different computing model infrastructures, such as network services, distributed computing, and grid computing infrastructures.
[0141] A computer program (also referred to as a program, software, software application, script, or code) may be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and it may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, object, or other unit suitable for use in a computing environment. A computer program may, but need not, correspond to a file in a file system. A program may be stored in a portion of a file that holds other programs or data (e.g., one or more scripts stored in a markup language document), in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files that store one or more modules, subroutines, or code portions). A computer program may be deployed to execute on one computer, or on multiple computers located at one site or distributed across multiple sites and interconnected by a communications network.
[0142] The processes and logic flows described in this specification can be performed by one or more programmable processors executing one or more computer programs to perform actions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuits (e.g., FPGAs (field programmable gate arrays) or (ASICs) application specific integrated circuits), and the apparatus can also be implemented as special purpose logic circuits.
[0143] For example, processors suitable for executing computer programs include general and special microprocessors, and any one or more processors of any kind of digital computer. Typically, the processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Typically, the computer will also include one or more mass storage devices (e.g., magnetic disks, magneto-optical disks, or optical disks) for storing data, or the computer will also be operably coupled to the one or more mass storage devices to receive data from the one or more mass storage devices, or to transfer data to the one or more mass storage devices, or both. However, the computer does not require such a device. In addition, the computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device (e.g., a universal serial bus (USB) flash drive), just to name a few. Devices suitable for storing computer program instructions and data include all forms of nonvolatile memory, media and storage devices, including, for example: semiconductor memory devices (e.g., EPROM, EEPROM and flash memory devices); magnetic disks (e.g., internal hard disks or removable disks); magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and memory may be supplemented by, or incorporated in, special purpose logic circuitry.
[0144] To provide interaction with a user, embodiments of the subject matter described in this specification may be performed using a computer having a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices may also be used to provide interaction with a user; for example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user may be received in any form, including sound, voice, or tactile input. In addition, a computer may interact with a user by sending documents to and receiving documents from a device used by the user; for example, by sending a web page to a web browser in response to a request received from a web browser on a user's client device.
[0145] Implementations of the subject matter described in this specification may be performed using a computing system that includes a back-end component (e.g., as a data server), or includes a middleware component (e.g., an application server), or includes a front-end component (e.g., a client computer with a graphical user interface or a web browser through which a user can interact with implementations of the subject matter described in this specification), or any combination of one or more such back-end, middleware, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), inter-networks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks).
[0146] A computing system may include a client and a server. The client and the server are usually remote from each other and usually interact through a communication network. The relationship between the client and the server occurs by means of computer programs running on each computer and having a client-server relationship with each other. In some embodiments, the server transmits data (e.g., an HTML page) to a client device (e.g., in order to display data to a user interacting with the client device and receive user input from it). Data generated at the client device (e.g., the result of the user interaction) can be received from the client device at the server.
[0147] In some illustrative embodiments, the features disclosed herein can be implemented on a smart TV module (or a connected TV module, a hybrid TV module, etc.), which may include a processing circuit configured to integrate an Internet connection with a more traditional TV program source (e.g., received via cable, satellite, air, or other signals). The smart TV module may be physically incorporated into a TV set, or may include separate devices such as a set-top box, a blue-ray or other digital media player, a game console, a hotel TV system, and other accompanying devices. The smart TV module may be configured to allow viewers to search and find videos, movies, photos, and other content on the web, on local cable TV channels, on satellite TV channels, or stored on a local hard drive. A set-top box (STB) or a set-top unit (STU) may include an information appliance device, which may include a tuner and be connected to a TV set and an external signal source, converting the signal into content, and then displaying it on a TV screen or other display device. The Smart TV module can be configured to provide a home screen or top-level screen that includes icons for multiple different applications, such as a web browser and multiple streaming services (e.g., Netflix, Vudu, Hulu, Disney+, etc.), connected cable or satellite media sources, other network "channels", etc. The Smart TV module can also be configured to provide an electronic program guide to the user. A companion application of the Smart TV module can be operated on the mobile computing device to provide the user with additional information about available programming, allow the user to control the Smart TV module, etc. In alternative embodiments, these features can be implemented on a laptop or other personal computer, a smartphone, other mobile phone, a handheld computer, a smart watch, a tablet PC, or other computing device.
[0148] Although this specification contains many specific implementation details, these details should not be interpreted as limitations on any invention or the scope that may be claimed, but should be interpreted as descriptions of features specific to specific embodiments of specific inventions. Certain features described in the context of separate embodiments in this specification may also be performed in combination in a single embodiment. On the contrary, the various features described in the context of a single embodiment can also be performed separately in multiple embodiments or in any suitable sub-combination. In addition, although the features may be described above as appearing in certain combinations and even required to be protected as such at the beginning, in some cases one or more features from the claimed combination may be deleted from the combination, and the claimed combination may point to a sub-combination or a variation of the sub-combination. In addition, the features described with respect to a specific title may be utilized with respect to and / or in combination with illustrative embodiments described under other titles; the titles provided are included only for readability purposes and should not be interpreted as limiting any features provided with respect to these titles.
[0149] Similarly, although operations are described in a particular order in the accompanying drawings, this should not be understood as requiring that the operations be performed in the particular order or sequence shown, or that all of the illustrated operations be performed, in order to achieve the desired results. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of various system components in the above-described embodiments should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product, or packaged into multiple software products contained on tangible media.
[0150] Thus, particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. In some cases, the actions recited in the claims can be performed in a different order and still achieve the desired results. Additionally, the processes described in the accompanying drawings do not necessarily require the particular order shown or sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing may be advantageous.
Claims
1. A method for training a machine learning model, the method comprising: receiving, by one or more processing circuits, a data set; determining, by the one or more processing circuits, a first portion of the data set associated with a plurality of entities, wherein the first portion of the data set comprises an open data set including features associated with the plurality of entities; training, by the one or more processing circuits, an entity model using the first portion of the data set, wherein the entity model includes one or more parameters associated with the plurality of entities and wherein the entity model is trained to recognize one or more patterns in subsequently received data; determining, by the one or more processing circuits, a second portion of the data set associated with a first subset of entities of the plurality of entities, wherein the second portion of the data set comprises a closed data set that is specific to the first subset of entities; determining, by the one or more processing circuits, a second subset of entities, wherein the second subset of entities does not include any entity in the first subset of entities, and wherein the second portion of the data set is not shared with the second subset of entities; freezing, by the one or more processing circuits, one or more parameters of the entity model associated with a first portion of the data set associated with the plurality of entities, such that the one or more parameters of the entity model remain fixed during subsequent training of the entity model using a second portion of the data set, and such that one or more non-frozen parameters of the trained entity model are trained to output predictions specific to the first subset of entities and not associated with the second subset of entities; and The entity model is trained, by the one or more processing circuits, using a second portion of the data set.
2. The method according to claim 1, further comprising: determining, by the one or more processing circuits, a third portion of the data set associated with a plurality of users, wherein the plurality of users includes a set of user identifiers and a plurality of user information; and training, by the one or more processing circuits, a user model using a third portion of the data set, wherein the user model is trained to output predictions from subsequently received data; Wherein training each of the user model and the entity model further comprises configuring at least one neural network.
3. The method according to claim 2, further comprising: receiving, by the one or more processing circuits, an input data set; inputting, by the one or more processing circuits, the input data set into the user model and the entity model; and generating, by the one or more processing circuits, an output prediction based on the trained user model and the trained entity model, wherein the output prediction is specific to the first subset of entities, and wherein the output prediction is an accuracy measure, the accuracy measure comprising a value; in, Generating the output prediction is also based on utilizing a user embedding vector generated by the user model and an entity embedding vector generated by the entity model.
4. The method according to claim 1 or 2, further comprising: determining, by the one or more processing circuits, a fourth portion of the data set associated with a third entity subset of the plurality of entities, the fourth portion of the data set comprising a closed data set specific to the third entity subset, wherein the third entity subset does not include any entity in the first entity subset or the second entity subset; and training, by the one or more processing circuits, a second entity model using a fourth portion of the data set, the second entity model being trained to output predictions specific to a third subset of entities, the second entity model being based on the entity model trained and frozen using the first portion of the data set, in, A fourth portion of the dataset utilized in training the second entity model does not include any data from the second portion of the dataset utilized in training the entity model.
5. The method according to claim 1 or 2, wherein: The first portion of the data set does not include any data from the second portion of the data set, such that each subset of the plurality of entities includes a particular data set.
6. The method according to claim 1 or 2, further comprising: configuring, by the one or more processing circuits, a first neural network associated with the user model; and A second neural network associated with the entity model is configured by the one or more processing circuits.
7. A system for training a machine learning model, comprising: At least one processing circuit configured to: Receive a dataset; determining a first portion of the data set associated with a plurality of entities, wherein the first portion of the data set comprises an open data set including features associated with the plurality of entities; training an entity model using the first portion of the data set, wherein the entity model includes one or more parameters associated with the plurality of entities and wherein the entity model is trained to recognize one or more patterns in subsequently received data; determining a second portion of the data set associated with a first subset of entities of the plurality of entities, wherein the second portion of the data set comprises a closed data set that is specific to the first subset of entities; determining a second subset of entities, wherein the second subset of entities does not include any entity in the first subset of entities, and wherein the second portion of the data set is not shared with the second subset of entities; freezing one or more parameters of the entity model associated with a first portion of the data set associated with the plurality of entities, such that the one or more parameters of the entity model remain fixed during subsequent training of the entity model using a second portion of the data set, and such that one or more non-frozen parameters of the trained entity model are trained to output predictions specific to the first subset of entities and not associated with the second subset of entities; and The entity model is trained using a second portion of the data set.
8. The system according to claim 7, wherein: The at least one processing circuit is further configured to: determining a third portion of the data set associated with a plurality of users, wherein the plurality of users includes a set of user identifiers and a plurality of user information; and A user model is trained using a third portion of the data set, wherein the user model is trained to output predictions from subsequently received data.
9. The system according to claim 8, wherein: The at least one processing circuit is further configured to: Receive an input data set; inputting the input data set into the user model and the entity model; and An output prediction is generated based on the trained user model and the trained entity model, wherein the output prediction is specific to the first subset of entities, and wherein the output prediction is an accuracy measure, the accuracy measure comprising a value.
10. The system according to claim 7 or 8, wherein: The at least one processing circuit is further configured to: determining a fourth portion of the data set associated with a third entity subset of the plurality of entities, the fourth portion of the data set comprising a closed data set specific to the third entity subset, wherein the third entity subset does not include any entity in the first entity subset or the second entity subset; and A second entity model is trained using a fourth portion of the data set, the second entity model being trained to output predictions specific to a third subset of entities, the second entity model being based on the entity model trained and frozen using the first portion of the data set.
11. The system according to claim 7 or 8, wherein: The at least one processing circuit is further configured to: determining a fifth portion of the data set associated with a fourth entity subset of the plurality of entities, wherein the fourth entity subset does not include any entity in the first entity subset, the second entity subset, or the third entity subset; and A third entity model is trained using a fifth portion of the data set, the third entity model being based on the entity model trained and frozen using the first portion of the data set.
12. The system according to claim 7 or 8, wherein: The first portion of the data set does not include any data from the second portion of the data set, such that each subset of the plurality of entities includes a particular data set.
13. One or more computer-readable storage media having stored thereon instructions that, when executed by at least one processing circuit, cause the at least one processing circuit to perform operations comprising: Receive a dataset; determining a first portion of the data set associated with a plurality of entities, wherein the first portion of the data set comprises an open data set including features associated with the plurality of entities; training an entity model using the first portion of the data set, wherein the entity model is trained to recognize one or more patterns in subsequently received data, wherein the entity model includes one or more parameters associated with the plurality of entities; determining a second portion of the data set associated with a first subset of entities of the plurality of entities, wherein the second portion of the data set comprises a closed data set that is specific to the first subset of entities; determining a second subset of entities, wherein the second subset of entities does not include any entity in the first subset of entities, and wherein the second portion of the data set is not shared with the second subset of entities; freezing one or more parameters of the entity model associated with a first portion of the data set associated with the plurality of entities, such that the one or more parameters of the entity model remain fixed during subsequent training of the entity model using a second portion of the data set, and such that one or more non-frozen parameters of the trained entity model are trained to output predictions specific to the first subset of entities and not associated with the second subset of entities; and The entity model is trained using a second portion of the data set.
14. The one or more computer-readable storage media of claim 13, the operations further comprising: determining a third portion of the data set associated with a plurality of users, wherein the plurality of users includes a set of user identifiers and a plurality of user information; and training a user model using a third portion of the data set, wherein the user model is trained to output predictions from subsequently received data; The operations also include: Receive an input data set; inputting the input data set into the user model and the entity model; and An output prediction is generated based on the trained user model and the trained entity model, wherein the output prediction is specific to the first subset of entities, and wherein the output prediction is an accuracy measure, the accuracy measure comprising a value.
15. The one or more computer-readable storage media of claim 13 or 14, the operations further comprising: configuring a first neural network associated with the user model; and A second neural network is configured to be associated with the entity model.
Citation Information
Patent Citations
Transfer learning without local data export in multi-node machine learning
US20190325350A1