Cooperative information generation adversarial network for efficient data co-clustering

By using the Collaborative Information Generative Adversarial Network (CI-GAN) model, the problem of processing unordered and unstructured data in recommendation systems is solved, enabling more accurate understanding of user preferences and personalized program recommendations.

CN114514537BActive Publication Date: 2025-11-04SAMSUNG ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080069028.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-24
Filing Date
2020-09-29
Publication Date
2025-11-04
Estimated Expiration
2040-09-29

AI Technical Summary

Technical Problem

Existing recommendation systems struggle to effectively handle users' unordered and unstructured TV program viewing habits, resulting in insufficient recommendation accuracy.

Method used

The Cooperative Information Generative Adversarial Network (CI-GAN) model is adopted. The first GAN and the second GAN synchronously and collaboratively cluster row vectors and column vectors to generate a collaborative clustering correlation matrix. The potential interrelationships in the data matrix are reconstructed by using the mutual information cost function and the distribution function.

Benefits of technology

This improves the accuracy of the recommendation system, enabling it to better understand user preferences and provide more personalized program recommendations and advertising content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114514537B_ABST
    Figure CN114514537B_ABST
Patent Text Reader

Abstract

A method implemented by one or more computing systems includes, by the one or more computing systems: accessing a first data matrix including a plurality of row data and a plurality of column data. The method further includes providing, to a first generative adversarial network (GAN), a first data input including a plurality of row vectors corresponding to the plurality of row data, and providing, to a second GAN, a second data input including a plurality of column vectors corresponding to the plurality of column data. The method further includes generating, by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN, a co-clustering correlation matrix based at least in part on the plurality of row vectors and the plurality of column vectors. The method further includes that the co-clustering correlation matrix includes co-clustering associations between the plurality of row data and the plurality of column data.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates generally to data co-clustering, and more particularly, to co-informatic generative adversarial networks (CI-GAN) for efficient data co-clustering. BACKGROUND

[0002] Recommendation systems typically include applications that facilitate users in locating items of interest for consumption, often in a quasi-personalized manner. One type of recommendation system can include a television (TV) recommendation system that can be used to provide recommendations for certain movies, TV series, news programs, sports telecasts, etc. that a particular user can be interested in watching, for example. Further, as an increasing number of program genres, program channels, video sharing platform publisher channels, and other features are increasingly available to users, deep learning techniques can be applied to recommendation systems to improve user viewing recommendations. In particular, while traditional recommendation systems can be suitable for modeling, for example, sequential and structured TV program data, actual TV program viewing habits of users can include non-sequential and / or unstructured data in many cases (e.g., due to user viewing preferences changing or evolving over time). It can be useful to provide one or more data co-clustering techniques to improve recommendation systems. In recent years, artificial intelligence systems are used in various fields. In particular, an artificial intelligence system is a system that trains, determines, and becomes intelligent by itself. As the use of artificial intelligence systems, the recognition rate is improved, the preference of a user can be more accurately understood, and thus, existing rule-based systems are gradually being replaced by artificial intelligence systems based on deep learning. SUMMARY

[0003] SOLUTION TO THE PROBLEM

[0004] According to one aspect of the present disclosure, a method is provided. A method includes, by one or more computing systems: accessing a first data matrix including a plurality of row data and a plurality of column data, providing, to a first generative adversarial network (GAN), a first data input including a plurality of row vectors corresponding to the plurality of row data from the first data matrix, providing, to a second GAN, a second data input including a plurality of column vectors corresponding to the plurality of column data from the first data matrix, and generating, by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN, a co-clustering correlation matrix based at least in part on the plurality of row vectors and the plurality of column vectors, wherein the co-clustering correlation matrix includes co-clustering associations between the plurality of row data and the plurality of column data of the first data matrix.

[0005] The first GAN and the second GAN can be part of a co-information generative adversarial network (CI-GAN) combiner model architecture, where the CI-GAN combiner model architecture can further include a distribution function configured to reconstruct one or more co-cluster labels provided as third data input to the first GAN and the second GAN.

[0006] The first GAN and the second GAN can be part of a co-information generative adversarial network (CI-GAN) autoencoder model architecture, where the CI-GAN autoencoder model architecture can further include a distribution function configured to reconstruct a plurality of row data and a plurality of column data based on a plurality of output row data generated by the first GAN and a plurality of output column data generated by the second GAN.

[0007] Generating the co-clustered correlation matrix by synchronously co-clustered the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN can include determining one or more latent interrelationships between the plurality of row data and the plurality of column data of the first data matrix based at least in part on a mutual information cost function.

[0008] Generating the co-clustered correlation matrix by synchronously co-clustered the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN can include determining a lower bound defining an objective function corresponding to mutual information between the co-clustered associations and the plurality of row data and the plurality of column data of the first data matrix.

[0009] Generating the co-clustered correlation matrix by synchronously co-clustered the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN can include maximizing mutual information between the co-clustered associations and the plurality of row data and the plurality of column data of the first data matrix.

[0010] Generating the co-clustered correlation matrix by synchronously co-clustered the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN can include separately determining one or more disentangled representations of the plurality of row data and the plurality of column data.

[0011] Generating the co-clustered correlation matrix by synchronously co-clustered the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN can include simultaneously determining one or more disentangled representations of the plurality of row data and the plurality of column data.

[0012] Simultaneously determining one or more disentangled representations of the plurality of row data and the plurality of column data can include determining a joint disentanglement representation for deciphering an interplay data structure in the co-clustered correlation matrix.

[0013] The first data matrix can include one or more of: user behavior data as a plurality of row data and user content viewing data as a plurality of column data, video data as a plurality of row data and video-caption data as a plurality of column data, biological genetic data as a plurality of row data and biological condition data as a plurality of column data, movie data as a plurality of row data and music data as a plurality of column data, or image data as a plurality of row data and audible data as a plurality of column data.

[0014] According to one aspect of the disclosure, a system is provided. A system comprising one or more non-transitory computer-readable storage media including instructions; and one or more processors coupled to the storage media, the one or more processors configured to execute the instructions to access a first data matrix comprising a plurality of row data and a plurality of column data, provide, to a first generative adversarial network (GAN), a first data input comprising a plurality of row vectors corresponding to the plurality of row data from the first data matrix, provide, to a second GAN, a second data input comprising a plurality of column vectors corresponding to the plurality of column data from the first data matrix, and generate, by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN, a co-clustering correlation matrix based at least in part on the plurality of row vectors and the plurality of column vectors, wherein the co-clustering correlation matrix comprises co-clustering associations between the plurality of row data and the plurality of column data of the first data matrix.

[0015] The first GAN and the second GAN can be part of a co-information generative adversarial network (CI-GAN) combiner model architecture, wherein the CI-GAN combiner model architecture further comprises a distribution function configured to reconstruct one or more co-clustering labels provided as a third data input to the first GAN and the second GAN.

[0016] The first GAN and the second GAN can be part of a co-information generative adversarial network (CI-GAN) autoencoder model architecture, wherein the CI-GAN autoencoder model architecture further comprises a distribution function configured to reconstruct the plurality of row data and the plurality of column data based on a plurality of output row data generated by the first GAN and a plurality of output column data generated by the second GAN.

[0017] The instructions to generate, by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN, the co-clustering correlation matrix can comprise instructions to determine one or more latent interrelationships between the plurality of row data and the plurality of column data of the first data matrix based at least in part on a mutual information cost function.

[0018] The instructions to generate a co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN can include instructions to determine a lower bound that defines an objective function corresponding to mutual information between the co-clustered associations and the plurality of row data and the plurality of column data of the first data matrix.

[0019] The instructions to generate a co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN can include instructions to maximize mutual information between the co-clustered associations and the plurality of row data and the plurality of column data of the first data matrix.

[0020] The instructions to generate a co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN can include instructions to separately determine one or more disentangled representations of the plurality of row data and the plurality of column data.

[0021] The instructions to generate a co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN can include instructions to simultaneously determine one or more disentangled representations of the plurality of row data and the plurality of column data.

[0022] The instructions to simultaneously determine one or more disentangled representations of the plurality of row data and the plurality of column data can include instructions to determine a joint disentanglement representation for decrypting an interaction data structure in the co-clustered correlation matrix.

[0023] According to one aspect of the disclosure, a non-transitory computer-readable medium is provided. A non-transitory computer-readable medium comprising instructions that, when executed by one or more processors of a computing system, cause the one or more processors to access a first data matrix comprising a plurality of row data and a plurality of column data, provide a first data input comprising a plurality of row vectors corresponding to the plurality of row data from the first data matrix to a first generative adversarial network (GAN), provide a second data input comprising a plurality of column vectors corresponding to the plurality of column data from the first data matrix to a second GAN, and generate a co-clustered correlation matrix based at least in part on the plurality of row vectors and the plurality of column vectors by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN, wherein the co-clustered correlation matrix comprises co-clustered associations between the plurality of row data and the plurality of column data of the first data matrix. BRIEF DESCRIPTION OF DRAWINGS

[0024] Figure 1 An example co-information generative adversarial network (CI-GAN) co-clustering system is shown.

[0025] Figure 2aEmbodiments of a combiner CI-GAN model for efficient data co-clustering are shown.

[0026] Figure 2b Embodiments of an autoencoder CI-GAN model for efficient data co-clustering are shown.

[0027] Figure 3a An example diagram showing efficient data clustering using CI-GAN is shown.

[0028] Figure 3b An example diagram showing efficient data clustering using CI-GAN is shown.

[0029] Figures 3c-3q An example training phase and additional run examples using CI-GAN for efficient data co-clustering are shown.

[0030] Figures 3r-3t An example inference phase and additional run examples using CI-GAN for efficient data co-clustering are shown.

[0031] Figure 4 A flowchart showing a method of providing a co-information generative adversarial network (CI-GAN) for data co-clustering is shown.

[0032] Figure 5 An example computer system is shown.

[0033] Figure 6 A diagram showing an example artificial intelligence (AI) architecture is shown. DETAILED DESCRIPTION

[0034] According to the presently disclosed embodiments, the present embodiments are directed to providing a co-information generative adversarial network (CI-GAN) for efficient data co-clustering. In particular embodiments, a CI-GAN model can access a first data matrix including a plurality of row data and a plurality of column data. In particular embodiments, the CI-GAN model can receive a data matrix of data that can include, for example, user behavior data as the plurality of row data and user content viewing data as the plurality of column data, video data as the plurality of row data and video description data as the plurality of column data, biological genetic data as the plurality of row data and biological condition data as the plurality of column data, movie data as the plurality of row data and music data as the plurality of column data, or image data as the plurality of row data and audible data as the plurality of column data. It should be understood that at least some dimensions (e.g., rows and columns) are interchangeable.

[0035] In particular embodiments, the CI-GAN model can receive, for a first GAN (e.g., a first InfoGAN), a first data input comprising a plurality of row vectors corresponding to a plurality of row data from a first data matrix. In particular embodiments, the CI-GAN model can also receive, for a second GAN (e.g., a second InfoGAN), a second data input comprising a plurality of column vectors corresponding to a plurality of column data from the first data matrix. In particular embodiments, the first GAN and the second GAN can be part of a CI-GAN combiner model architecture, where the CI-GAN combiner model architecture can also include a distribution function configured to reconstruct one or more co-clustering labels provided as a third data input to the first GAN and the second GAN. In particular embodiments, in combination or alternatively, the first GAN and the second GAN can be part of a CI-GAN autoencoder model architecture, where the CI-GAN autoencoder model architecture can also include a distribution function configured to reconstruct the plurality of row data and the plurality of column data based on a plurality of output row data generated by the first GAN and a plurality of output column data generated by the second GAN.

[0036] In particular, since the data clustering of row data and column data is implemented independently by the first GAN (e.g., a first InfoGAN) and the second GAN (e.g., a second InfoGAN) of the CI-GAN model, the CI-GAN model can allow maximizing the mutual information (e.g., a measure indicating how much one random variable can suggest another random variable) between the co-clustering associations and the input row data and column data in several distinct ways. In this way, as the present CI-GAN model for efficient data co-clustering is further enhanced with a mutual information objective function, it can be well suited for jointly learning dual data representations across a variety of technical applications. While the present embodiments can be discussed below primarily with respect to TV shows and recommendation systems, it should be understood that the present technology can be applied to any of a variety of systems or applications that can utilize two-dimensional data, such as TV shows and recommendation systems that can utilize multi-dimensional data, genomic data analysis, medical systems and analysis, GPS systems, indoor positioning and tracking systems, search engine systems, text document and web page search systems, oil and gas exploration and refining systems, database search systems, and / or other applications and systems.

[0037] Figure 1 An example co-information generative adversarial network (CI-GAN) co-clustering system 100 is shown. As Figure 1As shown, the CI-GAN system 100 can include a program analysis system 102, one or more databases 104, 106, and a TV program and recommended content subnetwork 108. In particular embodiments, the program analysis system 102 can include a cloud-based clustering computing architecture or other similar computing architecture that can receive one or more user automatic content recognition (ACR) user viewing data 110 that can be provided by first party or third party sources and can provide TV program content and advertising content to one or more client devices (e.g., a TV, a standalone monitor, a desktop computer, a laptop computer, a tablet computer, a mobile phone, a wearable electronic device, a voice-controlled personal assistant device, a car display, a gaming system, a home appliance, or other similar multimedia electronic device) that are suitable for displaying and / or playing back program and advertising content. In addition, the program analysis system 102 can be used to process and manage various analytics and / or data intelligence, such as TV program analytics, network analytics, user profile data, user payment data, user privacy preferences, and the like. For example, in particular embodiments, the program analysis system 102 can include a platform-as-a-service (PaaS) architecture, a software-as-a-service (SaaS) architecture, and an infrastructure-as-a-service (IaaS), or other various cloud-based clustering computing architectures.

[0038] In particular embodiments, as Figure 1 Further depicted, the program analysis system 102 can include a pre-processing function block 112, a CI-GAN modeling block 114, and a co-clustered correlation matrix function block 116. In particular embodiments, the pre-processing function block 112, the CI-GAN modeling block 114, and the co-clustered correlation matrix function block 116 can each include, for example, a computing engine. In particular embodiments, the pre-processing function block 112 can receive the ACR user viewing data 110, which can include, for example, particular program content (e.g., TV programs) that one or more particular users or a subgroup of users recently viewed. For example, the ACR user viewing data 110 can include metadata associated with the recently viewed program content (e.g., TV programs), a particular time slot (e.g., a daytime time) within which the recently viewed program content (e.g., TV programs) was viewed, and a program channel on which the program content (e.g., TV programs) was viewed, an identification of the recently viewed program content (e.g., TV programs).

[0039] In particular embodiments, the pre-processing function block 112 can then interface with the content database 104 to correlate the recently viewed program content included in the ACR user viewing data 110 with TV program content stored by the database 104. For example, the TV program content stored by the database 104 can include, for example, user or sub-group profile data, program genre data, program category data, program cluster category group data, or other TV program content or metadata that can be stored by the database 104. In particular embodiments, the ACR user viewing data 110 can include time series data expressed in hourly context, daily context, and / or day-hour context. For example, in particular embodiments, the time series ACR user viewing data 110 can be received, for example, at each predetermined time slot of each time period.

[0040] In particular embodiments, a CI-GAN modeling module 114 can be provided for performing efficient data co-clustering and generating one or more TV program recommendations or advertisements based thereon. In particular embodiments, as will be described below with respect to Figure 2a In more detail, the CI-GAN modeling module 114 can include a first GAN and a second GAN that can be part of a CI-GAN combiner model architecture for performing efficient data co-clustering. In particular embodiments, in conjunction or alternatively, and as will be described below with respect to Figure 2b In more detail, the first GAN and the second GAN can be part of a CI-GAN autoencoder model architecture that performs efficient data co-clustering. In particular embodiments, the CI-GAN modeling module 114 can then provide the determined co-clustering associations to a co-clustering correlation matrix function block 116 for further processing and management.

[0041] For example, as Figure 1 As further shown, the program analysis system 102 can provide the determined co-clustering associations based on the CI-GAN model to the database 106. In particular embodiments, as Figure 1 As further depicted, a network-based content orchestrator 118 can retrieve the determined co-clustering associations from the database 106. The content orchestrator 118 can then store the determined co-clustering associations along with TV programs and recommended and advertisement content to be viewed in the programs in a database 120. In particular embodiments, based on the determined co-clustering associations, the content orchestrator 118 can then provide the TV programs and recommended content 122 to, for example, end-user client devices for user viewing.

[0042] Figure 2aAn embodiment of a combiner CI-GAN model 200A for efficient data co-clustering is shown in accordance with the presently disclosed embodiments. In particular embodiments, the combiner CI-GAN model 200A can be provided for learning and determining underlying inter-relationships between row data and column data of a data matrix based on, for example, a mutual information cost function. For example, in particular embodiments, the combiner CI-GAN model 200A can co-cluster user behavior data and user content viewing data in sync to provide, for example, one or more program recommendations or advertisements to an end user. In particular embodiments, as shown in Figure 2a the combiner CI-GAN model 200A can include a first GAN 202A and a second GAN 202B. For example, the first GAN 202A can be provided to cluster row data of a data matrix while the second GAN 202B can be provided to cluster column data of the data matrix (or vice versa). In particular embodiments, one or more data matrices that can be co-clustered by the combiner CI-GAN model 200A can include, for example, user behavior data as row data and user content viewing data as column data, video data as row data and video description data as column data, biological genetic data as row data and biological condition data as column data, movie data as row data and music data as column data, or image data as row data and audible data as column data, or other similar data that exist in two disparate data dimensions. For example, in particular embodiments, the first GAN 202A and the second GAN 202B can be represented as:

[0043] Row: Min Gr,Qr Max Dr V(D r , G r )-L I (G r , Q r ) (Equation 1), and

[0044] Column: Min Gc,Qc Max Dc V(D r , G r )-L I (G c , Q c ) (Equation 2).

[0045] In particular embodiments, as Figure 2aFurther depicted and in view of Equation 1 and Equation 2, the first GAN 202A can include a row generator 204A (e.g., “Gr”), a row discriminator 206A (e.g., “Dr”), and a row data distribution 208A (e.g., “Qr”). Similarly, the second GAN 202B can include a column generator 204B (e.g., “Gc”), a column discriminator 206B (e.g., “Dc”), and a column data distribution 208B (e.g., “Qc”). In particular embodiments, the row generator 204A (e.g., “Gr”) and the column generator 204B (e.g., “Gc”) can each include, for example, a neural network that can model a transformation function in which the row generator 204A (e.g., “Gr”) and the column generator 204B (e.g., “Gc”) can be adapted to receive any row input data (e.g., row random variable data) and column input data (e.g., column random variable data) and generate row random variable output data and column random variable output data according to one or more target distribution functions. For example, in particular embodiments, the row generator 204A (e.g., “Gr”) can receive a row data input vector cr, a row random variable zr, and a co-cluster label d. Similarly, the column generator 204B (e.g., “Gc”) can receive a column data input vector cc, a column random variable zc, and the co-cluster label d. In particular embodiments, the co-cluster label d can include, for example, a categorical vector of size cr x cc (e.g., because the row data input vector cr and the column data input vector cc can represent a pair of clusters or co-clusters). In one example, the co-cluster label d can be represented as:

[0046]

[0047] In particular embodiments, referring to Equation 3 and Figure 2aThe row generator 204A (e.g., “Gr”) and the column generator 204B (e.g., “Gc”) can then generate a fake row data vector xr and a fake column data vector xc, respectively. For example, the fake row data vector xr and the fake column data vector xc can represent one or more fake data samples (e.g., fake unlabeled data samples) of the row data input vector cr and the column data input vector cc, respectively. In particular embodiments, the row generator 204A (e.g., “Gr”) can provide the fake row data vector xr to the row discriminator 206A (e.g., “Dr”), and the column generator 204B (e.g., “Gc”) can provide the fake column data vector xc to the column discriminator 206B (e.g., “Dc”). For example, as further depicted, the row discriminator 206A (e.g., “Dr”) can receive the fake row data vector xr and the true row data vector Xr, and the column discriminator 206B (e.g., “Dc”) can receive the fake column data vector xc and the true column data vector Xc. In particular embodiments, the row discriminator 206A (e.g., “Dr”) can be trained with the true row data vector Xr before receiving the fake row data vector xr, and the column discriminator 206B (e.g., “Dc”) can be trained with the true column data vector Xc before receiving the fake column data vector xc. For example, the true row data vector Xr and the true column data vector Xc can represent one or more true data samples (e.g., true unlabeled data samples) of the row data input vector cr and the column data input vector cc, respectively.

[0048] In particular, in particular embodiments, as part of the training of the first GAN 202A and the second GAN 202B, a row generator 204A (e.g., “Gr”) can be provided to “fool” the row discriminator 206A (e.g., “Dr”) and a column generator 204B (e.g., “Gc”) can be provided to “fool” the column discriminator 206B (e.g., “Dc”). For example, such “fooling” techniques can be utilized to maximize the final classification error between, for example, a fake row data vector xr- and a real row data vector Xr-, and between a fake column data vector xcand a real column data vector Xc, respectively. Conversely, for example, the row discriminator 206A (e.g., “Dr”) and the column discriminator 206B (e.g., “Dc”) can be provided to detect fake row data vector xr- and fake column data vector xc to minimize the final classification error between, for example, a fake row data vector xr- and a real row data vector Xr-, and between a fake column data vector xcand a real column data vector Xc, respectively. For example, in particular embodiments, the row generator 204A (e.g., “Gr”) and the column generator 204B (e.g., “Gc”) can generate fake data to try to “fool” the row discriminator 206A (e.g., “Dr”) and the column discriminator 206B (e.g., “Dc”), respectively. The row discriminator 206A (e.g., “Dr”) and the column discriminator 206B (e.g., “Dc”) can receive the fake data as input and then output true (e.g., if the fake data fools the row discriminator 206A (e.g., “Dr”) and the column discriminator 206B (e.g., “Dc”) to determine that the fake data is real data) or false (e.g., if the row discriminator 206A (e.g., “Dr”) and the column discriminator 206B (e.g., “Dc”) determine that the fake data is not real data). If the row discriminator 206A (e.g., “Dr”) and the column discriminator 206B (e.g., “Dc”) output false, then the row generator 204A (e.g., “Gr”) and the column generator 204B (e.g., “Gc”) are penalized and trained to produce better fake data for “fooling” the row discriminator 206A (e.g., “Dr”) and the column discriminator 206B (e.g., “Dc”). Thus, the row discriminator 206A (e.g., “Dr”) and the column discriminator 206B (e.g., “Dc”) can also provide a true / false (e.g., “T / F”) output that indicates the final classification error between, for example, a fake row data vector xr- and a real row data vector Xr-, and between a fake column data vector xcand a real column data vector Xc, respectively.

[0049] In particular embodiments, the row discriminator 206A (e.g., “Dr”) and the column discriminator 206B (e.g., “Dc”) can then provide outputs to the row clustering distribution model 208A (e.g., Qr), the column clustering distribution model 208B (e.g., Qc), and the auxiliary clustering distribution model 210 (e.g., Qco). In particular embodiments, the row clustering distribution model 208A (e.g., Qr) and the column clustering distribution model 208B (e.g., Qc) can generate a row data output vector cr’ and a column data output vector cc’. The row data output vector cr’ and the column data output vector cc’ can each include, for example, one-dimensional clusters of row data of the row data input vector cr and column data of the column data input vector cc, respectively. In particular embodiments, the auxiliary clustering distribution model 210 (e.g., Qco) can reconstruct the one or more co-clusters d and generate one or more co-cluster labels d’.

[0050] Figure 2b An embodiment of an autoencoder CI-GAN model 200B for efficient data co-clustering is shown in accordance with the presently disclosed embodiments. In particular embodiments, the autoencoder CI-GAN model 200B can be provided for determining a lower bound that defines an objective function corresponding to mutual information (e.g., I) between row data and column data of a data matrix and co-cluster associations generated by the CI-GAN. In particular, the autoencoder CI-GAN model 200B can maximize mutual information (e.g., I) between original data row data and column data and co-cluster associations generated by the CI-GAN for improved model optimization. In one example, the maximization of mutual information I can be represented as:

[0051] max I(X r′ , X c′ , c r′ , c c′ ) (Equation 4).

[0052] In particular, in particular embodiments, the autoencoder CI-GAN model 200B can synchronize co-cluster user behavior data and user content viewing data to provide, for example, one or more program recommendations or advertisements to an end user. In particular embodiments, as Figure 2bAs shown, the autoencoder CI-GAN model 200B can include a first GAN 212A and a second GAN 212B. For example, the first GAN 212A can be provided to cluster row data of a data matrix, while the second GAN 212A can be provided to cluster column data of the data matrix (or vice versa). In particular embodiments, one or more data matrices that can be co-clustered by the autoencoder CI-GAN model 200B can include, for example, user behavior data as row data and user content viewing data as column data, video data as row data and video description data as column data, biological genetic data as row data and biological condition data as column data, movie data as row data and music data as column data, or image data as row data and audible data as column data, or other similar data that exist in two disparate data dimensions.

[0053] In particular embodiments, as Figure 2b As further depicted, the first GAN 212A can include a row generator 214A (e.g., “Gr”), a row discriminator 216A (e.g., “Dr”), and a row data distribution 218A (e.g., “Qr”). Similarly, the second GAN 212B can include a column generator 214B (e.g., “Gc”), a column discriminator 216B (e.g., “Dc”), and a column data distribution 218B (e.g., “Qc”). In particular embodiments, the row generator 214A (e.g., “Gr”) and the column generator 214B (e.g., “Gc”) can each include, for example, a neural network that can model a transformation function in which the row generator 214A (e.g., “Gr”) and the column generator 214B (e.g., “Gc”) can be adapted to receive any row input data (e.g., row random variable data) and column input data (e.g., column random variable data) and generate row random variable output data and column random variable output data according to one or more target distribution functions. For example, in particular embodiments, the row generator 214A (e.g., “Gr”) can receive a row data input vector cr, a row random variable zr, and a co-cluster label d. Similarly, the column generator 214B (e.g., “Gc”) can receive a column data input vector cc, a column random variable zc, and the co-cluster label d. In particular embodiments, the co-cluster label d can include, for example, a category vector that can be designed to be of size cr x cc (e.g., because the row data input vector cr and the column data input vector cc can represent a pair of clusters or co-clusters).

[0054] In particular embodiments, the row generator 214A (e.g., “Gr”) and the column generator 214B (e.g., “Gc”) can then generate a fake row data vector xr and a fake column data vector xc, respectively. For example, the fake row data vector xr and the fake column data vector xc can represent one or more fake data samples (e.g., fake unlabeled data samples) of the row data input vector cr and the column data input vector cc, respectively. In particular embodiments, the row generator 214A (e.g., “Gr”) can provide the fake row data vector xr to the row discriminator 216A (e.g., “Dr”), and the column generator 214B (e.g., “Gc”) can provide the fake column data vector xc to the column discriminator 216B (e.g., “Dc”). For example, as further depicted, the row discriminator 216A (e.g., “Dr”) can receive the fake row data vector xr and the true row data vector Xr, and the column discriminator 216B (e.g., “Dc”) can receive the fake column data vector xc and the true column data vector Xc. In particular embodiments, the row discriminator 216A (e.g., “Dr”) can be trained with the true row data vector Xr before receiving the fake row data vector xr, and the column discriminator 216B (e.g., “Dc”) can be trained with the true column data vector Xc before receiving the fake column data vector xc. For example, the true row data vector Xr and the true column data vector Xc can represent one or more true data samples (e.g., true unlabeled data samples) of the row data input vector cr and the column data input vector cc, respectively.

[0055] In particular, in particular embodiments, as part of the training of the first GAN 212A and the second GAN 212B, the row generator 214A (e.g., “Gr”) can be provided to “fool” the row discriminator 216A (e.g., “Dr”), and the column generator 214B (e.g., “Gc”) can be provided to “fool” the column discriminator 216B (e.g., “Dc”). For example, this “fooling” technique can be utilized to maximize the final classification error between, for example, the fake row data vector xr and the true row data vector Xr, and between the fake column data vector xc and the true column data vector Xc, respectively. Conversely, for example, the row discriminator 216A (e.g., “Dr”) and the column discriminator 216B (e.g., “Dc”) can be provided to detect the fake row data vector xr and the fake column data vector xc to minimize the final classification error between, for example, the fake row data vector xr and the true row data vector Xr, and between the fake column data vector xc and the true column data vector Xc, respectively.

[0056] In particular embodiments, the row discriminator 216A (e.g., “Dr”) and the column discriminator 216B (e.g., “Dc”) can then provide outputs to the row clustering distribution model 218A (e.g., Qr) and the column clustering distribution model 218B (e.g., Qc). The row discriminator 216A (e.g., “Dr”) and the column discriminator 216B (e.g., “Dc”) can also provide true / false (e.g., “T / F”) outputs indicating the final classification errors between the false row data vector xr- and the true row data vector Xr-, and between the false column data vector xcand the true column data vector Xc, respectively. In particular embodiments, the row clustering distribution model 218A (e.g., Qr) and the column clustering distribution model 218B (e.g., Qc) can generate a row data output vector cr’ and a column data output vector cc’. The row data output vector cr’ and the column data output vector cc’ can each include, for example, one-dimensional clusters of row data and column data of the row data input vector cr and the column data input vector cc, respectively.

[0057] In particular embodiments, as Figure 2b Further as shown by the autoencoder CI-GAN model 200B, the reconstructor 220 (e.g., R) can receive the row data output vector cr’ and the column data output vector cc’ as inputs. In particular embodiments, the reconstructor 220 (e.g., R) can include a distribution model that can be used to learn and reconstruct the false row data vector xr- and the false column data vector xc- based on the row data output vector cr’ and the column data output vector cc’, and by extension, the original row data input vector cr and the original column data input vector cc. In particular embodiments, the row clustering distribution model 218A (e.g., Qr) and the column clustering distribution model (e.g., Qc) and the reconstructor 220 (e.g., R) can be defined by optimizing a lower bound of the autoencoder CI-GAN model 200B by, for example, increasing the mutual information (e.g., I) between the original row data input vector cr and the original column data input vector cc and the row data output vector cr’ and the column data output vector cc.

[0058] Figure 3a and Figure 3b Figures 300A and 300B are shown to depict utilization of a CI-GAN for efficient data co-clustering in accordance with the presently disclosed embodiments, as described above with respect to Figure 2a and Figure 2bThe described) running examples. For example, diagram 300A can include a data matrix, which can include, for example, user behavior data as row data 302 and user content viewing data as column data 304. To build useful and interesting two-dimensional data co-clusters for the purpose of, for example, providing program recommendations and advertising content that can be more appropriate to the desires, preferences, and interests of certain users, co-clustering of the data matrix can be performed by the combiner CI-GAN model 200A and / or the autoencoder CI-GAN model 200B. As shown by the user behavior row data 302 and the user content viewing column data 304, user 1 watches Korean dramas (e.g., K-Dramas) and changes program channels 8 times per hour; user 2 watches e-sports (e.g., video game videos) and changes program channels 12 times per hour; user 3 watches anime and changes program channels 5 times per hour; user 4 watches soccer and changes 1 time per hour; user 5 watches Korean pop music videos (e.g., K-Pop) and changes program channels 9 times per hour; and user 6 watches racing and changes program channels 2 times per hour.

[0059] Figure 3b Based on the described) running examples, it is shown that Figure 3aof the data matrix 306, where less advanced single dimension data clustering is performed as compared to the diagram 308 showing advanced techniques utilizing CI-GAN models (the combiner CI-GAN model 200A and / or the autoencoder CI-GAN model 200B) for efficient data co-clustering as described herein. As depicted by the diagram 306, the less advanced single dimension data clustering can only produce poor and non-specific user clusters 310, 312, 314, and 316 and TV show clusters 318, 320, and 322. Without the presently disclosed techniques, one or more co-clusters correlating the user behavior row data 302 and the user content viewing column data 304 cannot be achieved. In particular embodiments, as shown by the diagram 308, the CI-GAN models (the combiner CI-GAN model 200A and / or the autoencoder CI-GAN model 200B) can find one or more hidden and non-obvious underlying interrelationships (e.g., underlying relationships, similarities) among the user behavior row data 302 and the user content viewing column data 304. For example, the co-cluster 324 shows determined interrelationships between Korean dramas, K-Pop, e-sports, and anime, and thus, for example, clusters user 1, user 2, user 3, and user 5 together with Korean dramas, K-Pop, e-sports, and anime. Similarly, the co-cluster 326 shows determined interrelationships between soccer and racing, and thus, clusters user 4 and user 6 together with soccer and racing. In particular, as previously described, the one-dimensional clustering diagram 306 clusters columns independent of rows, and vice versa. Thus, when clustering rows in this example, the diagram 306 can cluster user 4 and user 6 as one cluster 316 (which changes channels 1-2 times), cluster user 3 as another cluster 314 (which changes channels 5 times), cluster user 2 as another cluster 312 (which changes channels 12 times), and cluster user 1 and user 5 as one cluster 310 (which changes channels 8-9 times). Continuing, the diagram 306 can cluster user 1 and user 5 as one cluster 318 (which watches K-pop and Korean dramas), cluster user 2 and user 3 as another cluster 320 (which watches e-sports and anime), and cluster user 4 and user 6 as another cluster 318 (which watches soccer and racing). However, as compared to the diagram 308, the diagram 306 shows that one-dimensional clustering techniques can miss one or more hidden and non-obvious underlying interrelationships. In contrast, the CI-GAN-based efficient data co-clustering techniques as disclosed herein and as shown by the diagram 308 can find such hidden and non-obvious underlying interrelationships. Thus, the present techniques can be used to provide both program recommendations and advertising content that can be more suitable to the desires, preferences, and interests of, for example, individual users and / or individual user subgroups.In practice, the present technology can improve program personalization / recommendation, targeted advertising (e.g., showing relevant ads to specific users), content demand prediction (e.g., predicting content that specific users want to watch), personalized device settings (e.g., adjusting device color settings if a specific user watches a nature show), and various other applications.

[0060] Figures 3c-3q An example training phase and further running examples are shown that utilize a CI-GAN for efficient data co-clustering (as described above with respect to Figure 2a and Figure 2b For example, Figure 3c An example running example is shown that utilizes, for example, the combiner CI-GAN model 200A. In this embodiment, assume that, for example, a data analyst wants to find 12 co-clusters (e.g., groups in two user dimensions). Thus, Figure 3c Example 300C shows that the data analyst can decide to have 12 co-clusters (e.g., 4 rows by 3 columns, as the analyst believes there can be more groups of behavior than groups of viewership). In Figure 3d Example 300D, there are 4 rows (e.g., 4 different types of behavior) and 3 columns (e.g., 3 different types of viewership). In this example, cc, cr, and d are uniformly distributed and randomly selected. For example, for all possible values of cr, each possible value represents a type / category of behavior, and for all possible values of cr, each possible value represents a type / category of viewership. In this example, all possible values of the co-cluster label d are also illustrated. Figure 3e A training example 300E is shown that illustrates that during the training phase, random values of cc, cr, and d can be selected. Figure 3f and Figure 3g Training examples 300F and 300G are shown that consider one particular real data sample as input. Figure 3f Training example 300F is shown that considers one particular real data sample. Based on the real data sample input, Figure 3h Example 300H shows that Gr generates a “fake” row data sample xr that attempts to be similar to the real row data sample Xr based on cr, zr, and d.

[0061] Figure 3i Example 300I is shown that illustrates a partial iteration of training one real data sample Xr based on one particular cr value, one particular d value, and one particular zr value. It should be understood that, in practice, training can be performed on all real data samples Xr, all possible cr values, all possible d values, and all possible zr values. Figure 3iExample 300J is shown, which illustrates continued training of a real data sample Xr based on one particular cr value, one particular d value, and one particular zr value for an additional n number of columns. Specifically, a real data sample Xr can be continued to be trained based on one d value, one Cr value, and all Zr values. The training iterations (loops of forward propagation and backpropagation) can then be repeated until xr is similar to Xr. In this example, approximately 7 million parameters can be used, but any suitable number of parameters are contemplated to be used with the disclosed technology. Figure 3k Example 300K is shown, which includes forward propagation and backpropagation examples. Forward propagation and backpropagation can be performed to make xr similar to Xr and maximize mutual information, which can be represented as:

[0062] min G,Q max D V InfoGAN (D, G, Q) = V(D, G) - λL I (G, Q) (Equation 5).

[0063] Figure 3l Example 300L is shown, which depicts a neural network (NN) 302L that can be implemented for row generator 214A (e.g., “Gr”). For example, this or another NN can be implemented for column generator 214B (e.g., “Gc”). Through forward propagation, one or more instances of “fake” row data sample xr can be generated based on input values cr, zr, and d. Figure 3m Example 300M is shown, which depicts various rows and / or columns for input values cr, zr, and d, and corresponding “fake” row data sample xr. Figure 3n Example 300N is shown, which depicts a NN that can be implemented on row clustering distribution model 208A (e.g., Qr), auxiliary clustering distribution model 210 (e.g., Qco), and / or row discriminator 206A (e.g., “Dr”). Through forward propagation, “fake” row data sample xr and real row data sample Xr can be input and compared, and the output can be true / false (T / F) and row data output vector cr’, and one or more collaborative clustering labels d’ are generated. In some embodiments, d’ takes into account information from Gr, Dr, Qco, and from Gc and Drc.

[0064] Figure 3oAn example 300O is shown that depicts a row discriminator 206A (e.g., “Dr”). In some embodiments, a column discriminator 206B (e.g., “Dc”) can be depicted in a similar manner. Specifically, the row discriminator 206A (e.g., “Dr”) and the column discriminator 206B (e.g., “Dc”) can provide a true / false (e.g., “T / F”) output that indicates a final classification error between the false row data vector xr- and the true row data vector Xr- and / or between the false column data vector xcand the true column data vector Xc, respectively. For example, the row discriminator 206A (e.g., “Dr”) can determine whether the false row data sample xr is sufficiently similar (e.g., true or false) to Xr. If it is determined that the false row data sample xr is not sufficiently similar (e.g., false) to the true data sample Xr, backpropagation can be performed. Figure 3p An example 300P is shown that includes a NN that can include one or more weights, for example, that can be used to increase the similarity between the false row data sample xr and the true data sample Xr. For example, the weights (e.g., wn) of the NN can be modified (e.g., increased or decreased) to increase the similarity between the false row data sample xr and the true data sample Xr. In one example, the similarity between the false row data sample xr and the true data sample Xr can be represented by V(D, G) in Equation 5 above. Moreover, the backpropagation (and the forward propagation) through the NN in example 300P can aim to learn the underlying relationship between the row data and the column data and maximize the mutual information, which can depend on λLI(G, Q) in the equation above. For example, the backpropagation can modify the weights (e.g., wn) that can be used to represent λLI(G, Q) (G,Q). The weights for the rows will affect the weights for the columns and vice versa. In one example, in the real world, d’ can affect both the weights of the row and column dimensions during backpropagation.

[0065] Figure 3qAn example 300Q is shown that includes a row clustering distribution model 208A (e.g., Qr), an auxiliary clustering distribution model 210 (e.g., Qco), and a row discriminator 206A (e.g., “Dr”). For example, the row discriminator 206A (e.g., “Dr”) can determine whether a fake row data sample xr is sufficiently similar to a real row data sample Xr (e.g., true or false). In particular embodiments, if a fake row data sample xr is sufficiently similar to a real row data sample (e.g., true), then training can be determined to be complete (e.g., converged) for one or more real row data samples Xr. Thus, once training is complete, a row data output vector cr’, a column data output vector cc’, and one or more co-clusters labels d’ can be generated. In particular embodiments, training can then be iteratively repeated for all fake row data samples xr (e.g., including all values for cr, d, and zr). The same training is then repeated for all fake column data samples xc (e.g., including all values for cc, d, and zc). In some embodiments, the disclosed techniques perform the foregoing computations simultaneously (e.g., computing all rows and columns at the same time), in parallel, and / or synchronously. However, it will be appreciated that other approaches are possible.

[0066] Figures 3r-3t An example inference phase utilizing a CI-GAN for efficient data co-clustering (as described above with respect to Figure 2a and Figure 2b the current disclosure is shown. Figure 3r A first inference phase example 300R is shown. In particular, based on the training phase, the disclosed techniques can optimize the application of the reconstructed latent code cr’ to the input row data and the reconstructed latent code cr’ to the input column data (e.g., considering maximizing mutual information). Moreover, based on the training phase, the disclosed techniques can generate a reconstructed latent code d’ that shows all possible co-clusters based on the initial user-made (e.g., data analyst) decisions for 4 rows and 3 columns (e.g., 12 co-clusters). Figure 3s An inference phase example 300S is shown. The inference phase example 300S shows that the reconstructed latent code d’ is used to flag each input real data sample and group it into a co-cluster. The example table in 300S represents the completed co-clusters. For example, in the real world, there can be vastly more data samples. Each co-cluster can include data samples that are similar or related to each other. In this way, a user (e.g., data analyst) can then examine the co-clusters to find relationships between two domains (e.g., behavior and viewership) or hidden or not obvious meaning. For example, Figure 3tA completed co-clustering example 300T is shown. For example, the present technology can provide personalized recommendations and personalized advertising because, for example, a developer can examine the co-clustering example 300U and ascertain that a user who changes channels at least 8 times per hour also watches action / adventure media content.

[0067] Figure 4 A flowchart of a method 400 for providing a co-information generating adversarial network (CI-GAN) for performing data co-clustering according to the presently disclosed embodiments is shown. The method 400 can be performed with one or more processing devices (e.g., the program analysis system 102), which can include hardware (e.g., a general-purpose processor, a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a system on a chip (SoC), a microcontroller, a field-programmable gate array (FPGA), a central processing unit (CPU), an application processor (AP), a visual processing unit (VPU), a neural processing unit (NPU), a neural decision processor (NDP), or any other processing device that can be suitable for processing image data), software (e.g., instructions running / executing on one or more processors), firmware (e.g., microcode), or some combination thereof.

[0068] The method 400 can begin with block 402 using one or more processing devices (e.g., the program analysis system 102) accessing a first data matrix including a plurality of row data and a plurality of column data. In particular embodiments, the program analysis system 102 can receive a data matrix of data, which can include, for example, user behavior data as the plurality of row data and user content viewing data as the plurality of column data, video data as the plurality of row data and video description data as the plurality of column data, biological genetic data as the plurality of row data and biological condition data as the plurality of column data, movie data as the plurality of row data and music data as the plurality of column data, image data as the plurality of row data and audible data as the plurality of column data.

[0069] The method 400 can then continue at block 404 using one or more processing devices (e.g., the program analysis system 102) providing a first data input to a first generative adversarial network (GAN), the first data input comprising a plurality of row vectors corresponding to a plurality of row data from the first data matrix. The method 400 can then continue at block 406 using one or more processing devices (e.g., the program analysis system 102) providing a second data input to a second GAN, the second data input comprising a plurality of column vectors corresponding to a plurality of column data from the first data matrix. In particular embodiments, the first GAN and the second GAN can be part of a co-informative generative adversarial network (CI-GAN) combiner model architecture, where the CI-GAN combiner model architecture can further comprise a distribution function configured to reconstruct one or more co-cluster labels as a third data input provided to the first GAN and the second GAN. In particular embodiments, in combination or alternatively, the first GAN and the second GAN can be part of a co-informative generative adversarial network (CI-GAN) autoencoder model architecture, where the CI-GAN autoencoder model architecture can further comprise a distribution function configured to reconstruct the plurality of row data and the plurality of column data based on a plurality of output row data generated by the first GAN and a plurality of output column data generated by the second GAN.

[0070] The method 400 can then terminate at block 408 using one or more processing devices (e.g., the program analysis system 102) generating a co-clustered correlation matrix based at least in part on the plurality of row vectors and the plurality of column vectors by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN, where the co-clustered correlation matrix comprises co-clustered associations between the plurality of row data and the plurality of column data of the first data matrix. In particular embodiments, generating the co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN can comprise determining one or more potential interrelationships between the plurality of row data and the plurality of column data of the first data matrix based at least in part on a mutual information cost function. For example, specifically, embodiments of the CI-GAN combiner model architecture can determine one or more potential interrelationships between the plurality of row data and the plurality of column data of the first data matrix based at least in part on a mutual information cost function.

[0071] In particular embodiments, generating the co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN can further include determining a lower bound that defines an objective function corresponding to the mutual information and co-clustered associations between the plurality of row data and the plurality of column data of the first data matrix. For example, embodiments of the CI-GAN autoencoder model architecture can specifically determine a lower bound that defines an objective function corresponding to the mutual information and co-clustered associations between the plurality of row data and the plurality of column data of the first data matrix. In particular embodiments, generating the co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN can further include separately determining one or more disentangled representations of the plurality of row data and the plurality of column data, or simultaneously determining one or more disentangled representations of the plurality of row data and the plurality of column data. In particular embodiments, simultaneously determining one or more disentangled representations of the plurality of row data and the plurality of column data can include determining a joint disentangled representation that is used to decrypt the interaction data structure in the co-clustered correlation matrix.

[0072] In particular, because the data clustering of the row data and the column data is implemented independently by the first GAN and the second GAN, the presently disclosed CI-GAN can allow for maximizing the mutual information (e.g., I, a measure of the degree to which one random variable can suggest another random variable) and co-clustered associations between the input row data and column data in several different ways. In this way, the present CI-GAN for efficient data co-clustering augmented with a mutual information objective function can be well suited for jointly learning dual data representations across a variety of technical applications. Indeed, the present technology can improve program personalization / recommendations, targeted advertising (e.g., showing relevant advertisements to specific users), content demand prediction (e.g., predicting content that a specific user wants to watch), personalized device settings (e.g., adjusting device color settings if a specific user watches a nature performance), and a variety of other applications.

[0073] Figure 5An example computer system 500 that can be used for providing a co-information generative adversarial network (CI-GAN) for data co-clustering in accordance with the presently disclosed embodiments is shown. In particular embodiments, the one or more computer systems 500 perform one or more steps of the one or more methods described or illustrated herein. In particular embodiments, the one or more computer systems 500 provide the functionality described or illustrated herein. In particular embodiments, software running on the one or more computer systems 500 performs one or more steps of the one or more methods described or illustrated herein, or provides the functionality described or illustrated herein. Particular embodiments include one or more portions of the one or more computer systems 500. Herein, reference to a computer system can encompass a computing device, and vice versa, where appropriate. Moreover, where appropriate, the reference can encompass one or more computer systems.

[0074] The present disclosure contemplates any suitable number of computer systems 500. The present disclosure contemplates computer systems 500 of varying size, varying maximum capacity, and varying maximum bandwidth. The present disclosure contemplates a computer system 500 having any suitable physical form. As example and not by way of limitation, computer system 500 can be embedded in a device, a system-on-chip (SoC), a single-board computer (SBC) (for example, a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile telephone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, computer system 500 can include one or more computer systems 500; be unitary or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud, which can include one or more cloud components in one or more networks.

[0075] Where appropriate, one or more computer systems 500 can perform one or more steps of the one or more methods described or illustrated herein in real time. As example and not by way of limitation, one or more computer systems 500 can perform one or more steps of one or more methods described or illustrated herein in batches. One or more computer systems 500 can perform one or more steps of one or more methods described or illustrated herein at different times and at different locations, where appropriate.

[0076] In particular embodiments, computer system 500 includes a processor 502, memory 504, storage 506, an input / output (I / O) interface 508, a communication interface 510, and a bus 512. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement. In particular embodiments, processor 502 includes hardware for executing instructions, such as those contained in the computer program 514. As an example and not by way of limitation, to execute instructions, processor 502 can retrieve (or fetch) the instructions from an internal register, an internal cache, memory 504, or storage 506; decode and execute them; and then write one or more results to the internal register, internal cache, memory 504, or storage 506. In particular embodiments, processor 502 can include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 502 with any suitable number of any suitable internal caches, which have any suitable arrangement. As an example and not by way of limitation, processor 502 can include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in the instruction caches can be copies of instructions in memory 504 or storage 506, and the instruction caches can speed up retrieval of those instructions by processor 502.

[0077] Data in the data caches can be copies of data in memory 504 or storage 506 for instructions executing at processor 502 to operate on; results of previous instructions executed at processor 502 for subsequent instructions executing at processor 502 to access or for writing to memory 504 or storage 506; or other suitable data. Data caches can speed up read or write operations by processor 502. TLBs can speed up virtual address translation for processor 502. In particular embodiments, processor 502 can include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 502 with any suitable number of any suitable internal registers, which have any suitable arrangement. Processor 502 can include one or more arithmetic logic units (ALUs); be a multi-core processor; or include one or more processors 502, as appropriate. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.

[0078] In particular embodiments, memory 504 includes main memory for storing instructions for execution by processor 502 or data used by processor 502 in its operation. As an example and not by way of limitation, computer system 500 can load instructions from storage 506 or another source (such as, for example, another computer system 500) into memory 504. Processor 502 can then load the instructions from memory 504 into internal registers or internal caches. To execute the instructions, processor 502 can retrieve the instructions from the internal registers or internal caches and decode them. During or after execution of the instructions, processor 502 can write one or more results (which can be intermediate or final results) to the internal registers or internal caches. Processor 502 can then write one or more of these results to memory 504. In particular embodiments, processor 502 only executes instructions in one or more internal registers or internal caches or memory 504 (as opposed to storage 506 or elsewhere) and only operates on data in one or more internal registers or internal caches or memory 504 (as opposed to storage 506 or elsewhere).

[0079] One or more memory buses, which can each include an address bus and a data bus, can couple processor 502 to memory 504. Bus 512 can include one or more memory buses, as described below. In particular embodiments, one or more memory management units (MMUs) reside between processor 502 and memory 504 and facilitate access to memory 504 by processor 502. In particular embodiments, memory 504 includes random access memory (RAM). This RAM can be volatile memory or can be non-volatile memory, as appropriate. Where volatile memory is used, this RAM can also be dynamic RAM (DRAM), which requires refreshes. Where such refreshes are used, the type of DRAM can be, but is not limited to, synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), SynchBurst SDRAM (SB-SDRAM), or extensible SDRAM (XSDRAM). Where non-volatile memory is used, this RAM can also be flash memory or magnetic block memory. Moreover, where non-volatile memory is used, this memory can be a storage device, such as storage 506. The RAM can be a

[0080] In particular embodiments, storage 506 includes mass storage for data or instructions. As an example and not by way of limitation, storage 506 can include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc (e.g., a compact disc or DVD, etc.), a solid-state drive (SSD), a tape drive, a USB drive, or a combination of two or more of these. Storage 506 can include removable or non-removable (or fixed) media, where appropriate. Storage 506 can be internal or external to computer system 500, where appropriate. In particular embodiments, storage 506 is nonvolatile, solid-state memory. In particular embodiments, storage 506 includes read-only memory (ROM). Where appropriate, this ROM can be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory or a combination of two or more of these. This disclosure contemplates mass storage 506 taking any suitable physical form.

[0081] In particular embodiments, I / O interface 508 includes hardware, software, or both, providing one or more interfaces for the transfer of information between computer system 500 and one or more I / O devices. Where appropriate, computer system 500 can include one or more of these I / O devices. One or more of these I / O devices can enable communication between a person and computer system 500. As an example and not by way of limitation, an I / O device can include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device or a combination of two or more of these. An I / O device can include one or more sensors. This disclosure contemplates any suitable I / O devices and any suitable I / O interfaces 506 for them. Where appropriate, I / O interface 508 can include one or more device or software drivers enabling processor 502 to drive one or more of these I / O devices. I / O interface 508 can include one or more I / O interfaces 506, where appropriate. Although this disclosure describes and illustrates a particular I / O interface, this disclosure contemplates any suitable I / O interface.

[0082] In particular embodiments, communication interface 510 includes hardware, software, or both providing one or more interfaces for the exchange of information between computer system 500 and one or more other computer systems 500 or one or more networks. As an example, and not by way of limitation, communication interface 510 can include a network interface controller (NIC) or network adapter for communicating with an Ethernet network or with a wireless network, such as a WI-FI network. The present disclosure contemplates any suitable network and any suitable communication interface 510 for it. As another example, communication interface 510 can include a wireless NIC (WNIC) or wireless adapter for communicating with a wireless PAN (WPAN) (e.g., a BLUETOOTH WPAN), a WI-FI network, a WI-MAX network, a cellular telephone network (e.g., a 4G 3G, or LTE network), or other suitable wireless network or combinations of two or more of these. The present disclosure contemplates any suitable network and any suitable communication interface 510 for it. As another example, communication interface 510 can include one or more input / output (I / O) devices 514, such as a mouse, a keyboard, a display, or a screen for communicating information to a user of computer system 500 and receiving input from the user.

[0083] As an example and not by way of limitation, computer system 500 can communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet or a combination of two or more of these. One or more portions of one or more of these networks can be wired or wireless. As an example, computer system 500 can communicate with a wireless PAN (WPAN) (such as, for example, a BLUETOOTH WPAN), a WI-FI network, a WI-MAX network, a cellular telephone network (such as, for example, a 4G, 3G, or LTE network), or other suitable wireless network or combinations of two or more of these. Computer system 500 can include any suitable communication interface 510 for any of these networks, where appropriate. The communication interface 510 can include one or more communication interfaces 510, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.

[0084] In particular embodiments, bus 512 includes a hardware, software, or both component that couples components of computer system 500 to one another. As an example and not by way of limitation, bus 512 can include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand (IB) interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI Express (PCIe) bus, a Serial Advanced Technology

[0085] In this document, where appropriate, one or more computer-readable non-transitory storage media may include one or more semiconductor-based integrated circuits (ICs) or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these. Where appropriate, computer-readable non-transitory storage media may be volatile storage media, non-volatile storage media, or a combination of volatile and non-volatile storage media.

[0086] Figure 6 Figure 600 illustrates an example artificial intelligence (AI) architecture 602 that can be used to provide a collaborative information generative adversarial network (CI-GAN) for collaborative data clustering, according to embodiments currently disclosed. In certain embodiments, the AI ​​architecture 602 may be implemented using, for example, one or more processing devices, which may include hardware (e.g., a general-purpose processor, graphics processing unit (GPU), application-specific integrated circuit (ASIC), system-on-a-chip (SoC), microcontroller, field-programmable gate array (FPGA), central processing unit (CPU), application processor (AP), vision processing unit (VPU), neural processing unit (NPU), neural decision processor (NDP), or any other processing device that may be suitable for processing various types of data and making one or more decisions based on various types of data), software (e.g., instructions that run / execute on one or more processors), firmware (e.g., microcode), or some combination thereof.

[0087] In certain embodiments, such as Figure 6 As shown, the AI ​​architecture 602 may include machine learning (ML) algorithms and functions 604, natural language processing (NLP) algorithms and functions 606, expert systems 608, computer-based vision algorithms and functions 610, speech recognition algorithms and functions 612, planning algorithms and functions 614, and robotics algorithms and functions 616. In a particular embodiment, the ML algorithms and functions 604 may include any statistical-based algorithm that may be suitable for finding patterns across large amounts of data (e.g., "big data," such as user click data or other user interactions, text data, image data, video data, audio data, voice data, digital data, etc.). For example, in a particular embodiment, the ML algorithms and functions 604 may include deep learning algorithms 618, supervised learning algorithms 620, and unsupervised learning algorithms 622.

[0088] In particular embodiments, the deep learning algorithm 618 can include any artificial neural network (ANN) that can be used to learn deep levels of representation and abstraction from large amounts of data. For example, the deep learning algorithm 618 can include ANNs such as a multilayer perceptron (MLP), an autoencoder (AE), a convolutional neural network (CNN), a recurrent neural network (RNN), a long short-term memory (LSTM), a gated recurrent unit (GRU), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a generative adversarial network (GAN), and a deep Q-network, neural autoregressive distribution estimator (NADE), an adversarial network (AN), an attention model (AM), deep reinforcement learning, etc.

[0089] In particular embodiments, the supervised learning algorithm 620 can include any algorithm that can be used to apply, e.g., past learning to new data for predicting future events using labeled examples. For example, starting from an analysis of a known training dataset, the supervised learning algorithm 620 can produce an inference function to make predictions about output values. The supervised learning algorithm 620 can also compare its output to the correct and expected output and find errors in order to modify the supervised learning algorithm 620 accordingly. On the other hand, the unsupervised learning algorithm 622 can include any algorithm that can be applied, e.g., when the data used to train the unsupervised learning algorithm 622 is not classified or labeled. For example, the unsupervised learning algorithm 622 can study and analyze how a system can infer a function that describes a hidden structure from unlabeled data.

[0090] In particular embodiments, the NLP algorithms and functions 606 can include any algorithms or functions that can be suitable for automatically manipulating natural language, such as speech and / or text. For example, in some embodiments, the NLP algorithms and functions 606 can include a content extraction algorithm or function 624, a classification algorithm or function 626, a machine translation algorithm or function 628, a question answering (QA) algorithm or function 630, and a text generation algorithm or function 632. In particular embodiments, the content extraction algorithm or function 624 can include a means for extracting, e.g., text or images, from an electronic document (e.g., a web page, a text editor document, etc.) to be used in other applications.

[0091] In particular embodiments, the classification algorithm or function 626 can include any algorithm that can utilize a supervised learning model (e.g., logistic regression, Naive Bayes, stochastic gradient descent (SGD), k-nearest neighbor, decision tree, random forest, support vector machine (SVM), etc.) to learn from data input to the supervised learning model and make new observations or classifications based on the data. The machine translation algorithm or function 628 can include any algorithm or function that can be suitable for automatically translating source text in one language to text in, for example, another language. The QA algorithm or function 630 can include any algorithm or function that can be suitable for automatically answering questions posed by humans in, for example, natural language (such as performed by voice-controlled personal assistant devices). The text generation algorithm or function 632 can include any algorithm or function that can be suitable for automatically generating natural language text.

[0092] In particular embodiments, the expert system 608 can include any algorithm or function that can be suitable for simulating the judgment and behavior of a human or organization with expertise and experience in a particular domain (e.g., stock trading, medicine, sports statistics, etc.). The computer-based vision algorithm and function 610 can include any algorithm or function that can be suitable for automatically extracting information from images (e.g., photographic images, video images). For example, the computer-based vision algorithm and function 610 can include an image recognition algorithm 634 and a machine vision algorithm 636. The image recognition algorithm 634 can include any algorithm that can be suitable for automatically recognizing and / or classifying objects, places, people, etc. that can be included in, for example, one or more image frames or other display data. The machine vision algorithm 636 can include any algorithm that can be suitable for allowing a computer to“see” or, for example, rely on an image sensor camera with specialized optics to acquire images for the purpose of processing, analyzing, and / or measuring various data characteristics for decision making.

[0093] In particular embodiments, the speech recognition algorithm and function 612 can include any algorithm or function that can be suitable for recognizing spoken language and converting the spoken language to text (such as through automatic speech recognition (ASR), computer speech recognition, speech-to-text (STT), or text-to-speech (TTS)) in order for, for example, a computer to communicate with one or more users via speech. In particular embodiments, the planning algorithm and function 614 can include any algorithm or function that can be suitable for generating sequences of actions, where each action can include its own set of preconditions that must be satisfied before the action is executed. Examples of AI planning can include classical planning, reduction to other problems, temporal planning, probabilistic planning, preference-based planning, conditional planning, etc. Finally, the robotic algorithm and function 616 can include any algorithm, function, or system that can enable one or more devices to replicate human behavior through, for example, motion, posture, performing tasks, making decisions, emotions, etc.

[0094] In this document, "or" is inclusive and not exclusive, unless expressly indicated otherwise or indicated otherw ise by context. Therefore, herein, "A or B" means "A, B, or both," unless expressly indicated otherwise or indicated otherw ise by context. Moreover, "and" is both joint and several, unless expressly indicated otherwise or indicated otherw ise by context. Therefore, herein, "A and B" means "A and B, jointly and severally," unles s expressly indicated otherw ise or indicated otherw ise by context.

[0095] In this document, "automatically" and its derivatives refer to the performance of an action or function without human intervention.

[0096] Embodiments disclosed herein are merely examples and the scope of the disclosure is not limited to them. Embodiments according to the disclosure are specifically disclosed in the appended claims directed to methods, storage media, systems and computer program products, wherein any feature mentioned in one claim category (e.g. method) can be claimed in another claim category (e.g. system etc.). The dependency or reference in the appended claims is chosen solely for formal reasons only. However, since any subject matter falling under any prior claim reference can be intended to be claimed (in particular multiple dependencies), any combination of features from the disclosure is disclosed and can be claimed, irrespective of the dependency chosen in the appended claims. The subject matter that can be claimed not only includes the combination of features as recited in the appended claims, but also any other combination of features from the claims, wherein each feature mentioned in a claim can be combined with any other feature or combination of other features from the claims. Furthermore, any embodiment and feature described or depicted herein can be claimed in separate claims and / or in any combination with any embodiment or feature described or depicted herein or in any combination with any feature of the appended claims.

[0097] The scope of the disclosure encompasses all changes, substitutions, variations, alterations, and modifications to the example embodiments described or illustrated herein that a person having ordinary skill in the art would comprehend. The scope of the disclosure is not limited to the example embodiments described or illustrated herein. Moreover, although the disclosure describes and illustrates respective embodiments herein as comprising specific components, elements, features, functions, operations or steps, any of these embodiments can include any combination or permutation of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that a person having ordinary skill in the art would comprehend. Moreover, references to devices or systems or components of systems or devices, adapted, arranged, capable, configured, enabled, operable, or otherwise, to perform a particular function refer to the same whether actually or inherently so adapted, arranged, capable, configured, enabled, operable, or otherwise, as long as the device, system, or component of system or device is so adapted, arranged, capable, configured, enabled, operable, or otherwise, whether or not activated, opened, or unlocked, to perform that particular function. Moreover, although the disclosure describes or illustrates particular embodiments as providing particular advantages, some embodiments can not provide one or more of those advantages, some embodiments can provide all of those advantages, and some embodiments can provide one, some, or all of those advantages.

Claims

1. A method comprising, by one or more computing systems: accessing a first data matrix comprising a plurality of row data and a plurality of column data; providing, to a first generative adversarial network (GAN), a first data input comprising a plurality of row vectors corresponding to the plurality of row data from the first data matrix; providing, to a second GAN, a second data input comprising a plurality of column vectors corresponding to the plurality of column data from the first data matrix; generating, by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN, a co-clustered correlation matrix based at least in part on the plurality of row vectors and the plurality of column vectors, wherein the co-clustered correlation matrix comprises co-clustered associations between the plurality of row data and the plurality of column data of the first data matrix; and based on the co-clustered associations, providing, by a network-based content orchestrator of the one or more computing systems, at least one of a TV program or a recommended content to an end-user client for user viewing.

2. The method of claim 1, wherein, the first GAN and the second GAN are part of a co-informative generative adversarial network (CI-GAN) combiner model architecture, wherein the CI-GAN combiner model architecture further comprises a distribution function configured to reconstruct one or more co-clustered labels as a third data input provided to the first GAN and the second GAN.

3. The method of claim 1, wherein, the first GAN and the second GAN are part of a co-informative generative adversarial network (CI-GAN) autoencoder model architecture, wherein the CI-GAN autoencoder model architecture further comprises a distribution function configured to reconstruct the plurality of row data and the plurality of column data based on a plurality of output row data generated by the first GAN and a plurality of output column data generated by the second GAN.

4. The method of claim 1, wherein, generating the co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN comprises: determining one or more latent interrelationships between the plurality of row data and the plurality of column data of the first data matrix based at least in part on a mutual information cost function.

5. The method of claim 1, wherein, generating the co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN comprises: determining a lower bound defining an objective function corresponding to mutual information between the co-clustered associations and the plurality of row data and the plurality of column data of the first data matrix.

6. The method of claim 1, wherein, generating the co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN comprises: maximizing mutual information between the co-clustered associations and the plurality of row data and the plurality of column data of the first data matrix.

7. The method of claim 1, wherein, generating the co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN comprises: determining one or more disentangled representations of the plurality of row data and the plurality of column data simultaneously.

8. The method of claim 1, wherein, generating the co-clustering correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN includes: determining one or more disentangled representations of the plurality of row data and the plurality of column data simultaneously.

9. The method of claim 8, wherein, determining the one or more disentangled representations of the plurality of row data and the plurality of column data simultaneously includes: determining a joint disentanglement representation for decrypting an interaction data structure in the co-clustering correlation matrix.

10. The method of claim 1, wherein, the first data matrix can include one or more of: user behavior data as the plurality of row data and user content viewing data as the plurality of column data; video data as the plurality of row data and video description data as the plurality of column data; bio-genetic data as the plurality of row data and bio-condition data as the plurality of column data; movie data as the plurality of row data and music data as the plurality of column data; or image data as the plurality of row data and audible data as the plurality of column data.

11. A system comprising: one or more non-transitory computer-readable storage media comprising instructions; and one or more processors coupled to the storage media, the one or more processors configured to execute the instructions to: access a first data matrix comprising a plurality of row data and a plurality of column data; provide, to a first generative adversarial network (GAN), a first data input comprising a plurality of row vectors corresponding to the plurality of row data from the first data matrix; provide, to a second GAN, a second data input comprising a plurality of column vectors corresponding to the plurality of column data from the first data matrix; generate, by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN, a co-clustering correlation matrix based at least in part on the plurality of row vectors and the plurality of column vectors, wherein the co-clustering correlation matrix comprises co-clustering associations between the plurality of row data and the plurality of column data of the first data matrix; and based on the co-clustering associations, provide, by a network-based content orchestrator of the one or more computing systems, at least one of a TV program or a recommended content to an end-user client for a user to view.

12. The system of claim 11, wherein, the first GAN and the second GAN are part of a co-informative generative adversarial network (CI-GAN) combiner model architecture, wherein the CI-GAN combiner model architecture further comprises a distribution function configured to reconstruct one or more co-clustering labels provided as a third data input to the first GAN and the second GAN.

13. The system of claim 11, wherein, The first GAN and the second GAN are part of a co-informatics generative adversarial network (CI-GAN) autoencoder model architecture, wherein the CI-GAN autoencoder model architecture further comprises a distribution function configured to reconstruct the plurality of row data and the plurality of column data based on a plurality of output row data generated by the first GAN and a plurality of output column data generated by the second GAN.

14. The system of claim 11, wherein, The instructions to generate the co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN comprise: instructions to determine one or more latent interrelationships between the plurality of row data and the plurality of column data of the first data matrix based at least in part on a mutual information cost function.

15. The system of claim 11, wherein, The instructions to generate the co-clustered correlation matrix by synchronously co-clustering the plurality of row vectors and the plurality of column vectors by the first GAN and the second GAN comprise: instructions to determine a lower bound defining an objective function corresponding to the mutual information between the co-clustered associations and the plurality of row data and the plurality of column data of the first data matrix.

Citation Information

Patent Citations

  • Method for learning cross-domain relations based on generative adversarial networks

    US20190205334A1