Electric vehicle intelligent cabin air conditioner user clustering method based on GAN network

Through the clustering method based on GAN network, the distribution of new user data is learned and the discriminator is used to filter historical user data, which solves the problems of distance measurement limitations and labeled data dependence in high-dimensional time-series data clustering, and achieves efficient and accurate user clustering.

CN120180167APending Publication Date: 2025-06-20CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510255110.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing clustering methods are difficult to accurately reflect the true distribution characteristics of the data in high-dimensional time series data, especially in the absence of labeled data, deep clustering methods also have limitations that rely on labeled data.

Method used

The user clustering method of electric vehicle smart cockpit air conditioner based on GAN network is adopted to learn new user data distribution through GAN network, and the discriminator is used to calculate indirect evaluation indicators to screen historical user data, overcoming the limitations of Euclidean distance equidistance measurement.

Benefits of technology

High-dimensional time-series data clustering without labeled data is realized, and historical user data that meets the distribution of new user data is accurately screened, which improves the accuracy and efficiency of clustering, and reduces the bandwidth pressure of data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180167A_ABST
    Figure CN120180167A_ABST
Patent Text Reader

Abstract

The invention relates to an electric vehicle intelligent cabin air conditioner user clustering method based on a GAN network, and belongs to the field of deep learning. The method comprises the following steps: collecting new user behavior data of an automobile, and clustering historical user data and new user data to obtain initial clustering data of the mixed historical user behavior data and new user behavior data; the GAN learns new user data distribution in any cluster; after learning is completed, calculating an indirect evaluation index through a GAN network discriminator according to new user data in the cluster; and evaluating the historical user data in the cluster through a GAN network discriminator, and screening the historical user data in the cluster according to a threshold value set by the indirect evaluation index. According to the method, the discriminator outputs the probability value to judge whether the probability value accords with data distribution or not so as to screen similar data, the similar data is introduced into a clustering evaluation system to provide an indirect index for the judgment of the clustering effect, and the problem that the clustering effectiveness cannot be judged due to the fact that no label data exists in clustering is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning, and relates to a method for clustering electric vehicle intelligent cockpit air conditioner users based on a GAN network. Background Art

[0002] In an intelligent cockpit system, air conditioner control has an important impact on passenger comfort. However, the personalized air conditioner solution faces many challenges, including individual differences in passenger preferences, dynamic changes in environmental factors, and the problem that general models are difficult to meet personalized needs. Personalization is actually population clustering. Currently, there are mainly two types of time series data clustering methods: (1) One is K-means, hierarchical clustering, etc., which rely on the calculation of distance or similarity, and cluster time series data through distance metrics (such as Euclidean distance). (2) The other is the combination of deep learning and clustering algorithms, using neural network models such as CNN as autoencoders, encoding the data and then using algorithms such as K-means and KNN clustering for clustering. These clustering algorithms mainly use the Euclidean distance, which mainly measures the straight-line distance between data points.

[0003] However, the Euclidean distance has certain limitations. For example, in a high-dimensional space, high-dimensional data usually shows sparsity, and data points are often distributed in a sub-region of a high-dimensional space. The Euclidean distance cannot well reflect the true distribution characteristics of the data. And in high-dimensional data, because the "clusters" of clustering are often not simple geometric shapes (such as circles or spheres), but complex distributions. These problems greatly affect the application of the clustered data in actual scenarios.

[0004] Previous research on high-dimensional data clustering is largely an extension of traditional clustering algorithms for low-dimensional data. Since the objects in a deterministic dataset are single points, traditional clustering algorithms do not consider the distribution of the objects themselves. Therefore, the research on extending traditional algorithms to high-dimensional data clustering is limited to using similarity metrics based on geometric distance and cannot capture the differences between objects with different distributions.

[0005] In the past decade, a large number of studies on improved clustering methods have been reported. Early studies proposed combining deep autoencoders with K-means to eliminate linearity, and a comprehensive solution called Deep Embedded Clustering (DEC) was proposed. In DEC, the autoencoder projects training samples into a lower-dimensional space through a bottleneck layer, which to some extent solves the problem of data representation in high-dimensional spaces. Although deep clustering is an unsupervised learning method, many deep clustering methods still rely to some extent on a large amount of labeled data to guide network training. For example, some deep clustering methods (such as Deep Embedded Clustering DEC) need to combine the supervision signal of clustering labels to enhance the training effect of the model. Although this is not strict supervised learning, it still relies on some labeled data. Therefore, deep clustering may be limited in the absence of labeled data. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide an electric vehicle intelligent cockpit air conditioner user clustering method based on a GAN network to improve the problem that geometric distances such as Euclidean distance cannot well reflect the true distribution characteristics of data when existing clustering methods such as K-means and hierarchical clustering are used for clustering unlabeled high-dimensional time series data.

[0007] To achieve the above purpose, the present invention provides the following technical solutions:

[0008] An electric vehicle intelligent cockpit air conditioner user clustering method based on a GAN network, which includes:

[0009] S1. Collect the behavior data of new car users, cluster the historical user data and the new user data, and obtain the initial clustering data of the mixed historical users and new user behavior data;

[0010] S2. The GAN network learns the distribution of new user data in any cluster;

[0011] S3. After the learning is completed, the GAN network discriminator calculates an indirect evaluation index according to the new user data in this cluster;

[0012] S4. The GAN network discriminator evaluates the historical user data in this cluster, and filters the historical user data in this cluster according to the threshold set by the indirect evaluation index.

[0013] Further, in step S1, clustering the historical user data and the new user data includes initializing K cluster centers {c1, c2,..., c K}, for each data point x i , calculate the distance between x i and each cluster center, and assign x i to the one that is the same as x iThe closest cluster; each data point x i After the allocation is completed for all, calculate the mean of the distances between the points within each cluster, and update the center of each cluster to the corresponding mean value.

[0014] Repeat the above operations until the cluster centers no longer change or the number of repeated operations reaches the maximum number of iterations.

[0015] Furthermore, in step S2, the GAN network learning the data distribution of new users in a certain cluster includes constructing a GAN network model to model the new user data, and the established model is shown as follows:

[0016]

[0017] In the formula, D represents the discriminator, D(x) represents the probability that the generated sample output by the discriminator comes from real data, G represents the generator, G(z) represents the generated sample output by the generator, V(D, G) represents the expected and game result of, p date represents real data, xvp date , p z represents the noise distribution, z represents the latent vector sampled from p z ;

[0018] The purpose of the discriminator is to maximize V(D, G) so that the discriminator can correctly distinguish real data and generated data; the purpose of the generator is to minimize V(D, G) so that the discriminator cannot distinguish generated data and real data;

[0019] During the training process, the loss functions of the generator and the discriminator are alternately updated until the model converges to obtain the optimal model, that is, to obtain the GAN network model that has learned the data distribution of new user behaviors.

[0020] Among them, the data used to train the GAN network model is new user behavior data and feature data.

[0021] Furthermore, the loss function of the generator is expressed as:

[0022]

[0023] Optimize the loss function L by updating the model weights G ;

[0024] The loss function of the discriminator is expressed as:

[0025]

[0026] Optimize the loss function L by updating the model weightsD 。

[0027] Further, in step S3, the discriminator of the GAN network calculates indirect evaluation indicators based on the new user data in this cluster, including: after the GAN network learns the data distribution of the new users in this cluster, the new user feature data in this cluster is input into the generator to generate behavior data, and then the discriminator evaluates the generated behavior data to obtain the indirect evaluation indicator log(1 - (D(G(z)))). This indirect evaluation indicator represents the fitting degree between the generated data of the generator and the real historical user behavior data.

[0028] Further, in step S4, screening the historical user data in this cluster according to the threshold set by the indirect evaluation indicator includes: inputting the historical user data in this cluster into the discriminator, and the discriminator outputs the probability that the historical user data conforms to the new user data distribution; using the indirect evaluation indicator as the threshold, comparing the probability value output by the discriminator for the historical user behavior data with the threshold, and screening out the historical user behavior data whose probability value does not meet the threshold.

[0029] The beneficial effects of the present invention are as follows: The present invention proposes an electric vehicle intelligent cockpit air conditioner user clustering method based on the GAN network. By judging whether the output probability value conforms to the data distribution to screen similar data, it overcomes the problem that the internal evaluation indicators are inaccurately evaluated due to the complexity of high-dimensional time-series data clustering. And the GAN network uses KL divergence to model the data, overcoming the problem that distance metrics such as Euclidean distance cannot well reflect the true distribution characteristics of the data. The similar data screening model only needs the label value of the data transmitted by the vehicle to the cloud and the model weights and hyperparameters trained by the vehicle local data, greatly reducing the pressure on the bandwidth for data transmission and enhancing the security of user information and vehicle driving. In addition, the method proposed by the present invention can be popularized and applied in other complex systems.

[0030] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings

[0031] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0032] Figure 1 It is a schematic flowchart of the user clustering method provided by the embodiment of the present invention;

[0033] Figure 2It is a flowchart for training a GAN network;

[0034] Figure 3 It is a flowchart for inferring a GAN network;

[0035] Figure 4 It is a schematic comparison between the data distribution map after optimizing the clustering effect of the GAN network and the data distribution maps of other clustering methods. Specific implementation manners

[0036] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0037] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation on the present invention; in order to better illustrate the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.

[0038] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the attached drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the attached drawings are only for illustrative purposes and should not be construed as a limitation on the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0039] As Figure 1 shown, it is an intelligent cockpit air conditioner user clustering method for electric vehicles based on a GAN network provided by an embodiment of the present invention. The specific implementation process of this method is as follows:

[0040] S1: Collect data of new car users and prepare enough historical user data.

[0041] S2: Since the driving times of users vary, the timestamps of data collection fluctuate, increasing the complexity of data processing. In addition, the inconsistent data sampling frequency may lead to data loss or inconsistent sampling intervals due to sensor failures or network problems in data transmission. These issues affect the integrity and continuity of the data, posing challenges to subsequent data analysis and model training.

[0042] To address these issues, in this embodiment, missing value filling and outlier deletion are performed on the collected raw data, and equidistant processing is carried out. Then, clustering methods (such as K-means, DBSCAN, etc.) are used to cluster the processed historical user data and new user data, thereby generating initial personalized data, which helps to reduce the complexity of subsequent calculations of the GAN network. In this embodiment, the optimal number of clusters for clustering is 3.

[0043] Among them, the specific process of clustering the historical user data and new user data is as follows:

[0044] Initialize K cluster centers {c1, c2,..., c K}, for each data point x i , calculate its distance from each cluster center, and assign x i to the cluster with the closest distance; after all data points are assigned, calculate the mean of the distances between points within each cluster, and update the center of the cluster to the mean. This process is repeated until the cluster centers no longer change or the number of repetitions reaches the set maximum number of iterations and then stops.

[0045] Calculate the distance between the data point and each cluster center through the following formula:

[0046]

[0047] In the formula, x = (x1, x2, x3..., x n ) and y = (y1, y2, y3..., y n ) are two n-dimensional points.

[0048] S3: After dividing the data into K clusters through K-means, for any i-th cluster, the GAN network learns the distribution of new user data in the i-th cluster, and the training data is the new user behavior data and feature data in this cluster. After appropriate epoch training, a GAN model that learns the distribution of new user behavior data is obtained. The historical user behavior data set is denoted as A = {a1, a2,..., a N}, where a i represents the feature vector of the i-th user.

[0049] For any cluster \(i\) (\(i = 1, 2, \ldots, K\)), the data within the cluster is used to train a generative adversarial network model. Assume the data in the \(i\)-th cluster is \(D\) i =\(\{x\) i1 , x\) i2 , \ldots, x\) iN \}, where \(i_N\) is the number of samples in the \(i\)-th cluster. The generative adversarial network consists of two main parts: a generator and a discriminator. The goal of the generator is to generate realistic samples \(G(Z)\) from a random noise \(z \sim N(0, 1)\); the goal of the discriminator is to distinguish between real samples and generated samples, that is, given a sample as input, it outputs the probability \(D(x)\) that it comes from real data. For cluster \(i\), the new user data \(P\) i within the cluster is used to train an independent GAN model. The specific process is as follows:

[0050] 1) Model the new user data and perform iterative processing. The model is as follows:

[0051]

[0052] where \(D\) represents the discriminator, \(D(x)\) represents the probability that the generated sample output by the discriminator comes from real data, \(G\) represents the generator, \(G(z)\) represents the generated sample output by the generator, \(V(D, G)\) represents the game result of the expectation and , \(p\) date represents real data, \(x \in p\) date , \(p\) z represents the noise distribution, and \(z\) represents the latent vector sampled from \(p\) z .

[0053] 2) During the training process of the GAN network, the discriminator \(D\) attempts to maximize \(V(D, G)\), that is, to make the discriminator able to correctly distinguish between real data and generated data as much as possible. The loss function of the discriminator is usually the cross-entropy loss, as shown in the following formula:

[0054]

[0055] The loss function \(L\) D is optimized by updating the model weights:

[0056]

[0057] The generator \(G\) attempts to minimize \(V(D, G)\), that is, to make the discriminator unable to distinguish between the generated data and the real data as much as possible. The loss function of the generator is expressed as:

[0058]

[0059] The loss function \(L\) G:

[0060]

[0061] 3) During the training process, when the loss functions of the generator and the discriminator are alternately updated, set an appropriate number of epochs until the model converges to obtain the optimal

[0062] S4: For the i-th cluster, after the GAN network generator learns the data distribution of the new users in the i-th cluster, the feature data of the new users in the i-th cluster is input into the generator to generate behavior data. The newly generated user behavior data is then evaluated by the discriminator to obtain the quantified indirect evaluation index log(1 - (D(G(z)))). This index log(1 - (D(G(z)))) represents the degree of fit between the generated data of the generator and the real historical user behavior data. A threshold is determined according to the new user probability value log(1 - (D(G(z)))). This method can ensure that the training effect of the generator does not affect the discriminator's objective screening of historical user behavior data and takes into account the influence of user feature data on behavior.

[0063] S5: Screen out the data in the historical user behavior data of the i-th cluster that does not conform to the data distribution of the new users in the i-th cluster through the indirect evaluation index and the set dynamic threshold.

[0064] Specifically, the historical user behavior data in the i-th cluster is input into the GAN network discriminator. The prediction result of the discriminator outputs the probability that the historical user data conforms to the new user data distribution through the sigmoid activation function. The form of the sigmoid function is as follows:

[0065]

[0066] For any input x, the output of σ(x) is always between 0 and 1.

[0067] Taking the indirect evaluation index as the threshold, screen out the historical user data whose probability value does not meet the threshold.

[0068] As Figure 4 shown is the comparison of the clustering data effects obtained by different clustering methods. From Figure 4 it can be seen that after introducing the indirect evaluation index using the method of the present invention, the obtained clustering data better conforms to the distribution of the required data.

[0069] In this embodiment, the parameters of the GAN model are shown in the following table:

[0070] Table 1 Parameters of the model

[0071]

[0072]

[0073] Using a user clustering method for the intelligent cockpit air conditioner of an electric vehicle based on a GAN network proposed by the present invention, the discriminator outputs a probability value to determine whether it conforms to the data distribution to screen similar data, overcoming the problem that the internal evaluation index evaluation is inaccurate due to the complexity of high-dimensional time-series data clustering. And the GAN network uses KL divergence to model the data, overcoming the problem that distance metrics such as Euclidean distance cannot well reflect the true distribution characteristics of the data. The similar data screening model only needs the label value of the data transmitted by the vehicle to the cloud and the model weights and hyperparameters trained by the local data of the vehicle, greatly reducing the pressure on the bandwidth for data transmission and enhancing the security of user information and vehicle driving. The method provided by the present invention can also be popularized and applied in other complex systems.

[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A GAN network-based electric vehicle intelligent cabin air conditioning user clustering method, characterized in that: The method includes: Collect new car user behavior data, cluster historical user data and new user data, and obtain initial clustering data of mixed historical user and new user behavior data; The GAN network learns the distribution of new user data in any cluster; After learning is completed, the indirect evaluation index is calculated based on the new user data in the cluster through the GAN network discriminator; The historical user data in the cluster is evaluated by a GAN network discriminator, and the historical user data in the cluster is screened according to a threshold set by the indirect evaluation indicator.

2. The method according to claim 1, characterized in that: Clustering the historical user data and new user data includes initializing K cluster centers {c1, c2, ..., c K }, for each data point x i , calculate x i The distance from the center of each cluster, x i Assigned to x i The closest cluster; for each data point x i After the distribution is completed, the mean of the distances between the points in each cluster is calculated, and the center of each cluster is updated to the corresponding mean; Repeat the above steps until the cluster center no longer changes, or the number of repeated operations reaches the maximum number of iterations.

3. The method according to claim 1, characterized in that The GAN network learning the distribution of new user data in a certain cluster includes building a GAN network model to model the new user data. The established model is shown in the following formula: In the formula, D represents the discriminator, D(x) represents the probability that the generated sample output by the discriminator comes from the real data, G represents the generator, G(z) represents the generated sample output by the generator, and V(D,G) represents the expected and The game result, p date represents real data, x∈p date , p z represents the noise distribution, z represents the distribution from p z The latent vector sampled in ; The purpose of the discriminator is to maximize V(D, G) so that the discriminator can correctly distinguish between real data and generated data; the purpose of the generator is to minimize V(D, G) so that the discriminator cannot distinguish between generated data and real data; During the training process, the loss functions of the generator and the discriminator are updated alternately until the model converges and the optimal model is obtained, that is, the GAN network model that learns the distribution of new user behavior data is obtained.

4. The method according to claim 3, characterized in that: The data used to train the GAN network model are new user behavior data and feature data.

5. The method according to claim 3, characterized in that: The loss function of the generator is expressed as: Optimize the loss function L by updating the model weights G ; The loss function of the discriminator is expressed as: Optimize the loss function L by updating the model weights D .

6. The method according to claim 1, characterized in that The indirect evaluation index is calculated according to the new user data in the cluster by using the GAN network discriminator, including: after the GAN network learns the distribution of the new user data in the cluster, the feature data of the new user in the cluster is input into the generator to generate behavior data, and then the generated behavior data is evaluated by the discriminator to obtain the indirect evaluation index log(1-(D(G(z)))), which represents the fit between the generated data of the generator and the real historical user behavior data.

7. The method according to claim 1, characterized in that Filtering the historical user data in the cluster according to the threshold set by the indirect evaluation indicator includes: inputting the historical user data in the cluster into a discriminator, and the discriminator outputs the probability that the historical user data conforms to the distribution of new user data; using the indirect evaluation indicator as a threshold, comparing the probability value of the historical user behavior data output by the discriminator with the threshold, and filtering out the historical user behavior data whose probability value does not conform to the threshold.