Electronic device, recording medium, and method for clustering data related to battery charging pattern thereof

By generating features from battery charging data, classifying them into sets, and using clustering algorithms and indices to determine optimal clusters, the method addresses data mixing and subjectivity in conventional clustering, achieving efficient and objective battery charging pattern classification.

WO2025244215A1PCT designated stage Publication Date: 2025-11-27LG ENERGY SOLUTION LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/019879
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-25
Filing Date
2024-12-05
Publication Date
2025-11-27

AI Technical Summary

Technical Problem

Conventional clustering methods for battery charging patterns face challenges such as data mixing across different types, leading to reduced performance and subjectivity in determining the appropriate number of clusters, and the lack of objective metrics for selecting the optimal number of clusters.

Method used

A method involving a processor to generate features from battery charging data, classify them into data sets, determine clusters using clustering algorithms and indices like Silhouette, Calinski-Harabasz, and Davies-Bouldin, and set a reference number of clusters based on optimal values to improve clustering efficiency and objectivity.

Benefits of technology

This approach allows for more accurate identification of battery charging patterns by minimizing data boundary influence and preventing overfitting, resulting in efficient and objective data classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024019879_27112025_PF_FP_ABST
    Figure KR2024019879_27112025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an electronic device for clustering data related to a battery charging pattern. The electronic device may be configured to: acquire a plurality of pieces of battery charging data for a plurality of vehicles; generate a plurality of features related to a battery charging pattern on the basis of the plurality of pieces of battery charging data; classify the plurality of features into a preset number of data sets by sampling the plurality of features; determine the number of clusters corresponding to each of the preset number of data sets on the basis of a clustering algorithm and an index for evaluating a clustering result; set a reference number of clusters corresponding to the plurality of features on the basis of the optimal number of clusters of each of the preset number of data sets; and cluster the plurality of features into clusters of the reference number of clusters.
Need to check novelty before this filing date? Find Prior Art

Description

Method for clustering data relating to electronic devices, recording media and their battery charging patterns

[0001] The present disclosure relates to a method for clustering data related to an electronic device, a recording medium, and a battery charging pattern thereof, and more particularly, to a technique for clustering battery charging pattern data for defining types related to charging patterns of electric vehicle users.

[0002] This application claims the benefit of priority to Republic of Korea Patent Application No. 2024-0066357, dated May 22, 2024, and Republic of Korea Patent Application No. 2024-0098789, dated July 25, 2024, the entire contents of which are incorporated herein by reference.

[0003] Clustering algorithms, an unsupervised learning method, have been proposed to standardize and categorize time-series data. However, clustering excessively large amounts of time-series data can lead to data contiguously existing on the boundaries of different data types, resulting in an excessively small number of clusters. This can lead to data with different characteristics being categorized into a single type. Considering that the purpose of clustering is to identify the characteristics of each target data type and identify valid types, clustering performance can be significantly reduced when data within each type exhibits different characteristics due to data mixing.

[0004] Therefore, conventional methods have prevented this mixing problem by clustering into a sufficiently large number of clusters, and have utilized the characteristics of probability distributions to discover and define data types through a process of overlapping and artificially merging data. However, this approach suffers from limitations such as subjectivity and the significant time required.

[0005] Furthermore, although clustering is generally an unsupervised learning method, evaluation metrics (indices) can be utilized to evaluate clustering performance, allowing for the optimal number of clusters to be selected. However, these metrics are not convex and exhibit monotonically decreasing or increasing trends, making it difficult to select an appropriate elbow point (the point where the rate of decrease of a given cost function for clustering performance evaluation sharply decreases). Therefore, there is a growing need for a method that can derive an objective metric for determining the appropriate number of clusters for clustering for type discovery.

[0006] An embodiment of the present disclosure is proposed to solve the above-described problem, and provides a method for clustering data related to an electronic device and its battery charging pattern.

[0007] The technical task to be achieved by this embodiment is not limited to the task described above, and other technical tasks can be inferred from the following examples.

[0008] An electronic device according to one embodiment includes a transceiver; a processor; and one or more memories storing one or more instructions, wherein the one or more instructions, when executed, cause the processor to obtain a plurality of battery charging data for a plurality of vehicles, generate a plurality of features regarding battery charging patterns based on the plurality of battery charging data, classify the plurality of features into a preset number of data sets by sampling the plurality of features, determine a number of clusters corresponding to each of the preset number of data sets based on a clustering algorithm and an index for evaluating a clustering result, set a reference number of clusters corresponding to the plurality of features based on an optimal number of clusters for each of the preset number of data sets, and perform clustering of the plurality of features into clusters of the reference number of clusters.

[0009] According to one embodiment, the plurality of battery charging data may include information on the state of charge (SoC) of the batteries over time, which is data collected for each of the plurality of batteries included in the plurality of vehicles over a preset period of time.

[0010] According to one embodiment, one or more instructions may be configured to cause the processor, when executed, to identify one or more charging sections based on first battery charging data for a first vehicle among a plurality of battery charging data for a plurality of vehicles, identify, for each of the one or more charging sections, a SoC of the battery at a start time of charging and a SoC of the battery at a end time of charging, and generate a first feature indicating a frequency of the one or more charging sections corresponding to a range of the SoC of the battery at the start time of charging and a range of the SoC of the battery at the end time of charging.

[0011] According to one embodiment, each of the predetermined number of data sets includes at least some features, and the number of at least some features can be determined proportionally to the number of the plurality of features.

[0012] According to one embodiment, one or more instructions may be configured to cause the processor, when executed, to perform clustering multiple times on a first data set among a preset number of data sets to have different numbers of clusters according to a clustering algorithm, and to determine an optimal number of clusters for the first data set based on the results of the multiple clusterings and an index for evaluating the results of the clusterings.

[0013] According to one embodiment, the clustering algorithm may include at least one of a hierarchical clustering algorithm that hierarchizes at least some features included in each data set based on similarity and a K-means clustering algorithm that clusters at least some features included in each data set based on centroids.

[0014] According to one embodiment, the index includes a first index, a second index, and a third index, and one or more instructions may be configured to cause the processor, when executed, to calculate, based on clustering results performed multiple times, each of a first index according to the number of clusters, a second index according to the number of clusters, and a third index according to the number of clusters, and to determine an optimal number of clusters for the first data set based on a first number of clusters corresponding to a maximum or minimum value among the first index according to the number of clusters calculated, a second number of clusters corresponding to a maximum or minimum value among the second index according to the number of clusters calculated, and a third number of clusters corresponding to a maximum or minimum value among the third index according to the number of clusters calculated.

[0015] According to one embodiment, each of the first index, the second index, and the third index may include at least one of a Silhouette Index technique, a Calinski-Harabasz Index technique, and a Davies-Bouldin Index technique that evaluate clustering results based on the cohesion of each cluster and the separation between clusters, for each clustering result performed according to a clustering algorithm.

[0016] According to one embodiment, one or more instructions, when executed, may be configured to cause the processor to set the average of the optimal number of clusters for each of a predetermined number of data sets as the reference number of clusters.

[0017] A method for clustering data regarding a battery charging pattern according to one embodiment may include the steps of: obtaining a plurality of battery charging data for a plurality of vehicles; generating a plurality of features regarding the battery charging pattern based on the plurality of battery charging data; classifying the plurality of features into a preset number of data sets by sampling the plurality of features; determining a number of clusters corresponding to each of the preset number of data sets based on a clustering algorithm and an index for evaluating a clustering result; setting a reference number of clusters corresponding to the plurality of features based on an optimal number of clusters for each of the preset number of data sets; and performing clustering of the plurality of features into clusters having the reference number of clusters.

[0018] A computer-readable, non-transitory recording medium having recorded thereon a program for executing a method for clustering data regarding a battery charging pattern according to one embodiment of the present invention on a computer, the method for clustering data regarding a battery charging pattern comprising: acquiring a plurality of battery charging data for a plurality of vehicles; generating a plurality of features regarding the battery charging pattern based on the plurality of battery charging data; classifying the plurality of features into a preset number of data sets by sampling the plurality of features; determining a number of clusters corresponding to each of the preset number of data sets based on a clustering algorithm and an index for evaluating a clustering result; setting a reference number of clusters corresponding to the plurality of features based on an optimal number of clusters for each of the preset number of data sets; and performing clustering of the plurality of features into clusters having the reference number of clusters.

[0019] According to the present disclosure, when discovering types of multiple battery charging data, by selecting an appropriate number of clusters, types of charging patterns that can more efficiently identify data characteristics can be discovered.

[0020] In addition, according to the present disclosure, optimal types can be defined by minimizing the influence of data existing on the boundary between each cluster and preventing overfitting that may occur in a specific cluster.

[0021] The effects of the invention are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art from the description of the claims.

[0022] Figure 1 illustrates a block diagram of an electronic device according to one embodiment.

[0023] FIG. 2 illustrates a flowchart of a method for clustering data regarding battery charging patterns according to one embodiment.

[0024] Figure 3 illustrates a feature generation process according to one embodiment.

[0025] Figure 4 illustrates a feature classification process according to one embodiment.

[0026] Figure 5 illustrates a process for determining the optimal number of clusters corresponding to a data set according to one embodiment.

[0027] Figure 6a shows a first index according to the number of clusters, according to one embodiment.

[0028] Figure 6b shows a second index according to the number of clusters, according to one embodiment.

[0029] Figure 6c shows a third index according to the number of clusters, according to one embodiment.

[0030] Figure 7 illustrates a process for setting the number of reference clusters according to one embodiment.

[0031] The terms used in the examples have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, in which case their meanings will be described in detail in the relevant description. Therefore, the terms used in this disclosure should not be defined simply as names, but rather based on the meanings of the terms and the overall content of the present disclosure.

[0032] When a part of the specification is said to "include" a component, this does not exclude other components, but rather implies the inclusion of other components, unless otherwise specifically stated. Furthermore, terms such as "part" and "module" used in the specification mean a unit that processes at least one function or operation, which may be implemented in hardware, software, or a combination of hardware and software.

[0033] The expression "at least one of a, b, and c" described throughout the specification may encompass 'a alone', 'b alone', 'c alone', 'a and b', 'a and c', 'b and c', or 'all of a, b, and c'.

[0034] Below, embodiments of the present disclosure are described in detail with reference to the attached drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein.

[0035]

[0036] Hereinafter, embodiments of the present disclosure relating to an electronic device for clustering data regarding a battery charging pattern are described in detail with reference to the drawings.

[0037] Figure 1 illustrates a block diagram of an electronic device according to one embodiment.

[0038] Referring to FIG. 1, an electronic device (100) may include, according to one embodiment, a transceiver (110), a processor (120), and a memory (130). The electronic device (100) illustrated in FIG. 1 only includes components related to the present embodiment. Therefore, it will be understood by those skilled in the art related to the present embodiment that other general components may be included in addition to the components illustrated in FIG. 1.

[0039] For example, the electronic device (100) may include a communication device including one or more transceivers (110), an input unit, and an output unit. The communication unit is a device for performing wired / wireless communication and may communicate with an external electronic device. The external electronic device may be a terminal or a server. In addition, communication technologies used by the communication unit may include GSM (Global System for Mobile communication), CDMA (Code Division Multi Access), LTE (Long Term Evolution), 5G, WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Bluetooth, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), ZigBee, NFC (Near Field Communication), etc. The input unit may be, for example, a traditional keypad or keyboard, a mouse, a microphone for inputting voice signals, a camera, and various other input means for detecting or receiving various types of user input. The output unit may be, for example, a display that outputs images, a speaker that outputs sounds, a haptic device that generates vibrations, and various other forms of output means.

[0040] According to one embodiment, the electronic device (100) may be a server that acquires and processes data for multiple vehicles. Specifically, the data for multiple vehicles may include multiple battery charge data for the multiple vehicles. The multiple battery charge data may be acquired, for example, from at least one of an On-Board Diagnostics (OBD) device mounted on each vehicle, a battery management system (BMS), and a device (e.g., a database) in which battery charge data is previously stored, via the transceiver (110) of the electronic device (100). The manner in which the electronic device (100) acquires the multiple battery charge data is not limited to the above example, and it will be clearly understood by those skilled in the art that the electronic device (100) may acquire the multiple battery charge data from various devices with which it can communicate via the transceiver (110). The type of the electronic device (100) is not limited thereto, and various embodiments of the present disclosure may be applied to various devices capable of acquiring and processing data for vehicles.

[0041] The processor (120) can control the overall operation of the electronic device (100) and process data and signals. The processor (120) can be composed of at least one hardware unit. In addition, the processor (120) can operate by one or more software modules generated by executing program codes stored in the memory (130). The processor (120) can include a memory, and the processor (120) can control the overall operation of the electronic device (100) and process data and signals by executing the program codes stored in the memory.

[0042] The processor (120) may be implemented as a computer or similar device based on hardware, software, or a combination thereof. In terms of hardware, the processor (120) may be implemented in the form of an electronic circuit that processes electrical signals to perform control functions, and in terms of software, the processor (120) may be implemented in the form of a program that drives the hardware processor (120). Meanwhile, unless otherwise specified in the following description, the operation of the electronic device may be interpreted as being performed under the control of the processor (120). That is, when modules implemented in the clustering system for data regarding battery charging patterns are executed, the modules may be interpreted as controlling the processor (120) to perform the following operations of the electronic device (100).

[0043] The memory (130) can store various types of information. The memory (130) can store data temporarily or semi-permanently. For example, the memory (130) of the electronic device (100) can store data related to an operating program (OS: Operating System) for operating the electronic device (100). Examples of the memory (130) may include a hard disk drive (HDD: Hard Disk Drive), a solid state drive (SSD), flash memory, read-only memory (ROM: Read-Only Memory), random access memory (RAM: Random Access Memory), etc. The memory (130) may be provided as a built-in type or a detachable type.

[0044] In summary, the various embodiments may be implemented through various means. For example, the various embodiments may be implemented through hardware, firmware, software, or a combination thereof.

[0045] In the case of hardware implementation, the methods according to various embodiments may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, etc.

[0046] When implemented via firmware or software, the methods according to various embodiments may be implemented in the form of modules, procedures, or functions that perform the functions or operations described above. For example, software code may be stored in memory and executed by a processor. The memory may be located within or external to the processor and may exchange data with the processor via various known means.

[0047] FIG. 2 illustrates a flowchart of a method for clustering data regarding battery charging patterns according to one embodiment.

[0048] In step S210, the electronic device (100) may acquire multiple battery charging data for multiple vehicles. Here, the multiple battery charging data is data collected for each of the multiple batteries included in the multiple vehicles over a preset period of time, and may include information on the state of charge (SoC) of the batteries over time. The preset period may be a period during which battery charging and discharging may be repeated multiple times, and may be, for example, one month. Accordingly, the multiple battery charging data may be time-series data including information on changes in SoC over time during multiple repetitions of charging and discharging of each of the multiple batteries included in the multiple vehicles. A more specific embodiment of the battery charging data will be described in detail with reference to FIG. 3 below.

[0049] At step S220, the electronic device (100) may generate multiple features regarding a battery charging pattern based on multiple battery charging data. A more specific process by which the electronic device (100) generates the features will be described in detail with reference to FIG. 3 below.

[0050] At step S230, the electronic device (100) can classify the plurality of features into a preset number of data sets by sampling the plurality of features. The sampling of the features can be performed randomly, for example. Furthermore, according to an embodiment, the electronic device (100) can perform the sampling evenly so that the composition of each data set is not biased toward features related to a specific vehicle. A more specific embodiment of classifying the plurality of features into a preset number of data sets will be described in detail with reference to FIG. 3 below.

[0051] In step S240, the electronic device (100) may determine the number of clusters corresponding to each of a preset number of data sets based on a clustering algorithm and an index for evaluating the clustering result. The clustering algorithm may be an algorithm that divides the data to be performed into clusters with similar characteristics, and the index may be an index that quantitatively evaluates the performance of clustering based on the cohesion of each cluster (a measure of how closely data within the same cluster are grouped) and the separation between clusters (a measure of how clearly different clusters are distinguished) according to the clustering. A more specific embodiment in which the electronic device (100) determines the number of clusters for each data set based on the clustering algorithm and the index will be described in detail with reference to FIG. 4 below.

[0052] At step S250, the electronic device (100) can set a reference number of clusters corresponding to multiple features based on the optimal number of clusters for each of a preset number of data sets. That is, the electronic device (100) can finally set a reference number of clusters that can more efficiently cluster multiple features based on the optimal number of clusters for each data set. A more specific embodiment in which the electronic device (100) sets the reference number of clusters using the optimal number of clusters for each data set will be described in detail with reference to FIG. 6 below.

[0053] At step S260, the electronic device (100) can cluster multiple features into clusters based on a reference cluster count. By performing this clustering, types that enable more detailed and objective classification of battery charging patterns can be identified.

[0054] Figure 3 illustrates a feature generation process according to one embodiment.

[0055] According to one embodiment, the electronic device (100) may identify one or more charging sections based on first battery charging data for a first vehicle among a plurality of battery charging data for a plurality of vehicles, and may identify the SoC of the battery at the start time of charging and the SoC of the battery at the end time of charging for each of the one or more charging sections. Subsequently, the electronic device (100) may generate a first feature indicating the frequency of one or more charging sections corresponding to the range of the SoC of the battery at the start time of charging and the range of the SoC of the battery at the end time of charging as a feature regarding the charging pattern of the first battery charging data for the first vehicle. More specifically, the electronic device (100) may analyze the frequency of the charging sections based on the first battery charging data, and estimate a probability distribution function regarding the frequency of each charging section. Subsequently, the electronic device (100) may generate an image-type feature indicating the frequency of the battery charging sections of the corresponding vehicle based on the estimated probability distribution function. The electronic device (100) can generate multiple features regarding battery charging patterns by performing this process for each of multiple battery charging data.

[0056] Each of the multiple features in image form may represent a specific battery charging pattern extracted from battery charging data for a specific vehicle. Specifically, if the users of the vehicles for which battery charging data is collected are different for each vehicle, each feature may represent a charging pattern for each user of the corresponding vehicle.

[0057] Referring to FIG. 3, the first battery charging data for the first vehicle may be in the form of a graph (300) with time as the x-axis and SoC as the y-axis. Specifically, the graph (300) may include SoC information of the first battery of the first vehicle for a preset period of time, and more specifically, may include information about a charging section confirmed each time the first battery is charged. Accordingly, the graph (300) may indicate at what SoC the charging starts and at what SoC the charging ends each time the first battery is charged. The electronic device (100) may analyze the frequency of a plurality of charging sections confirmed based on the first battery charging data graph (300) to estimate a probability distribution function regarding the frequency of the charging sections, and may generate a feature (310) regarding the first battery charging data based on the estimated probability distribution function.

[0058] The feature (310) generated in response to the first battery charging data may be a two-dimensional image with the charging start SoC as the x-axis and the charging end SoC as the y-axis, and the shading of the feature (310) may indicate the frequency of the charging section. That is, the darker the shading in the feature (310), the more frequently the first battery is charged. Accordingly, the feature (310) may indicate a charging habit or tendency regarding the SoC at which charging mainly starts and ends in the corresponding vehicle, and as a more specific example, the feature (310) may indicate that the user of the first vehicle has a charging habit of mainly starting charging at a low SoC and ending charging at a low SoC.

[0059] Figure 4 illustrates a feature classification process according to one embodiment.

[0060] Referring to FIG. 4, based on a plurality of features (410) regarding a battery charging pattern, a preset number (n) of data sets (420) including data set 1 (422), data set 2 (424), ..., data set n (426) may be generated. Each of the preset number of data sets may include, for example, at least some features, and the number of at least some features included in each data set may be determined in proportion to the number of the plurality of features. As a specific example, if a total of 10,000 features are generated and the n value is set to 10, the number of at least some features included in each data set may be set to 800, which is 8,000, which corresponds to 80% of 10,000, divided by 10. In order to obtain data sets in which the generated features are evenly sampled, according to an embodiment, the number of at least some features included in each data set may be determined to a more appropriate value by considering the total number of generated features.

[0061] Figure 5 illustrates a process for determining the optimal number of clusters corresponding to a data set according to one embodiment.

[0062] According to one embodiment, the electronic device (100) may perform clustering multiple times on a first data set among a preset number of data sets so that the first data set has different numbers of clusters according to a clustering algorithm, and may determine an optimal number of clusters for the first data set based on the results of the multiple clusterings and an index for evaluating the results of the clusterings. Specifically, the electronic device (100) may perform clustering multiple times on the first data set so that the number of clusters is k (k∈{2, 3, ..., k}) according to the clustering algorithm. The electronic device (100) may cluster the first data set multiple times so that the number of clusters varies each time clustering is performed.

[0063] In one embodiment, the clustering algorithm performed on the first data set may be any type of clustering algorithm that allows the number of clusters to be set as a parameter (i.e., hyperparameterized). For example, it may be a hierarchical clustering algorithm that hierarchizes at least some features included in each data set based on similarity. The hierarchical clustering algorithm may have the advantage of very small computational load because it iterates the process of calculating an initial distance matrix and then merging or dividing clusters with close distances into sub-clusters based on the calculated distance matrix. In another example, the clustering algorithm may be a K-means clustering algorithm that clusters at least some features included in each data set based on centroids. The K-means clustering algorithm performs clustering by assigning data points to the closest cluster based on an initially selected centroid, calculating the centroids of the clusters, and repeating the process until the centroids are the same. Since this iterative operation can be processed in parallel, it may have the advantage of enabling faster computation on large data sets.

[0064] Next, the electronic device (100) can determine the optimal number of clusters for the first data set based on the indexes for each of the clustering results performed multiple times. The indexes may include, for example, a first index, a second index, and a third index, and the electronic device (100) can calculate a first index according to the number of clusters, a second index according to the number of clusters, and a third index according to the number of clusters, based on the clustering results performed multiple times. The index according to the number of clusters may refer to a performance index that evaluates the results of performing clustering with different k values, and may be organized in the form of a graph, for example. Meanwhile, the types of indexes that the electronic device (100) calculates based on the clustering results performed multiple times are not limited to the three types of indexes, and it will be clearly understood by those skilled in the art that more types of indexes may be calculated depending on the embodiment.

[0065] According to one embodiment, each of the first index, the second index, and the third index may include at least one of a Silhouette Index technique, a Calinski-Harabasz Index technique, and a Davies-Bouldin Index technique, which evaluate clustering results based on the cohesion of each cluster and the separation between clusters, for each clustering result performed according to a clustering algorithm. The techniques have differences in the process of calculating cohesion and separation. For example, according to the Silhouette Index technique, cohesion may be calculated based on the average of the distances between each data point and all other data points in the same cluster, and separation may be calculated based on the average of the distances between each cluster and all other data points in the closest other cluster. In contrast, according to the Kalinsky-Harabats index technique, cohesion is calculated based on the average square of the distance between each data point and its cluster centroid, while separability can be calculated based on the average square of the distance between each cluster's centroid and the entire data set's centroid. Furthermore, according to the Davis-Boldin index technique, cohesion is calculated based on the average of the distance between each data point and its cluster centroid, while separability is calculated based on the average of the distance between each cluster's centroid.

[0066] Next, the electronic device (100) may determine an optimal number of clusters for the first data set based on a first cluster number corresponding to a maximum or minimum value among the first indexes according to the calculated number of clusters, a second cluster number corresponding to a maximum or minimum value among the second indexes according to the calculated number of clusters, and a third cluster number corresponding to a maximum or minimum value among the third indexes according to the calculated number of clusters. Here, each of the maximum or minimum value among the first indexes, the maximum or minimum value among the second indexes, and the maximum or minimum value among the third indexes may correspond to an optimal value among the maximum and minimum values ​​set based on the characteristics of the indexes. For example, when the index is a silhouette index technique or a Kalinsky-Harabats index technique, the maximum value may correspond to the optimal value, and when the index is a Davis-Boldin index technique, the minimum value may correspond to the optimal value. Meanwhile, the optimal value for each index is not limited to the maximum or minimum value, and can be set to an appropriate value depending on the characteristics of the index.

[0067] In one embodiment, the optimal number of clusters for the first data set may be set to an overlapping value among the first number of clusters, the second number of clusters, and the third number of clusters. That is, the number of clusters where the maximum or minimum values ​​intersect as evaluated by each of the three indices may be set as the optimal number of clusters. For example, if the first number of clusters and the second number of clusters are 9, and the third number of clusters is 8, the optimal number of clusters may be set to 9. In this way, by setting the number of clusters evaluated as having the optimal value according to various index techniques as the optimal number of clusters for the first data set, clustering can be performed with guaranteed high performance. In another embodiment, if the first number of clusters, the second number of clusters, and the third number of clusters all have different values, their average value may be set as the optimal number of clusters for the first data set.

[0068] Referring to FIG. 5, which illustrates a process (500) for setting the optimal number of clusters for data set 1 (422), the electronic device (100) can perform multiple clustering with k cluster numbers for data set 1 (422). As described above, the multiple clustering can be performed with different numbers of clusters for different k values. Subsequently, the electronic device (100) can calculate a first index (522) according to the k value, a second index (524) according to the k value, and a third index (526) according to the k value, and can determine the optimal number of clusters for data set 1 (422) based on the first index (522), the second index (524) according to the k value, and the third index (526) according to the k value.

[0069] The electronic device (100) can repeatedly perform the same process for each of a preset number of data sets, and this repetition may be performed sequentially or simultaneously in parallel, depending on the embodiment. In the past, in order to achieve higher performance clustering, the intervention of a worker was required to identify data judged to be similar and merge them one by one to reduce the number of clusters. However, according to the present invention, the number of reference clusters can be set as a converged value through the above-described repeated process utilizing a clustering algorithm and index without the intervention of a worker, so that more efficient and objective data classification can be possible.

[0070] Figure 6a shows a first index according to the number of clusters, according to one embodiment.

[0071] Here, the first index may be a silhouette index. The graph (600) represents the silhouette index evaluated for each value of the number of clusters k ranging from 2 to 30, with the number of clusters (k) on the x-axis and the silhouette index on the y-axis. SI JS For graphs, the silhouette index is calculated using Jensen-Shannon Divergence, which measures the similarity between data points based on probability distribution, and SI EUCFor graphs, the silhouette index is calculated using the Euclidean distance, which calculates the distance between data points based on the straight-line distance between data points. In the silhouette index technique, considering that the maximum value is the optimal value, the number of clusters corresponding to the optimal value among the first indices is SI JS Graphs and SI EUC For the graph, it can all be 7.

[0072] Figure 6b shows a second index according to the number of clusters, according to one embodiment.

[0073] Here, the second index may be the Kalinsky-Harabatz index. The graph (610) represents the Kalinsky-Harabatz index (CHI) evaluated for each value of the number of clusters k ranging from 2 to 30, with the number of clusters (k) on the x-axis and the Kalinsky-Harabatz index on the y-axis. Considering that the maximum value is the optimal value in the Kalinsky-Harabatz index technique, the number of clusters corresponding to the optimal value among the second indices may be 7, as in FIG. 6a.

[0074] Figure 6c shows a third index according to the number of clusters, according to one embodiment.

[0075] Here, the third index may be the Davis-Boldin index. Graph (620) represents the Davis-Boldin index (DBI) evaluated for each value of the number of clusters k ranging from 2 to 30, with the number of clusters (k) on the x-axis and the Davis-Boldin index on the y-axis. Considering that the minimum value is the optimal value in the Davis-Boldin index technique, the number of clusters corresponding to the optimal value among the third indices may be 7, as in FIGS. 6a and 6b.

[0076] That is, in the examples of FIGS. 6A to 6C, the optimal number of clusters of the target data set determined based on the produced indices may be 7.

[0077] Figure 7 illustrates a process for setting the number of reference clusters according to one embodiment.

[0078] According to one embodiment, the electronic device (100) may set the average value of the optimal number of clusters for each of a preset number of data sets as the reference number of clusters. Referring to FIG. 7, the electronic device (100) may determine the optimal number of clusters corresponding to each data set by repeatedly performing the process described in detail with respect to FIGS. 4 and 5 for each of data set 1 (422), data set 2 (424), ..., data set n (426). The optimal number of clusters corresponding to data set 1 (422) is K1, the optimal number of clusters corresponding to data set 2 (424) is K2, ..., the optimal number of clusters corresponding to data set n (426) is K n When this is said, the number of reference clusters set to perform clustering on multiple features is K. 1, K2, ..., K n That is, the electronic device (100) can cluster multiple features into a more appropriate number by setting the optimal reference number of clusters corresponding to multiple features based on the optimal number of clusters corresponding to each data set generated by sampling multiple features.

[0079]

[0080] The electronic device according to the above-described embodiments may include a processor, a memory for storing and executing program data, permanent storage such as a disk drive, a communication port for communicating with an external device, a user interface device such as a touch panel, a key, a button, etc. The methods implemented as software modules or algorithms may be stored on a computer-readable recording medium as computer-readable codes or program instructions executable on the processor. Here, the computer-readable recording medium includes a magnetic storage medium (e.g., read-only memory (ROM), random-access memory (RAM), floppy disk, hard disk, etc.) and an optical reading medium (e.g., CD-ROM, DVD: Digital Versatile Disc)). The computer-readable recording medium may be distributed to computer systems connected to a network, so that the computer-readable code may be stored and executed in a distributed manner. The medium may be readable by a computer, stored in a memory, and executed by a processor.

[0081] The present embodiment may be represented by functional block configurations and various processing steps. These functional blocks may be implemented by various hardware and / or software configurations that perform specific functions. For example, the embodiment may employ direct circuit configurations such as memory, processing, logic, look-up tables, etc., which may perform various functions under the control of one or more microprocessors or other control devices. Similarly, the present embodiment may be implemented in a programming or scripting language such as C, C++, Java, assembler, etc., including various algorithms implemented as a combination of data structures, processes, routines, or other programming configurations. Functional aspects may be implemented as algorithms that execute on one or more processors. Furthermore, the present embodiment may employ conventional techniques for electronic configuration, signal processing, and / or data processing. Terms like "mechanism," "element," "means," and "composition" can be used broadly and are not limited to mechanical or physical components. These terms can also encompass a series of software routines, such as those associated with a processor.

[0082] The above-described embodiments are merely examples, and other embodiments may be implemented within the scope of the claims set forth below.

Claims

1. In electronic devices, transceiver; processor; and Contains one or more memories that store one or more instructions, The one or more instructions, when executed, cause the processor to: Obtain multiple battery charging data for multiple vehicles, Based on the above plurality of battery charging data, a plurality of features regarding a battery charging pattern are generated, By sampling the above plurality of features, the plurality of features are classified into a preset number of data sets, Based on the clustering algorithm and the index for evaluating the clustering results, the number of clusters corresponding to each of the above-described number of data sets is determined, Based on the optimal number of clusters for each of the above-described number of data sets, a reference number of clusters corresponding to the plurality of features is set, An electronic device configured to perform clustering of the above plurality of features into clusters having the above reference number of clusters.

2. In paragraph 1, The above multiple battery charging data is, An electronic device comprising, for each of a plurality of batteries included in the plurality of vehicles, data collected over a preset period of time, information on the state of charge (SoC) of the battery over time.

3. In paragraph 2, The one or more instructions, when executed, cause the processor to: Based on the first battery charging data for the first vehicle among the plurality of battery charging data for the plurality of vehicles, one or more charging sections are identified, For each of the above one or more charging sections, check the SoC of the battery at the start of charging and the SoC of the battery at the end of charging, An electronic device configured to generate a first feature indicating a frequency of one or more charging intervals corresponding to a range of SoC of the battery at a start time of charging and a range of SoC of the battery at a end time of charging.

4. In paragraph 1, Each of the above-described number of data sets contains at least some features, An electronic device, wherein the number of at least some of the features is determined in proportion to the number of the plurality of features.

5. In paragraph 1, The one or more instructions, when executed, cause the processor to: For the first data set among the above-described number of data sets, clustering is performed multiple times to have different numbers of clusters according to the clustering algorithm, An electronic device configured to determine the optimal number of clusters for the first data set based on the results of the clustering performed multiple times and an index for evaluating the results of the clustering.

6. In paragraph 5, The above clustering algorithm is, An electronic device comprising at least one of a hierarchical clustering algorithm that hierarchizes at least some features included in each data set based on similarity and a K-means clustering algorithm that clusters at least some features included in each data set based on center points.

7. In paragraph 5, The above index includes a first index, a second index, and a third index, The one or more instructions, when executed, cause the processor to: Based on the results of the clustering performed multiple times, the first index according to the number of clusters, the second index according to the number of clusters, and the third index according to the number of clusters are each calculated, An electronic device configured to determine an optimal number of clusters for the first data set based on a first number of clusters corresponding to a maximum or minimum value among the first indexes according to the calculated number of clusters, a second number of clusters corresponding to a maximum or minimum value among the second indexes according to the calculated number of clusters, and a third number of clusters corresponding to a maximum or minimum value among the third indexes according to the calculated number of clusters.

8. In paragraph 7, Each of the first index, the second index and the third index, An electronic device comprising at least one of the Silhouette Index technique, the Calinski-Harabasz Index technique, and the Davies-Bouldin Index technique, which evaluates the clustering results based on the cohesion of each cluster and the separation between clusters, for each clustering result performed according to the above clustering algorithm.

9. In paragraph 1, The one or more instructions, when executed, cause the processor to: An electronic device configured to set the average value of the optimal number of clusters for each of the above-described number of data sets as the reference number of clusters.

10. A method for clustering data regarding battery charging patterns performed by an electronic device, A step of acquiring multiple battery charging data for multiple vehicles; A step of generating a plurality of features regarding a battery charging pattern based on the plurality of battery charging data; A step of classifying the plurality of features into a preset number of data sets by sampling the plurality of features; A step of determining the number of clusters corresponding to each of the preset number of data sets based on a clustering algorithm and an index for evaluating the clustering results; A step of setting a reference number of clusters corresponding to the plurality of features based on the optimal number of clusters for each of the above-described number of data sets; and A method for clustering data regarding battery charging patterns, comprising the step of performing clustering of the plurality of features into clusters having the number of reference clusters.

11. A non-transitory computer-readable recording medium having recorded thereon a program for executing the method of Article 10 on a computer.

Citation Information

Patent Citations

  • Electronic apparatus, recording medium, and method for clustering data regarding battery charging pattern thereof

    KR1020250167451A

  • Improved functional pillow

    KR1020250059684A

  • A food composition for preventing or improving osteoporosis comprising propolis bioconversion extract as an active ingredient

    KR1020250161974A

  • Method and apparatus for detecting anomaly of time series power data

    KR102614798B1

  • Battery Detachment Structure and Battery Replacement Method of Unmanned Aircraft

    KR102677782B1