Charging station data processing method and device, electronic equipment and storage medium

By setting a similarity threshold and using a connected graph method, charging station data is separated and fused, solving the problem of fusion errors in data with medium similarity and achieving more accurate and efficient charging station site selection planning.

CN121959034APending Publication Date: 2026-05-01ZHEJIANG XIAOJU GREEN ENERGY TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG XIAOJU GREEN ENERGY TECHNOLOGY CO LTD
Filing Date
2024-10-24
Publication Date
2026-05-01

Smart Images

  • Figure CN121959034A_ABST
    Figure CN121959034A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a charging station data processing method and device, electronic equipment and a storage medium. The method comprises the steps that a station data set and an undetermined data set are determined according to the relation between the similarity between different charging station data sets from different data sources and a first similarity threshold value and the relation between the similarity between the different charging station data sets from different data sources and a second similarity threshold value; fusing the attribute information in each charging station data group in the station data set of which the similarity reaches a second similarity threshold, and manually fusing the attribute information in each charging station data group in the undetermined data set of which the similarity is between the first similarity threshold and the second similarity threshold, and the charging station distribution parameters are determined according to the fused actual charging station data, and the supply distribution of the charging stations is optimized according to the charging station distribution parameters, so that the accuracy and efficiency of data fusion can be considered at the same time, and convenience is provided for arrangement optimization of the charging stations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and more specifically to a data processing method, apparatus, electronic device, and storage medium for charging stations. Background Technology

[0002] In existing multi-source data fusion methods, when calculating data similarity, the results are simply divided into high similarity or low similarity. High similarity means that the two data actually refer to the same entity, while low similarity means that the two data actually refer to different entities.

[0003] However, this method has a problem of incorrectly treating two data points with moderate similarity. For example, two data points with moderate similarity may actually refer to different entities, but because their similarity is slightly higher than a set threshold, the fusion method incorrectly identifies them as referring to the same entity. Or, two data points with moderate similarity may actually refer to the same entity, but because their similarity is slightly lower than a set threshold, the fusion method incorrectly identifies them as referring to different entities. Furthermore, while this error is acceptable in some scenarios, it can have a significant negative impact in scenarios where high accuracy is required, such as charging station site selection and planning. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a charging station data processing method that simultaneously considers the accuracy and efficiency of data fusion, thereby facilitating the optimization of charging station layout.

[0005] In a first aspect, embodiments of the present invention aim to provide a charging station data processing method, the method comprising:

[0006] Obtain an initial dataset, which includes multiple charging station data sets from different data sources;

[0007] Based on the attribute information of each charging station data group, at least one charging station data group and at least one undetermined data group corresponding to the initial dataset are determined. The similarity between different charging station data groups in the charging station data group reaches a second similarity threshold. The similarity between different charging station data groups in the undetermined data group is between a first similarity threshold and a second similarity threshold. The first similarity threshold is less than the second similarity threshold.

[0008] According to a preset attribute priority order, the attribute information in each charging station data group in each of the station datasets is fused to determine the first fused data, which includes the actual charging station data of the charging station represented by the corresponding station dataset.

[0009] The charging station data group in the pending dataset is sent to the manual processing terminal to determine the second fused data based on the manual processing results. The second fused data includes the actual charging station data of at least one charging station.

[0010] The charging station distribution parameters are determined based on the first fused data and the second fused data, so as to optimize the supply distribution of charging stations according to the charging station distribution parameters.

[0011] Further, determining at least one charging station dataset and at least one pending dataset corresponding to the initial dataset based on the attribute information of each of the charging station data groups includes:

[0012] The data groups of each charging station in the initial dataset are combined in pairs to determine at least one filtered data pair;

[0013] The overall similarity result of each of the filtered data pairs is determined based on the attribute information of each of the charging station data groups. The overall similarity result is used to characterize the similarity between the charging station data groups in the corresponding filtered data pairs.

[0014] The overall similarity results of each of the selected data pairs are compared with the first similarity threshold to determine at least one similar data pair in the initial dataset. The similar data pair is the selected data pair whose overall similarity results reach the first similarity threshold.

[0015] The overall similarity results of each of the similar data pairs are compared with the second similarity threshold to determine at least one site dataset and at least one undetermined dataset corresponding to the initial dataset.

[0016] Further, the step of comparing the overall similarity result of each of the similar data pairs with the second similarity threshold to determine at least one site dataset and at least one pending dataset corresponding to the initial dataset includes:

[0017] The overall similarity result of each of the similar data pairs is compared with the second similarity threshold to filter and determine the matching data pairs and the data pairs to be matched. The matching data pairs are the similar data pairs whose overall similarity result reaches the second similarity threshold, and the data pairs to be matched are the similar data pairs whose overall similarity result reaches the first similarity threshold and is less than the second similarity threshold.

[0018] Based on the connected graph method, the charging station data groups in each of the matched data pairs are clustered to determine at least one station dataset.

[0019] Based on the connected graph method, the charging station data groups in each of the data pairs to be matched are clustered to determine at least one undetermined dataset.

[0020] Furthermore, the charging station data group includes at least one attribute information, and determining the overall similarity result of each filtered data pair based on the attribute information of each charging station data group includes:

[0021] Determine the attribute similarity results between attribute information of the same type in the filtered data pairs;

[0022] The overall similarity result of the corresponding filtered data pairs is determined based on the attribute similarity results of different types of attribute information.

[0023] Furthermore, the charging station data group includes the station name, operator name, city name, address name, latitude and longitude information, number of fast charging guns and / or number of slow charging guns.

[0024] Furthermore, determining the attribute similarity results between attribute information of the same type in the filtered data pairs includes:

[0025] The names of charging stations in each data group of the filtered data pair are encoded to determine the word vectors of the corresponding station names;

[0026] The word vectors are similarity calculated to determine the word vector similarity between corresponding word vectors, so as to determine the attribute similarity between corresponding station names based on the word vector similarity.

[0027] Furthermore, determining the attribute similarity results between attribute information of the same type in the filtered data pairs includes:

[0028] The distance between charging stations is determined based on the latitude and longitude information in each charging station data group in the filtered data pair. The distance between charging stations is used to represent the distance between charging stations represented by different charging station data groups.

[0029] The similarity of attributes between corresponding latitude and longitude information is determined based on the relationship between the distance and the preset distance threshold.

[0030] Furthermore, determining the attribute similarity results between attribute information of the same type in the filtered data pairs includes:

[0031] Determine the difference in the number of fast charging guns in each charging station data group within the filtered data pair;

[0032] The similarity of attributes between the corresponding fast charging gun quantities is determined based on the relationship between the quantity difference and the preset quantity threshold.

[0033] Furthermore, determining the overall similarity result of the corresponding filtered data pairs based on the attribute similarity results of different types of attribute information includes:

[0034] The similarity results of each attribute are compared with a preset matching rule table to determine the overall similarity result of the corresponding filtered data pairs.

[0035] Furthermore, determining the overall similarity result of the corresponding filtered data pairs based on the attribute similarity results of different types of attribute information includes:

[0036] The weights corresponding to the attribute information representing similarity results are accumulated and summed to determine the summation result;

[0037] The summation result is compared with the preset weight intervals of each similarity level to determine the overall similarity result of the corresponding filtered data pair.

[0038] Secondly, embodiments of the present invention aim to provide a charging station data processing device, the device comprising:

[0039] An acquisition unit is used to acquire an initial dataset, which includes multiple charging station data groups from different data sources;

[0040] The processing unit is configured to determine at least one charging station dataset and at least one pending dataset corresponding to the initial dataset based on the attribute information of each charging station data group. The similarity between different charging station data groups in the charging station dataset reaches a second similarity threshold. The similarity between different charging station data groups in the pending dataset is between a first similarity threshold and the second similarity threshold. The first similarity threshold is less than the second similarity threshold.

[0041] An analysis unit is configured to fuse attribute information in each charging station data group within each of the aforementioned station datasets according to a preset attribute priority order, to determine first fused data, wherein the first fused data includes actual charging station data of the charging station represented by the corresponding station dataset; and to send the charging station data group in the pending dataset to a manual processing terminal to determine second fused data based on the manual processing results, wherein the second fused data includes actual charging station data of at least one charging station.

[0042] An optimization unit is used to determine charging station distribution parameters based on the first fused data and the second fused data, so as to optimize the supply distribution of charging stations based on the charging station distribution parameters.

[0043] Thirdly, embodiments of the present invention aim to provide a computer program product, the computer program product including a computer program / instruction, which, when executed by a processor, implements the method described in any of the preceding claims.

[0044] Fourthly, embodiments of the present invention aim to provide an electronic device, including a memory and a processor, wherein the memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any of the preceding claims.

[0045] Fifthly, embodiments of the present invention aim to provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in any of the preceding claims.

[0046] The technical solution of this invention obtains an initial dataset comprising multiple charging station data groups from different data sources. Based on the attribute information of each charging station data group and the relationship between the similarity between different charging station data groups and a first similarity threshold and a second similarity threshold, it determines at least one charging station dataset and at least one undetermined dataset corresponding to the initial dataset. It then fuses the attribute information of each charging station data group in the charging station dataset where the similarity reaches the second similarity threshold, and manually fuses the attribute information of each charging station data group in the undetermined dataset where the similarity is between the first and second similarity thresholds. This allows for the classification of charging station data groups from different data sources according to different similarity thresholds. Highly similar charging station data is automatically and quickly fused, while moderately similar charging station data is accurately fused through manual processing. This achieves both accuracy and efficiency in data fusion when fusing data from different data sources. Furthermore, by using the fused first and second fused data to determine charging station distribution parameters, the supply distribution of charging stations can be optimized based on these parameters, facilitating the optimization of charging station layout. Attached Figure Description

[0047] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:

[0048] Figure 1 This is a flowchart of a charging station data processing method according to an embodiment of the present invention;

[0049] Figure 2 This is a flowchart illustrating the process of determining the site dataset and the dataset to be determined according to an embodiment of the present invention;

[0050] Figure 3 This is a flowchart illustrating the determination of overall similarity results according to an embodiment of the present invention;

[0051] Figure 4 This is a schematic diagram of the matching rule table according to an embodiment of the present invention;

[0052] Figure 5This is a flowchart illustrating the process of determining the site dataset according to an embodiment of the present invention;

[0053] Figure 6 This is a schematic diagram illustrating the determination of the site dataset and the dataset to be determined according to an embodiment of the present invention;

[0054] Figure 7 This is a charging station data processing device according to an embodiment of the present invention;

[0055] Figure 8 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0056] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.

[0057] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0058] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".

[0059] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0060] The solutions described in this specification and embodiments, if involving information acquisition, will collect data under legal and compliant conditions, ensuring the legality of the data source, and will take appropriate technical and management measures to ensure data security. If involving personal information processing, processing will be carried out under legal grounds (e.g., obtaining the consent of the personal information subject, or being necessary for contract performance), and will only be conducted within the prescribed or agreed scope. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0061] This embodiment uses the data fusion processing of charging station sites during site selection and planning as an example to illustrate the data processing method, aiming to balance accuracy and efficiency during data fusion. However, it should be understood that the data processing method in this embodiment can also be applied to other application scenarios that require the fusion of similar data; the application scenarios of the method are not limited here.

[0062] When planning the site selection for charging stations, it is crucial to understand the current distribution of charging station supply in the city to avoid constructing stations in areas of oversupply or undersupply, which could lead to investment losses. Charging station information, such as station name, operator name, address, and the number of fast and slow charging guns, may come from different data sources. The way this information is expressed may also vary; for example, the name of the same charging station may differ across data sources, address descriptions may vary depending on the level of detail, and even station names may differ from their actual locations. Therefore, a method is needed to quickly and accurately integrate multi-source data for the same station to accurately assess the supply distribution of charging stations and facilitate optimized station deployment.

[0063] Figure 1 This is a flowchart of a charging station data processing method according to an embodiment of the present invention. Figure 1 As shown in the figure, the data processing method in this embodiment includes the following steps.

[0064] In step S110, the initial dataset is obtained.

[0065] In this embodiment, the initial dataset includes multiple charging station data groups from different data sources. Each charging station data group has a corresponding data source and charging station, and includes at least one attribute information obtained from the corresponding data source to characterize the corresponding charging station. For example, the initial dataset includes data from data source A and data source B. Data source A includes charging station data groups A1, A2, and A3. Charging station data group A1 includes relevant data (i.e., attribute information) characterizing charging station M, and charging station data groups A2 and A3 are relevant data characterizing charging station N. Data source B includes charging station data groups B1 and B2. Charging station data group B1 is relevant data characterizing charging station M, and charging station data group B2 is relevant data characterizing charging station N.

[0066] Optionally, the charging station data group in this embodiment includes the station name, operator name, city name, address name, latitude and longitude information, number of fast charging guns, number of slow charging guns and / or other attribute information related to the charging station.

[0067] Meanwhile, in this embodiment, different data sources can correspond to different charging station locations. In this case, the initial dataset includes charging station data from different cities, and each city has multiple charging station data sets. Different data sources can also correspond to different charging station data acquisition methods. In this case, the initial dataset includes charging station data acquired through different data acquisition methods, and each data acquisition method has multiple charging station data sets. Optionally, the data sources under different data acquisition methods include equipment databases provided by charging station operators, charging station maintenance databases, or other sources from which charging station data can be reasonably obtained, such as equipment databases provided by equipment suppliers.

[0068] Optionally, considering that charging stations are usually deployed in specific geographical areas (e.g., a city), and that there are significant differences between charging station data in different geographical areas, the initial dataset in this embodiment corresponds to a single geographical area (e.g., a city). It includes multiple charging station data from multiple sources obtained through different data acquisition methods. This facilitates the rapid determination of the actual charging station data for each city, improving data fusion efficiency.

[0069] Furthermore, to improve the accuracy of data fusion, this embodiment first acquires an original dataset containing charging station data groups from multiple cities when obtaining the initial dataset. The acquisition methods for each city's charging station data groups can be the same or different. Then, based on the city labels of each charging station data group, the original dataset's charging station data groups are clustered to obtain the corresponding initial dataset for each city. This ensures the comprehensiveness of the data in the initial dataset and improves the accuracy of subsequent data fusion results. Simultaneously, this embodiment can determine the original dataset based on existing data sources for each city, thereby avoiding unnecessary data collection costs.

[0070] Optionally, taking a single city as an example, in this embodiment, after obtaining the original dataset, the charging station data group in the original dataset can be preprocessed first, and then the initial dataset for a single city can be determined from the preprocessed original dataset, so as to execute subsequent data processing methods based on the initial dataset. Alternatively, in this embodiment, the initial dataset for a single city can also be determined first from the original dataset, and then the initial dataset can be preprocessed, so as to execute subsequent data processing methods based on the preprocessed initial dataset. The specific choice can be made according to the actual use scenario.

[0071] Furthermore, when preprocessing charging station data, for attribute information containing characters such as station name, operator name, and address name in the charging station data group, this embodiment will remove special characters (such as separators) that have no obvious meaning from the attribute information, and only retain the Chinese and English names and Arabic numerals. For charging station data that only includes address name and does not have latitude and longitude information, the preprocessing operation in this embodiment also includes parsing the address name and supplementing geographical location information such as latitude and longitude and standardized city, so as to facilitate the subsequent use of latitude and longitude information to determine the similarity between different charging station data groups, and further improve the accuracy of data fusion.

[0072] In step S120, at least one station dataset and at least one pending dataset are determined based on the attribute information of each charging station data group.

[0073] In this embodiment, the similarity between different charging station data groups in the charging station dataset all reach the second similarity threshold. The similarity between different charging station data groups in the undetermined dataset lies between the first and second similarity thresholds, with the first similarity threshold being less than the second. Therefore, by setting different values ​​for the first and second similarity thresholds and determining the charging station dataset and the undetermined dataset corresponding to the initial dataset, this embodiment can cluster charging station data groups with high similarity representing the same charging station into the charging station dataset, and cluster charging station data groups with high similarity that may represent the same charging station into the undetermined dataset. This provides a reliable data foundation for better identification and fusion of data from the same charging station in the future, improving the accuracy of data fusion.

[0074] Optionally, in this embodiment, when determining at least one charging station dataset and at least one undetermined dataset corresponding to the initial dataset, the overall similarity result between different charging station datasets is determined based on the attribute information of each charging station dataset group. Then, the charging station dataset and the undetermined dataset are determined by the relationship between the overall similarity result between different charging station dataset groups and the first similarity threshold and the second similarity threshold, respectively.

[0075] Furthermore, in this embodiment, by means of... Figure 2 The method shown determines the site dataset and the dataset to be determined, and specifically includes the following steps.

[0076] In step S210, the data groups of each charging station in the initial dataset are combined in pairs to determine at least one filtered data pair.

[0077] In this embodiment, by traversing each charging station data group in the initial dataset, the pairwise combination of charging station data groups in the initial dataset is determined, thereby determining at least one filter data pair corresponding to the initial dataset. Moreover, the pairwise combination of charging station data usually corresponds to different data sources (i.e., data acquisition methods), that is, the data sources of each charging station data in the filter data pair are different.

[0078] Optionally, considering that duplicate or erroneous recording operations may occur during the generation of charging station data groups, resulting in multiple charging station data groups corresponding to the same charging station from the same data source, the data sources for each charging station in the filtered data pairs in this embodiment can be the same or different.

[0079] In step S220, the overall similarity result of each selected data pair is determined based on the attribute information of each charging station data group. The overall similarity result is used to characterize the similarity between the charging station data groups in the corresponding selected data pair.

[0080] In one optional implementation, this embodiment uses all attribute information from the charging station data groups simultaneously to determine the overall similarity. Optionally, when determining the overall similarity result of the selected data pairs, this embodiment can first encode the attribute information of each charging station data group in different charging station data groups in a certain order using a preset encoding method to determine the attribute vector corresponding to each charging station data group; then, perform similarity calculation on the attribute vectors of each charging station data group to determine the similarity value between different charging station data groups in the corresponding selected data pair; finally, determine the overall similarity result based on the similarity value between different charging station data groups.

[0081] Optionally, to intuitively describe similarity, this embodiment pre-sets multiple similarity levels, each with a corresponding similarity value range. When the similarity value between charging station data groups falls within a certain similarity value range, the overall similarity result between the corresponding charging station data groups is determined as the similarity level corresponding to that similarity value range.

[0082] Optionally, when encoding the attribute information, different types of attribute information can use different encoding methods. For example, the station name and operator name can use one-hot encoding, the address information can use hash encoding, and the number of fast charging guns and slow charging guns can use binary encoding. Furthermore, after the attribute information is encoded, the attribute information can be converted into a unified format based on the encoding conversion tool to generate the attribute vector of the corresponding charging station data group.

[0083] In another alternative implementation, this embodiment can first determine the attribute similarity results of each attribute information in the filtered data pair, and then determine the overall similarity result of the filtered data pair based on the attribute similarity results.

[0084] Figure 3 This is a flowchart illustrating the determination of overall similarity results according to an embodiment of the present invention. For example... Figure 3 As shown, in this embodiment, the similarity between different charging station data groups in the filtered data pair is determined through the following steps.

[0085] In step S310, the attribute similarity results between attribute information of the same type in the filtered data pairs are determined.

[0086] In this embodiment, when the charging station data group includes multiple attribute information, the attribute similarity result can be determined according to the type of attribute information. The attribute similarity result between different types of attribute information can be determined using different similarity calculation methods. Furthermore, the attribute similarity result in this embodiment can be a quantitatively described similarity probability or a qualitatively described similarity result (e.g., similar or dissimilar, consistent or inconsistent).

[0087] Optionally, the attribute similarity result in this embodiment is a similarity result. The following describes the method for determining the attribute similarity result for different types of attribute information.

[0088] For charging station names or other descriptive attribute information in the charging station data group, the same method can be used to determine the attribute similarity results. Taking the attribute similarity results of charging station names as an example, in this embodiment, when determining the attribute similarity results of charging station names in different charging station data groups, the charging station names in each charging station data group in the filtered data pair will be encoded separately to determine the word vectors of the corresponding charging station names; then, the similarity of each word vector will be calculated to determine the word vector similarity between the corresponding word vectors, so as to determine the attribute similarity results between the corresponding charging station names based on the word vector similarity.

[0089] Optionally, in this embodiment, the names of charging stations in each data group of the filtered data pair can be segmented into words, and word vector models (such as Word2Vec, GloVe, FastText, BERT, etc.) can be used to determine the word vectors of the corresponding charging station names; then, the word vector similarity between each word vector can be determined according to the similarity calculation model (such as Vector Control Model (VSM), Word Embedding Model, Encoder-Decoder Model, etc.), so as to determine the attribute similarity results between the corresponding charging station names based on the word vector similarity.

[0090] Furthermore, in this embodiment, when determining the attribute similarity results between corresponding charging station names based on word vector similarity, the word vector similarity *s* between charging station names in different charging station data groups is compared with a preset threshold *S*. When the word vector similarity *s* reaches the preset threshold *S*, the attribute similarity results between the corresponding charging station names are determined to be that the station names are similar; conversely, when the word vector similarity *s* is less than the preset threshold *S*, the attribute similarity results between the corresponding charging station names are determined to be that the station names are not similar. The preset threshold *S* can be determined based on the coverage of the charging station names. For example, the preset threshold *S* can cover cases where the charging station names contain prefixes such as "region" (e.g., "xx district xx building") or "operator" (e.g., "xx operator").

[0091] For the operator names in the charging station data group, this embodiment can use the same method as the method used to determine the attribute similarity of different operator names when determining the attribute similarity of charging station names. Furthermore, considering the limited number of operators and the fact that the name of the same operator in the dataset generally only has a limited number of variations, this embodiment can also pre-establish an operator alias mapping table. This table includes at least one operator and at least one corresponding operator name. Then, it is only necessary to match the operator names in each charging station data group with the operator alias mapping table to determine whether different operator names map to the same operator.

[0092] Furthermore, when determining that different operator names are mapped to the same operator, the attribute similarity result between the corresponding operator names is determined to be 1 or the operator names are consistent; or when determining that different operator names are mapped to different operators, the attribute similarity result between the corresponding operator names is determined to be 0 or the operator names are inconsistent.

[0093] For location information in charging station data groups, when determining the attribute similarity results of location information in different charging station data groups, this embodiment will determine the distance between them based on the latitude and longitude information in each charging station data group in the filtered data pair. The distance is used to characterize the distance between charging stations represented by different charging station data groups; then, the attribute similarity results between the corresponding latitude and longitude information will be determined based on the relationship between the distance and the preset distance threshold.

[0094] Furthermore, in this embodiment, when determining the attribute similarity result between corresponding latitude and longitude information based on the relationship between the distance d and preset distance thresholds D and E, the distance d between charging stations represented by different charging station data is compared with the preset distance thresholds D and E. When the distance d is less than or equal to the preset distance threshold D, the attribute similarity result between the corresponding location information is determined to be that the station locations are consistent; while when the distance d is greater than the preset distance threshold D but less than the preset distance threshold E, the attribute similarity result between the corresponding station names is determined to be that the station locations are inconsistent. The distance threshold D is determined based on the level of detail in the address description of the location information of the same charging station. For example, for multiple charging station data groups referring to the same charging station, the different levels of detail in the address description may result in a coverage area corresponding to the location of the same charging station, such as x meters; therefore, the distance threshold D is set to x meters. The distance threshold E is determined based on the maximum distance between adjacent charging stations when the same operator deploys charging stations. For example, it is rare for two different charging stations to appear within a y-meter range under normal circumstances for the same operator; therefore, the distance threshold is set to y meters.

[0095] The same method can be used to determine the similarity of attributes for the number of fast charging guns and slow charging guns in the charging station data group. Taking the similarity of attributes for the number of fast charging guns as an example, in this embodiment, when determining the similarity of attributes for the number of fast charging guns in different charging station data groups, the difference in the number of fast charging guns in each charging station data group in the filtered data pair will be determined; then, the similarity of attributes between the corresponding number of fast charging guns will be determined based on the relationship between the difference in the number of fast charging guns and the preset number threshold.

[0096] Furthermore, in this embodiment, when determining the attribute similarity result between corresponding fast charging gun quantities based on the relationship between the quantity difference fd and the preset quantity threshold FD, the quantity difference fd can be compared with the preset quantity threshold FD. If the quantity difference fd is less than or equal to the preset quantity threshold FD, the attribute similarity result between corresponding fast charging gun quantities is determined to be that the number of fast charging guns is consistent; while if the quantity difference fd is greater than the preset quantity threshold FD, the attribute similarity result between corresponding fast charging gun quantities is determined to be that the number of fast charging guns is inconsistent. The preset quantity threshold FD is typically set to 0. Meanwhile, since there may be situations where charging bays are under maintenance and not counted when acquiring charging station data sets, resulting in a discrepancy between the number of charging guns in the charging station data set and the actual number of charging guns, the preset quantity threshold FD in this embodiment can also be set to a value greater than 0.

[0097] Meanwhile, for the difference sd between the number of slow charging guns, the similarity of attributes between the number of slow charging guns can be determined based on the relationship between the difference sd and the corresponding preset quantity threshold SD. The specific determination process is the same as the method for determining the similarity of attributes between the number of fast charging guns, and will not be repeated here.

[0098] Additionally, it should be noted that the various thresholds involved in this embodiment can be set according to the actual usage scenario. The examples provided here are merely examples and do not limit the specific values ​​of the thresholds.

[0099] In step S320, the overall similarity result of the corresponding filtered data pairs is determined based on the attribute similarity results of each type of attribute information.

[0100] In this embodiment, the overall similarity result is used to characterize the degree of similarity between different charging station data groups, and the overall similarity result can be determined by the similarity level corresponding to the degree of similarity.

[0101] Furthermore, when determining the overall similarity result of the corresponding filtered data pair based on the attribute similarity results of different types of attribute information, optionally, in this embodiment, a corresponding weight is set for each attribute information according to the importance of different attribute information in the charging station data group in advance. When determining the overall similarity result of each charging station data group in the filtered data pair, the weights corresponding to the attribute information that represent similarity (including positive similarity results such as similarity and consistency) are accumulated and summed to determine the summation result; then the summation result is compared with the preset weight range of each similarity level to determine the overall similarity result of the corresponding filtered data pair.

[0102] Alternatively, in this embodiment, a similarity result matching rule is preset. By comparing the similarity results of each attribute with the preset matching rule table, the overall similarity result of the corresponding filtered data pair is determined.

[0103] Figure 4 This is a schematic diagram of the matching rule table according to an embodiment of the present invention. Figure 4 As shown in the figure, the overall similarity result in this embodiment includes 6 similarity levels. The higher the similarity level value, the higher the degree of similarity. Specifically, when determining the overall similarity result of the filtered data pair, the similarity result of each attribute information in the filtered data pair is matched with the similarity result of the corresponding attribute information in the matching rule table to determine the overall similarity of the filtered data pair. For example, if the charging station data groups in the filtered pair have similar station names (i.e., s≥S), consistent station locations (i.e., d≤D), consistent number of fast charging guns (i.e., fd≤FD), and consistent number of slow charging guns (i.e., sd≤SD), regardless of whether the operators are the same, the overall similarity result is determined to be 6. If the charging station data groups in the filtered pair have dissimilar station names (i.e., s<S), consistent station locations (i.e., d≤D), consistent number of fast charging guns (i.e., fd≤FD), and consistent number of slow charging guns (i.e., sd≤SD), the overall similarity result is determined to be 5.

[0104] It should be understood that the matching rule table given in this embodiment is only an example, and the specific settings can be made according to the actual use scenario. There are no restrictions on how the matching rule table is set.

[0105] In step S230, the overall similarity results of each selected data pair are compared with the first similarity threshold to determine at least one similar data pair in the initial dataset. The similar data pair is the selected data pair whose overall similarity results reach the first similarity threshold.

[0106] In this embodiment, after determining the overall similarity result of each selected data pair, each overall similarity result is compared with a first similarity threshold. When the overall similarity is less than the first similarity threshold, it indicates that the different charging station data in the corresponding selected data pair are dissimilar, meaning that the probability of different charging station data groups in the selected data pair representing the same charging station is very low. Conversely, when the overall similarity is greater than or equal to the first similarity threshold, it indicates that the different charging station data groups in the corresponding selected data pair are similar, and the probability of them representing the same charging station is relatively high. In this case, the corresponding selected data pair is determined as a similar data pair. Therefore, in this embodiment, by filtering all selected data pairs using the first similarity threshold and determining the similar data pairs among all selected data pairs, charging station data groups with similarity in the initial dataset can be selected.

[0107] In step S240, the overall similarity results of each similar data pair are compared with the second similarity threshold to determine at least one site dataset and at least one undetermined dataset corresponding to the initial dataset.

[0108] In this embodiment, for each similar data pair in the initial dataset, different charging station data groups within the similar data pair may represent the same charging station or different charging stations. To improve the accuracy of data fusion among different charging station data groups of the same charging station, this embodiment sets a second similarity threshold in addition to the first similarity threshold. This second similarity threshold is used to cluster all similar data pairs corresponding to the initial dataset, thereby determining at least one charging station data group representing the same charging station, and at least one pending dataset that may represent the same charging station. That is, one charging station data set corresponds to one actual charging station, and one pending dataset corresponds to one or more actual charging stations.

[0109] Optionally, the first and second similarity thresholds in this embodiment can be determined based on the overall similarity between charging station data groups. Similar to the overall similarity, both the first and second similarity thresholds are represented using similarity levels to characterize the degree of similarity between corresponding charging station data groups. For example, when using the overall similarity value from the aforementioned preset rule table, the first similarity threshold in this embodiment can be set to 1, and the second similarity threshold can be set to 6. That is, when the overall similarity between different charging station data groups reaches the first similarity threshold of 1, it indicates that the corresponding charging station data groups are similar; when the overall similarity between different charging station data groups reaches the second similarity threshold of 6, it indicates that the corresponding charging station data groups are highly similar, representing the same charging station; and when the overall similarity between different charging station data groups is 2-5, i.e., between the first similarity threshold of 1 and the second similarity threshold of 6, it indicates that the corresponding charging station data groups are relatively similar, possibly representing the same charging station.

[0110] Optionally, to more clearly determine whether a group of charging station data represents the same charging station, this embodiment uses methods such as... Figure 5 The method shown determines at least one site dataset and at least one pending dataset corresponding to the initial dataset.

[0111] In step S510, the overall similarity result of each similar data pair is compared with the second similarity threshold to filter and determine the matching data pairs and the data pairs to be matched.

[0112] In this embodiment, the matching data pair is a similar data pair whose overall similarity result reaches the second similarity threshold, and the data pair to be matched is a similar data pair whose overall similarity result reaches the first similarity threshold and is less than the second similarity threshold.

[0113] When comparing the overall similarity result of each similar data pair with a second similarity threshold, if the overall similarity result of the similar data pair reaches the second similarity threshold, then the similar data pair is determined to be a matched data pair; if the overall similarity result of the similar data pair does not reach (i.e., is less than) the second similarity threshold, then the similar data pair is determined to be a data pair to be matched. Therefore, in this embodiment, the above method is used to filter each similar data pair corresponding to the initial dataset, and to determine matched data pairs and data pairs to be matched from all similar data pairs.

[0114] In step S520, the charging station data groups in each matching data pair are clustered based on the connected graph method to determine at least one station dataset.

[0115] In this embodiment, when determining at least one charging station dataset based on the connected graph method, each charging station data group in each matched data pair is used as a node, and the overall similarity of each matched data pair represents the edges between different charging station data groups in the matched data pair. By traversing each charging station data group in each matched data pair and the overall similarity between each charging station data in each matched data pair, at least one complete connected graph is constructed, ensuring that the overall similarity between any two charging station data groups in each connected graph is greater than a second similarity threshold. After the connected graph is constructed, each connected graph corresponds to one charging station dataset, and the charging station data group corresponding to each node in a connected graph is a charging station data group in the corresponding charging station dataset.

[0116] In step S530, the charging station data groups in each pair of data to be matched are clustered based on the connected graph method to determine at least one undetermined dataset.

[0117] In this embodiment, the method for determining at least one pending dataset based on the connected graph approach is similar to the method for determining at least one charging station dataset. The difference lies in that the nodes in the connected graph corresponding to the pending dataset are all charging station data groups within the data pairs to be matched. Therefore, by constructing a connected graph to determine at least one charging station dataset and at least one pending dataset corresponding to the initial dataset, this embodiment can fully and comprehensively utilize the similarity relationships between charging station data groups from different data sources, providing an accurate data foundation for subsequent data fusion and thus improving the accuracy of data fusion.

[0118] To facilitate understanding, this embodiment uses a specific example to illustrate the process of determining the field dataset and the dataset to be determined corresponding to the initial dataset.

[0119] Figure 6 This is a schematic diagram illustrating the determination of the site dataset and the pending dataset according to an embodiment of the present invention. Figure 6 As shown in the diagram, the initial dataset is assumed to include charging station data sets from three data sources: data source A, data source B, and data source C. Each data source is assumed to contain two charging station data sets: A1 and A2 from data source A, B1 and B2 from data source B, and C1 and C2 from data source C.

[0120] After obtaining the initial dataset, the data from each charging station in the initial dataset are combined in pairs to determine the corresponding 15 filter data pairs: {A1, A2}, {A1, B1}, {A1, B2}, {A1, C1}, {A1, C2}, {A2, B1}, {A2, B2}, {A2, C1}, {A2, C2}, {B1, B2}, {B1, C1}, {B1, C2}, {B2, C1}, {B2, C2}, and {C1, C2}.

[0121] Then, based on the attribute information in the charging station data group of each selected data pair, the overall similarity result of the corresponding selected data pair is determined, and the overall similarity result of each selected data pair is compared with a preset first similarity threshold 'a'. Similar data pairs whose overall similarity results reach the first similarity threshold 'a' are selected: {A1, A2}, {A1, B1}, {A2, B1}, and {B2, C1}. Simultaneously, all other selected data pairs whose overall similarity results are less than the first similarity threshold 'a', except for {A1, A2}, {A1, B1}, {A2, B1}, and {B2, C1}, are removed.

[0122] Furthermore, the overall similarity of each similar data pair is compared with the second similarity threshold b, and matching data pairs with an overall similarity result reaching the second similarity threshold b are selected: {A1, A2}, {A1, B1}, {A2, B1}, and unmatched data pairs with an overall similarity result between the first similarity threshold a and the second similarity threshold b: {B2, C1}, where the second similarity threshold b is greater than the first similarity threshold a.

[0123] Subsequently, connectivity was established for each charging station data group in {A1, A2}, {A1, B1}, and {A2, B1} based on the matching data. Figure 1 , connect Figure 1 The nodes in the data include charging station data groups A1, A2 and B1, and the overall similarity between any two pairs of charging station data groups A1, A2 and B1 reaches the second similarity threshold b. At this time, the charging station dataset {A1, A2, B1} can be determined.

[0124] Simultaneously, connectivity is constructed for the charging station data group in {B2, C1} based on the data to be matched. Figure 2 , connect Figure 2 The nodes in the dataset include charging station data groups B2 and C1, and the overall similarity between charging station data groups B2 and C1 is between the first similarity threshold a and the second similarity threshold b. At this point, the dataset to be determined {B2, C1} can be determined.

[0125] Furthermore, after determining at least one charging station dataset and at least one pending dataset corresponding to the initial dataset based on the above method, the charging station data groups in each charging station dataset and each pending dataset will be fused to determine the actual charging station data represented by each charging station dataset and each pending dataset.

[0126] In step S130, the attribute information in each charging station data group in each station dataset is fused according to the preset attribute priority order to determine the first fused data. The first fused data includes the actual charging station data of the charging station represented by the corresponding station dataset.

[0127] In this embodiment, for each station dataset, since each charging station data group in the station dataset represents the same charging station, the attribute information in the charging station data group in each station dataset can be automatically fused after determining each station dataset, thereby determining the actual charging station data of the corresponding charging station represented by each station dataset.

[0128] Optionally, since the types of attribute information contained in the charging station data group are known, in order to improve the accuracy and efficiency of data fusion, when fusing different attribute information in different charging station data groups in the station dataset, this embodiment will fuse the attribute information in each charging station data group in the station dataset according to a preset attribute priority order.

[0129] Furthermore, in this embodiment, the number of fast charging guns can be considered the highest priority attribute, with other attributes listed in descending order. The priority order of these attributes can be set according to the actual application scenario. For example, the priority order of the attributes from highest to lowest could be: number of fast charging guns, number of slow charging guns, site name, latitude and longitude information, city name, address name, and operator name, among other attributes.

[0130] Optionally, during data fusion, this embodiment can sequentially fill in the true values ​​of each attribute information in the actual charging station data corresponding to the charging station dataset according to the attribute priority order. Specifically, when filling in each attribute information, this embodiment first determines the number of fast charging guns in the charging station data group with the largest number of fast charging guns in the charging station dataset as the actual number of fast charging guns in the corresponding charging station; then, it determines the number of slow charging guns in the charging station data group with the largest number of slow charging guns in the charging station dataset as the actual number of slow charging guns in the corresponding charging station; then, it uses other attribute information such as the most comprehensive station name, the most frequently occurring latitude and longitude information, the most standard city name, the most comprehensive address name, and the operator name in the charging station dataset as the true values ​​of the corresponding type of attribute information in the actual charging station data.

[0131] It should be understood that this embodiment is an example of a method for fusing attribute information in each charging station data group in the charging station dataset according to a preset attribute priority order. Different attribute information in different charging station data groups in the charging station dataset can also be fused according to other orders or methods. There is no restriction on the specific method of data fusion.

[0132] Therefore, in this embodiment, the above method can automatically fuse attribute information from multiple charging station data groups in a highly similar charging station dataset to generate actual charging station data that corresponds to the actual charging station, thereby improving the fusion efficiency of real charging station data.

[0133] In step S140, the charging station data group in the pending dataset is sent to the manual processing terminal to determine the second fused data based on the manual processing results. The second fused data includes the actual charging station data of at least one charging station.

[0134] In this embodiment, for each pending dataset, since each charging station data group in the pending dataset may represent the same charging station, in order to ensure the accuracy of data fusion, after determining each pending dataset, the charging station data group in each pending dataset and the attribute information in each charging station data group can be sent to the manual processing terminal, so as to manually determine the charging station represented by each pending dataset, and fuse the attribute information in different charging station data groups corresponding to the same charging station.

[0135] Therefore, in this embodiment, the above method can be used to manually fuse attribute information from multiple charging station data groups in a highly similar dataset, achieving more refined judgment and fusion processing. This allows attribute information from charging station data groups with different attribute descriptions but actually belonging to the same charging station to be fused into one, thereby better realizing the identification and fusion of actual charging station data from the same charging station, improving the completeness of the fusion results and the overall accuracy of data fusion.

[0136] In step S150, charging station distribution parameters are determined based on the first fusion data and the second fusion data, so as to optimize the supply distribution of charging stations according to the charging station distribution parameters.

[0137] In this embodiment, after determining the first fused data and the second fused data, the actual charging station data corresponding to the first fused data and the second fused data will be integrated to determine the existing charging stations and attribute information of each charging station in the geographical area (i.e., the target area) corresponding to the initial dataset.

[0138] Optionally, to further facilitate the subsequent optimization of the charging station supply distribution based on the charging station distribution parameters, in this embodiment, after integrating the actual charging station data corresponding to the first fused data and the second fused data, and determining the attribute information of the existing charging stations in the geographical area (i.e., the target area) corresponding to the initial dataset, the existing charging stations will be visualized on the map in the form of points using their latitude and longitude location information. This allows for a quick understanding of the charging station supply point distribution within the geographical area based on the points displayed on the map. It also facilitates further processing based on the map to obtain the supply competition pattern of local areas within the geographical area, thereby aiding in site selection planning.

[0139] Subsequently, in this embodiment, the distribution parameters of charging stations within the corresponding geographical area are determined based on the attribute information of these existing charging stations, so as to optimize the supply distribution of charging stations according to the charging station distribution parameters. Optionally, the charging station distribution parameters in this embodiment include total charging supply, regional charging supply, and / or other parameters that can affect the layout of charging stations. The total charging supply is used to characterize the charging supply capacity of the corresponding geographical area, which can be represented by the sum of the number of fast charging guns within the geographical area. The regional charging supply is used to characterize the charging supply capacity of a sub-region within the corresponding geographical area, which can be represented by the sum of the number of fast charging guns within the corresponding sub-region.

[0140] Furthermore, when optimizing the supply distribution of charging stations based on the distribution parameters of charging stations, this embodiment can determine whether to add new charging stations in a geographical area based on the total charging supply within that area. For example, if the total charging supply is greater than the demand or a preset value, it indicates that the number of charging stations in the current geographical area is already saturated, and there is no need to continue deploying charging stations in that area; conversely, if the total charging supply is less than the demand or a preset value, it indicates that the number of charging stations in the current geographical area is insufficient, and charging stations can continue to be deployed in that area.

[0141] Furthermore, in this embodiment, the layout of charging stations within a geographical area can be adjusted based on the regional charging supply corresponding to sub-regions within the geographical area. For example, if the total charging supply within the geographical area can meet the overall charging demand, but the regional charging supply in a sub-region cannot meet the regional charging demand, the number of charging stations in that sub-region can be increased. Alternatively, if the regional charging supply in a sub-region far exceeds the charging demand, the charging equipment in the charging stations within that sub-region can be moved to other sub-regions with insufficient charging supply, thereby optimizing the distribution of charging station supply.

[0142] The technical solution of this embodiment obtains an initial dataset comprising multiple charging station data groups from different data sources. Based on the attribute information of each charging station data group and the relationship between the similarity between different charging station data groups and a first similarity threshold and a second similarity threshold, it determines at least one charging station dataset and at least one undetermined dataset corresponding to the initial dataset. It then fuses the attribute information of each charging station data group in the charging station dataset where the similarity reaches the second similarity threshold, and manually fuses the attribute information of each charging station data group in the undetermined dataset where the similarity is between the first and second similarity thresholds. This allows for the classification of charging station data groups from different data sources according to different similarity thresholds. Highly similar charging station data is automatically and quickly fused, while moderately similar charging station data is accurately fused through manual processing. This achieves both accuracy and efficiency in data fusion when fusing data from different data sources. Furthermore, by using the fused first and second fused data to determine charging station distribution parameters, the supply distribution of charging stations can be optimized based on these parameters, facilitating the optimization of charging station layout.

[0143] Figure 7 This is a charging station data processing device according to an embodiment of the present invention, such as... Figure 7 As shown, the charging station data processing device in this embodiment includes: an acquisition unit 1, a processing unit 2, an analysis unit 3, and an optimization unit 4. The acquisition unit 1 acquires an initial dataset, which includes multiple charging station data groups from different data sources. The processing unit 2 determines at least one charging station dataset and at least one pending dataset corresponding to the initial dataset based on the attribute information of each charging station data group. The similarity between different charging station data groups in the charging station dataset reaches a second similarity threshold, while the similarity between different charging station data groups in the pending dataset is between a first similarity threshold and a second similarity threshold, with the first similarity threshold being less than the second similarity threshold. The analysis unit 3 fuses the attribute information of each charging station data group in each charging station dataset according to a preset attribute priority order to determine first fused data, which includes the actual charging station data of the charging station represented by the corresponding charging station dataset. The analysis unit 3 sends the charging station data groups in the pending dataset to a manual processing end to determine second fused data based on the manual processing results. The second fused data includes the actual charging station data of at least one charging station. The optimization unit 4 is used to determine the charging station distribution parameters based on the first fusion data and the second fusion data, so as to optimize the supply distribution of charging stations according to the charging station distribution parameters.

[0144] Optionally, the processing unit 2 in this embodiment includes: a combination subunit, a determination subunit, and a comparison subunit. The combination subunit is used to combine each charging station data group in the initial dataset in pairs to determine at least one filtered data pair. The determination subunit is used to determine the overall similarity result of each filtered data pair based on the attribute information of each charging station data group. The overall similarity result is used to characterize the similarity between the charging station data groups in the corresponding filtered data pair. The comparison subunit is used to compare the overall similarity result of each filtered data pair with the first similarity threshold to determine at least one similar data pair in the initial dataset. The similar data pair is the filtered data pair whose overall similarity result reaches the first similarity threshold; and to compare the overall similarity result of each similar data pair with a second similarity threshold to determine at least one charging station dataset and at least one pending dataset corresponding to the initial dataset.

[0145] Furthermore, the determining subunit in this embodiment is also used to determine the attribute similarity results between attribute information of the same type in the filtered data pair; and to determine the overall similarity result of the corresponding filtered data pair based on the attribute similarity results of attribute information of different types.

[0146] Specifically, the charging station data group includes the station name, operator name, city name, address name, latitude and longitude information, number of fast charging guns, and / or number of slow charging guns. When determining the attribute similarity results between similar attribute information in the filtered data pair, the determining subunit specifically performs word segmentation on the station names in each charging station data group of the filtered data pair, determines the word vectors of the corresponding station names, and determines the cosine similarity between each word vector using the cosine similarity method, so as to determine the attribute similarity results between the corresponding station names based on the cosine similarity; determines the distance based on the latitude and longitude information in each charging station data group of the filtered data pair, the distance being used to characterize the distance between charging stations represented by different charging station data groups, and determines the attribute similarity results between the corresponding latitude and longitude information based on the relationship between the distance and a preset distance threshold; and determines the quantity difference of the number of fast charging guns in each charging station data group of the filtered data pair, and determines the attribute similarity results between the corresponding number of fast charging guns based on the relationship between the quantity difference and a preset quantity threshold.

[0147] When determining the overall similarity result of the corresponding filtered data pair based on the attribute similarity results of different types of attribute information, the specific sub-unit is used to compare each attribute similarity result with the preset matching rule table to determine the overall similarity result of the corresponding filtered data pair; or to sum up the weights corresponding to the attribute information that represent similarity in the attribute similarity results, determine the summation result, and compare the summation result with the preset weight range of each similarity level to determine the overall similarity result of the corresponding filtered data pair.

[0148] Furthermore, the comparison subunit in this embodiment is also used to compare the overall similarity result of each similar data pair with the second similarity threshold, and filter and determine the matching data pairs and the data pairs to be matched. The matching data pairs are similar data pairs whose overall similarity result reaches the second similarity threshold, and the data pairs to be matched are similar data pairs whose overall similarity result reaches the first similarity threshold and is less than the second similarity threshold. The charging station data groups in each matching data pair are clustered based on the connected graph method to determine at least one charging station dataset. The charging station data groups in each data pair to be matched are clustered based on the connected graph method to determine at least one undetermined dataset.

[0149] Figure 8 This is a schematic diagram of an electronic device according to an embodiment of the present invention. (For example...) Figure 8 As shown, Figure 8 The illustrated electronic device is a general-purpose data processing device, comprising a general-purpose computer hardware architecture, including at least a processor 81 and a memory 82. The processor 81 and memory 82 are connected via a bus 83. The memory 82 is adapted to store instructions or programs executable by the processor 81. The processor 81 can be a standalone microprocessor or a collection of one or more microprocessors. Thus, the processor 81 executes the instructions stored in the memory 82, thereby performing the method flow of the embodiments of the present invention as described above to process data and control other devices. The bus 83 connects the aforementioned components together, and also connects these components to a display controller 84, a display device, and an input / output (I / O) device 85. The input / output (I / O) device 85 can be a mouse, keyboard, modem, network interface, touch input device, motion-sensing input device, printer, and other devices known in the art. Typically, the input / output device 85 is connected to the system via an input / output (I / O) controller 86.

[0150] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus (devices), or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0151] This application is described with reference to flowchart illustrations of methods, apparatus (devices), and computer program products according to embodiments of this application. It should be understood that each step in the flowchart can be implemented by computer program instructions.

[0152] These computer program instructions may be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction means, the implementation process of which is described in the instruction means. Figure 1 The function specified in one or more processes.

[0153] These computer program instructions may also be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, produce instructions for implementing processes. Figure 1 A device for a function specified in one or more processes.

[0154] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program for use by a computer to execute some or all of the above-described method embodiments.

[0155] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program specifying the relevant hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0156] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A data processing method for charging stations, characterized in that, The method includes: Obtain an initial dataset, which includes multiple charging station data sets from different data sources; Based on the attribute information of each charging station data group, at least one charging station data group and at least one undetermined data group corresponding to the initial dataset are determined. The similarity between different charging station data groups in the charging station data group reaches a second similarity threshold. The similarity between different charging station data groups in the undetermined data group is between a first similarity threshold and a second similarity threshold. The first similarity threshold is less than the second similarity threshold. According to a preset attribute priority order, the attribute information in each charging station data group in each of the station datasets is fused to determine the first fused data, which includes the actual charging station data of the charging station represented by the corresponding station dataset. The charging station data group in the pending dataset is sent to the manual processing terminal to determine the second fused data based on the manual processing results. The second fused data includes the actual charging station data of at least one charging station. The charging station distribution parameters are determined based on the first fused data and the second fused data, so as to optimize the supply distribution of charging stations according to the charging station distribution parameters.

2. The method according to claim 1, characterized in that, The step of determining at least one charging station dataset and at least one pending dataset corresponding to the initial dataset based on the attribute information of each charging station data group includes: The data groups of each charging station in the initial dataset are combined in pairs to determine at least one filtered data pair; The overall similarity result of each of the filtered data pairs is determined based on the attribute information of each of the charging station data groups. The overall similarity result is used to characterize the similarity between the charging station data groups in the corresponding filtered data pairs. The overall similarity results of each of the selected data pairs are compared with the first similarity threshold to determine at least one similar data pair in the initial dataset. The similar data pair is the selected data pair whose overall similarity results reach the first similarity threshold. The overall similarity results of each of the similar data pairs are compared with the second similarity threshold to determine at least one site dataset and at least one undetermined dataset corresponding to the initial dataset.

3. The method according to claim 2, characterized in that, The step of comparing the overall similarity result of each of the similar data pairs with the second similarity threshold to determine at least one site dataset and at least one pending dataset corresponding to the initial dataset includes: The overall similarity result of each of the similar data pairs is compared with the second similarity threshold to filter and determine the matching data pairs and the data pairs to be matched. The matching data pairs are the similar data pairs whose overall similarity result reaches the second similarity threshold, and the data pairs to be matched are the similar data pairs whose overall similarity result reaches the first similarity threshold and is less than the second similarity threshold. Based on the connected graph method, the charging station data groups in each of the matched data pairs are clustered to determine at least one station dataset. Based on the connected graph method, the charging station data groups in each of the data pairs to be matched are clustered to determine at least one undetermined dataset.

4. The method according to claim 2, characterized in that, The charging station data group includes at least one attribute information, and determining the overall similarity result of each filtered data pair based on the attribute information of each charging station data group includes: Determine the attribute similarity results between attribute information of the same type in the filtered data pairs; The overall similarity result of the corresponding filtered data pairs is determined based on the attribute similarity results of different types of attribute information.

5. The method according to claim 4, characterized in that, The charging station data group includes the station name, operator name, city name, address name, latitude and longitude information, number of fast charging guns and / or number of slow charging guns.

6. The method according to claim 5, characterized in that, The determination of attribute similarity results between attribute information of the same type in the filtered data pairs includes: The names of charging stations in each data group of the filtered data pair are encoded to determine the word vectors of the corresponding station names; The word vectors are similarity calculated to determine the word vector similarity between corresponding word vectors, so as to determine the attribute similarity between corresponding station names based on the word vector similarity.

7. The method according to claim 5, characterized in that, The determination of attribute similarity results between attribute information of the same type in the filtered data pairs includes: The distance between charging stations is determined based on the latitude and longitude information in each charging station data group in the filtered data pair. The distance between charging stations is used to represent the distance between charging stations represented by different charging station data groups. The similarity of attributes between corresponding latitude and longitude information is determined based on the relationship between the distance and the preset distance threshold.

8. The method according to claim 5, characterized in that, The determination of attribute similarity results between attribute information of the same type in the filtered data pairs includes: Determine the difference in the number of fast charging guns in each charging station data group within the filtered data pair; The similarity of attributes between the corresponding fast charging gun quantities is determined based on the relationship between the quantity difference and the preset quantity threshold.

9. The method according to claim 4, characterized in that, The determination of the overall similarity result of the corresponding filtered data pair based on the attribute similarity results of different types of attribute information includes: The similarity results of each attribute are compared with a preset matching rule table to determine the overall similarity result of the corresponding filtered data pairs.

10. The method according to claim 4, characterized in that, The determination of the overall similarity result of the corresponding filtered data pair based on the attribute similarity results of different types of attribute information includes: The weights corresponding to the attribute information representing similarity results are accumulated and summed to determine the summation result; The summation result is compared with the preset weight intervals of each similarity level to determine the overall similarity result of the corresponding filtered data pair.

11. A data processing device for charging stations, characterized in that, The device includes: An acquisition unit is used to acquire an initial dataset, which includes multiple charging station data groups from different data sources; The processing unit is configured to determine at least one charging station dataset and at least one pending dataset corresponding to the initial dataset based on the attribute information of each charging station data group. The similarity between different charging station data groups in the charging station dataset reaches a second similarity threshold. The similarity between different charging station data groups in the pending dataset is between a first similarity threshold and the second similarity threshold. The first similarity threshold is less than the second similarity threshold. An analysis unit is configured to fuse attribute information in each charging station data group within each of the aforementioned station datasets according to a preset attribute priority order, to determine first fused data, wherein the first fused data includes actual charging station data of the charging station represented by the corresponding station dataset; and to send the charging station data group in the pending dataset to a manual processing terminal to determine second fused data based on the manual processing results, wherein the second fused data includes actual charging station data of at least one charging station. An optimization unit is used to determine charging station distribution parameters based on the first fused data and the second fused data, so as to optimize the supply distribution of charging stations based on the charging station distribution parameters.

12. A computer program product, characterized in that, The computer program product includes a computer program / instruction that, when executed by a processor, implements the method of any one of claims 1-10.

13. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method of any one of claims 1-10.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-10.