A power customer evaluation data processing method, device, equipment and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-21
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本申请提供了一种电力客户评价数据处理方法、装置、设备及介质,能够解决现有技术中分布式环境下电力客户评价数据处理效率低的问题
[0008]相比现有技术,上述实施例具有以下有益效果:通过在同一存储节点组内选择第一存储节点并将其他目标存储节点中的数据分片迁移至该第一存储节点,实现了对分散数据的集中管理,使得后续数据处理能够在单一节点上完成,从而减少跨节点数据交互,降低数据传输次数及网络开销,提升数据处理的整体效率。
Smart Images

Figure CN122547518A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and in particular to a method, apparatus, equipment and medium for processing electricity customer evaluation data. Background Technology
[0002] As the power system develops towards intelligence and service orientation, power companies are gradually acquiring customer evaluation data through various channels (such as business hall services, online platforms, and customer follow-ups) to reflect users' satisfaction with power supply reliability, service quality, and response efficiency. Therefore, processing power customer evaluation data has become an important part of the refined management of power services.
[0003] In existing technologies, the processing of electricity customer evaluation data mostly adopts centralized processing methods or simple distributed processing architectures. With the continuous growth of data volume, these methods have gradually revealed insufficient processing efficiency in practical applications. On the one hand, the frequent transmission of large amounts of data between different nodes during processing easily increases network bandwidth consumption, thus affecting the overall processing speed. On the other hand, the execution methods of data processing tasks are relatively fixed, making it difficult to flexibly adjust according to data distribution. This results in some nodes facing heavy processing pressure while others underutilize resources, further restricting data processing efficiency. Therefore, how to improve the overall processing efficiency of electricity customer evaluation data in a distributed environment has become a technical problem that needs to be solved. Summary of the Invention
[0004] This application provides a method, apparatus, equipment, and medium for processing electricity customer evaluation data, which can solve the problem of low efficiency in processing electricity customer evaluation data in a distributed environment in the prior art.
[0005] Some embodiments of this application provide a method for processing electricity customer evaluation data, including: The system acquires operational data for each distributed node and generates a node score representing the resource sufficiency of the distributed node based on the operational data; the distributed node includes storage nodes and computing nodes. When an evaluation data processing task is received, several target storage nodes are determined from the storage nodes according to the data required by the evaluation data processing task. Obtain the first network communication distance between each target storage node, and group the target storage nodes according to the first network communication distance to obtain at least one storage node group; wherein, the first network communication distance between each target storage node in the same storage node group is less than a first preset distance threshold or there is only one target storage node in the same storage node group; Based on the scores of each node, the target computing node corresponding to each storage node group is determined from all the computing nodes; The target computing node processes customer evaluation data corresponding to each target storage node in the storage node group.
[0006] Compared with existing technologies, the above embodiments have the following beneficial effects: This application quantifies the sufficiency of node resources by acquiring the operating data of each distributed node and generating node scores, thereby providing a basis for selecting computing nodes and avoiding performance bottlenecks caused by task concentration; simultaneously, after receiving processing tasks, grouping is performed based on the first network communication distance between target storage nodes, so that spatially similar data is processed together, thereby reducing cross-node data interaction and reducing transmission latency; furthermore, target computing nodes are matched to each storage node group according to the node scores, so that data processing is preferentially executed on nodes with more abundant resources, improving resource utilization and processing efficiency. Through the synergistic effect of the above technical means, this application improves the overall processing efficiency of power customer evaluation data in a distributed environment while reducing data transmission overhead.
[0007] Further, determining the target computing node corresponding to each storage node group from all the computing nodes based on the node scores includes: When the storage node group includes multiple target storage nodes, a first storage node is selected from the storage node group, and data fragments from other target storage nodes in the storage node group are migrated to the first storage node. When there is only one target storage node in the storage node group, the target storage node shall be designated as the first storage node. For each of the storage node groups, the target computing node is determined from all the computing nodes based on the first storage node and the scores of each node in the storage node group.
[0008] Compared with the prior art, the above embodiments have the following beneficial effects: by selecting a first storage node within the same storage node group and migrating data fragments from other target storage nodes to the first storage node, centralized management of dispersed data is achieved, enabling subsequent data processing to be completed on a single node, thereby reducing cross-node data interaction, reducing data transmission frequency and network overhead, and improving the overall efficiency of data processing.
[0009] Further, determining the corresponding target computing node from all the computing nodes based on the first storage node and the node scores of the storage node group includes: The network latency and second network communication distance between each computing node and the first storage node are obtained, and the data transmission overhead between each computing node and the first storage node is calculated based on the network latency and the second network communication distance. Based on the second network communication distance, a number of candidate computing nodes whose second network communication distance with the target storage node is lower than a preset transmission threshold are determined. Among the candidate computing nodes, the target computing node is determined based on the corresponding node score.
[0010] Compared with the prior art, the above embodiments have the following beneficial effects: by comprehensively considering the network latency and distance between the computing node and the first storage node, and screening candidate computing nodes accordingly, and then combining the node score to determine the target computing node, the computing task can be executed on the node with low data transmission overhead and sufficient resources, thereby reducing data transmission costs while improving computing efficiency and system response performance.
[0011] Further, when the storage node group includes multiple target storage nodes, selecting a first storage node from the storage node group includes: Several candidate storage nodes are determined based on the first network communication distance between each target storage node; the candidate storage node is a storage node whose sum of the first network communication distances with other target storage nodes in the storage node group is less than a second preset distance threshold; Among the candidate storage nodes, the target storage node with the highest node score is selected as the first storage node.
[0012] Compared with the prior art, the above embodiments have the following beneficial effects: by screening candidate storage nodes based on the first network communication distance between each target storage node, and further combining the node score to determine the first storage node, the selected node has both a better spatial location advantage and a higher resource processing capability, thereby reducing the overall overhead of data migration and improving the efficiency of centralized data processing.
[0013] Furthermore, upon receiving customer feedback data, it includes: The customer review data is divided into several data segments; each data segment corresponds to a popularity tag that represents the frequency of data reading. The data segments are classified according to the popularity tags corresponding to each data segment to obtain a first popularity data segment and a second popularity data segment; wherein, the data reading frequency of the first popularity data segment is greater than the data reading frequency of the second popularity data segment; Based on the node score corresponding to each of the storage nodes, determine the first storage node whose node score is higher than a preset score threshold; The first heat data shard is allocated to the first storage node, and the second heat data shard is allocated to the remaining storage nodes.
[0014] Compared with the prior art, the above embodiments have the following beneficial effects: by assigning popularity tags to data shards and classifying them, high-frequency access data and low-frequency access data can be distinguished, and by combining node scores, high-popularity data can be allocated to storage nodes with more abundant resources, which helps to alleviate the access pressure of hot data, avoid node overload, and thus improve data access efficiency and system stability.
[0015] Furthermore, upon receiving customer feedback data, the following is also included: The customer review data is denoised to obtain the first review data; Identify the source of the first evaluation data and determine the scoring mapping rule for the first evaluation data based on the source of the first evaluation data; The first evaluation data is normalized according to the scoring mapping rules.
[0016] Compared with existing technologies, the above embodiments have the following beneficial effects: by performing noise reduction, channel identification, and rating normalization on customer evaluation data, evaluation data from different channels have a consistent data format and evaluation standards, thereby improving data quality and comparability, providing a reliable foundation for subsequent data sharding and processing, and improving the accuracy and effectiveness of overall data processing.
[0017] The operational data includes: processing load, access latency, resource utilization, and data read / write request frequency; generating a node score characterizing the resource sufficiency of the distributed nodes based on the operational data includes: Based on the processing load, determine the load index that characterizes the resource consumption of the distributed nodes; Determine a latency metric characterizing the data access performance of the distributed nodes based on the access latency; Based on the resource utilization rate, resource indicators characterizing the available capacity of the distributed nodes are determined; The access pressure index, which characterizes the data access pressure of the distributed node, is determined based on the frequency of data read and write requests. The node score is obtained by weighted summation based on the load metric, the latency metric, the resource metric, and the access pressure metric.
[0018] Compared with existing technologies, the above embodiments have the following beneficial effects: by decomposing the running data into load indicators, latency indicators, resource indicators and access pressure indicators, and performing weighted summation to generate node scores, the resource status of distributed nodes can be comprehensively characterized from multiple dimensions, thereby improving the accuracy of node scores, providing a more reasonable basis for subsequent data distribution and selection of computing nodes, and further improving system resource utilization and data processing efficiency.
[0019] Another embodiment of this application provides an electricity customer evaluation data processing device, including: a scoring module, a first screening module, a grouping module, and a processing module; The scoring module is used to acquire the operating data corresponding to each distributed node, and generate a node score representing the resource sufficiency of the distributed node based on the operating data; the distributed node includes: storage node and computing node; The first filtering module is used to determine several target storage nodes from the storage nodes according to the data required by the evaluation data processing task when an evaluation data processing task is received. The grouping module is used to obtain a first network communication distance between each target storage node, and to group the target storage nodes according to the first network communication distance to obtain at least one storage node group; wherein, the first network communication distance between each target storage node in the same storage node group is less than a first preset distance threshold or there is only one target storage node in the same storage node group; The second filtering module is used to determine the target computing node corresponding to each storage node group from all the computing nodes based on the scores of each node. The processing module is used to process customer evaluation data corresponding to each target storage node in the storage node group through the target computing node.
[0020] Another embodiment of this application provides a terminal device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the steps of the power customer evaluation data processing method of this application.
[0021] Another embodiment of this application also provides a computer-readable storage medium item, including: a stored computer program, which, when the computer program is running, controls the device where the computer-readable storage medium is located to perform the steps of the power customer evaluation data processing method of this application. Attached Figure Description
[0022] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating a method for processing electricity customer evaluation data provided in some embodiments of this application; Figure 2 This is a schematic diagram of the structure of a power customer evaluation data processing device provided in some embodiments of this application. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0026] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0027] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0028] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0029] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).
[0030] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.
[0031] In existing technologies, the processing of electricity customer evaluation data mostly adopts centralized processing methods or simple distributed processing architectures. With the continuous growth of data volume, these methods have gradually revealed insufficient processing efficiency in practical applications. On the one hand, the large amount of data needs to be frequently transmitted between different nodes during processing, which easily increases network bandwidth consumption and thus affects the overall processing speed. On the other hand, the execution method of data processing tasks is relatively fixed, making it difficult to flexibly adjust according to data distribution. This results in some nodes facing heavy processing pressure while others underutilize resources, further restricting data processing efficiency.
[0032] Please refer to Figure 1 To address the problem of low overall processing efficiency of electricity customer evaluation data in a distributed environment in existing technologies, this application provides an electricity customer evaluation data processing method, including the following steps S101 to S105: S101: Obtain the running data corresponding to each distributed node, and generate a node score representing the resource sufficiency of the distributed node based on the running data; the distributed node includes: storage node and computing node.
[0033] Furthermore, in some embodiments of this application, upon receiving customer review data, the process includes: The customer review data is divided into several data segments; each data segment corresponds to a popularity tag that represents the frequency of data reading. The data segments are classified according to the popularity tags corresponding to each data segment to obtain a first popularity data segment and a second popularity data segment; wherein, the data reading frequency of the first popularity data segment is greater than the data reading frequency of the second popularity data segment; Based on the node score corresponding to each of the storage nodes, determine the first storage node whose node score is higher than a preset score threshold; The first heat data shard is allocated to the first storage node, and the second heat data shard is allocated to the remaining storage nodes.
[0034] Specifically, in the actual implementation, upon receiving customer review data, the system first integrates the raw data from different channels (such as branch office reviews, online service reviews, and follow-up records) and segments the data according to preset data partitioning rules. For example, the data can be segmented based on time windows, user identifiers, or data size, ensuring each data segment contains a relatively independent set of review records. After data segmentation, the system combines historical access log statistics to calculate the number of reads or access frequency of each data segment within a preset time period. Based on this, each data segment is assigned a corresponding popularity tag to characterize its access activity level. Furthermore, all data segments are categorized based on the popularity tags. For example, by setting a popularity threshold or using a quantile interval division method, data segments with higher access frequencies are classified as the first popularity data segment, and data segments with lower access frequencies are classified as the second popularity data segment. Meanwhile, the system calculates a node score for each storage node based on its operational status information (including processing load, resource usage, and access response status), and selects storage nodes with relatively abundant resources as the first storage node according to a preset score threshold. Based on this, the first batch of frequently accessed data shards is preferentially allocated to the first storage node for storage, ensuring that high-frequency access data can be quickly read from nodes with sufficient resources. The second batch of frequently accessed data shards is allocated to the remaining storage nodes, thereby ensuring overall storage balance while reducing the impact of hot access on system performance, thus providing an efficient data access foundation for subsequent data processing tasks.
[0035] This application distinguishes between high-frequency and low-frequency access data by assigning popularity tags to data shards and classifying them. Combined with node scoring, high-popularity data is allocated to storage nodes with sufficient resources, which helps to alleviate the access pressure of hot data, avoid node overload, and thus improve data access efficiency and system stability.
[0036] Furthermore, in some embodiments of this application, upon receiving customer review data, the method further includes: The customer review data is denoised to obtain the first review data; Identify the source of the first evaluation data and determine the scoring mapping rule for the first evaluation data based on the source of the first evaluation data; The first evaluation data is normalized according to the scoring mapping rules.
[0037] Specifically, upon receiving customer review data from different channels, the raw data is first preprocessed to improve data quality. Noise reduction processing may include: removing duplicate submissions, eliminating obviously abnormal or invalid data (e.g., missing ratings, incomplete fields, or ratings significantly deviating from a reasonable range), and cleaning up irrelevant symbols, garbled text, or interfering information in the text reviews, thus obtaining the first set of review data. Based on this, the channel source of the first set of review data is identified, for example, by determining whether it belongs to a branch office review, online platform review, or telephone follow-up review based on data source identifiers, interface types, or data field characteristics. Since different channels may differ in their rating methods—for example, some channels use a percentage rating system, while others use a five-point system or a satisfaction / dissatisfaction rating system—a corresponding rating mapping rule library is pre-established based on the channel source, and the rating mapping rule corresponding to the current data is matched from it. Subsequently, the rating information in the first set of review data is uniformly converted according to the rating mapping rules, for example, mapping ratings of different dimensions to a unified standard rating range, thereby achieving data normalization and making the review data from different channels comparable on the same scale. The above processing not only improves the standardization and consistency of the data, but also provides a unified and reliable data foundation for subsequent data sharding, heat analysis and scheduling.
[0038] Preferably, in some embodiments of this application, before receiving customer evaluation data, the method further includes unified access processing of multi-channel evaluation data, specifically including: receiving customer evaluation data from online business halls, mobile applications, mini-programs, SMS follow-ups, telephone follow-ups, and counter evaluation terminals through a unified access gateway; performing identity authentication processing on access requests to verify the legality of the data source and interface call permissions; after successful authentication, performing rate limiting control and queued caching processing on the customer evaluation data to avoid data congestion in high-concurrency scenarios; and simultaneously, supplementing each piece of customer evaluation data with a unified timestamp, channel identifier, and access batch identifier, and generating an association key based on the customer identifier, business acceptance number, and time window for subsequent association matching between evaluation data and business record data.
[0039] This application improves data quality and comparability by denoising, identifying channels, and normalizing ratings in customer review data. This results in consistent data format and evaluation standards for review data from different channels, providing a reliable foundation for subsequent data sharding and processing, and enhancing the accuracy and effectiveness of overall data processing.
[0040] Furthermore, in some embodiments of this application, the operational data includes: processing load, access latency, resource utilization, and data read / write request frequency; generating a node score characterizing the resource sufficiency of the distributed nodes based on the operational data includes: Based on the processing load, determine the load index that characterizes the resource consumption of the distributed nodes; Determine a latency metric characterizing the data access performance of the distributed nodes based on the access latency; Based on the resource utilization rate, resource indicators characterizing the available capacity of the distributed nodes are determined; The access pressure index, which characterizes the data access pressure of the distributed node, is determined based on the frequency of data read and write requests. The node score is obtained by weighted summation based on the load metric, the latency metric, the resource metric, and the access pressure metric.
[0041] Specifically, during the operation of the distributed system, the system continuously collects operational data from each distributed node to characterize the current resource status and access pressure of the nodes. This operational data includes at least processing load, access latency, resource utilization, and data read / write request frequency. Processing load reflects the computational and storage pressure a node bears at any given moment, preferably obtained by weighted fusion of CPU utilization, memory utilization, and disk I / O usage. Access latency reflects the latency performance of a node when responding to data read or write requests, and can be obtained by statistically analyzing request response time or network round-trip latency. Resource utilization characterizes the remaining available resources of a node, preferably measured by the percentage of available storage capacity or remaining computing resources. Data read / write request frequency reflects the data access pressure a node bears within a certain time window, and can be obtained based on the number of data reads and writes per unit time.
[0042] Furthermore, after acquiring the aforementioned operational data, corresponding indicators are constructed to characterize the different resource dimensions of the nodes. Specifically, a load indicator is determined based on processing load to describe the resource occupancy of the node; a latency indicator is determined based on access latency to describe the performance of data access; a resource indicator is determined based on resource utilization to describe the current available capacity of the node, where a higher resource utilization corresponds to a larger resource indicator value, indicating less remaining resources for the node; and an access pressure indicator is determined based on the frequency of data read / write requests to describe the intensity of data access currently being handled by the node and the concentration of hot data access. To ensure comparability between different indicators, each indicator can be normalized to map it to a uniform numerical range.
[0043] Furthermore, based on load, latency, resource, and access pressure metrics, a weighted summation is performed to calculate a node score. The weight parameters for each metric are pre-set according to the system operation strategy or adaptively adjusted using historical operational data. The node score comprehensively reflects the suitability of a node for currently handling new data shards or executing computational tasks. A higher node score indicates higher load, greater access latency, or resource constraints, making it unsuitable for further allocation of high-intensity data or computational tasks. Conversely, a lower node score indicates relatively abundant resources and better access performance, making it more suitable as a target node for data allocation or task execution. Based on this node score, it can be further used in conjunction with data shard popularity tags and distributed storage indexes to provide quantitative data allocation, shard migration, and task scheduling, thereby achieving dynamic balancing and efficient utilization of system resources.
[0044] S102: When an evaluation data processing task is received, several target storage nodes are determined from the storage nodes according to the data required by the evaluation data processing task.
[0045] Specifically, upon receiving an evaluation data processing task, the task is first parsed to determine its processing objectives and corresponding data requirement scope. This scope includes at least a time window range, a business type range, a customer region range, and the data types involved. These data types include one or more of the following: evaluation data shards, business record shards, and customer profile shards. Subsequently, based on a pre-built distributed storage index, a matching query is performed on the data requirement scope. The distributed storage index records the shard identifier, node identifier, primary key or time range, data version information, and popularity tag for each data shard. By searching the distributed storage index, the set of data shards that satisfy the evaluation data processing task is determined, and the current storage node of each data shard is obtained and designated as the target storage node.
[0046] S103: Obtain the first network communication distance between each target storage node, and group the target storage nodes according to the first network communication distance to obtain at least one storage node group; wherein, the first network communication distance between each target storage node in the same storage node group is less than a first preset distance threshold or there is only one target storage node in the same storage node group.
[0047] Specifically, after identifying the target storage nodes, to further reduce cross-node data interaction overhead and improve task execution efficiency, the system acquires and models the first network communication distance between each target storage node. Specifically, based on the network topology information of the distributed deployment environment and node operation monitoring results, the system calculates the first network communication distance between any two target storage nodes. The first network communication distance can be comprehensively characterized by one or more indicators among network communication latency, inter-node bandwidth overhead, hop count, or physical / logical topology hierarchy differences. Preferably, node access latency is used as the basic metric, and weighted correction is applied by combining information on the node's rack location, the same data center, or cross-regional deployment, thereby obtaining the first network communication distance between nodes that reflects the actual data transmission cost. The first network communication distance between each node can be represented by a network communication distance matrix.
[0048] Furthermore, after obtaining the first network communication distance between each pair of target storage nodes, the target storage nodes are clustered and grouped according to a first preset distance threshold. Specifically, any two target storage nodes whose first network communication distance is less than the preset distance threshold are divided into the same storage node group. Based on this, several storage node groups are formed through connectivity expansion, so that any node within the same storage node group satisfies the distance constraint or indirectly satisfies the distance constraint through an intermediate node. For target storage nodes that fail to satisfy the distance threshold condition with other nodes, a separate storage node group containing only one node is formed.
[0049] Furthermore, during the grouping process, the grouping results can be optimized by combining the node score of each node and the popularity label of the data shards it carries: preferably, nodes with lower node scores and carrying the most popular data shards are aggregated into the same node group, so that subsequent tasks can complete the data reading and centralized processing within this node group; at the same time, multiple high-load nodes are avoided from being assigned to the same node group, thereby preventing the intensification of local resource competition. Through the above grouping process, each storage node group has strong proximity in spatial topology and good availability in terms of resource status.
[0050] Furthermore, based on the grouping results, in the subsequent task scheduling and feature extraction process, execution nodes can be selected first within the same storage node group, or the local aggregation and processing of data fragments can be completed within the node group. Cross-group data interaction is only performed when necessary, thereby forming a synergy with the aforementioned scheduling mechanism based on distributed storage index and dynamic allocation scheme, further reducing cross-node data transmission, reducing network overhead, and improving overall processing efficiency.
[0051] S104: Based on the scores of each node, determine the target computing node corresponding to each storage node group from all the computing nodes.
[0052] Furthermore, in some embodiments of this application, determining the target computing node corresponding to each storage node group from all the computing nodes based on the node scores includes the following steps S1041 to S1043: S1041: When the storage node group includes multiple target storage nodes, select a first storage node from the storage node group and migrate the data fragments of the other target storage nodes in the storage node group to the first storage node; Furthermore, in some embodiments of this application, when the storage node group includes multiple target storage nodes, selecting a first storage node from the storage node group includes: Several candidate storage nodes are determined based on the first network communication distance between each target storage node; the candidate storage node is a storage node whose sum of the first network communication distances with other target storage nodes in the storage node group is less than a second preset distance threshold; Among the candidate storage nodes, the target storage node with the highest node score is selected as the first storage node.
[0053] Specifically, when a storage node group includes multiple target storage nodes, to determine the first storage node for subsequent data aggregation or task execution within the node group, the system first further filters each target storage node based on the first network communication distance between nodes. Specifically, using the constructed inter-node network communication distance matrix, the system calculates the sum of the first network communication distances between each target storage node and the remaining target storage nodes in the same storage node group, characterizing the overall proximity or communication cost centrality of the node within the group. Subsequently, the sum of distances corresponding to each target storage node is compared with a second preset distance threshold, and nodes with a sum of distances less than this threshold are selected as candidate storage nodes. After obtaining the set of candidate storage nodes, the node scores of each node are compared, and the target storage node with the highest node score is selected as the first storage node. Since a higher node score usually indicates a higher current load or greater access pressure, in the specific implementation, the node score can be directionally processed according to the system scheduling strategy, such as reverse normalization of the score or exclusion of overloaded nodes during the threshold filtering stage, thereby ensuring that the finally selected first storage node, while satisfying topological proximity, still possesses relatively superior resource carrying capacity.
[0054] This application filters candidate storage nodes based on the first network communication distance between each target storage node, and further determines the first storage node by combining node scores, so that the selected node has both a better spatial location advantage and a higher resource processing capability, thereby reducing the overall overhead of data migration and improving the efficiency of centralized data processing.
[0055] S1042: When there is only one target storage node in the storage node group, the target storage node shall be used as the first storage node; S1043: For each of the storage node groups, determine the corresponding target computing node from all the computing nodes based on the first storage node and the node scores of the storage node group.
[0056] This application achieves centralized management of dispersed data by selecting a first storage node within the same storage node group and migrating data fragments from other target storage nodes to that first storage node. This enables subsequent data processing to be completed on a single node, thereby reducing cross-node data interaction, decreasing the number of data transfers and network overhead, and improving the overall efficiency of data processing.
[0057] Furthermore, in some embodiments of this application, determining the corresponding target computing node from all the computing nodes based on the first storage node and the node scores of the storage node group includes: The network latency and second network communication distance between each computing node and the first storage node are obtained, and the data transmission overhead between each computing node and the first storage node is calculated based on the network latency and the second network communication distance. Based on the second network communication distance, a number of candidate computing nodes whose second network communication distance with the target storage node is lower than a preset transmission threshold are determined. Among the candidate computing nodes, the target computing node is determined based on the corresponding node score.
[0058] Specifically, after determining the first storage node corresponding to each storage node group, to ensure that subsequent feature extraction or analysis calculations are completed as close to the data as possible, the system further filters target computing nodes within the entire computing node range. Specifically, it first acquires network status information and topology information between each computing node and the first storage node. The network status information includes at least the data access latency between nodes, and the topology information characterizes the physical or logical distance between nodes, i.e., the second network communication distance. This second network communication distance can be determined based on information such as the data center location, rack location, network hop count, or regional division of the node. Based on this, the data transmission cost between the two nodes is comprehensively measured by combining network latency and the second network communication distance. Preferably, a weighted method is used to calculate the data transmission overhead, reflecting the communication cost required to read data from the first storage node when performing a task on that computing node.
[0059] Furthermore, after obtaining the data transmission overhead corresponding to each computing node, it is compared with a preset transmission threshold, and computing nodes with transmission overhead lower than the threshold are selected as candidate computing nodes. This ensures that the candidate computing nodes are as close as possible to the first storage node or in the same area in the network topology, thereby reducing cross-node data transmission and network bandwidth consumption.
[0060] This application comprehensively considers the network latency and distance between the computing node and the first storage node, and selects candidate computing nodes accordingly. Then, it combines node scores to determine the target computing node, so that computing tasks can be executed on nodes with low data transmission overhead and sufficient resources. This reduces data transmission costs while improving computing efficiency and system response performance.
[0061] S105: Process customer evaluation data corresponding to each target storage node in the storage node group through the target computing node.
[0062] Specifically, during data processing, the target computing nodes first perform unified data parsing and structuring processing on the evaluation data in each target storage node, converting the evaluation data into a computable structured data format. They then extract features from the structured data, generating feature data that includes rating features, text features, business process features, customer profile features, and channel features. Afterward, the target computing nodes can perform statistical analysis, correlation analysis, anomaly detection, or predictive modeling operations according to task requirements. For example, they can perform dimensional aggregation statistics and trend analysis on the feature data, or execute satisfaction prediction model calculations based on the feature data and output the corresponding analysis or prediction results.
[0063] Furthermore, in some embodiments of this application, during data processing and result generation, the transmission of evaluation data, feature data, and analysis results is encrypted, and sensitive fields in storage nodes and computing nodes are subjected to hierarchical desensitization or encrypted storage. Simultaneously, based on preset access control policies, access permissions for different roles are limited, ensuring that users can only access data and analysis results within their authorized scope. Furthermore, the entire process of data reading, processing, exporting, and model invocation is recorded, generating corresponding audit logs to facilitate tracing and alerting in the event of abnormal access or unauthorized operations.
[0064] Furthermore, in some embodiments of this application, after completing the evaluation data analysis and outputting the results, the satisfaction analysis results are published through visualization or reports, and the analysis results and their corresponding business execution feedback are fed back to form feedback data. Based on the feedback data, the parameters and strategy configurations of the models embedded in the computing nodes are updated or optimized, for example, by continuously training the models through incremental training or by adaptively adjusting the threshold parameters according to the model performance indicators. When a decline in model performance or a change in data distribution is detected, a model retraining or strategy adjustment process is triggered, thereby enabling the system to continuously adapt to business changes during long-term operation and ensuring the accuracy and stability of the evaluation data analysis results. Through the above-mentioned security control and feedback iteration mechanism, the evaluation data processing flow forms a complete analysis chain while possessing traceability, compliance, and continuous optimization capabilities, and forms a closed-loop collaboration with the aforementioned data processing, node scheduling, and computation execution processes.
[0065] In summary, the power customer evaluation data processing method provided in this application has the following advantages compared to the prior art: This application quantifies the sufficiency of node resources by acquiring the operational data of each distributed node and generating node scores, thereby providing a basis for selecting computing nodes and avoiding performance bottlenecks caused by task concentration. Simultaneously, after receiving a processing task, it groups the data based on the first network communication distance between the target storage nodes, allowing spatially similar data to be processed together, thus reducing cross-node data interaction and transmission latency. Furthermore, it matches target computing nodes to each storage node group according to the node scores, prioritizing data processing on nodes with more abundant resources, improving both resource utilization and processing efficiency. Through the synergistic effect of the above technical means, this application improves the overall processing efficiency of power customer evaluation data in a distributed environment while reducing data transmission overhead.
[0066] like Figure 2 As shown, based on the above method embodiments, an embodiment of this application provides an electricity customer evaluation data processing device, including: a scoring module 201, a first screening module 202, a grouping module 203, a second screening module 204, and a processing module 205; The scoring module 201 is used to acquire the running data corresponding to each distributed node, and generate a node score representing the resource sufficiency of the distributed node based on the running data; the distributed node includes: storage node and computing node; The first filtering module 202 is used to determine several target storage nodes from the storage nodes according to the data required by the evaluation data processing task when an evaluation data processing task is received. The grouping module 203 is used to obtain a first network communication distance between each target storage node, and to group the target storage nodes according to the first network communication distance to obtain at least one storage node group; wherein, the first network communication distance between each target storage node in the same storage node group is less than a first preset distance threshold or there is only one target storage node in the same storage node group; The second filtering module 204 is used to determine the target computing node corresponding to each storage node group from all the computing nodes based on the scores of each node. The processing module 205 is used to process customer evaluation data corresponding to each target storage node in the storage node group through the target computing node.
[0067] Further, in some embodiments of this application, the second filtering module 204 includes: a first filtering unit, a second filtering unit, and a first calculation unit; the second filtering module 204 is used to determine the target computing node corresponding to each storage node group from all the computing nodes according to the node scores, including: The first filtering unit is used to select a first storage node from the storage node group when the storage node group includes multiple target storage nodes, and to migrate data fragments from other target storage nodes in the storage node group to the first storage node; The second filtering unit is used to select the target storage node as the first storage node when there is only one target storage node in the storage node group. The first computing unit is configured to, for each of the storage node groups, determine the corresponding target computing node from all the computing nodes based on the first storage node and the node scores of the storage node group.
[0068] Further, in some embodiments of this application, the first computing unit is configured to determine the corresponding target computing node from all the computing nodes based on the first storage node and the node scores of the storage node group, including: The network latency and second network communication distance between each computing node and the first storage node are obtained, and the data transmission overhead between each computing node and the first storage node is calculated based on the network latency and the second network communication distance. Based on the second network communication distance, a number of candidate computing nodes whose second network communication distance with the target storage node is lower than a preset transmission threshold are determined. Among the candidate computing nodes, the target computing node is determined based on the corresponding node score.
[0069] Furthermore, in some embodiments of this application, the first filtering unit is configured to select a first storage node from the storage node group when the storage node group includes multiple target storage nodes, including: Several candidate storage nodes are determined based on the first network communication distance between each target storage node; the candidate storage node is a storage node whose sum of the first network communication distances with other target storage nodes in the storage node group is less than a second preset distance threshold; Among the candidate storage nodes, the target storage node with the highest node score is selected as the first storage node.
[0070] Furthermore, in some embodiments of this application, upon receiving customer review data, the process includes: The customer review data is divided into several data segments; each data segment corresponds to a popularity tag that represents the frequency of data reading. The data segments are classified according to the popularity tags corresponding to each data segment to obtain a first popularity data segment and a second popularity data segment; wherein, the data reading frequency of the first popularity data segment is greater than the data reading frequency of the second popularity data segment; Based on the node score corresponding to each of the storage nodes, determine the first storage node whose node score is higher than a preset score threshold; The first heat data shard is allocated to the first storage node, and the second heat data shard is allocated to the remaining storage nodes.
[0071] Furthermore, in some embodiments of this application, upon receiving customer review data, the method further includes: The customer review data is denoised to obtain the first review data; Identify the source of the first evaluation data and determine the scoring mapping rule for the first evaluation data based on the source of the first evaluation data; The first evaluation data is normalized according to the scoring mapping rules.
[0072] Further, in some embodiments of this application, the operational data includes: processing load, access latency, resource utilization, and data read / write request frequency; the scoring module 201 includes: a second computing unit, a third computing unit, a fourth computing unit, a fifth computing unit, and a sixth computing unit; the scoring module 201 is used to generate a node score characterizing the resource sufficiency of the distributed node based on the operational data, including: The second computing unit is used to determine a load index that characterizes the resource consumption of the distributed nodes based on the processing load. The third computing unit is used to determine a latency index characterizing the data access performance of the distributed node based on the access latency; The fourth calculation unit is used to determine resource indicators that characterize the available capacity of the distributed nodes based on the resource occupancy rate. The fifth computing unit is used to determine an access pressure index that characterizes the data access pressure of the distributed node based on the frequency of data read and write requests. The sixth calculation unit is used to perform a weighted summation based on the load index, the latency index, the resource index, and the access pressure index to obtain the node score.
[0073] In summary, the power customer evaluation data processing device provided in this application has the following advantages compared to the prior art: This application quantifies the sufficiency of node resources by acquiring the operating data of each distributed node and generating node scores, thereby providing a basis for selecting computing nodes and avoiding performance bottlenecks caused by task concentration. Simultaneously, after receiving a processing task, it groups the data based on the first network communication distance between the target storage nodes, allowing spatially similar data to be processed together, thus reducing cross-node data interaction and transmission latency. Furthermore, it matches target computing nodes to each storage node group according to the node scores, prioritizing data processing on nodes with sufficient resources, improving both resource utilization and processing efficiency. Through the synergistic effect of the above technical means, this application improves the overall processing efficiency of power customer evaluation data in a distributed environment while reducing data transmission overhead.
[0074] It is understood that the above-described device embodiments correspond to the method embodiments of this application, and can implement the power customer evaluation data processing method provided by any of the above-described method embodiments of this application.
[0075] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided in this application, the connection relationships between modules indicate that they have communication connections, which can specifically be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0076] Based on the above embodiments of the electricity customer evaluation data processing method, another embodiment of this application provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the electricity customer evaluation data processing method of any embodiment of this application.
[0077] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete this application. The one or more module units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the terminal device.
[0078] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0079] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.
[0080] Based on the above-described method embodiments, another embodiment of this application provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the power customer evaluation data processing method described in any of the above-described method embodiments of this application.
[0081] The modules / units integrated in the device / terminal equipment, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
Claims
1. A power customer evaluation data processing method characterized by, include: Obtain the operational data corresponding to each distributed node, and generate a node score representing the resource sufficiency of the distributed node based on the operational data; The distributed nodes include: storage nodes and computing nodes; When an evaluation data processing task is received, several target storage nodes are determined from the storage nodes according to the data required by the evaluation data processing task. Obtain the first network communication distance between each target storage node, and group the target storage nodes according to the first network communication distance to obtain at least one storage node group; wherein, the first network communication distance between each target storage node in the same storage node group is less than a first preset distance threshold or there is only one target storage node in the same storage node group; Based on the scores of each node, the target computing node corresponding to each storage node group is determined from all the computing nodes; The target computing node processes customer evaluation data corresponding to each target storage node in the storage node group.
2. The power customer evaluation data processing method of claim 1, wherein, The step of determining the target computing node corresponding to each storage node group from all the computing nodes based on the scores of each node includes: When the storage node group includes multiple target storage nodes, a first storage node is selected from the storage node group, and data fragments from other target storage nodes in the storage node group are migrated to the first storage node. When there is only one target storage node in the storage node group, the target storage node shall be designated as the first storage node. For each of the storage node groups, the target computing node is determined from all the computing nodes based on the first storage node and the scores of each node in the storage node group.
3. The electric power customer evaluation data processing method of claim 2, wherein, The step of determining the corresponding target computing node from all the computing nodes based on the first storage node and the scores of each node in the storage node group includes: Obtain the second network communication distance between each computing node and the first storage node; Based on the second network communication distance, a number of candidate computing nodes whose second network communication distance with the target storage node is lower than a preset transmission threshold are determined. Among the candidate computing nodes, the target computing node is determined based on the corresponding node score.
4. The power customer evaluation data processing method of claim 2, wherein, When the storage node group includes multiple target storage nodes, selecting a first storage node from the storage node group includes: Several candidate storage nodes are determined based on the first network communication distance between each target storage node; the candidate storage node is a storage node whose sum of the first network communication distances with other target storage nodes in the storage node group is less than a second preset distance threshold; Among the candidate storage nodes, the target storage node with the highest node score is selected as the first storage node.
5. The electric power customer evaluation data processing method of claim 1, wherein, Upon receiving customer review data, including: The customer review data is divided into several data segments; each data segment corresponds to a popularity tag that represents the frequency of data reading. The data segments are classified according to the popularity tags corresponding to each data segment to obtain a first popularity data segment and a second popularity data segment; wherein, the data reading frequency of the first popularity data segment is greater than the data reading frequency of the second popularity data segment; Based on the node score corresponding to each of the storage nodes, determine the first storage node whose node score is higher than a preset score threshold; The first heat data shard is allocated to the first storage node, and the second heat data shard is allocated to the remaining storage nodes.
6. The electric power customer evaluation data processing method of claim 5, wherein, After receiving customer review data, the following is also included: The customer review data is denoised to obtain the first review data; Identify the source of the first evaluation data and determine the scoring mapping rule for the first evaluation data based on the source of the first evaluation data; The first evaluation data is normalized according to the scoring mapping rules.
7. The electric power customer evaluation data processing method of claim 1, wherein, The operational data includes: processing load, access latency, resource utilization, and data read / write request frequency; the generation of a node score characterizing the resource sufficiency of the distributed nodes based on the operational data includes: Based on the processing load, determine the load index that characterizes the resource consumption of the distributed nodes; Determine a latency metric characterizing the data access performance of the distributed nodes based on the access latency; Based on the resource utilization rate, resource indicators characterizing the available capacity of the distributed nodes are determined; The access pressure index, which characterizes the data access pressure of the distributed node, is determined based on the frequency of data read and write requests. The node score is obtained by weighted summation based on the load metric, the latency metric, the resource metric, and the access pressure metric.
8. An electric power customer evaluation data processing apparatus characterized by comprising: include: The system includes a scoring module, a first filtering module, a grouping module, a second filtering module, and a processing module. The scoring module is used to obtain the running data corresponding to each distributed node and generate a node score that characterizes the resource sufficiency of the distributed node based on the running data. The distributed nodes include: storage nodes and computing nodes; The first filtering module is used to determine several target storage nodes from the storage nodes according to the data required by the evaluation data processing task when an evaluation data processing task is received. The grouping module is used to obtain a first network communication distance between each target storage node, and to group the target storage nodes according to the first network communication distance to obtain at least one storage node group; wherein, the first network communication distance between each target storage node in the same storage node group is less than a first preset distance threshold or there is only one target storage node in the same storage node group; The second filtering module is used to determine the target computing node corresponding to each storage node group from all the computing nodes based on the scores of each node. The processing module is used to process customer evaluation data corresponding to each target storage node in the storage node group through the target computing node.
9. A terminal device, comprising: The device includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements a power customer evaluation data processing method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform a power customer evaluation data processing method as described in any one of claims 1 to 7.