A Market Participant Behavior Prediction System for Electricity Trading Platforms Based on Big Data Analysis

CN122089463APending Publication Date: 2026-05-26SHENZHEN COMTOP INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN COMTOP INFORMATION TECH
Filing Date
2025-12-23
Publication Date
2026-05-26

Smart Images

  • Figure CN122089463A_ABST
    Figure CN122089463A_ABST
Patent Text Reader

Abstract

This invention relates to the field of intelligent big data analysis and decision support technology in the power market, specifically a big data analysis-based system for predicting the behavior of market participants in a power trading platform. The system includes: a data acquisition module, used to configure and execute real-time acquisition strategies for data filtering and parsing based on multi-source heterogeneous data feature analysis from the power trading platform, third-party data service providers, and enterprise internal management systems, to obtain a preliminary processed data stream; and to implement parallel transmission aggregation and transmission path optimization control on the preliminary processed data stream to obtain a raw dataset containing market transaction data, basic data of market participants, and external influence data. This invention achieves efficient processing and accurate decision support for the entire chain of power market transaction data through a systematic process of integrated data acquisition, intelligent quality optimization, multi-dimensional behavior prediction, and dynamic risk warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent analysis and decision support technology for big data in the power market, and more specifically, to a system for predicting the behavior of market participants in a power trading platform based on big data analysis. Background Technology

[0002] Provincial-level power trading centers bring together various entities, including power generation companies, electricity sales companies, and large power users. Their trading activities are highly complex and dynamic, making accurate forecasting crucial for optimizing trading decisions, ensuring stable grid operation, and implementing effective regulation. However, traditional forecasting methods primarily rely on statistical analysis of historical data, fixed rules, or single-model extrapolation. These methods struggle to effectively address the distinct characteristics of power market data, such as its multi-source (trading, meteorological, fuel, policy), heterogeneous (structured data, unstructured text, real-time streaming data), strong time-series nature, and high real-time requirements. This leads to the following core deficiencies in data processing: First, existing systems lack a unified analysis and adaptive access mechanism for diverse data characteristics. Data fusion relies on manual rules, making it difficult to automatically resolve semantic conflicts and format standardization issues, potentially leading to a loss of control over data quality at the source. Second, predictive models lack the ability to adapt to complex market dynamics. Mainstream models are typically based on stationarity assumptions, making it difficult to effectively capture non-stationary and nonlinear market characteristics caused by extreme weather, policy adjustments, etc. More importantly, models generally ignore the game interaction and collaborative behavior between market participants, potentially resulting in low efficiency when predicting key risk scenarios such as price manipulation and joint bidding. Third, there are architectural bottlenecks in real-time data processing and high-concurrency access. Spot trading requires decision-making responses at the minute or even second level, while traditional batch processing or micro-batch processing architectures have inherent latency, which may not meet the needs of rapid calculation and simulation of massive real-time data during peak trading periods, resulting in weak system concurrency support capabilities. Summary of the Invention

[0003] The technical problem to be solved by this invention is to overcome the shortcomings of the existing technology and provide a power trading platform market participant behavior prediction system based on big data analysis. Through a systematic process of integrated data collection, intelligent quality optimization, multi-dimensional behavior prediction and dynamic risk warning, it realizes efficient processing and accurate decision support for the entire chain of power market transaction data.

[0004] To solve the above-mentioned technical problems, the basic concept of the technical solution adopted by the present invention is as follows: Firstly, a market participant behavior prediction system for an electricity trading platform based on big data analysis includes: The data acquisition module is used to configure and execute real-time acquisition strategies to filter and parse data based on the multi-source heterogeneous data characteristics analysis of the power trading platform, third-party data service providers and enterprise internal management systems, and to obtain a preliminary processed data stream. Parallel transmission aggregation and transmission path optimization control are implemented on the preliminary processed data stream to obtain the raw dataset containing market transaction data, main basic data and external impact data. The optimization module performs multidimensional quality assessments on the original dataset in terms of data integrity, consistency, and timeliness, resulting in a quality assessment matrix. Based on the quality assessment matrix, feature extraction and pattern recognition are performed to form a core quality feature vector set. A strategy configuration vector is generated based on the core quality feature vector set. The strategy configuration vector is used to dynamically adapt data cleaning rules, standardization thresholds, and feature extraction parameters to construct a standardized dataset with unified semantics and format. The prediction module is used to input standardized datasets into a prediction model trained on historical regular data to obtain the distribution of trading strategy tendencies and risk assessment indicators of collaborative behavior, and to integrate them to obtain comprehensive prediction data. The early warning module is used to obtain a market state evolution matrix based on comprehensive forecast data; integrate the market state evolution matrix with collaborative behavior risk assessment indicators to obtain a risk feature vector set; perform dynamic threshold judgment on the risk feature vector set to obtain a list of abnormal risk behaviors; generate structured early warning prompts with embedded priority identifiers based on the list of abnormal risk behaviors, and perform differentiated information push through the early warning information routing and distribution device.

[0005] Secondly, a control method for a power trading platform market participant behavior prediction system based on big data analysis includes the following steps: Based on the multi-source heterogeneous data feature analysis of power trading platforms, third-party data service providers and enterprise internal management systems, a real-time acquisition strategy is configured and executed to filter and parse the data to obtain a preliminary processed data stream. Parallel transmission aggregation and transmission path optimization control are implemented on the preliminary processed data stream to obtain a raw dataset containing market transaction data, main basic data and external impact data. A multidimensional quality assessment is performed on the original dataset in terms of data integrity, consistency, and timeliness to obtain a quality assessment matrix. Feature extraction and pattern recognition are performed based on the quality assessment matrix to form a core quality feature vector set. A strategy configuration vector is generated based on the core quality feature vector set. The strategy configuration vector is used to dynamically adapt data cleaning rules, standardization thresholds, and feature extraction parameters to construct a standardized dataset with unified semantics and format. The standardized dataset is input into a prediction model trained on historical regular data to obtain the distribution of trading strategy tendencies and risk assessment indicators of collaborative behavior, and then integrated to obtain comprehensive prediction data. Based on comprehensive forecast data, a market state evolution matrix is ​​obtained; by integrating the market state evolution matrix with collaborative behavior risk assessment indicators, a risk feature vector set is obtained; dynamic threshold judgment is performed on the risk feature vector set to obtain a list of abnormal risk behaviors; a structured early warning prompt with embedded priority identifier is generated based on the list of abnormal risk behaviors, and differentiated information is pushed through an early warning information routing and distribution device.

[0006] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art.

[0007] By analyzing the characteristics of heterogeneous multi-source data, the adaptability of data acquisition is improved; the configuration and execution of real-time acquisition strategies enable accurate data filtering and parsing, reducing interference from irrelevant data; parallel transmission aggregation and transmission path optimization control improve data transmission efficiency and aggregation timeliness, ensuring efficient acquisition of raw datasets; multi-dimensional quality assessment covers data integrity, consistency, and timeliness, providing a comprehensive basis for data quality control; feature extraction and pattern recognition based on the quality assessment matrix accurately extract core quality features; strategy configuration vectors dynamically adapt to data processing rules and parameters, improving the flexibility and adaptability of data processing; and the construction of standardized datasets with unified semantics and formats enhances the universality of data. The system ensures compatibility with subsequent processing; standardized dataset input ensures the uniformity of model data input and reduces the impact of data format differences on the processing; the integration of trading strategy tendency distribution and collaborative behavior risk assessment indicators realizes the systematic integration of multi-dimensional predictive information, providing comprehensive data support for subsequent stages; the construction of the market state evolution matrix realizes the structured representation of market operation state data, facilitating the mining of risk characteristics; the integration of market state and risk assessment indicators enhances the comprehensiveness and correlation of risk feature vector sets; dynamic threshold judgment adapts to the dynamic change characteristics of data, enhancing the adaptability of anomaly identification; structured early warning prompts and differentiated push notifications improve the organization and standardization of early warning information and the targeted delivery, ensuring information transmission efficiency.

[0008] The specific embodiments of the present invention will now be described in further detail with reference to the accompanying drawings. Attached Figure Description

[0009] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. Some specific embodiments of this application will be described in detail below with reference to the accompanying drawings in an exemplary and non-limiting manner. The same reference numerals in the drawings designate the same or similar parts or components. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings: Figure 1This is a schematic diagram of the market participant behavior prediction system for the power trading platform based on big data analysis, as described in this invention.

[0010] Figure 2 This is a schematic diagram of the control method of the power trading platform market participant behavior prediction system based on big data analysis, according to the present invention.

[0011] It should be noted that these accompanying drawings and textual descriptions are not intended to limit the scope of the invention in any way, but rather to illustrate the concept of the invention to those skilled in the art by referring to specific embodiments. The elements in the drawings are schematic and not drawn to scale. Detailed Implementation

[0012] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort should fall within the scope of protection of the present application.

[0013] The following embodiments of this application use a big data analysis-based power trading platform market participant behavior prediction system as an example to illustrate the solution of this application in detail. However, this embodiment does not limit the scope of protection of this application.

[0014] like Figure 1 As shown, this invention provides a market participant behavior prediction system for power trading platforms based on big data analysis, comprising: The acquisition module 11 is used to configure and execute real-time acquisition strategies to perform data filtering and parsing based on the multi-source heterogeneous data feature analysis of the power trading platform, third-party data service providers and enterprise internal management systems, to obtain a preliminary processed data stream; and to implement parallel transmission aggregation and transmission path optimization control on the preliminary processed data stream to obtain a raw dataset containing market transaction data, main basic data and external impact data. Optimization module 12 is used to perform multi-dimensional quality assessment on the original dataset in terms of data integrity, consistency and timeliness to obtain a quality assessment matrix; based on the quality assessment matrix, feature extraction and pattern recognition are performed to form a core quality feature vector set; a strategy configuration vector is generated based on the core quality feature vector set; and the strategy configuration vector is used to dynamically adapt data cleaning rules, standardization thresholds and feature extraction parameters to construct a standardized dataset with unified semantics and format. Prediction module 13 is used to input standardized datasets into a prediction model trained with historical regular data to obtain the distribution of trading strategy tendencies and risk assessment indicators of collaborative behavior, and to integrate them to obtain comprehensive prediction data. The early warning module 14 is used to obtain a market state evolution matrix based on comprehensive forecast data; integrate the market state evolution matrix with collaborative behavior risk assessment indicators to obtain a risk feature vector set; perform dynamic threshold judgment on the risk feature vector set to obtain a list of abnormal risk behaviors; generate a structured early warning prompt with embedded priority identifier based on the list of abnormal risk behaviors, and perform differentiated information push through the early warning information routing and distribution device.

[0015] In this embodiment of the invention, the following features are adapted to improve the targeting and adaptability of data collection: parallel data transmission and convergence with path optimization are achieved, enhancing data transmission efficiency and convergence timeliness; the original dataset is ensured to cover multi-dimensional key information, providing comprehensive data support for subsequent processing; multi-dimensional quality assessment enables comprehensive control of data quality and extracts core quality features; data cleaning rules, standardized thresholds, and feature extraction parameters are dynamically adapted, improving the flexibility and adaptability of data processing; a standardized dataset with unified semantics and format is constructed, reducing the complexity of subsequent data processing and improving data reusability; the consistency of data input is ensured by relying on the standardized dataset, assisting the model in efficiently processing data; the simultaneous acquisition of trading strategy tendency distribution and collaborative behavior risk assessment indicators is achieved, enriching prediction dimensions; comprehensive prediction data is formed through data fusion, improving the comprehensiveness of prediction results; market state evolution matrix and collaborative behavior risk assessment indicators are integrated to comprehensively extract risk features; dynamic threshold judgment improves the accuracy of identifying abnormal risk behaviors; structured early warning prompts with embedded priority identifiers are generated to determine early warning priorities; differentiated information push improves the targeting of early warning information delivery, ensuring key entities receive important early warning information in a timely manner.

[0016] In the power trading platform market participant behavior prediction system based on big data analysis described in this embodiment of the invention, the aforementioned data acquisition module 11, based on the multi-source heterogeneous data feature analysis of the power trading platform, third-party data service providers, and enterprise internal management systems, configures and executes a real-time data acquisition strategy to perform data filtering and parsing, obtaining a preliminary processed data stream; parallel transmission aggregation and transmission path optimization control are implemented on the preliminary processed data stream to obtain an original dataset containing market transaction data, basic participant data, and external influence data, including: Step 1101 involves performing data source feature analysis on multi-source heterogeneous data from power trading platforms, third-party data service providers, and enterprise internal management systems. Data source feature vectors are extracted from each data source. These feature vectors must contain quantified values ​​for data type, data update frequency, and data scale. Specifically, this includes: conducting data source feature analysis on the multi-source heterogeneous data from power trading platforms, third-party data service providers, and enterprise internal management systems; determining the data type attribute for each data source; counting the number of data updates per unit time to determine the data update frequency; measuring the data storage capacity or number of data records to define the data scale; converting the attribute information corresponding to the data type, data update frequency, and data scale into quantifiable values; and then extracting data source feature vectors containing the aforementioned quantified values ​​from the multi-source heterogeneous data of each data source.

[0017] Step 1102: Based on the feature vectors of the data source, calculate the data sampling rate, key field filtering conditions, and parsing template by querying the mapping table, retrieving the rule base, and matching the parsing template. Based on the data sampling rate, key field filtering conditions, and parsing template, obtain the real-time acquisition strategy parameter set. Specifically, this includes: retrieving a preset feature vector and acquisition parameter mapping table. This preset mapping table is based on feature vector samples from historical multi-source heterogeneous data sources in the power trading field. It combines the acquisition effect verification results corresponding to different feature vectors to select the optimal acquisition parameters, establishes the correlation between feature vectors and optimal acquisition parameters through statistical analysis, and pre-forms a mapping table by organizing various feature vector entries and their corresponding basic data sampling rates. Compare the extracted data source feature vectors with the feature vector entries in the mapping table to obtain the corresponding basic data sampling rate. Retrieve a preset acquisition rule base. This preset rule base is pre-formed by combining the core business needs of power trading data processing. This process involves identifying core data fields valuable for subsequent predictive analysis under different data types and update frequencies, developing corresponding key field filtering rules for various combinations of data types and update frequencies, and pre-setting all filtering rules into a data collection rule library. Based on the data type and update frequency information in the data source feature vector, corresponding key field filtering rules are matched from the data collection rule library to determine the key field filtering conditions. Based on the data type information in the data source feature vector, suitable format parsing templates are retrieved and matched from a pre-set parsing template library. The pre-set parsing templates are designed by identifying common format types of multi-source heterogeneous data in power trading scenarios, designing corresponding format parsing logic and conversion rules for each format type, generating standardized parsing templates, and pre-setting all parsing templates into the parsing template library. Finally, the basic data sampling rate, matched key field filtering conditions, and parsing templates obtained from the above queries are integrated and, after parameter verification, form a real-time data collection strategy parameter set.

[0018] Step 1103: Based on the real-time acquisition strategy parameter set, perform real-time conditional filtering and format parsing on the raw data streams of each data source to filter out irrelevant data and unify the data format, obtaining a preliminary processed data stream. Specifically, this includes: filtering the raw data streams output by each data source in real time based on key field filtering conditions such as subject identifier, data type, and timestamp in the real-time acquisition strategy parameter set; verifying the key field values ​​of each data record in the raw data stream line by line to determine whether they conform to the preset field value range or matching rules, and removing irrelevant data that does not meet the conditions in real time; according to the format type of each data source, calling the pre-matched exclusive parsing template in the real-time acquisition strategy parameter set to parse the format of the raw data streams of each data source after filtering; converting data streams of different formats such as JSON, XML, and CSV into preset standardized data table formats through the field extraction and format conversion rules built into the template, ensuring that the encoding rules and numerical precision of each field are consistent; sorting and integrating the data streams of each data source after real-time conditional filtering and format unification by data acquisition timestamp in ascending order, simultaneously verifying the integrity and time consistency of data records, removing duplicate data records, and finally obtaining a preliminary processed data stream with a standardized structure and complete information.

[0019] Step 1104: Perform dynamic data segmentation on the preliminary processed data stream to obtain a data segment set consisting of market transaction data, entity basic data, and external influence data; obtain real-time network status information, specifically including: retrieving preset data classification rules. The preset method for these data classification rules is to combine the business objectives of predicting the behavior of market participants on the power trading platform, sort out the business attributes corresponding to various types of data in the preliminary processed data stream, and determine the definition standards for market transaction data, entity basic data, and external influence data. Among them, market transaction data is defined as data related to power trading quotations, transactions, and settlements; entity basic data is defined as basic information data such as the qualifications, installed capacity, and user scope of market participants such as power generation companies and power sales companies; and external influence data is defined as external related data such as weather, fuel prices, and policy regulation that affect trading behavior. Based on this, data classification rules are formulated and preset; according to the above data classification rules and the corresponding data... Based on business attributes, the initial data stream is divided into three data subsets: market transaction data, main basic data, and external impact data. For each data subset, dynamic splitting is performed according to preset splitting thresholds and data transmission requirements. The preset splitting thresholds are based on historical data transmission efficiency statistics and the average data volume of different data subsets to set splitting thresholds suitable for different data transmission scenarios. The preset data transmission requirements are determined based on the real-time requirements of power trading, and preset requirements such as transmission timeliness and reliability levels for each data subset. Based on the preset splitting thresholds and data transmission requirements, each data subset is dynamically split to form a set of fragments containing multiple data fragments. At the same time, the system's built-in network status monitoring unit collects real-time status information of each available transmission link in the current network environment. This real-time network status information includes at least link bandwidth, transmission delay, and link occupancy rate.

[0020] Step 1105: Based on the data size of each fragment in the data fragment set and the preset fragment priority, combined with the bandwidth and latency indicators of each available path in the real-time network status, calculate the path adaptability metric for each data fragment using a preset path evaluation function. Specifically, this includes: extracting the specific data size of each data fragment in the data fragment set. This data fragment set is formed based on preset data classification rules, fragmentation thresholds, and data transmission requirements. The preset data classification rules are formulated by combining the business objectives of predicting the behavior of market participants on the power trading platform, sorting out data business attributes, and defining standards for market transaction data, basic data of the main body, and external influence data. The preset fragmentation thresholds are set based on historical data transmission efficiency statistics combined with the average data size of different types of data subsets to set a splitting threshold suitable for the transmission scenario. The preset data transmission requirements are... To meet the real-time requirements of power trading, the required indicators such as transmission timeliness and reliability levels of various data subsets are determined and preset. Preset fragmentation priority division rules are retrieved to determine the priority level corresponding to each data fragment. The actual bandwidth values ​​and real-time transmission delay indicators of each available transmission path are extracted from the real-time network status information obtained in step 1104. A preset path evaluation function is retrieved. This function integrates the key attributes of data fragments with the performance indicators of transmission paths to achieve a quantitative evaluation of the compatibility between data fragments and each available transmission path. The data size, priority level of each data fragment, and the bandwidth and transmission delay indicators of each available transmission path are substituted into the preset path evaluation function. The function calculates the path compatibility metric between each data fragment and each available transmission path. This path compatibility metric provides a quantitative basis for the subsequent dynamic allocation of suitable transmission paths to each data fragment.

[0021] Step 1106: Based on the path fit metric, dynamically allocate suitable transmission paths to each data fragment according to the load balancing strategy, and generate a routing table containing path identifiers and fragment sequence numbers. Specifically, this includes: sorting the path fit metric values ​​of each data fragment and each available transmission path; defining a convex optimization objective function to minimize the load variance of each transmission path and maximize the sum of the fit between the data fragment and the path; constructing a feasible region using the maximum load threshold of a single transmission path as a constraint; initializing the load weight parameters of each transmission path and the data fragment path allocation variables; setting the iteration step size and convergence threshold; and calculating the gradient value of the objective function based on the current path fit metric value and the real-time load data of each transmission path, and updating the path allocation variables along the gradient descent direction. The updated path allocation variables are projected into a preset feasible region to ensure that the load of each transmission path does not exceed the maximum threshold, thus avoiding excessive load on a single transmission path. Gradient calculation, variable update, and projection operations are repeated until the change in path allocation variables between two adjacent iterations is less than the convergence threshold, completing the iterative calculation of the convex optimization projection gradient descent algorithm. Based on the path allocation results after iterative convergence, the transmission path with the highest fitness and a reasonable load range is dynamically selected for each data fragment. The path identifier of the selected transmission path corresponding to each data fragment is recorded, and a unique fragment sequence number is assigned to each data fragment. The fragment sequence number, the corresponding path identifier, and the data fragment category information of each data fragment are associated and integrated to generate a corrected routing table containing the above information.

[0022] Step 1107: Based on the routing schedule table, perform multi-path parallel transmission and data aggregation to obtain a preliminary aggregated data stream. Specifically, this includes: first, parsing the routing schedule table to extract the unique path identifier, target node address, and transmission priority information corresponding to each data fragment, establishing a one-to-one mapping relationship between data fragments and transmission paths; second, allocating corresponding transmission channel resources to each data fragment according to the mapping relationship, synchronously initializing the multi-path parallel transmission control module, setting the bandwidth threshold, retransmission mechanism, and transmission status monitoring frequency for each transmission path, starting the multi-path parallel data transmission process, and monitoring the transmission rate, latency, and data packet loss of each path in real time. In case of data loss, at the data stream receiving end of each transmission path, a real-time receiving buffer unit is deployed to receive and verify the integrity of each data fragment arriving in real time, discarding data fragments with damaged or missing fields, and marking the received timestamp of the data fragments that pass the verification; extracting the fragment sequence number and data category information from the data fragments that pass the verification, grouping and classifying them according to data category, and sorting and completing them in ascending order of fragment sequence number within the same category, checking for any missing fragments, and triggering a retransmission request if any exist; integrating and splicing the sorted and completed data fragments of each data category to generate a preliminary aggregated data stream containing complete data content and divided by category.

[0023] Step 1108: During the execution of multi-path parallel transmission, real-time performance monitoring is performed on each transmission link carrying market transaction data, main body basic data, and external impact data fragments to obtain transmission delay and packet loss rate indicators, thus obtaining transmission performance monitoring data. Specifically, this includes: initiating a link performance monitoring process throughout the entire multi-path parallel transmission process, and performing real-time performance monitoring on each transmission link carrying market transaction data fragments, main body basic data fragments, and external impact data fragments; real-time collection of transmission delay data for data fragments in each transmission link through the monitoring unit, where the transmission delay is the time difference between the data fragment being sent from the sending end to the receiving end; simultaneously collecting the number of data packets lost on each transmission link and calculating the packet loss rate indicator for each link; and organizing and recording the collected transmission delay data and packet loss rate indicators to form transmission performance monitoring data.

[0024] Step 1109: Based on transmission performance monitoring data, dynamically adjust the weight parameters of the transmission path and the size parameters of the data fragments to generate an updated routing table. Specifically, this includes: first, extracting the core indicators of each transmission link from the transmission performance monitoring data, including transmission delay, packet loss rate, bandwidth utilization, and transmission stability coefficient; validating each indicator and removing invalid monitoring data with abnormal fluctuations; analyzing the verified monitoring data one by one, comparing the transmission delay of each link with a preset delay threshold and the packet loss rate with a preset packet loss rate threshold; if the transmission delay of a transmission link exceeds the preset delay threshold or the packet loss rate exceeds the preset packet loss rate threshold, then the weight parameter of that transmission path in subsequent transmissions is reduced according to the degree of performance deviation, with a larger deviation resulting in a larger reduction in weight; if the transmission delay of a transmission link is less than 80% of the preset delay threshold and the packet loss rate is less than 50% of the preset packet loss rate threshold, meaning the transmission performance is better than the preset standard, then its weight parameter is increased according to the degree of performance superiority, with a larger increase in weight depending on the degree of superiority; combining the link carrying capacity reflected in the transmission performance monitoring data... First, the bandwidth redundancy and transmission fluctuation of each link are quantitatively assessed, and the links are divided into three carrying capacity levels: high, medium, and low. Based on the carrying capacity level, the data fragment size parameters are dynamically adjusted. For links with low carrying capacity and poor transmission performance, the corresponding data fragment size is appropriately reduced, with the reduction controlled between 20% and 40% of the original fragment size, to ensure compatibility with the link's transmission capacity. For links with high carrying capacity and good transmission performance, the corresponding data fragment size can be appropriately increased, with the increase controlled between 10% and 30% of the original fragment size, to improve transmission efficiency. Based on the adjusted transmission path weight parameters and data fragment size parameters, the old parameters in the original routing table are first cleared. Then, according to the correspondence between data fragment identifiers and path identifiers, the updated weight parameters and fragment size parameters are re-entered, and the parameter adjustment timestamps and adjustment basis are added simultaneously. The updated routing table is then checked for completeness and logical consistency to ensure that there are no conflicts between each data fragment and the adjusted path and fragment size. Finally, the updated routing table is generated.

[0025] Step 1110: Based on the updated routing table, perform retransmission and supplementary transmission control on data fragments marked as having transmission anomalies in the initial aggregated data stream to obtain a set of data fragments after optimized transmission. Specifically, this includes: first, verifying each data fragment in the initial aggregated data stream, including the uniqueness of the fragment identifier, the validity of the transmission status code, and the data integrity check code, removing invalid fragment records found during the verification process; second, classifying and filtering data fragments based on their built-in transmission status identifiers to accurately locate data fragments marked as having transmission anomalies, specifically defining three types of transmission anomalies: 1) fragments that have not been received after exceeding the preset reception timeout threshold; 2) incomplete fragments found to have missing fields or incorrect lengths after reception and verification; and 3) verification error fragments whose integrity check codes do not match the preset standard code; third, grouping and statistically analyzing the filtered abnormal fragments according to anomaly type and data category, recording the identifier, original transmission path, and anomaly occurrence timestamp for each group of abnormal fragments; and fourth, extracting the target path identifier, adjusted weight parameters, and fragment type corresponding to each abnormal fragment based on the updated routing table. Based on the size parameters and the real-time performance status of each transmission link, the optimal transmission path is determined for each abnormal fragment, prioritizing links with increased weight and superior transmission performance. A differentiated retransmission or supplementary transmission process for the abnormal data fragments is initiated. For fragments that are not received, they are repackaged according to the updated fragment size parameters and then retransmitted completely, while setting a maximum number of retransmission attempts to avoid infinite retransmissions. For fragments with incomplete reception, only the data packets corresponding to missing fields or missing segments are extracted and targeted supplementary transmission is performed to reduce redundant data transmission. During retransmission and supplementary transmission, the transmission status is monitored in real time. If a single retransmission or supplementary transmission fails, the system switches to the suboptimal adapted path and tries again until the maximum number of retransmissions is reached or the transmission is successful. After all abnormal data fragments have been retransmitted or supplemented, a secondary integrity check is performed on all transmitted fragments, and those that pass the check are selected. The verified fragments are then systematically integrated according to the correspondence between data category and fragment sequence number, and the optimized transmission status identifier and completion timestamp are added, ultimately resulting in a set of fragments with optimized transmission and complete information.

[0026] Step 1111 involves verifying and reassembling the optimized data fragment set to generate the original dataset. Specifically, this includes: performing integrity and accuracy checks on each data fragment in the optimized data fragment set; removing invalid data fragments that fail the checks; and retaining the data fragments that pass the checks. A preset fragment reassembly rule is retrieved. This rule is preset by combining the data fragment generation logic with the sorting criteria for fragment sequence numbers, the integration priority of data categories, and the matching standards for data association relationships. The sorting rule is set in ascending order of fragment sequence numbers. The integration priority is set according to the business importance of market transaction data, main body basic data, and external impact data. At the same time, the matching rules of the correlation fields between different types of data are preset to ensure the logical consistency of data. Based on this, the fragmentation and reorganization rules are formulated and preset. According to the fragmentation sequence number, data category information and data correlation relationship of each data fragment, the data fragments are systematically spliced ​​and integrated. Through the reorganization operation, the effective data corresponding to the market transaction data fragment, main body basic data fragment and external impact data fragment are integrated into one, and finally the original dataset containing market transaction data, main body basic data and external impact data is generated.

[0027] In this embodiment of the invention, multi-dimensional quantitative features of each data source are extracted to clearly grasp the core attributes of multi-source heterogeneous data, providing data support for the configuration of subsequent acquisition strategies. Based on data source features, accurate matching and calculation of acquisition strategy parameters are achieved, improving the adaptability and rationality of acquisition strategy parameters and ensuring the targeted nature of real-time acquisition strategies. Irrelevant data is filtered out to reduce the amount of data processed subsequently, and a unified data format reduces processing obstacles caused by format differences, improving the efficiency of data preprocessing and laying the foundation for subsequent transmission and processing. Dynamic data sharding enables refined data management, facilitating efficient parallel transmission, and synchronously acquiring real-time network status provides immediate reference for transmission path optimization. Combining data sharding characteristics with real-time network status to calculate path adaptability improves the matching accuracy between transmission paths and data shards, providing support for efficient transmission path selection. Dynamically allocating adaptable transmission paths and generating a routing schedule table realizes transmission... Rational resource planning ensures the adaptability of data fragments with different characteristics for transmission, providing clear guidance for parallel transmission; multi-path parallel transmission improves data transmission efficiency, accelerates data aggregation, shortens data acquisition cycles, and ensures the timeliness of data transmission; real-time monitoring of transmission link performance obtains key indicators, promptly grasps status changes during transmission, provides real-time data basis for subsequent transmission parameter adjustments, and ensures controllable transmission processes; dynamic adjustment of transmission parameters and updating of routing tables enhance the adaptability of transmission strategies to network changes, ensuring the stability and reliability of transmission performance; retransmission and supplementation of abnormal data fragments compensate for data loss during transmission, ensuring the integrity of transmitted data and improving the quality of transmitted data; verification ensures data accuracy, and reorganization and integration of multiple types of data form a complete original dataset, ensuring the standardization and integrity of the original dataset and providing high-quality basic data for subsequent data processing.

[0028] In the power trading platform market participant behavior prediction system based on big data analysis described in this embodiment of the invention, the optimization module 12 performs multi-dimensional quality assessment on the original dataset in terms of data integrity, consistency, and timeliness to obtain a quality assessment matrix; based on the quality assessment matrix, it performs feature extraction and pattern recognition to form a core quality feature vector set; it generates a strategy configuration vector based on the core quality feature vector set; and it uses the strategy configuration vector to dynamically adapt data cleaning rules, standardization thresholds, and feature extraction parameters to construct a standardized dataset with unified semantics and format, including: Step 1201: For each data unit in the original dataset, calculate the evaluation values ​​for data integrity, data consistency, and data timeliness dimensions to obtain the quality feature vector of each data unit; integrate the quality feature vectors of all data units to form a quality evaluation matrix, specifically including: traversing all data units in the original dataset to determine the field composition and data attributes corresponding to each data unit; for the data integrity dimension, calculating the proportion of missing fields in each data unit to the total number of fields in that data unit, and converting this proportion into the corresponding integrity evaluation value based on a preset integrity scoring standard; for the data consistency dimension, verifying whether the format of each data unit conforms to the preset data format specification, and verifying... Check whether the logical relationships between fields within a data unit meet preset logical rules, and determine the consistency assessment value based on the format compliance and logical rationality results; for the data timeliness dimension, obtain the generation timestamp and the current system time for each data unit, calculate the time difference between the two, and convert the time difference into a timeliness assessment value according to the preset timeliness scoring gradient; combine the integrity assessment value, consistency assessment value, and timeliness assessment value corresponding to each data unit in a fixed order to form a unique quality feature vector for each data unit; collect the quality feature vectors of all data units in the original dataset, integrate and consolidate them according to the row or column arrangement rules, and finally form a quality assessment matrix with data units as rows and quality assessment dimensions as columns.

[0029] Step 1202: Based on the quality assessment matrix, a clustering algorithm is used to group and classify data units according to the similarity between their score vectors in the dimensions of completeness, consistency, and timeliness, resulting in a set of data quality clusters. Specifically, this includes: loading the quality assessment matrix; firstly, standardizing and preprocessing the assessment values ​​within the matrix to uniformly map assessment values ​​of different dimensions to a numerical range of 0 to 1, eliminating the interference of differences in magnitude between dimensions on subsequent calculations; determining the data unit identifier corresponding to each row and the quality assessment dimension corresponding to each column in the matrix, establishing a one-to-one mapping relationship between data units and assessment values; and calling a preset clustering algorithm, first setting the range of the number of clusters based on the data size of the original dataset and historical clustering experience, typically taking a value of 3 to 8 clusters, while also referring to historical data quality... A reasonable similarity threshold is set for the quality assessment criteria, generally within the range of 0.7 to 0.8. The similarity calculation adopts a method based on Euclidean distance in vector space, and the specific operation is as follows: First, extract the quality feature vectors corresponding to any two data units from the quality assessment matrix. Extract the assessment values ​​of the two vectors in the three dimensions of completeness, consistency, and timeliness. Calculate the difference between the two assessment values ​​in each dimension and square the difference. Add the squared differences of the three dimensions and take the square root. The result is the Euclidean distance between the two quality feature vectors. Compare the calculated Euclidean distance with the preset similarity threshold. If the Euclidean distance is less than the preset similarity threshold, the two quality feature vectors are considered to be similar; otherwise, they are considered not to be similar.

[0030] Based on the similarity determination results, the iterative grouping process of the clustering algorithm is initiated: First, several quality feature vectors are randomly selected as initial cluster centers. The Euclidean distance between each quality feature vector and all initial cluster centers is calculated, and the vector is assigned to the temporary data quality cluster corresponding to the nearest initial cluster center. After the first round of grouping is completed, the mean of the evaluation values ​​of each dimension of all quality feature vectors in each temporary data quality cluster is calculated. The mean values ​​are combined to form new cluster centers. The Euclidean distance between each quality feature vector and the new cluster center is recalculated and the clusters are regrouped. The cluster center update and regrouping operation is repeated until the change in cluster centers in two adjacent iterations is less than the preset convergence threshold. The iteration is stopped and the final grouping result is determined. Quality feature vectors with matching similarity are assigned to the same data quality cluster, and those that do not match are assigned to different data quality clusters. After grouping all quality feature vectors, a unique cluster identifier is assigned to each data quality cluster. The identifier information includes the cluster number and the number of data units in the cluster. All identified data quality clusters are summarized to form a data quality cluster set.

[0031] Step 1203: For each data unit in the data quality cluster set, principal component analysis is used to calculate the covariance matrix of the feature space formed by the data unit in the cluster across the three quality dimensions. Principal component eigenvectors representing the main directions of its quality distribution are extracted through eigenvalue decomposition. Specifically, this includes: traversing the data quality cluster set and extracting the quality feature vectors of all data units within each data quality cluster; using a single data quality cluster as a processing unit, constructing a three-dimensional feature space based on all quality feature vectors within the cluster, where the three dimensions correspond to the integrity, consistency, and timeliness evaluation values, respectively; and calculating the... The mean of all quality feature vectors in the three-dimensional feature space is used to construct a covariance matrix based on the deviation between the mean and each quality feature vector. This covariance matrix is ​​used to characterize the linear correlation between the three quality dimensions. Eigenvalue decomposition is performed on the constructed covariance matrix to extract the eigenvalues ​​and corresponding eigenvectors. The eigenvectors are sorted according to the magnitude of the eigenvalues, and eigenvectors with eigenvalues ​​greater than a preset eigenvalue threshold are selected. These eigenvectors are the principal component feature vectors that can characterize the main direction of the quality distribution of the data unit in the cluster. Each principal component feature vector retains the corresponding cluster identifier.

[0032] Step 1204: Summarize the principal component feature vectors extracted from various data units to form a core quality feature vector set. This specifically includes: collecting the principal component feature vectors extracted from each data quality cluster in step 1203, verifying the cluster identifier corresponding to each principal component feature vector to ensure no omissions or duplications; performing a unified format conversion on all collected principal component feature vectors to ensure consistency in the number of dimensions and numerical precision of principal component feature vectors from different data quality clusters; sorting all principal component feature vectors according to the classification order of data quality clusters or the importance order of principal component feature vectors, naming the sorted set of principal component feature vectors the core quality feature vector set, and establishing a mapping table between the core quality feature vector set and each data quality cluster.

[0033] Step 1205: Based on the core quality feature vector set, a strategy configuration vector is obtained through preset mapping rules and weighted calculations. Each component in the strategy configuration vector corresponds to the threshold standard for data cleaning, the benchmark range used for standardization, and the weights of key fields of interest in feature extraction. Specifically, this includes: retrieving the core quality feature vector set and loading preset feature and configuration mapping rules. The preset method for these mapping rules is as follows: First, the representational meaning of each dimension component of the core quality feature vector is analyzed to determine the control requirements for data missing and anomalies corresponding to the integrity dimension component, the standardization requirements for data format and logic corresponding to the consistency dimension component, and the filtering requirements for data time validity corresponding to the time validity dimension component. Second, the core control objectives of the data processing strategy parameters are determined, and the specific parameter types for the data cleaning threshold standard, the standardization benchmark range, and the weights of key fields of feature extraction are analyzed. Finally, the historical experience of power trading data processing, core business needs, and data quality assessment are combined. The evaluation standard establishes the correlation between each dimension of feature components and their corresponding strategy parameters, determines the corresponding logic between the value range of different feature components and the parameter values, and finally solidifies the above correlation and corresponding logic into a preset feature-configuration mapping rule. This rule predefines the correspondence between each dimension of the core quality feature vector and the data processing strategy parameters. A preset importance weight is assigned to each principal component feature vector in the core quality feature vector set. Based on this weight, the weighted calculation of each dimension of each principal component feature vector is performed to obtain the weighted feature component values. According to the preset feature-configuration mapping rule, the weighted feature component values ​​are converted into corresponding initial values ​​of strategy parameters, including the initial values ​​of the data cleaning threshold standard, the initial values ​​of the standardized benchmark range, and the initial values ​​of the key field weights for feature extraction. The above three types of initial values ​​of strategy parameters are combined in a preset order to form a strategy configuration vector, where each component uniquely corresponds to a type of data processing strategy parameter.

[0034] Step 1206 involves converting the data cleaning components in the strategy configuration vector into specific missing value handling rules and outlier judgment thresholds to generate a data cleaning configuration set; converting the corresponding standardized components into normalized intervals and baseline values ​​for each data field to generate a standardized configuration set; and converting the corresponding feature extraction components into weight parameters in a feature selection algorithm to generate a feature extraction configuration set. Specifically, this includes: parsing the strategy configuration vector and extracting the components corresponding to the data cleaning threshold standard; converting these component values ​​into specific missing value handling rules, determining the imputation method used when the proportion of missing values ​​is below the threshold and the removal method used when the proportion of missing values ​​is above the threshold, and simultaneously converting the component values ​​into outlier... A threshold for missing value determination is established, and data exceeding this threshold is classified as outlier data. This missing value handling rule is integrated with the outlier determination threshold to generate a data cleaning configuration set. The components corresponding to the standardized baseline value range in the strategy configuration vector are extracted and converted into normalized intervals for each data field. The baseline reference values ​​used in the normalization process are determined, forming a standardized configuration set. The components corresponding to the weights of key feature extraction fields in the strategy configuration vector are extracted and mapped to weight parameters in the feature selection algorithm. The priority of different data fields in the feature extraction process is defined, generating a feature extraction configuration set. Corresponding configuration activation conditions are added to each configuration set to ensure that the configuration set can adaptively match data types.

[0035] Step 1207: By comprehensively applying the data cleaning configuration set, standardization configuration set, and feature extraction configuration set, corresponding cleaning, transformation, and feature extraction operations are performed on the original dataset to construct a standardized dataset with unified semantics and format. Specifically, this includes: calling the data cleaning configuration set to perform data cleaning operations on the original dataset; processing missing fields of each data unit according to the missing value handling rules in the configuration set; and removing or correcting abnormal data based on the outlier judgment threshold. After cleaning, the standardization configuration set is loaded, and field-level standardization transformation is performed on the cleaned dataset according to the normalization interval and benchmark value in the configuration set to unify the numerical magnitude and format of all data fields. After standardization, the feature extraction configuration set is activated, and core data fields are selected according to the weight parameters in the configuration set to extract the core feature information of each data unit. Semantic verification is performed on the data after cleaning, standardization, and feature extraction to ensure that the semantic representation of data from different sources and of different types is unified, and semantic mapping transformation is performed on data with inconsistent semantics. Finally, all processed data are integrated according to a preset data structure to construct a standardized dataset with unified semantics and format.

[0036] In this embodiment of the invention, a multi-dimensional quantitative assessment of data quality is achieved, covering the core dimensions of data integrity, consistency, and timeliness, ensuring that the quality status of each data unit is accurately characterized. An integrated quality assessment matrix provides structured and refined basic data support for subsequent quality feature analysis, ensuring the comprehensiveness and accuracy of the quality analysis. Clustering algorithms are used to group and classify data with similar quality features, simplifying data processing complexity and allowing data units with similar quality features to adopt a unified and appropriate processing logic. This improves the targeting of data processing, avoiding inefficiency or imbalance caused by mixing data with different quality features, and ensuring consistency in processing similar data. Principal component analysis is used to extract the core quality features of each cluster, eliminating redundant information in the quality features, reducing feature dimensionality while retaining key quality distribution information. Focusing on the core distribution direction of data quality provides a feature basis for subsequent strategy configuration. The principal component characteristics of various data units are summarized. The system constructs a unified set of core quality feature vectors, integrating core quality information from different data clusters; it forms a comprehensive and refined quality feature system, avoiding the limitations of features from a single data cluster; it establishes the relationship between core quality features and strategy configurations through preset mapping rules and weighted calculations, ensuring that each component of the strategy configuration vector is generated based on data quality features; it ensures the relevance and adaptability of strategy configurations, providing data support for parameter settings in data cleaning, standardization, and feature extraction; it transforms abstract strategy configuration vectors into concrete, executable data cleaning rules, standardization parameters, and feature extraction weights, determining the operational standards and thresholds for each processing step; it guarantees the standardization and repeatability of data processing, providing clear execution guidelines for cleaning, standardization, and feature extraction operations; it comprehensively applies three types of configuration sets to carry out data processing, achieving collaborative adaptation of cleaning, standardization, and feature extraction steps; and it constructs a standardized dataset with unified semantics and format, improving the structure and reusability of the data.

[0037] In the power trading platform market participant behavior prediction system based on big data analysis described in this embodiment of the invention, the prediction module 13 inputs a standardized dataset into a prediction model trained with historical conventional data to obtain trading strategy tendency distribution and collaborative behavior risk assessment indicators, and integrates them to obtain comprehensive prediction data, including: Step 1301 involves inputting the standardized dataset into the behavioral prediction sub-model of the prediction model to perform time-series feature analysis on the historical transaction sequences of different types of market participants in the power trading platform. These market participants include power generation companies, power sales companies, and large users. The aim is to obtain the distribution of trading strategy preferences for each participant in future trading periods. Specifically, this includes: loading the standardized dataset and selecting a subset of transaction data related to market participants in the power trading platform; classifying the data subset according to market participant type and extracting historical transaction data corresponding to power generation companies, power sales companies, and large users respectively; and organizing the classified historical transaction data into continuous historical transaction sequences according to the time dimension. The behavioral prediction sub-model used here is constructed by combining the strong time-series characteristics of power trading data, selecting a time-series prediction network as the basic network architecture, and building a sub-model structure including an input layer, a time-series feature extraction layer, a feature fusion layer, and an output layer. The input layer receives standardized historical transaction sequence data, and the time-series feature extraction layer captures the long-term and short-term dependencies in the transaction sequences. The system relies on relationships, with a feature fusion layer integrating temporal features from different dimensions and an output layer outputting predictions of trading strategy tendencies for each entity. It also embeds an attention mechanism adapted to the power trading scenario to enhance focus on key trading periods and core features. The training process involves collecting historical regular trading data from the power trading platform, classifying it by market entity type as the training sample set, performing data augmentation on the sample set, dividing it into training, validation, and test sets, and setting training parameters such as the number of training iterations, learning rate, and loss function for the sub-model. The training set is input into the sub-model for iterative training. After each iteration, the model performance is verified using the validation set, and the model parameters are dynamically adjusted based on the validation results until the model loss function converges and the prediction error on the validation set is below a preset threshold. Finally, the model performance is verified using the test set. After training, the model parameters are solidified, forming a usable behavioral prediction sub-model. The implementation process involves the sub-model loading the solidified training parameters, receiving the input historical trading sequence, and completing temporal feature analysis and strategy tendency deduction through built-in feature extraction and prediction inference logic.

[0038] The processed historical transaction sequences are input into the behavioral prediction sub-model. The behavioral prediction sub-model performs time series feature analysis on the historical transaction sequences of various types of market participants, specifically extracting time interval features, transaction frequency features, price fluctuation trend features, and transaction volume change features from the transaction sequences. Combined with the industry attributes and operating characteristics of each participant, the extracted time series features are screened and integrated. Based on the screened and integrated time series features, the sub-model's built-in time series prediction logic is used to deduce the pricing tendency, transaction volume tendency, and transaction timing selection tendency of each market participant in different future trading periods. The above tendency information is structured and organized according to participant type and trading period to form the distribution of trading strategy tendency of each participant in future trading periods.

[0039] Step 1302: The standardized dataset is simultaneously input into the game analysis sub-model in the prediction model. By constructing the interaction relationship topology between market participants, the bidding correlation and electricity synergy among market participants are quantitatively analyzed to obtain a risk assessment index for quantifying potential joint manipulation risk. Specifically, this includes: synchronously retrieving the standardized dataset and extracting related data such as transaction data, cooperation agreement data, and bidding synchronization data between market participants. The game analysis sub-model in the prediction model is constructed as follows: combining the interaction relationship characteristics between market participants, a graph neural network is selected as the basic network architecture to build a sub-model structure including an input layer, a topology construction layer, a correlation quantification layer, and a risk assessment layer. The input layer is used to receive standardized subject correlation data, the topology construction layer is used to transform the correlation data into a subject interaction relationship topology, the correlation quantification layer is used to calculate quantitative indicators such as bidding correlation and electricity synergy among subjects, and the risk assessment layer is used to output the quantitative results of joint manipulation risk. An adaptive adjustment module for subject correlation weights is also embedded. To improve adaptability to interactions with varying degrees of correlation strength, the training process involves collecting historical entity correlation data, historical joint manipulation risk event data, and corresponding handling results data from the power trading platform. This data is categorized by entity combination type and used as the training sample set. The sample set undergoes data cleaning and enhancement, and is divided into training, validation, and test sets. Training parameters such as the number of training iterations, learning rate, and loss function are set for the sub-model. The training set is input into the sub-model for iterative training. After each iteration, the validation set is used to verify the model's correlation quantification accuracy and risk assessment accuracy. Based on the validation results, the model parameters are dynamically adjusted until the model's loss function converges and the evaluation error on the validation set is below a preset threshold. Finally, the model performance is verified using the test set. After training, the model parameters are solidified, forming a usable game analysis sub-model. The implementation process involves the sub-model loading the solidified training parameters, receiving input entity correlation data, and using built-in topology construction logic, correlation quantification logic, and risk assessment logic to complete the topology construction of interaction relationships, the quantification of correlation characteristics, and the generation of risk indicators.

[0040] The extracted correlation data is input into the game analysis sub-model. The sub-model constructs an interaction relationship topology among market participants, using nodes to represent market participants and edge weights to represent the degree of correlation between participants. Based on the correlation data, the correlation strength value between each participant is calculated and assigned to the corresponding edge, forming a visualized interaction relationship topology. Based on this interaction relationship topology, the price correlation among market participants is quantitatively analyzed, calculating the price similarity and price change synchronization rate of different participants in the same or adjacent trading periods. At the same time, the electricity volume synergy is quantitatively analyzed, calculating the complementarity coefficient and adjustment synchronization coefficient of the trading volume between participants. Based on the quantitative results of price correlation and electricity volume synergy, a risk assessment system is constructed, setting correlation strength thresholds and synergy coefficient thresholds. Participant combinations exceeding the thresholds are marked, and the probability value of joint manipulation risk of the marked participant combinations is calculated. This probability value is correlated and integrated with the participant combination information to form a synergistic behavior risk assessment index that quantifies the potential joint manipulation risk.

[0041] Step 1303: Based on a preset fusion weight matrix, the strategy data of each subject in the trading strategy tendency distribution and the corresponding risk quantification values ​​in the collaborative behavior risk assessment index are weighted and fused using a dot product to generate preliminary fused data. Specifically, this includes: retrieving the preset fusion weight matrix. The preset method for this fusion weight matrix is ​​as follows: First, identify the core types of trading strategy data and risk assessment data, determine the various strategy data dimensions of different subjects in the trading strategy tendency distribution, and the risk quantification value dimensions of different subject combinations in the collaborative behavior risk assessment index; then, combine the regulatory needs of the power trading market, the core objectives of trading decisions, and the historical data fusion effects to determine the importance level of each data dimension, assigning higher weight coefficients to dimensions with higher importance levels and lower weight coefficients to dimensions with lower importance levels; subsequently, collect historical data fusion cases to verify the initially set weight coefficients, calculate the degree of fit between the data fusion results under different weight combinations and the actual market situation, based on... The matching degree dynamically adjusts the weight coefficients; finally, the adjusted weight coefficients of each dimension are arranged according to the preset matrix arrangement rules and solidified into a preset fusion weight matrix; this weight matrix pre-sets the weight coefficients of each dimension of data according to the importance of trading strategy data and risk assessment data, where the various strategy data of different subjects in the trading strategy tendency distribution and the risk quantification values ​​of different subject combinations in the collaborative behavior risk assessment indicators all correspond to specific weight coefficients in the matrix; the correlation between the strategy data of each subject in the trading strategy tendency distribution and the corresponding risk quantification values ​​in the collaborative behavior risk assessment indicators is determined, and the strategy data and risk quantification values ​​corresponding to the same subject or subject combination are determined; the corresponding data and quantification values ​​are used as vector elements respectively, and strategy data vectors and risk quantification value vectors are constructed in a preset order; based on the weight coefficients in the fusion weight matrix, a dot product weighted calculation is performed on the two vectors, and the calculation results are classified and integrated according to subject type, subject combination and trading time to generate preliminary fusion data.

[0042] Step 1304 involves normalizing and reducing the dimensionality of the preliminary fused data to generate comprehensive prediction data. Specifically, this includes: normalizing the preliminary fused data by using a preset normalization method to map fused data of different dimensions and magnitudes to a unified numerical range, eliminating the interference of data magnitude differences on subsequent applications; reducing the dimensionality of the normalized preliminary fused data by using a feature selection algorithm built into the sub-model to select key features with a core impact on predicting market participant behavior, eliminating redundant features and features with excessive correlation, thus reducing data dimensionality and processing complexity; and structurally integrating the dimensionality-reduced key feature data by classifying and archiving it according to market participant type, trading period, and risk level to generate comprehensive prediction data that is both complete and concise.

[0043] In this embodiment of the invention, differentiated time-series characteristic analysis is conducted for different types of market participants to accurately capture the time-series evolution patterns of each participant's historical transaction sequences; the transaction behaviors of multiple market participants are classified and sorted, providing targeted data support for subsequent strategy preference analysis; potential time correlation information of transaction behaviors is mined through time-series characteristic analysis, enriching the data source dimensions of strategy preference distribution; an interactive relationship topology is constructed to clearly present the relationship structure between market participants, providing intuitive relationship support for quantitative analysis; implicit collaborative behaviors between participants are transformed into measurable indicators by quantifying price correlation and electricity synergy; and the potential risks of joint manipulation are visualized. This method supplements the data dimensions for single-behavior prediction, improving the prediction information system; it establishes a correlation between strategy data and risk quantification values ​​based on a preset fusion weight matrix, ensuring the orderly nature of the fusion process; it integrates two types of core prediction information through dot product weighted fusion, achieving complementary advantages of data from different dimensions; it provides a weighted basis for the fusion process, enhancing the relevance and rationality of the initial fusion data; it normalizes the scale range of the initial fusion data, avoiding interference caused by differences in the scale of data from different dimensions; it uses feature dimensionality reduction to remove redundant information from the data, simplifying the data structure and improving data processing efficiency; and it generates refined comprehensive prediction data, providing a standardized and efficient processing foundation for subsequent data applications.

[0044] In the power trading platform market participant behavior prediction system based on big data analysis described in this embodiment of the invention, the aforementioned early warning module 14 obtains a market state evolution matrix based on comprehensive prediction data; integrates the market state evolution matrix with collaborative behavior risk assessment indicators to obtain a risk feature vector set; performs dynamic threshold judgment on the risk feature vector set to obtain a list of abnormal risk behaviors; generates a structured early warning prompt with embedded priority identifier based on the list of abnormal risk behaviors, and performs differentiated information push through an early warning information routing and distribution device, including: Step 1401 involves using the trading strategy preference distribution contained in the comprehensive forecast data as a simulation input parameter to drive market clearing simulation and obtain market supply and demand curves reflecting the power supply and demand balance in future time periods. Specifically, this includes: loading the comprehensive forecast data; analyzing and filtering the data to extract the trading strategy preference distribution data of various market participants, which covers the bidding preferences, transaction volume preferences, and trading timing preferences of power generation companies, electricity sales companies, and large users in different future trading periods; using the extracted trading strategy preference distribution data as the core input parameter for market clearing simulation, while simultaneously loading preset market clearing rules, historical market supply and demand benchmark data, and power commodity trading attribute parameters to construct a market clearing simulation model; starting the simulation model to perform multi-time period simulations, simulating the game process between market supply and demand under different trading strategy combinations, and recording power supply and demand data for each future trading period; and drawing supply and demand relationship curves based on the supply and demand data for each time period to form market supply and demand curves that intuitively reflect the power supply and demand balance and supply and demand gap trends in future time periods.

[0045] Step 1402: Based on the market supply and demand curve and combined with the power network topology and transmission constraints, perform iterative solutions to obtain the future marginal electricity price and network congestion distribution of each node. Specifically, this includes: retrieving the market supply and demand curve and loading the power system network topology data, which includes transmission line distribution, node connection relationships, and key equipment parameters such as transformers; extracting the power network transmission constraints, specifically including transmission line transmission capacity thresholds, node voltage stability thresholds, allowable frequency fluctuation ranges, and power supply reliability constraints; and inputting the market supply and demand curve, network topology data, and transmission constraints into a preset... The power grid operation analysis model sets a convergence accuracy threshold and a maximum number of iterations for iterative solutions. It iteratively solves the problem by first calculating the initial electricity price and power flow direction for each node based on the initial supply and demand state. Then, it verifies the calculation results by combining transmission constraints. If constraints are violated, the supply and demand matching scheme is adjusted and recalculated until the calculation results satisfy all constraints and reach convergence accuracy. After iterative convergence, it outputs the marginal electricity price data for different power grid nodes in future trading periods. Simultaneously, it analyzes the matching relationship between power flow direction and transmission capacity, determines the congestion status and degree of each transmission line, and forms the marginal electricity price and network congestion distribution data for each node in the future.

[0046] Step 1403: Integrate the market supply and demand curves, nodal marginal electricity prices, and network congestion distribution, and parse the unit output status from them. Organize the data in a structured manner according to the time series to obtain a market state evolution matrix. Specifically, this includes: collecting market supply and demand curves, nodal marginal electricity prices, and network congestion distribution data; standardizing the format of these data to ensure consistency in the statistical caliber of the time dimension; based on the standardized market supply and demand curves, nodal marginal electricity prices, and network congestion distribution data, combined with preset unit output analysis rules, and by associating the marginal electricity price fluctuations with the supply and demand balance relationship and network congestion constraints, inferring the planned output values ​​of each generating unit at different future time periods, and parsing the unit output status data; structuring the market supply and demand curve nodal marginal electricity price network congestion distribution and unit output status data according to the time series, constructing a matrix with the time dimension as rows and each state parameter as columns. Each element in the matrix corresponds to a specific market state parameter value at a specific time point, ultimately forming a market state evolution matrix that can completely characterize the dynamic changes in the future market operation status.

[0047] Step 1404 involves using the collaborative behavior risk assessment indicators as influencing parameters and weighting them into the corresponding market entity association dimension of the market state evolution matrix to generate a fused risk-enhanced state matrix. Specifically, this includes: retrieving the collaborative behavior risk assessment indicators, which contain the quantitative values ​​of price correlation, electricity coordination, and joint manipulation risk probability for each market entity combination; establishing a mapping relationship between the collaborative behavior risk assessment indicators and the market entity association dimension in the market state evolution matrix based on market entity identity identifiers, and determining the specific position of each risk assessment indicator's corresponding entity combination in the matrix; pre-setting weight coefficients for the risk indicators, which are determined based on the degree of impact of the risk indicators on market stability and regulatory priority, with the joint manipulation risk probability value having a higher weight than the price correlation and electricity coordination quantitative values; weighting each collaborative behavior risk assessment indicator according to the pre-set weight coefficients to obtain a weighted risk value; and injecting the weighted risk value into the corresponding market entity association dimension column in the market state evolution matrix, forming an association mapping with the original state parameters of that dimension to generate a fused risk-enhanced state matrix that combines market operation status information and risk warning information.

[0048] Step 1405: Extract pattern features reflecting the correlation between systemic fluctuations and local anomalies from the fused risk-enhanced state matrix to obtain a risk feature vector set. Specifically, this includes: preprocessing the fused risk-enhanced state matrix by removing isolated abnormal data and filling in missing values ​​to ensure data integrity and continuity; using a preset pattern feature extraction algorithm to focus on key information in the matrix that reflects the correlation between market systemic fluctuations and local anomalies, specifically extracting pattern features including abrupt changes in market supply and demand curves, abnormal fluctuations in marginal electricity prices, diffusion of network congestion, coordinated changes in unit output, and clustering of risk values; reducing and integrating the extracted pattern features, removing redundant features and features with excessive correlation, and combining the retained core features into feature vectors in a preset order; summarizing all feature vectors to form a risk feature vector set that can characterize the market risk state, and simultaneously establishing an association index between the feature vectors and market participants in the corresponding time period.

[0049] Step 1406: Train a high-dimensional hyperspherical model based on historical risk data to determine the baseline distribution center and dynamic radius of the risk feature vector set in the high-dimensional space. Specifically, this includes: collecting market operation data, risk feature data, and risk disposal result data corresponding to historical risk events from the power trading platform to construct a historical risk dataset; cleaning and filtering the historical risk dataset, retaining valid data and standardizing it, and dividing the processed dataset into a training set and a validation set. The construction, training, and implementation process of the high-dimensional hyperspherical model used here is as follows: The construction process involves building a model structure with the hyperspherical decision boundary as the core, taking into account the high-dimensional characteristics of the risk feature vectors. The model's input layer, output layer, and core computation layer are determined. The input layer receives the standardized risk feature vectors, the core computation layer constructs the high-dimensional space and calculates the spatial distance between the feature vectors and the baseline center, and the output layer outputs the distribution assignment results of the feature vectors. A dimension adaptive adjustment unit is configured to adapt to risk feature vectors of different dimensions. The training process involves using historical risk feature vectors from the training set... Using the core training samples, training parameters such as the number of iterations and convergence threshold are set for the model. After inputting the training samples into the model, the center coordinates and radius parameters of the hypersphere are adjusted through iterative optimization, enabling the model to gradually fit the distribution pattern of historical risk feature vectors in high-dimensional space. The process is as follows: the model loads the parameters after training and receives the input risk feature vectors. Through built-in spatial distance calculation logic and distribution attribution judgment logic, the distribution position of the feature vectors is determined. The high-dimensional hypersphere model is trained based on the training set, using historical risk feature vectors as model input. The model parameters are adjusted through iterative optimization, enabling the model to accurately fit the distribution pattern of historical risk feature vectors in high-dimensional space. The performance of the trained model is verified using the validation set, and the model parameters are adjusted to improve the fitting accuracy. After the model training is completed, the baseline distribution center of the risk feature vectors in high-dimensional space is determined. This center is the mean vector of historical normal risk feature vectors. At the same time, the distance distribution of historical risk feature vectors to the baseline center is calculated, and the dynamic radius that can cover the distribution range of normal risk feature vectors is determined.

[0050] Step 1407 involves flexibly adjusting the dynamic radius based on real-time market volatility to obtain adaptive judgment boundaries for each feature dimension. This includes: acquiring real-time transaction data, price fluctuation data, and supply-demand change data of the current electricity market through a real-time data acquisition unit; calculating real-time market volatility based on this data, specifically by statistically analyzing the standard deviation of price fluctuation amplitude and supply-demand gap change rate within a unit of time; pre-setting the mapping relationship between volatility and radius adjustment coefficients; increasing the radius adjustment coefficient to expand the dynamic radius when the real-time market volatility is higher than the preset benchmark value; decreasing the radius adjustment coefficient to shrink the dynamic radius when the real-time market volatility is lower than the preset benchmark value; flexibly adjusting the dynamic radius based on the adjustment coefficient corresponding to the real-time market volatility; and calculating the adjusted dynamic radius threshold for each feature dimension of the risk feature vector set to form adaptive judgment boundaries for each feature dimension, ensuring that the judgment boundaries can adapt to the volatility state of the real-time market.

[0051] Step 1408: Calculate the distance from each vector in the risk feature vector set to the baseline distribution center; compare the distance with the adaptive judgment boundary, and filter out vector combinations whose distance exceeds the boundary and characterize multi-dimensional collaborative anomalies to obtain a list of abnormal risk behaviors. Specifically, this includes: using a preset Euclidean distance calculation method, specifically extracting the component values ​​of each feature vector and the baseline distribution center in each dimension of the risk feature vector set, calculating the difference between the two component values ​​dimension by dimension and squaring the difference, summing the squared differences of all dimensions and taking the square root to obtain the spatial distance from each feature vector to the baseline distribution center; comparing the distance value of each feature vector in each dimension with the adaptive judgment boundary of the corresponding feature dimension according to the dimensional order of the feature vectors, recording the comparison results of each dimension, i.e., whether the distance value exceeds the corresponding judgment boundary; filtering out the existing... Feature vectors whose distance exceeds the corresponding adaptive judgment boundary in at least one dimension are included in the candidate anomaly vector set. Multi-dimensional correlation analysis is performed on the vectors in the candidate anomaly vector set to specifically verify the temporal correlation, subject correlation, and feature correlation of anomalies in different dimensions. It is determined whether there are coordinated anomalies in multiple dimensions that simultaneously point to the same market subject, the same trading period, or the same risk type, and the vector combination representing the multi-dimensional coordinated anomaly is determined. Based on the correlation index corresponding to the anomaly vector combination, the corresponding market subject information, specific trading period, risk feature type, and risk quantification level are traced. This information is integrated according to a preset structured format to determine that the information entries include the name of the anomaly subject, the type of anomaly, the time of anomaly occurrence, the description of risk features, and the level of risk, and finally a list of anomaly risk behaviors containing the above complete information is formed.

[0052] Step 1409: Based on the list of abnormal risk behaviors, generate a structured early warning prompt with embedded priority identifiers. Specifically, according to the subject type, risk characteristic value, and potential impact range associated with each abnormal behavior in the list of abnormal risk behaviors, assign risk level identifiers to them through preset priority calculation rules, and organize the generation of structured early warning information containing risk subject, risk type, risk level, and recommended measures. Specifically, this includes: loading the list of abnormal risk behaviors and retrieving the preset priority calculation rules. The preset method for these priority calculation rules is to first sort out the core dimensions affecting the priority of abnormal risks, and determine the subject type, risk characteristic value, and potential impact range as the three core assessment dimensions; then, combined with the key points of power market supervision, power grid safety operation requirements, and historical risk handling experience, determine the weight ratio of each core dimension. Among them, the weight of key subject types such as power generation enterprises is higher than that of ordinary subjects, the weight of serious risk characteristic values ​​such as joint manipulation is higher than that of general risk characteristic values, and the weight of the potential impact range affecting the power supply of the entire network is higher than that of the local impact range; subsequently, set quantitative standards for the subdivision level of each dimension. The subject type dimension is divided into three levels—core, important, and ordinary—according to the subject's influence on the market, and assigns weights to each level. The risk characteristic value dimension is divided into three levels—extremely high, relatively high, and moderate—based on the severity of the risk, and each level is assigned a corresponding quantitative score. The potential impact scope dimension is divided into three levels—network-wide, regional, and local—based on the scope of impact, and each level is assigned a corresponding quantitative score. Finally, the calculation logic for the comprehensive priority score is established, determining that the comprehensive score is obtained by weighted summation of the quantitative scores of each dimension and their corresponding weights. This calculation logic, along with the weight proportions and quantitative standards of each dimension, forms a preset priority calculation rule. This rule defines the quantitative standards for the subject type weight, risk characteristic value weight, and potential impact scope weight. Based on the comprehensive priority score of each abnormal behavior in the rule calculation list, the abnormal behavior is divided into three risk levels—high, medium, and low—based on the score, and a corresponding risk level identifier is assigned to each abnormal behavior. Warning prompts are generated according to a preset structured warning information template. The template includes fields for risk subject, risk type, risk level, risk period, risk characteristic description, and suggested measures. The suggested measures field matches a preset response and handling plan based on the abnormal risk type. The warning information for all abnormal behaviors is then organized to form a set of structured warning prompts with embedded priority identifiers.

[0053] Step 1410: The structured early warning prompts are input into the early warning information routing and distribution device. The device, based on the risk level identifier embedded in the early warning information and a preset push strategy, pushes different levels of early warning information to the corresponding market entity terminals and regulatory terminals in a differentiated manner. Specifically, this includes: inputting the set of structured early warning prompts into the early warning information routing and distribution device; loading the preset push strategy into the device, which defines the push channels, push timelines, and receiving entities for early warning information of different risk levels. High-level early warning information is pushed in real-time, simultaneously to the corresponding market entity terminals, the core regulatory terminals of regulatory departments, and the operation and maintenance terminals; medium- and low-level early warning information is aggregated and pushed to the corresponding receiving terminals according to preset time periods; the device parses the priority identifier and associated entity information of each structured early warning prompt, and determines the corresponding push path and push method based on the push strategy; initiating a differentiated push process, pushing high-level early warning information immediately and providing feedback on the push status, and aggregating and organizing medium- and low-level early warning information before pushing it; recording the push time, receiving terminal, and receiving status of all early warning information to form a push log for retention, ensuring that early warning information can accurately and efficiently reach the corresponding entities.

[0054] In this embodiment of the invention, the distribution of trading strategy preferences is transformed into simulation input parameters. A correlation between future trading strategies and supply-demand balance is established through market clearing simulation. The generated market supply-demand curves intuitively represent the power supply and demand status at various future time periods, providing concrete data support for subsequent market state analysis and effectively transforming predictive strategy data into basic market state data. The market supply-demand curves are integrated with the core constraints of the power network, and a deep correlation between supply-demand status and network operation status is achieved through iterative solutions. Precise output of node marginal electricity prices and network congestion distribution supplements key quantitative indicators of market operation, enhancing the comprehensiveness of market state analysis and making the data processing results more closely aligned with the actual operation scenarios of the power network. Multiple types of market operation data are structurally integrated. The system synchronously analyzes core status information of unit output; organizes market status evolution matrix according to time series to achieve temporal and systematic representation of market status, facilitating subsequent tracking of dynamic changes in market status and providing an orderly and comprehensive basic data system for risk analysis; organically integrates collaborative behavior risk assessment indicators with the market status evolution matrix through weighted injection; the generated risk-enhanced status matrix combines market status and risk dimension information, enriching the data representation dimensions of the matrix and realizing the correlation between the two types of core data, laying the foundation for subsequent risk feature extraction; focusing on the correlation characteristics between systemic fluctuations and local anomalies in the risk-enhanced status matrix, the system achieves accurate screening of core risk information through pattern feature extraction; and the generated risk feature vector set eliminates redundancy. This approach simplifies the data structure, enhances the relevance of risk analysis, and provides refined feature data support for subsequent risk assessment. It trains models based on historical risk data to construct a high-dimensional benchmark distribution framework for risk feature vectors. By defining the benchmark distribution center and dynamic radius, it provides a quantifiable historical reference standard for risk assessment, ensuring solid historical data support. The dynamic radius is flexibly adjusted based on real-time market volatility, allowing the judgment boundary to adaptively match real-time market changes. This avoids the limitations of fixed judgment standards, improves the adaptability of risk judgment boundaries to the current market state, and ensures the real-time nature and flexibility of risk assessment. Distance calculations enable quantitative comparison between risk feature vectors and the benchmark distribution, identifying anomalies exceeding the adaptive boundary. Vector-based filtering of abnormal risk behavior lists focuses on multi-dimensional collaborative anomalies, enabling precise positioning of abnormal risks and improving the accuracy and targeting of risk identification. Based on multi-dimensional anomaly information, risk level identifiers are assigned according to preset rules, structurally organizing the core elements of early warning information. The generated embedded priority identifier early warning prompts are standardized, defining the risk subject, type, level, and response suggestions, improving the readability and usability of early warning information and providing clear guidance for subsequent handling. Relying on an early warning information routing and distribution system, differentiated pushes are achieved according to risk level and preset strategies. High-priority early warning information is ensured to reach the corresponding subject first, optimizing the efficiency of early warning information transmission, ensuring accurate matching of early warning needs at different levels, and improving the effectiveness of early warning information application.

[0055] like Figure 2 As shown, a control method for a power trading platform market participant behavior prediction system based on big data analysis is disclosed. The control method includes: Based on the multi-source heterogeneous data feature analysis of power trading platforms, third-party data service providers and enterprise internal management systems, a real-time acquisition strategy is configured and executed to filter and parse the data to obtain a preliminary processed data stream. Parallel transmission aggregation and transmission path optimization control are implemented on the preliminary processed data stream to obtain a raw dataset containing market transaction data, main basic data and external impact data. A multidimensional quality assessment is performed on the original dataset in terms of data integrity, consistency, and timeliness to obtain a quality assessment matrix. Feature extraction and pattern recognition are performed based on the quality assessment matrix to form a core quality feature vector set. A strategy configuration vector is generated based on the core quality feature vector set. The strategy configuration vector is used to dynamically adapt data cleaning rules, standardization thresholds, and feature extraction parameters to construct a standardized dataset with unified semantics and format. The standardized dataset is input into a prediction model trained on historical regular data to obtain the distribution of trading strategy tendencies and risk assessment indicators of collaborative behavior, and then integrated to obtain comprehensive prediction data. Based on comprehensive forecast data, a market state evolution matrix is ​​obtained; by integrating the market state evolution matrix with collaborative behavior risk assessment indicators, a risk feature vector set is obtained; dynamic threshold judgment is performed on the risk feature vector set to obtain a list of abnormal risk behaviors; a structured early warning prompt with embedded priority identifier is generated based on the list of abnormal risk behaviors, and differentiated information is pushed through an early warning information routing and distribution device.

[0056] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the system as described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0057] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the system as described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.

[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A system for predicting the behavior of market participants in a power trading platform based on big data analysis, characterized in that, include: The data acquisition module is used to configure and execute real-time acquisition strategies to filter and parse data based on the multi-source heterogeneous data characteristics analysis of power trading platforms, third-party data service providers and enterprise internal management systems, and to obtain a preliminary processed data stream. Parallel transmission aggregation and transmission path optimization control are implemented on the initially processed data stream to obtain the raw dataset containing market transaction data, main basic data and external impact data; The optimization module performs multidimensional quality assessments on the original dataset in terms of data integrity, consistency, and timeliness, resulting in a quality assessment matrix. Based on the quality assessment matrix, feature extraction and pattern recognition are performed to form a core quality feature vector set. Generate strategy configuration vectors based on the core quality feature vector set; use the strategy configuration vectors to dynamically adapt data cleaning rules, standardized thresholds and feature extraction parameters to construct a standardized dataset with unified semantics and format; The prediction module is used to input standardized datasets into a prediction model trained on historical regular data to obtain the distribution of trading strategy tendencies and risk assessment indicators of collaborative behavior, and to integrate them to obtain comprehensive prediction data. The early warning module is used to obtain a market state evolution matrix based on comprehensive forecast data; and to integrate the market state evolution matrix with collaborative behavior risk assessment indicators to obtain a risk feature vector set. Perform dynamic threshold judgment on the risk feature vector set to obtain a list of abnormal risk behaviors; Structured early warning prompts with embedded priority identifiers are generated based on the list of abnormal risk behaviors, and differentiated information is pushed through the early warning information routing and distribution device.

2. The power trading platform market participant behavior prediction system based on big data analysis according to claim 1, characterized in that, The acquisition module includes: Data source feature analysis is performed on multi-source heterogeneous data from power trading platforms, third-party data service providers and enterprise internal management systems. Data source feature vectors are extracted from each data source. The data source feature vectors include at least quantitative values ​​in the dimensions of data type, data update frequency and data scale. Based on the feature vector of the data source, the data sampling rate, key field filtering conditions, and parsing template are calculated by querying the mapping table, retrieving the rule base, and matching the parsing template, respectively; based on the data sampling rate, key field filtering conditions, and parsing template, the real-time acquisition strategy parameter set is obtained. Based on the real-time acquisition strategy parameter set, real-time conditional filtering and format parsing are performed on the raw data streams of each data source to filter out irrelevant data and unify the data format, resulting in a preliminarily processed data stream.

3. The power trading platform market participant behavior prediction system based on big data analysis according to claim 2, characterized in that, The acquisition module also includes: The preliminary processed data stream is dynamically segmented to obtain a data segment set consisting of market transaction data, main entity basic data, and external influence data; real-time network status information is obtained. Based on the data size of each fragment in the data fragment set and the preset fragment priority, combined with the bandwidth and latency indicators of each available path in the real-time network status, a preset path evaluation function is used to calculate the path adaptability metric for each data fragment. Based on the path adaptability metric, each data fragment is dynamically allocated an appropriate transmission path according to the load balancing strategy, and a routing table containing path identifiers and fragment sequence numbers is generated. Based on the routing schedule table, multi-path parallel transmission and data aggregation are performed to obtain a preliminary aggregated data stream; During the execution of multi-path parallel transmission, the performance of each transmission link carrying market transaction data, main basic data and external impact data fragments is monitored in real time to obtain transmission delay and packet loss rate indicators and obtain transmission performance monitoring data. Based on transmission performance monitoring data, the weight parameters of the transmission path and the size parameters of the data fragments are dynamically adjusted to generate an updated routing table. Based on the updated routing table, retransmission and supplementary transmission control is performed on the data fragments marked as transmission anomalies in the initial aggregated data stream to obtain a set of data fragments after optimized transmission. The optimized data fragment set is verified and reassembled to generate the original dataset.

4. The power trading platform market participant behavior prediction system based on big data analysis according to claim 3, characterized in that, The optimization module includes: For each data unit in the original dataset, the evaluation values ​​of its data integrity, data consistency, and data timeliness are calculated to obtain the quality feature vector of each data unit; the quality feature vectors of all data units are integrated to form a quality evaluation matrix. Based on the quality assessment matrix, a clustering algorithm is used to group and classify data units according to the similarity between their score vectors in the dimensions of completeness, consistency, and timeliness, thereby obtaining a set of data quality clusters. For each data unit in the data quality cluster set, principal component analysis is used to calculate the covariance matrix of the feature space formed by the data unit in the three quality dimensions, and principal component eigenvectors representing the main directions of its quality distribution are extracted by eigenvalue decomposition. The principal component feature vectors extracted from various data units are summarized to form the core quality feature vector set.

5. The power trading platform market participant behavior prediction system based on big data analysis according to claim 4, characterized in that, The optimization module further includes: Based on the core quality feature vector set, a strategy configuration vector is obtained through preset mapping rules and weighted calculation. Each component in the strategy configuration vector corresponds to the threshold standard for data cleaning, the benchmark value range used for standardization, and the weight of the key fields of interest for feature extraction. The data cleaning components in the strategy configuration vector are converted into specific rules for handling missing data values ​​and thresholds for judging outliers, generating a data cleaning configuration set; the standardized components are converted into normalized intervals and baseline values ​​for each data field, generating a normalization configuration set; and the feature extraction components are converted into weight parameters in the feature selection algorithm, generating a feature extraction configuration set. By comprehensively applying the data cleaning configuration set, standardization configuration set, and feature extraction configuration set, corresponding cleaning, transformation, and feature extraction operations are performed on the original dataset to construct a standardized dataset with unified semantics and format.

6. The power trading platform market participant behavior prediction system based on big data analysis according to claim 5, characterized in that, The prediction module includes: The standardized dataset is input into the behavioral prediction sub-model in the prediction model to perform time series feature analysis on the historical transaction sequences of different types of market participants in the power trading platform. The market participants include power generation companies, power sales companies and large users, so as to obtain the transaction strategy tendency distribution of each participant in future trading periods. The standardized dataset is simultaneously input into the game analysis sub-model in the prediction model. By constructing the topology of the interaction relationship between market participants, the bidding correlation and electricity coordination among market participants are quantitatively analyzed to obtain a risk assessment index for the coordinated behavior that quantifies the potential risk of joint manipulation. Based on a preset fusion weight matrix, the strategy data of each subject in the trading strategy tendency distribution and the corresponding risk quantification value in the collaborative behavior risk assessment index are fused by dot product weighting to generate preliminary fusion data. The preliminary fused data is normalized and its features are reduced to generate comprehensive prediction data.

7. The power trading platform market participant behavior prediction system based on big data analysis according to claim 6, characterized in that, The early warning module includes: The trading strategy preference distribution contained in the comprehensive forecast data is used as the simulation input parameter to drive the market clearing simulation and obtain the market supply and demand curves that reflect the power supply and demand balance in each future period. Based on the market supply and demand curves, and combined with the power network topology and transmission constraints, an iterative solution is performed to obtain the marginal electricity price and network congestion distribution of each node in the future. By integrating the market supply and demand curves, nodal marginal electricity prices, and network congestion distribution, and analyzing the unit output status, the data are structured according to time series to obtain the market state evolution matrix. The risk assessment indicators of collaborative behavior are used as influencing parameters and weighted and injected into the corresponding market entity association dimension of the market state evolution matrix to generate a fused risk-enhanced state matrix. From the fused risk-enhanced state matrix, pattern features reflecting the correlation between systemic fluctuations and local anomalies are extracted to obtain a risk feature vector set.

8. The power trading platform market participant behavior prediction system based on big data analysis according to claim 7, characterized in that, The early warning module also includes: A high-dimensional hyperspherical model is trained based on historical risk data to determine the baseline distribution center and dynamic radius of the risk feature vector set in high-dimensional space. By combining real-time market volatility with flexible adjustment of the dynamic radius, adaptive judgment boundaries for each feature dimension are obtained. Calculate the distance from each vector in the risk feature vector set to the center of the baseline distribution; compare the distance with the adaptive judgment boundary, and filter out vector combinations that exceed the boundary and represent multi-dimensional collaborative anomalies to obtain a list of abnormal risk behaviors; Based on the list of abnormal risk behaviors, a structured early warning prompt with embedded priority identifiers is generated. Specifically, according to the subject type, risk characteristic value and potential impact range associated with each abnormal behavior in the list of abnormal risk behaviors, a risk level identifier is assigned to it through a preset priority calculation rule, and a structured early warning information containing the risk subject, risk type, risk level and suggested measures is generated. The structured early warning prompts are input into the early warning information routing and distribution device. The early warning information routing and distribution device pushes early warning information of different levels to the corresponding market entity terminals and regulatory terminals in a differentiated manner according to the risk level identifier embedded in the early warning information and the preset push strategy.

9. A control method for a power trading platform market participant behavior prediction system based on big data analysis, characterized in that, Applied to the system as described in any one of claims 1 to 8, the method comprises: Based on the multi-source heterogeneous data feature analysis of power trading platforms, third-party data service providers and enterprise internal management systems, a real-time acquisition strategy is configured and executed to filter and parse the data to obtain a preliminary processed data stream. Parallel transmission aggregation and transmission path optimization control are implemented on the preliminary processed data stream to obtain a raw dataset containing market transaction data, main basic data and external impact data. A multidimensional quality assessment is performed on the original dataset in terms of data integrity, consistency, and timeliness to obtain a quality assessment matrix. Feature extraction and pattern recognition are performed based on the quality assessment matrix to form a core quality feature vector set. A strategy configuration vector is generated based on the core quality feature vector set. The strategy configuration vector is used to dynamically adapt data cleaning rules, standardization thresholds, and feature extraction parameters to construct a standardized dataset with unified semantics and format. The standardized dataset is input into a prediction model trained on historical regular data to obtain the distribution of trading strategy tendencies and risk assessment indicators of collaborative behavior, and then integrated to obtain comprehensive prediction data. Based on comprehensive forecast data, a market state evolution matrix is ​​obtained; by integrating the market state evolution matrix with collaborative behavior risk assessment indicators, a risk feature vector set is obtained; dynamic threshold judgment is performed on the risk feature vector set to obtain a list of abnormal risk behaviors; a structured early warning prompt with embedded priority identifier is generated based on the list of abnormal risk behaviors, and differentiated information is pushed through an early warning information routing and distribution device.

10. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to perform the method as described in any one of claims 1 to 8.