Railway power supply equipment online data processing system based on Internet of Things coding
Through the Internet of Things coding and dynamic adaptive reinforcement learning model, the adaptability problem of the online data processing system of railway power supply equipment in complex environments is solved, real-time monitoring and stability improvement of the equipment operation status is achieved.
Patent Information
- Application Number
- CN202510354509.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
When the existing online data processing system of railway power supply equipment faces complex load changes and environmental interference, the model is poorly adaptable, lacks flexibility and targeted, resulting in unstable training process.
The Internet of Things encoding module is used to generate enhanced coding labels with non-domain correlation characteristics, combine the hierarchical structure model to filter the optimal features, and build a dynamic adaptive reinforcement learning model to adjust the strategies in real time to adapt to changes in the operating state of the device.
It improves the pertinence and efficiency of data processing, enhances the adaptability to dynamic environments, and ensures the stable operation of railway power supply equipment and fault prediction capabilities.
Smart Images

Figure CN120296496A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to an online data processing system for railway power supply equipment based on Internet of Things coding. Background Art
[0002] With the continuous growth and high-speed development of railway transportation volume, power supply equipment faces more complex load changes (such as the start and stop of multiple unit trains, line switching) and environmental interferences (such as sudden temperature changes, electromagnetic interference), and railway power supply equipment involves multi-type sensor data such as current, machinery, and temperature. Traditional systems lack a unified data coding and fusion mechanism.
[0003] Currently, in existing online data processing systems for railway power supply equipment, most data analysis models are established based on historical data. When the operating environment, load conditions, etc. of the equipment change, the adaptability of the models is poor. For example, with the increase in railway transportation volume, the load of power supply equipment will change. When using traditional reinforcement learning models to process dynamic environments, it is difficult to quickly adapt to environmental changes, the strategy adjustment is not flexible enough, and there is a lack of in-depth customization and optimization for specific application scenarios. When applied to specific scenarios such as railway power supply equipment, a large amount of adjustment and adaptation work is required, which will lead to unstable situations in the training process. Therefore, an online data processing system for railway power supply equipment based on Internet of Things coding is proposed here. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art and to achieve the above object, the present invention proposes the following technical solutions:
[0005] An online data processing system for railway power supply equipment based on Internet of Things coding, comprising:
[0006] An Internet of Things coding module: collecting sensor data and obtaining corresponding Internet of Things coding data, coupling the pseudo-random sequence data based on the Internet of Things coding data with the sensor data to generate an enhanced coding label with non-local correlation characteristics;
[0007] An optimal feature screening module: extracting coding features for the enhanced coding label, evaluating the importance index of data types through a hierarchical structure model according to the coding features, and obtaining optimal data features according to the importance index;
[0008] A dynamic adaptation optimization module: obtaining the behavioral feature sequence of the optimal data features, and constructing a dynamic adaptive reinforcement learning model according to the behavioral feature sequence;
[0009] A data result output module: outputting an action mode based on the dynamic adaptive reinforcement learning model, the action mode including a normal mode and an abnormal mode, and performing an intervention operation on the data in the abnormal mode.
[0010] The process of obtaining the Internet of Things (IoT) encoded data: The operating parameters of the power supply device are sampled in real time by sensors at a sampling frequency f and binary encoded.
[0011] A multi-source power supply data sequence V = {v1, v2, v3} is obtained, where v1 represents the data collected by the current sensor, v2 represents the data collected by the mechanical sensor, and v3 represents the data collected by the temperature sensor.
[0012] The process of the binary encoding includes multiplying the normalized sensor stream data v i by the amplification factor M and converting it into an 8-bit binary number.
[0013] The process of obtaining the enhanced encoding tag is as follows:
[0014] A pseudo-random sequence of the IoT encoded data is generated using a pseudo-random number generator.
[0015] The binary form R of the pseudo-random sequence is obtained.
[0016] The binary sequence S after encoding the sensor data is obtained.
[0017] The enhanced encoding tag is obtained by coupling the binary sequence S with the binary form R of the pseudo-random sequence.
[0018] The process of coupling the binary sequence S with the binary form R of the pseudo-random sequence is as follows:
[0019] The binary sequence S after encoding the sensor data and the pseudo-random sequence R are divided into subsequences according to the length k.
[0020] An exclusive OR operation is performed on each pair of subsequences to obtain an intermediate result
[0021] The non-local correlation characteristics of the intermediate result are restricted to obtain the coupled enhanced encoding tag C = (c1, c2... c n )
[0022] The process of restricting the non-local correlation characteristics is as follows:
[0023] The correlation function R C (τ) of the enhanced encoding tag C is calculated to evaluate the non-local correlation characteristics, and the formula is expressed as:
[0024]
[0025] where n represents the number of elements in the enhanced encoding tag C, that is, the length of the enhanced encoding tag C, τ is the time delay, is the mean value of C.
[0026] The process of extracting coding features for enhanced coding labels is as follows:
[0027] Process the binary sequence of the enhanced coding label C;
[0028] Perform numerical conversion on the binary sequence to obtain the numerical vector X = [x1, x2,..., x n ;
[0029] Perform standardization processing on the numerical vector of the enhanced coding label conversion to construct an enhanced data matrix D, and calculate the covariance matrix Cov;
[0030] After obtaining the covariance matrix Cov, perform eigenvalue decomposition, that is, Cov = UΛU T , where U is a matrix composed of eigenvectors,, U T is the transpose of U, and Λ is a diagonal matrix composed of eigenvalues;
[0031] Select the first m eigenvectors with larger eigenvalues, and the principal components corresponding to the m eigenvectors constitute the coding feature F = [f1, f2,..., f m ;
[0032] The process of obtaining the optimal data features is as follows:
[0033] Construct a hierarchical structure model;
[0034] The hierarchical structure model includes an objective layer, a criterion layer, and a scheme layer;
[0035] Based on the criterion layer, construct a judgment matrix A = (a ij ) p×p , where a ij represents the importance score of the i-th data type relative to the j-th data type for the stable operation of the device. p is the matrix dimension, that is, the number of data types, and satisfies reciprocity;
[0036] Solve the maximum eigenvalue λ max and the eigenvector W of the judgment matrix A, and normalize W to obtain the data type weight w = [w1, w2,..., w p , where, The weight w i is the importance index;
[0037] The process of obtaining the optimal data features according to the importance index is as follows:
[0038] For the coding feature F, based on the formula Calculate the score;
[0039] Among them, Z represents the importance score, and δ ij is an indicator function, indicating f iWhen it belongs to the j-th type, δ ij = 1, otherwise δ ij = 0;
[0040] Obtain the importance score of the data corresponding to the coding feature according to the calculation formula, and select the optimal coding data feature set based on the importance score, denoted as where is the quantity index of the optimal coding feature data.
[0041] The construction process of the hierarchical structure model is as follows:
[0042] The construction process of the target layer is as follows:
[0043] The device operating state is The normal operating condition is Cnormal, and the target layer is denoted as ∈Cnormal, where Cnormal includes the condition set of the current within the rated range, the displacement of mechanical components within the allowable interval, and the temperature not exceeding the preset threshold;
[0044] The criterion layer is constructed based on the data acquisition dimension of the railway power supply equipment. Suppose the criterion layer includes p data types, and is represented by the set G: G = {g1, g2...g p};
[0045] The solution layer is a process of feature extraction based on enhanced coding tags.
[0046] The construction process of the dynamic adaptive reinforcement learning model is as follows:
[0047] Collect the time series of the optimal data features in real time, and form the behavior feature sequence through filtering and normalization processing as the input state of the reinforcement learning model;
[0048] Define the model elements including:
[0049] Define the state space, and use the standardized feature vector as the state to represent the real-time operating state of the device;
[0050] Define the action space H = {h1, h2, h3}, where h1 is to maintain the current power supply parameters, h2 is to adjust the voltage output, and h3 is to switch to the backup power supply;
[0051] Design the reward function
[0052] The device operates normally and the features are normal, U = +1;
[0053] Anomaly is detected but no effective intervention is taken, U = -1;
[0054] After taking intervention, the device returned to normal, U = +3;
[0055] Based on a basic deep Q-network, its input layer is set to receive the behavior feature sequence as the state;
[0056] The hidden layer combines a multi-layer perceptron;
[0057] The output layer outputs the Q values of each action where θ is the network parameter;
[0058] Set up a buffer to store the state transition tuples During training, randomly sample data to break the order correlation, where is the next action;
[0059] Regularly copy the main network parameters to the target network and calculate the target Q value: where γ is the discount factor, represents using the target network parameter θ to calculate the next state all possible actions of the maximum Q value, representing the estimation of the future state reward;
[0060] Optimize the model structure so that it adaptively iteratively updates the model parameters based on new data;
[0061] When the importance of the current feature changes dynamically due to load changes in the device, after new data is input, the model adjusts the parameters through adaptive iterative update of the model parameters and re-evaluates the action value;
[0062] Obtain an adaptive reinforcement learning model.
[0063] The process of optimizing the model structure is as follows:
[0064] Using the experience replay mechanism, store the state transition tuples into the buffer;
[0065] During training, randomly sample, calculate the target Q value through the target network, and optimize the loss function Update the main network parameter θ with the help of stochastic gradient descent.
[0066] The present invention has the following beneficial effects:
[0067] In the present invention, first, the Internet of Things coding module collects data through multiple sensors and couples it with a pseudo-random sequence. The generated enhanced coding tags incorporate non-local correlation characteristics, retaining the key information of device operation. At the same time, the optimal feature screening module evaluates the importance of data types with the help of a hierarchical structure model and selects the optimal data features. This process fully considers the impact of different data types on the stable operation of the device. In the railway power supply scenario, the features corresponding to key data types such as current and temperature are clarified, focusing on the core data, which improves the pertinence and efficiency of data processing;
[0068] Secondly, in the dynamic adaptive reinforcement learning model, the input state data is the quantum coding tags coupled with the quantum bit state after being processed by the Internet of Things coding module and the optimal data features obtained by the optimal feature screening module. These features not only contain the information of the railway power supply device sensor data but also incorporate the non-local correlation characteristics of the quantum bits. Compared with the traditional reinforcement learning model that directly uses the original data or simply processed features, the data source and features of this improved model are more unique and complex, and can capture more subtle and deep information during the device operation process;
[0069] Finally, the overall system can capture the changes in device operation in real time by inputting the processed optimal data features into the dynamic adaptive reinforcement learning model. In the face of situations such as device load changes, the model uses the experience replay and target network mechanisms to iteratively update the parameters based on new data and re-evaluate the action value. Compared with the traditional reinforcement learning model, this improved model has stronger adaptability to the dynamic environment, more flexible strategy adjustment, and more stable training process. Brief Description of the Drawings
[0070] Figure 1 It is a system block diagram of the on-line data processing system for railway power supply equipment based on Internet of Things coding proposed by the present invention. Detailed Embodiments
[0071] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0072] Embodiment 1
[0073] As Figure 1 shown, the on-line data processing system for railway power supply equipment based on Internet of Things coding proposed by the present invention includes:
[0074] Internet of Things Coding Module: Collect sensor data and obtain corresponding Internet of Things coding data, couple the pseudo-random sequence data based on the Internet of Things coding data with the sensor data to generate an enhanced coding label with non-local correlation characteristics;
[0075] Multiple types of sensors are distributed on railway power supply equipment, including current sensors, mechanical sensors, and temperature sensors;
[0076] The sensors sample the operating parameters of the power supply equipment in real time at a sampling frequency f. Let the multi-source power supply data sequence collected by the sensors be V = {v1, v2, v3}, where v1 represents the data collected by the current sensor, v2 represents the data collected by the mechanical sensor, and v3 represents the data collected by the temperature sensor;
[0077] The data collected by the sensors is binary-coded through a coding algorithm, and the normalized sensor stream data v i is multiplied by an amplification factor M (for example, M = 256) and converted into an 8-bit binary number to obtain the Internet of Things coding data;
[0078] The process of obtaining the enhanced coding label is as follows:
[0079] Use a pseudo-random number generator to generate a pseudo-random sequence of the Internet of Things coding data;
[0080] Obtain the binary form R = (r1, r2... r n ) of the pseudo-random sequence;
[0081] Obtain the binary sequence S = (s1, s2... s n ) after encoding the sensor data;
[0082] Specifically, the process of obtaining the binary sequence after encoding the sensor data is as follows:
[0083] The data collected by different types of sensors is preprocessed and feature-extracted, and then the final encoded binary sequence s is generated through a weighted fusion method;
[0084] The process of obtaining the binary sequence after encoding the sensor data is as follows:
[0085] Suppose there are m sensors participating in the fusion, and the binary sequence after encoding the data of the jth sensor (different sensors correspond to different data types) is The fusion weight is After fusion where represents the fusion operation;
[0086] The coupling process is as follows:
[0087] The binary sequence S and the pseudo-random sequence R after encoding the sensor data are divided into subsequences according to the length k, that is Among them, since the number of data points of the pseudo-random sequence is generated based on the sensor data, the number of data points of the pseudo-random sequence is equal to the number of data points n of the sensor data;
[0088] Perform an exclusive OR operation on each pair of subsequences to obtain an intermediate result Then, perform a non-local correlation property restriction on these intermediate results. The final coupled enhanced encoding label C=(c1, c2...c n ), and the coupled data is equal to the original number of data points n of the sensor;
[0089] Specifically, the formula By dividing the binary sequence S and the pseudo-random sequence R after encoding the sensor data into subsequences according to the length k, and then performing an exclusive OR operation to obtain an intermediate result, the coupling of the sensor data and the pseudo-random sequence is realized. The characteristics of the exclusive OR operation enable the generated enhanced encoding label C to have a unique information fusion effect, and the introduction of the pseudo-random sequence increases the randomness and complexity of the data;
[0090] The process of restricting the non-local correlation property is as follows:
[0091] Calculate the correlation function R C (τ) of the enhanced encoding label C to evaluate the non-local correlation property. The formula is expressed as:
[0092]
[0093] Among them, n represents the number of elements in the enhanced encoding label C, that is, the length of the enhanced encoding label C, τ is the time delay, is the mean value of C;
[0094] Specifically, by combining the Internet of Things encoding technology, pseudo-random sequence generation, and hierarchical exclusive OR coupling algorithm, and applying it to the data processing of railway power supply equipment, the abstract physical characteristics (non-local correlation) are transformed into computable and optimizable engineering indicators, ensuring that the final result covers the key information of equipment operation.
[0095] Optimal feature screening module: Extract the encoding features for the enhanced encoding label, and evaluate the importance index of the data type according to the encoding features through the hierarchical structure model, and obtain the optimal data features according to the importance index;
[0096] Process the binary sequence of the enhanced encoding label C;
[0097] Perform a numerical conversion on the binary sequence to obtain the numerical vector X=[x1, x2,..., xn , where n is the length of the enhanced coding tag, that is, the number of elements therein;
[0098] Standardize the numerical vector obtained by converting the enhanced coding tag to construct an enhanced data matrix D, and calculate the covariance matrix Cov;
[0099] Among them, the covariance matrix Cov is used to measure the correlation between variables, and its element Cov ij The calculation formula is: Among them, here is the mean value of the data in the i-th column;
[0100] After obtaining the covariance matrix Cov, perform eigenvalue decomposition, that is, Cov = UΛU T , where U is a matrix composed of eigenvectors,, U T is the transpose of U, and Λ is a diagonal matrix composed of eigenvalues;
[0101] Specifically, in the covariance matrix, the larger the eigenvalue of the eigenvalue decomposition, the more information the corresponding eigenvector contains. Select the first m eigenvectors with larger eigenvalues, and the principal components corresponding to these eigenvectors constitute the coding feature F = [f1, f2,... f m , and the process of extracting this coding feature is used as the solution layer of the subsequent hierarchical structure model and at the same time as the core data feature for subsequent analysis;
[0102] Construct a hierarchical structure model;
[0103] The hierarchical structure model includes an objective layer, a criterion layer, and a solution layer;
[0104] The construction process of the objective layer is as follows:
[0105] The device operating state is The normal operating condition is Cnormal, and the objective layer is expressed as where Cnormal includes a set of conditions where the current is within the rated range, the displacement of mechanical components is within the allowable interval, and the temperature does not exceed the preset threshold;
[0106] The criterion layer is constructed based on the data collection dimensions of railway power supply equipment. These data types are the key dimensions for evaluating the device operating state. Suppose the criterion layer contains p data types, and is represented by the set G: G = {g1, g2... g p};
[0107] The solution layer is the process of feature extraction based on the enhanced coding tag;
[0108] Construct a judgment matrix based on the criterion layer;
[0109] Based on the criterion layer, construct the judgment matrix A = (a ij ). p×p , where a ij represents the importance score of the i-th data type relative to the j-th data type for the stable operation of the device. p is the matrix dimension, that is, the number of data types, and it satisfies reciprocity
[0110] Solve the judgment matrix to obtain the importance weights of the data types;
[0111] Solve the maximum eigenvalue λ max of the judgment matrix A and the eigenvector W, and normalize W to obtain the data type weights The weight w i is the importance index. The higher the weight, the more critical the data type is to the stable operation of the device;
[0112] Specifically, the target layer is used to clarify the ultimate goal of the entire evaluation, that is, to ensure the stable operation of the railway power supply equipment. By solving the maximum eigenvalue λ max of the judgment matrix A and the eigenvector W, and normalizing W to obtain the data type weights. This method combines expert experience and mathematical calculations to quantify the importance of different data types to the stable operation of the device;
[0113] The criterion layer is used to determine the criteria to be evaluated, that is, different data types, such as current data, mechanical data, temperature data, etc. These data types come from various sensors on the railway power supply equipment and reflect the operating status of different aspects of the equipment;
[0114] The scheme layer includes specific coding features F, which are extracted from the enhanced coding labels and are further abstract and quantitative representations of the data types;
[0115] The hierarchical structure model comprehensively considers the impacts of multiple factors (different data types and coding features) on the stable operation of the railway power supply equipment, avoiding the limitations of single-factor analysis. By analyzing the factors at each level and calculating the weights, it can comprehensively evaluate the equipment operating status, discover potential risk factors in advance, and thus enhance the reliability and stability of the entire railway power supply system;
[0116] The process of obtaining the optimal data features according to the importance index is as follows:
[0117] For the coding feature F, calculate the score based on the formula ;
[0118] Among them, Z represents the importance score, and δ ij is the indicator function, indicating that when f i belongs to the j-th type, δ ij = 1, otherwise δ ij = 0;
[0119] Obtain the importance scores of the data corresponding to the coding features according to the calculation formula, and select the optimal coding data feature set based on the importance scores, denoted as where is the quantity index of the optimal coding feature data;
[0120] Through this formula, the data type weight w j is assigned to the corresponding features to quantify the feature importance. The coding features are sorted in descending order according to the score S, and the top features are selected as the optimal data features according to the requirements. These features focus on key information, provide core data support for subsequent equipment status analysis, fault diagnosis, etc., and improve the pertinence and effectiveness of railway power supply data processing.
[0121] Dynamic adaptation optimization module: Obtain the behavioral feature sequence of the optimal data features, and construct a dynamic adaptive reinforcement learning model according to the behavioral feature sequence;
[0122] As the railway power supply equipment operates, the time series of the optimal data features is collected in real time. After filtering and normalization processing, a behavioral feature sequence is formed as the input state of the reinforcement learning model, and the latest information of the equipment operation is transmitted in real time;
[0123] Construct an adaptive reinforcement learning model;
[0124] Define the model elements including:
[0125] State: The standardized feature vector is used as the state to represent the real-time operation state of the equipment.
[0126] Action: Define the action space H = {h1, h2, h3}, where h1 is "maintain the current power supply parameters", h2 is "adjust the voltage output", and h3 is "switch to the standby power supply";
[0127] Reward: Design the reward function
[0128] The equipment is operating normally and the features are normal, U = +1;
[0129] Anomaly is detected but no effective intervention is taken, U = -1;
[0130] The equipment returns to normal after taking intervention, U = +3;
[0131] Model architecture and training:
[0132] Based on a basic deep Q-network, set its input layer to receive the behavioral feature sequence as the state;
[0133] The hidden layer combines a multi-layer perceptron;
[0134] The output layer outputs the Q-values of each action where θ are the network parameters;
[0135] Set up a buffer to store state transition tuples During training, randomly sample data to break the sequential correlation. Among them, is the next action;
[0136] Regularly copy the main network parameters to the target network and calculate the target Q-value: where γ is the discount factor, represents using the target network parameters θ to calculate the next state all possible actions in the next the maximum Q-value, representing the estimation of future state rewards;
[0137] Furthermore, by introducing a target network and regularly copying the main network parameters, when calculating the target Q-value, use the relatively stable parameters θ of the target network. At the same time, through the formula provides a stable training target. Compared with traditional methods, it avoids the interference of the dynamic change of the main network parameters on the target calculation, and significantly improves the stability and convergence speed of the training process;
[0138] Further optimize the model structure so that it can adaptively update the model parameters based on new data;
[0139] Using the experience replay mechanism, store the state transition tuple into the buffer;
[0140] During training, randomly sample, calculate the target Q-value through the target network, and optimize the loss function Update the main network parameters θ with the help of stochastic gradient descent;
[0141] Specifically, the parameters of traditional reinforcement learning models are updated with a lag and are difficult to handle dynamic scenarios such as changes in the load of railway power supply equipment and environmental interference. The improved reinforcement learning model can iterate in real time based on new data. When the importance of current characteristics changes due to load changes in the equipment, after new data is input, optimize the loss function Update the main network parameters θ with the help of stochastic gradient descent;
[0142] This mechanism enables the improved reinforcement learning model to respond in real time to changes in the operating state of the equipment, re-evaluate the action value, break through the limitation of the "static training" of traditional models, and achieve precise adaptation to the dynamic scenarios of railway power supply;
[0143] When the importance of the current characteristics changes dynamically due to load changes in the device, after new data is input, the model adjusts the parameters through the above mechanism and re-evaluates the action value;
[0144] Finally, an adaptive reinforcement learning model that has completed training is obtained;
[0145] Specifically, in the dynamic adaptive reinforcement learning model, this method of selecting features based on the importance of data types makes the state information entering the reinforcement learning model more targeted. As the operating state of the railway power supply equipment changes continuously and new data is continuously generated, the model can adjust its own parameters and strategies in real time according to the behavior feature sequence.
[0146] Data result output module: Based on the dynamic adaptive reinforcement learning model, it outputs action modes, including normal mode and abnormal mode, and performs intervention operations on the data in the abnormal mode;
[0147] The dynamic adaptive reinforcement learning model outputs the Q-value distribution when the equipment is operating normally in the historical data, and sets a threshold When It is determined to be in the normal mode;
[0148] If Then it is in the abnormal mode;
[0149] Furthermore, the model calculates the Q-values of all actions according to the current state and selects the action with the largest Q-value In the normal mode, it outputs the action to maintain the operation of the equipment. In the abnormal mode, it outputs the corresponding processing action, such as "voltage data is abnormal, track and record the abnormal voltage data";
[0150] For example, a specific implementation example of the Q-value and the action H with the largest Q-value is as follows:
[0151] In the dynamic adaptive reinforcement learning model, assume the current state The corresponding behavior feature sequence contains some of the encoded features extracted above. The model calculates the Q-values of all actions according to the current state For example, in the current state, the Q-value of the action of "maintaining the current power supply parameters" is 0.5, the Q-value of the action of "adjusting the voltage output" is 0.3, and the Q-value of the action of "switching to the standby power supply" is 0.1. If the equipment is operating normally and the features are normal, the reward function U = +1; if an abnormality is detected but no effective intervention is taken, U = -1; if the equipment returns to normal after taking intervention, U = +3. Assume the current state is normal and the reward U = +1;
[0152] According to the target Q-value calculation formula (assuming the discount factor γ = 0.9), calculate the target Q-value, assuming the next state The Q value of the "maintain current power supply parameters" action is 0.6, the Q value of the "adjust voltage output" action is 0.4, and the Q value of the "switch to backup power supply" action is 0.2. Then the maximum action The target Q value y = 1 + 0.9×0.6 = 1.54.
[0153] In the application, several formulas involved are calculated by taking their numerical values after dimensionless. The establishment of the formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the real situation. Some coefficients or weights in the formula are set by those skilled in the art according to the actual situation, so no more details will be given here.
[0154] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution.
[0155] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An online data processing system for railway power supply equipment based on Internet of Things coding, characterized in that Including: IoT Coding Module: Collect sensor data and obtain corresponding IoT coding data, couple the pseudo-random sequence data based on the IoT coding data with the sensor data to generate enhanced coding tags with non-local correlation characteristics; Optimal Feature Screening Module: Extract coding features for the enhanced coding tags, evaluate the importance indicators of data types according to the coding features through a hierarchical structure model, and obtain the optimal data features according to the importance indicators; Dynamic Adaptation Optimization Module: Obtain the behavior feature sequence of the optimal data features, and construct a dynamic adaptive reinforcement learning model according to the behavior feature sequence; Data Result Output Module: Output action modes based on the dynamic adaptive reinforcement learning model. The action modes include normal mode and abnormal mode, and perform intervention operations on the data in the abnormal mode.
2. The on-line data processing system for railway power supply equipment based on Internet of Things coding according to claim 1, characterized in that, The process of obtaining the IoT coding data: Real-time sample the operating parameters of the power supply equipment through a sensor at a sampling frequency f and perform binary coding; Obtain a multi-source power supply data sequence V = {v1, v2, v3}, where v1 represents the data collected by the current sensor, v2 represents the data collected by the mechanical sensor, and v3 represents the data collected by the temperature sensor; The process of the binary encoding includes multiplying the normalized sensor stream data v i by the amplification factor M and converting it into an 8-bit binary number.
3. The online data processing system for railway power supply equipment based on Internet of Things coding according to claim 1, wherein The process of obtaining the enhanced coding tags is: Use a pseudo-random number generator to generate a pseudo-random sequence of IoT coding data; Obtain the binary form R of the pseudo-random sequence; Obtain the binary sequence S after encoding the sensor data; Couple the binary sequence S with the binary form R of the pseudo-random sequence to obtain enhanced coding tags.
4. The online data processing system for railway power supply equipment based on Internet of Things coding according to claim 3, wherein The process of coupling the binary sequence S with the binary form R of the pseudo-random sequence is: Divide the binary sequence S after encoding the sensor data and the pseudo-random sequence R into subsequences according to the length k; Perform an exclusive OR operation on each pair of subsequences to obtain an intermediate result Qualify the non-local correlation characteristics of the intermediate result to obtain the enhanced encoded label C = (c1, c2... c n ).
5. The on-line data processing system for railway power supply equipment based on Internet of Things coding according to claim 4, characterized in that The process of defining the non-local correlation characteristics is; Calculate the correlation function R C (τ) of the enhanced coding label C to evaluate the non-local correlation characteristics, which is expressed by the formula as follows: where n represents the number of elements in the enhanced coding label C, i.e., the length of the enhanced coding label C, and τ is the time delay, is the mean value of C.
6. The online data processing system for railway power supply equipment based on Internet of Things coding according to claim 1, wherein The process of extracting coding features for the enhanced coding tags is: Process the binary sequence of the enhanced coding tag C; Perform a numerical conversion on the binary sequence to obtain a numerical vector X = [x1, x2, …, x n ; Perform normalization processing on the numerical vector converted from the enhanced coding tag to construct an enhanced data matrix D, and calculate the covariance matrix Cov; After obtaining the covariance matrix Cov, perform eigenvalue decomposition, i.e., Cov = UΛU T , where U is a matrix composed of eigenvectors, and U T is the transpose of U, and Λ is a diagonal matrix composed of eigenvalues; Select the first m eigenvectors with larger eigenvalues, and the principal components corresponding to the m eigenvectors constitute the encoded feature F = [f1, f2,..., f m .
7. The online data processing system for railway power supply equipment based on Internet of Things coding according to claim 1, wherein The process of obtaining the optimal data features is: Construct a hierarchical structure model; The hierarchical structure model includes an objective layer, a criterion layer, and a scheme layer; Based on the criterion layer, construct the judgment matrix A = (a ij ) p×p , where a ij represents the importance score of the i-th data type relative to the j-th data type for the stable operation of the device. p is the matrix dimension, that is, the number of data types, and it satisfies reciprocity; Solve for the maximum eigenvalue λ of the judgment matrix A max And the eigenvector W, normalize W to obtain the weight w of the data type, w = [w1, w2,..., w p , where The weight w i That is, the importance index; The process of obtaining the optimal data features according to the importance indicators is: For the coding feature F, calculate the score based on the formula ; where Z represents the importance score, and δ ij is an indicator function, indicating that when f i belongs to the j-th type, δ ij = 1; otherwise, δ ij = 0; Obtain the importance score of the data corresponding to the coding feature according to the calculation formula, and select the optimal coding data feature set based on the importance score, denoted as wherein is the quantity index of the optimal coding feature data.
8. The online data processing system for railway power supply equipment based on Internet of Things coding according to claim 7, characterized in that The construction process of the hierarchical structure model is: The construction process of the objective layer is: The operating state of the device is The normal operating condition is Cnormal, and the target layer is represented as where Cnormal includes a set of conditions where the current is within the rated range, the displacement of mechanical components is within the allowable range, and the temperature does not exceed the preset threshold; The criterion layer is constructed based on the data collection dimensions of railway power supply equipment. Suppose the criterion layer contains p types of data, which are represented by the set G: G = {g1, g2... g p}; The scheme layer is the process of feature extraction based on the enhanced coding tags.
9. The online data processing system for railway power supply equipment based on IoT coding according to claim 1, the construction process of the dynamic adaptive reinforcement learning model is: Collect the time series of the optimal data features in real time, and through filtering and normalization processing, form a behavior feature sequence As the input state of the reinforcement learning model; Define the model elements including: Define the state space and use the normalized eigenvector as the state to represent the real-time operating state of the device; Define the action space H = {h1, h2, h3}, where h1 is to maintain the current power supply parameters, h2 is to adjust the voltage output, and h3 is to switch to the standby power supply; Design reward function When the equipment runs normally and the features are normal, U = +1; When an anomaly is detected but no effective intervention is taken, U = -1; When the equipment returns to normal after taking intervention, U = +3; Based on a basic deep Q-network, its input layer is set to receive a sequence of behavioral features as the state; The hidden layer combines a multi-layer perceptron; The output layer outputs the Q-values of each action where θ is the network parameter; Set up a buffer to store state transition tuples During training, randomly sample data to break the sequential relationship, where is the next action Regularly copy the main network parameters to the target network and calculate the target Q value: where γ is the discount factor, means using the target network parameters θ to calculate the next state for all possible actions of the maximum Q value, representing the estimation of future state rewards; Optimize the model structure to iteratively update the model parameters adaptively based on new data; When the importance of the current feature changes dynamically due to load changes in the device, after new data is input, the model iteratively updates and adjusts the parameters by adapting the model parameters, and re-evaluates the action value; An adaptive reinforcement learning model is obtained.
10. For the online data processing system of railway power supply equipment based on Internet of Things coding according to claim 9, the process of optimizing the model structure is as follows: Using the experience replay mechanism, store the state transition tuple into the buffer; During training, randomly sample, calculate the target Q-value through the target network, and optimize the loss function Update the parameters θ of the main network by means of stochastic gradient descent.