Large-scale data feature extraction method, system and equipment based on average field
By dividing large-scale data into local subproblems through feature correlation analysis and dynamic coupling mechanism, the communication and computing bottlenecks in ultra-large-scale data processing are solved, and a fast and flexible global optimal solution is achieved, which can adapt to changes in data distribution.
Patent Information
- Application Number
- CN202511535187.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-10-27
AI Technical Summary
Existing technologies suffer from communication bottlenecks, high computational complexity, and insufficient flexibility when processing ultra-large-scale data, making it difficult to meet the demand for efficient processing.
By using feature correlation analysis, large-scale data is divided into multiple local subproblems that can be processed independently and in parallel. A dynamic coupling mechanism based on information entropy and variance is introduced to coordinate the solution process of local subproblems and achieve rapid convergence of the global optimal solution.
It effectively overcomes the bottleneck of network communication overhead, reduces computational complexity, and improves the flexibility of the algorithm to adapt to dynamic changes in data distribution.
Smart Images

Figure CN120995088A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of mass data processing, and particularly relates to a large-scale data feature extraction method, system and device based on an average field. BACKGROUND
[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute prior art.
[0003] At present, the popularity of Internet of Things devices, high-definition sensors and social media continues to increase, directly leading to explosive growth in data size. These data not only have a large volume, but also often exhibit complex characteristics of high dimensionality and non-structure, bringing significant challenges to data feature extraction. Under this background, how to quickly and accurately extract valuable information and patterns from mass data has become a core problem for promoting the landing of artificial intelligence, optimizing business decisions and assisting scientific research.
[0004] In view of this problem, two types of mainstream technical solutions have been formed: One type is a distributed learning framework based on gradient descent, the core of which is to complete model optimization on a computing cluster by means of data or model parallel strategy and stochastic gradient descent (SGD) algorithm; The other type is a technical path based on a probabilistic graphical model, including Bayesian networks, Markov random fields, etc., which is mainly used to describe complex dependency relationships between data. In order to reduce the computational cost of the probabilistic graphical model, mean field variational inference has become a common approximation method, which simplifies the solving process of the posterior distribution by assuming variable independence.
[0005] However, it should be noted that the existing technology still has inherent defects that are difficult to avoid when dealing with truly super large-scale data tasks, and cannot meet the efficient processing needs. From a practical application point of view, the existing technology mainly faces three bottlenecks: First, the communication bottleneck problem: in the distributed framework based on SGD, the worker nodes need to frequently exchange gradient information with the parameter server to realize model synchronization. This process will produce huge network communication overhead as the model parameter quantity increases or the cluster size expands, often making the communication time much longer than the computation time, leading to a sharp decline in the expansion efficiency of the entire system, and even possibly appearing the paradox of performance decreasing with size increasing; Second, traditional methods are generally restricted by computational complexity: many traditional mean field variational inference methods need to perform global computation on the entire data set, resulting in computational complexity proportional to the square of the data volume or feature dimension, which seriously limits their applicability in super large-scale scenarios. The root cause of this problem lies in the fact that traditional methods lack a mechanism to effectively decompose high-dimensional complex problems into low-dimensional sub-problems; In addition, the lack of flexibility and adaptability is also a significant shortcoming of the prior art: in many existing schemes, the key parameters are often statically preset or rely on costly global statistical calculations, which are difficult to adapt to the dynamic changes of data distribution. This rigid parameter setting method is due to the failure to fully utilize real-time feedback information generated during local calculation in algorithm design. SUMMARY
[0006] To solve the technical problems in the background art, the present application provides a large-scale data feature extraction method, system and device based on mean field, which intelligently divides large-scale data through feature correlation analysis, decouples the complex global problem into multiple independent and parallel local sub-problems, and introduces a dynamic coupling mechanism based on information entropy and variance to coordinate the solution process of each local sub-problem, realizes the fast convergence of global optimal solution with minimum communication overhead, and effectively breaks through the network communication overhead bottleneck of super large-scale data feature extraction.
[0007] To achieve the above purpose, the present application adopts the following technical solutions: The first aspect of the present application provides a large-scale data feature extraction method based on mean field, applied to a parameter server, comprising: obtaining a data set, dividing the feature correlation graph of the data set into a plurality of subgraphs, dividing the data set into a same number of subsets as the subgraphs, each subset containing all data samples and only containing the features of the corresponding subgraph; sending each subset to a computing node, so that the computing node combines the coupling coefficient to calculate the probability distribution, expectation and entropy of each feature through a distributed mean field equation; Based on the probability distribution, expectation and entropy of each feature calculated by all computing nodes, update the coupling coefficient, and then issue the calculation instruction and the updated coupling coefficient to the computing nodes until the iteration converges. For each feature, based on the probability distribution calculated by the computing node containing the feature, calculate the global probability distribution.
[0008] Further, the updating step of the coupling coefficient comprises: for each feature, calculating the global entropy based on the entropy calculated by the computing node containing the feature; calculating the global expectation based on the expectation calculated by the computing node containing the feature; calculating the variance based on the global expectation and the expectation calculated by the computing node containing the feature; updating the coupling coefficient through the entropy-variance weighting mechanism based on the global entropy and the variance.
[0009] Further, the entropy-variance weighting mechanism is:
[0010] wherein, Ct represents the coupling coefficient obtained in the tth iteration; Representation of features global entropy; Representation of features global entropy; Representation of features Variance between different computation nodes; Representation of features Variance between different computation nodes; It's a hyperparameter.
[0011] Furthermore, the construction steps of the feature association graph include: calculating the feature mutual information matrix based on the dataset using the mutual information formula; treating each feature in the dataset as a graph node, and constructing the feature association graph using the value of the mutual information matrix as the edge weight.
[0012] Furthermore, the condition for the iteration to converge is that the Jensen-Shannon divergence between the global feature field distribution of the current round and the previous round is less than a threshold, and the global feature field distribution is the product of the global probability distributions of all features.
[0013] Furthermore, the threshold is: ;in, As the initial threshold, The attenuation rate, This represents the number of iterations.
[0014] Furthermore, the distributed mean-field equation is: ; in, Indicates the first During the round of iteration, the first Features computed by each computing node The probability distribution; It is a temperature hyperparameter; It is the first Normalization factor for each computing node; Representative characteristics The set of indices of the neighbor features; Representing the Features of all neighbors on each computing node Features The average resultant force produced; partial derivatives Energy function Features The gradient.
[0015] A second aspect of the present invention provides a method for large-scale data feature extraction based on mean field, applied to computing nodes, comprising: Receive a subset of parameters sent by the server; in response to the calculation instruction sent by the parameter server, and receiving the coupling coefficient sent by the parameter server, for each feature in the subset, combining the coupling coefficient, calculating the probability distribution, expectation and entropy through the distributed mean field equation, and sending to the parameter server, so that the parameter server updates the coupling coefficient based on the probability distribution, expectation and entropy of each feature calculated by all the computing nodes, until the iteration converges, for each feature, calculating the global probability distribution based on the probability distribution calculated by the computing node containing the feature; The parameter server associates feature correlation graphs of the data set into a plurality of subgraphs, divides the data set into a same number of subsets as the subgraphs, each subset containing all data samples and only containing features of a corresponding subgraph.
[0016] The third aspect of the present application provides a large-scale data feature extraction system based on mean field, comprising a parameter server and a plurality of computing nodes. The parameter server is configured to obtain a data set, associate feature correlation graphs of the data set into a plurality of subgraphs, divide the data set into a same number of subsets as the subgraphs, each subset containing all data samples and only containing features of a corresponding subgraph. The computing node is configured to receive the subset sent by the parameter server, in response to the calculation instruction sent by the parameter server, and receiving the coupling coefficient sent by the parameter server, for each feature in the subset, combining the coupling coefficient, calculating the probability distribution, expectation and entropy through the distributed mean field equation. The parameter server is further configured to, after updating the coupling coefficient based on the probability distribution, expectation and entropy of each feature calculated by all the computing nodes, issue the calculation instruction and the updated coupling coefficient to the computing node until the iteration converges, for each feature, calculating the global probability distribution based on the probability distribution calculated by the computing node containing the feature.
[0017] The fourth aspect of the present application provides a computer device, comprising a computer readable storage medium, a processor and a computer program stored on the computer readable storage medium and executable on the processor, wherein the processor implements the steps of the above-mentioned large-scale data feature extraction method based on mean field when executing the program.
[0018] Compared with the prior art, the present application has the following advantages: The present application intelligently divides large-scale data through feature correlation analysis, decouples the complex global solving problem into a plurality of local sub-problems which can be independently and parallelly processed, and introduces a dynamic coupling mechanism based on information entropy and variance to coordinate the solving process of each local sub-problem, realizes the fast convergence of the global optimal solution with minimum communication overhead, and effectively breaks through the network communication overhead bottleneck of super large-scale data feature extraction.
[0019] The application can utilize local real-time feedback information in time by setting a dynamic change threshold, can adapt to dynamic changes of data distribution, and makes the algorithm more flexible. BRIEF DESCRIPTION OF DRAWINGS
[0020] The accompanying drawings, which form a part of this specification, are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification. The embodiments of these drawings are set to explain the application, and do not constitute an improper limitation to the application.
[0021] Figure 1 is a flow chart of a large-scale data feature extraction method based on an average field of embodiment one of the application; Figure 2 is a schematic diagram of the working interaction of a parameter server cluster and each computing node of embodiment four of the application; Figure 3 is a structural schematic diagram of a computer device of embodiment five of the application. DETAILED DESCRIPTION
[0022] To make the purpose, technical scheme and advantages of the embodiments of the application more clear, the technical scheme in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application.
[0023] It should be pointed out that the following detailed description is exemplary and is intended to provide further explanation of the application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as generally understood by those skilled in the art to which the application belongs.
[0024] Embodiment one The large-scale data feature extraction method based on an average field provided in this embodiment innovatively embeds the average field idea into a distributed computing architecture, can break through the performance limit of massive data feature extraction, significantly optimizes the computing bottleneck and computing complexity of the feature extraction task, and provides technical support for massive data driven application scenarios.
[0025] The large-scale data feature extraction method based on an average field provided in this embodiment first intelligently divides massive data through feature correlation analysis, decouples the complex solving problem at the global level into multiple local sub-problems that can be independently and parallelly processed; secondly, introduces a dynamic coupling mechanism based on information entropy and variance for coordinating the solving process of each local sub-problem; finally, realizes the fast convergence of the global optimal solution with minimum communication overhead, and effectively breaks through the performance bottleneck of super large-scale data feature extraction.
[0026] The large-scale data feature extraction method based on an average field provided in this embodiment, as shown in Figure 1 contains five stages, and the specific steps are as follows: Step 1: Obtain the original large-scale data and construct a feature association graph.
[0027] To process large-scale datasets, the first step is to construct a feature association graph corresponding to them. The main construction steps are as follows: (1) The large-scale data to be processed Convert to data matrix ,in Represents the number of data samples. This represents the number of data features.
[0028] Specifically, for structured data, numerical features can be directly standardized or categorical features can be one-hot encoded; for unstructured text, image, and audio data, word embedding, CNN (convolutional neural network) feature extraction, spectrogram analysis, and other methods can be used to transform them into fixed-length feature vectors.
[0029] (2) Based on data matrix Calculate the feature mutual information matrix .
[0030] Specifically, the statistical dependencies between features can be calculated using the mutual information formula, thus enabling the calculation of the feature mutual information matrix of the data. : ; in, It is a random variable, representing the sum of all samples in the dataset at the _th ... The distribution of values across each feature dimension; Indicates the first The marginal probability distribution of a feature, that is, the distribution pattern of the value of the feature itself; Indicates the first The first feature and the first The joint probability distribution of features, that is, the pattern of the co-occurrence of values between two features; and Through data matrix Calculated.
[0031] (3) Based on the feature mutual information matrix Construct a feature association graph.
[0032] Specifically, each feature dimension Treat them as graph nodes and use the mutual information matrix The value is the edge weight, and a weighted feature association graph is constructed. .
[0033] Furthermore, in order to make the diagram Removing a large number of weak or even meaningless connections requires threshold filtering. The specific steps are as follows: set a threshold... For the graph Medium weight Remove the edges and keep the rest to form a new feature association graph. .
[0034] Furthermore, threshold The selection can be made based on the domain and characteristics of the data to be processed.
[0035] Step 2: Intelligent data segmentation.
[0036] Specifically, in order to decompose the high-dimensional feature space, graph partitioning algorithms such as spectral clustering can be used to partition the feature association graph. Divided into A tightly connected subgraph There may be feature overlap between different subgraphs, meaning that a feature may belong to two or more classes at the same time.
[0037] Furthermore, to reduce the data dimensionality that each computing node needs to process, the original dataset can be divided according to the feature partitioning results. Divided into the same number Subset And let each subset (data sub-block) Includes all data samples, but only retains those belonging to the subgraph. The feature dimensions.
[0038] Furthermore, the main purpose of this step is to make the features within each subset highly correlated, while the features between different subsets are as independent as possible. By implementing this step, high-dimensional complex problems can be decomposed into low-dimensional simple subproblems, thus significantly reducing computational complexity.
[0039] Step 3: Update the distributed mean-field equations in parallel.
[0040] Specifically, each computing node receives a subset of data. and the current global coupling coefficient Then, in a GPU-accelerated environment, each computing node independently and in parallel updates the following mean-field equation: ; in, Indicates the first During the round of iteration, the first Each computing node has its assigned features The best local estimate of the probability distribution; It is a temperature hyperparameter used to control the smoothness of the optimization; is the normalization factor for the th computing node to ensure is a valid probability distribution; represents the current feature 's neighbor features' index set; is given based on the mean field idea, representing the average force generated by all neighbor features on the th computing node on the current feature ; denotes the coupling coefficient between the current feature and neighbor features obtained at the t-1th iteration; denotes the expectation of neighbor features on the kth computing node at the t-1th iteration; the partial derivative originates from the gradient of the energy function with respect to , driving the distribution of to evolve towards the direction of reducing the reconstruction error.
[0041] Further, the purpose of this equation is to update the probability distribution of each feature dimension in each computing node , after obtaining the probability distribution, various statistics of the feature can be easily calculated: the expectation of the computing node, i.e. the best local estimate of the feature in the node; the entropy , i.e. the degree of confusion of the distribution in the computing node, where represents the Shannon entropy.
[0042] Further, the energy function in this equation measures the error between the reconstructed dataset with the feature and the subset , for example, it can be set as:
[0043] where, represents the reconstruction.
[0044] Further, it is worth noting that due to the different data perspectives seen by each computing node, different computing nodes may obtain slightly different distribution estimates for the same feature .
[0045] Further, by decoupling the complex global equation solving problem into multiple local mean field sub-problems that can be independently and parallelly processed, the network communication overhead in the data feature extraction process can be greatly reduced.
[0046] Step 4: Global coupling coefficient aggregator, update global coupling coefficient.
[0047] (1) Compute global expectation.
[0048] Specifically, after the parameter server receives the local statistics sent by all workers (computing nodes) in each iteration, for each feature , its global expectation is computed by averaging the local expectations of all computing nodes, i.e. where denotes the number of computing nodes containing feature .
[0049] Further, in the process of computing global expectation, if the size difference between data subsets is too large, the computation can be weighted according to the size of the data subsets.
[0050] (2) Compute global entropy.
[0051] Specifically, for each feature , its global entropy can be obtained by averaging the local entropies, i.e. where denotes the global authority distribution of the probability distribution of feature in the th iteration, which can also be considered as the most likely distribution after all local information is fused in the th iteration.
[0052] Further, is the global knowledge of the feature formed by aggregating the statistical characteristics of all computing nodes containing the feature about feature , i.e., the global probability distribution, through the parameter server.
[0053] (3) Compute variance.
[0054] Specifically, for each feature , the variance of its expectation value among different computing nodes is computed, i.e. .
[0055] (4) Update global dynamic coupling coefficient .
[0056] Specifically, the global coupling coefficient is updated using the following entropy-variance weighting mechanism:
[0057] where is a hyperparameter used to adjust the overall importance of the entropy component.
[0058] Step 5: Adaptive convergence determinator, making adaptive determination convergence.
[0059] This step mainly gives how to intelligently determine whether the distributed average field iteration has converged, so as to terminate the calculation and output the final feature extraction result, and its main process is as follows: (1) Calculate the change amount.
[0060] Specifically, according to the collected global field distribution information, the change amount of the current iteration round Global feature field distribution and the last round Global feature field distribution Jensen-Shannon (Jensen-Shannon) divergence between them: ; Wherein, ( ) represents the KL distance; is the average distribution of two rounds, that is: .
[0061] Further, the in the above formula represents the joint probability distribution of all features in the round iteration, which describes the relationship between all features and the state of the entire system.
[0062] Further, according to the mean field approximation theory, the complex joint distribution can be approximated as the product of all individual edge distributions, so the update equation of is: .
[0063] Further, the numerical value of Jensen-Shannon divergence can reflect the overall update amplitude of the system in the round iteration, and the smaller the value, the more stable and the weaker the change, so the numerical value of in each iteration can be used to determine the size of the update amplitude.
[0064] (2) Calculate the current threshold.
[0065] Specifically, according to the current iteration number , calculate , wherein, is the initial threshold, which should be set at a higher value (for example, set to 0.5), which ensures that the distribution has a larger change in the early stage of iteration, encourages the algorithm to "explore", and avoids falling into local optimum too early; is the decay rate, which controls the speed of threshold decline.
[0066] Further, with the increase of the number of iterations , the threshold value is automatically and gradually reduced, which represents the optimization strategy of "coarse first and fine later": that is, the early threshold value is large, allowing the algorithm to take a big step to explore and quickly approach the optimal solution area, and the late threshold value is small, requiring the algorithm to be fine and delicate, and only when the distribution improves extremely slightly will the iteration continue, thus, the local real-time feedback information can be used in time, the dynamic changes of the data distribution can be adapted, and the algorithm is more flexible.
[0067] (3) Comparison and convergence determination.
[0068] Specifically, the following comparison is made: ; The above formula shows that if the comparison is True (Yes), the parameter server sends a termination signal to all computing nodes and outputs the final result; otherwise, that is, the comparison is False (No), the next round of iteration is started.
[0069] Further, after the iteration is completed and convergence is achieved, the final output is a stable, global feature field distribution , specifically the probability distribution of each feature and its statistics, and these outputs constitute a new and efficient representation of the original data set, which can be used for subsequent higher-level tasks.
[0070] The large-scale data feature extraction method based on mean field provided in this embodiment does not use the existing SGD technology, but solves the complex global equation problem by decoupling it into multiple local mean field sub-problems that can be independently and parallelly processed, so as to greatly reduce the network communication overhead bottleneck problem in the data feature extraction process.
[0071] The large-scale data feature extraction method based on mean field provided in this embodiment greatly reduces the computational complexity by intelligently dividing the high-dimensional feature space and decomposing the high-dimensional complex problem into low-dimensional simple sub-problems.
[0072] The large-scale data feature extraction method based on mean field provided in this embodiment can use local real-time feedback information in time by setting a dynamically changing threshold value, can adapt to the dynamic changes of the data distribution, and makes the algorithm more flexible.
[0073] Embodiment Two The large-scale data feature extraction method based on mean field provided in this embodiment is applied to a parameter server and includes the following steps: obtaining a dataset, partitioning a feature correlation graph of the dataset into a plurality of subgraphs, and partitioning the dataset into a same number of subsets as the subgraphs, each of the subsets containing all data samples and only containing features of a corresponding subgraph; sending each of the subsets to a computing node, to enable the computing node to calculate a probability distribution, an expectation, and an entropy of each of the features by a distributed mean field equation in combination with the coupling coefficients; updating the coupling coefficients based on the probability distribution, the expectation, and the entropy of each of the features calculated by all of the computing nodes, and then issuing a calculation instruction and the updated coupling coefficients to the computing nodes until an iteration converges, and for each of the features, calculating a global probability distribution based on the probability distribution calculated by the computing node containing the feature.
[0074] Further, the updating of the coupling coefficients comprises, for each of the features, calculating a global entropy based on the entropy calculated by the computing node containing the feature, calculating a global expectation based on the expectation calculated by the computing node containing the feature, calculating a variance based on the global expectation and the expectation calculated by the computing node containing the feature, and updating the coupling coefficients by an entropy-variance weighting mechanism based on the global entropy and the variance.
[0075] Further, the entropy-variance weighting mechanism is:
[0076] wherein, denotes the coupling coefficients obtained in the tth iteration; denotes a global entropy of the feature denotes a global entropy of the feature denotes a variance of the feature between different computing nodes; denotes a variance of the feature between different computing nodes; denotes a variance of the feature between different computing nodes; is a hyperparameter.
[0077] Further, the constructing of the feature correlation graph comprises, based on the dataset, calculating a feature mutual information matrix by a mutual information formula, and constructing the feature correlation graph by taking each feature in the dataset as a graph node and taking a value of the mutual information matrix as an edge weight.
[0078] Further, the condition of the iteration convergence is that a Jensen-Shannon divergence between a global feature field distribution of a current round and a global feature field distribution of a previous round is less than a threshold value, and the global feature field distribution is a product of global probability distributions of all features.
[0079] Further, the threshold value is: ; wherein, is an initial threshold value, is a decay rate, This represents the number of iterations.
[0080] Furthermore, the distributed mean-field equation is: ; in, Indicates the first During the round of iteration, the first Features computed by each computing node The probability distribution; It is a temperature hyperparameter; It is the first Normalization factor for each computing node; Representative characteristics The set of indices of the neighbor features; Representing the Features of all neighbors on each computing node Features The average resultant force produced; partial derivatives Energy function Features The gradient.
[0081] It should be noted that each step in this embodiment corresponds one-to-one with each step in Embodiment 1, and their specific implementation process is the same, so it will not be repeated here.
[0082] Example 3 This embodiment provides a large-scale data feature extraction method based on mean field, applied to computing nodes, including: Receive a subset of parameters sent by the server; In response to the computation instructions sent by the parameter server and receiving the coupling coefficients sent by the parameter server, for each feature in the subset, the probability distribution, expectation, and entropy are calculated by combining the coupling coefficients and the distributed mean field equations, and sent to the parameter server so that the parameter server updates the coupling coefficients based on the probability distribution, expectation, and entropy of each feature calculated by all computing nodes until the iteration converges. For each feature, the global probability distribution is calculated based on the probability distribution calculated by the computing nodes containing that feature. The parameter server divides the feature association graph of the dataset into several subgraphs, and divides the dataset into the same number of subsets as the subgraphs. Each subset contains all data samples and only contains the features of the corresponding subgraph.
[0083] It should be noted that each step in this embodiment corresponds one-to-one with each step in Embodiment 1, and their specific implementation process is the same, so it will not be repeated here.
[0084] Example 4 This embodiment provides a large-scale data feature extraction system based on mean field, such as... Figure 2 As shown, it includes a parameter server cluster (abbreviated as parameter server) and multiple computing nodes (computing node 1, computing node 2, ..., computing node K). The parameter server is used to acquire the dataset, divide the feature association graph of the dataset into several subgraphs, and divide the dataset into subsets (data blocks) of the same number as the subgraphs. Each subset contains all data samples and only contains the features of the corresponding subgraph; The computing node is used to receive a subset sent by the parameter server; respond to the computing instructions sent by the parameter server and receive the coupling coefficient sent by the parameter server; for each feature in the subset, in combination with the coupling coefficient, calculate the probability distribution, expectation and entropy through the distributed mean field equation; The parameter server is also used to update the coupling coefficients based on the probability distribution, expectation, and entropy of each feature calculated by all computing nodes, and then issue calculation instructions (convergence instructions) and updated coupling coefficients to the computing nodes until iterative convergence. For each feature, the global probability distribution is calculated based on the probability distribution calculated by the computing nodes containing that feature.
[0085] The parameter server includes a global coupling coefficient aggregator, which is used to update the coupling coefficients: for each feature, the global entropy is calculated based on the entropy calculated by the computing nodes containing the feature; the global expectation is calculated based on the expectation calculated by the computing nodes containing the feature; the variance is calculated based on the global expectation and the expectation calculated by the computing nodes containing the feature; and the coupling coefficients are updated based on the global entropy and variance through an entropy-variance weighting mechanism and synchronized to the adaptive convergence arbiter.
[0086] The parameter server includes an adaptive convergence determiner, which is used to determine whether the convergence condition is met. When the iteration converges, for each feature, the global probability distribution is calculated based on the probability distribution calculated by the computing nodes containing that feature. The condition for iterative convergence is that the Jensen-Shannon divergence between the global feature field distribution of the current round and the previous round is less than a threshold. The global feature field distribution is the product of the global probability distributions of all features.
[0087] Furthermore, the entropy-variance weighting mechanism is as follows:
[0088] in, Represents the coupling coefficient obtained in the t-th iteration; Representation of features global entropy; Representation of features global entropy; Representation of features variance between different computing nodes; representing features variance between different computing nodes; is a hyperparameter.
[0089] Further, the constructing step of the feature correlation graph comprises: calculating a feature mutual information matrix based on the data set by a mutual information formula; and constructing the feature correlation graph by taking each feature in the data set as a graph node and taking the value of the mutual information matrix as an edge weight.
[0090] Further, the condition of the iterative convergence is that the Jensen-Shannon divergence between the global feature field distribution of the current round and the global feature field distribution of the last round is less than a threshold value, and the global feature field distribution is the product of the global probability distribution of all features.
[0091] Further, the threshold value is: ; wherein, is an initial threshold value, is a decay rate, is the number of iterations.
[0092] Further, the distributed mean field equation is: ; wherein, denotes the probability distribution of the feature calculated by the th computing node at the th iteration; is a temperature hyperparameter; is a normalization factor of the th computing node; represents the serial number set of the neighbor features of the feature ; represents the average resultant force generated by all neighbor features on the feature on the th computing node; the partial derivative is the gradient of the energy function with respect to the feature .
[0093] It should be noted that each execution end in the embodiment corresponds to each step in Embodiment One, and the specific implementation process is the same, which will not be repeated here.
[0094] Embodiment Five The embodiment provides a computer device, which comprises: Figure 3As shown, it comprises a computer readable storage medium 1003, a processor 1001, a communication interface 1002 and a computer program stored in the computer readable storage medium 1003 and capable of running on the processor 1001, wherein the processor 1001, the communication interface 1002 and the computer readable storage medium 1003 are connected through a bus or other means. The communication interface 1002 is used for receiving and sending data, and the processor 1001 implements the steps of the large-scale data feature extraction method based on the average field when executing the program.
[0095] The above only describes the preferred embodiments of the present application and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for feature extraction from large-scale data based on mean field, characterized in that, Applied to parameter servers, including: Obtain the dataset, divide the feature association graph of the dataset into several subgraphs, divide the dataset into the same number of subsets as the subgraphs, each subset contains all data samples and only contains the features of the corresponding subgraph; Each subset is sent to a computing node, which then uses the coupling coefficients to compute the probability distribution, expectation, and entropy of each feature through the distributed mean-field equation. Based on the probability distribution, expectation, and entropy of each feature calculated by all computing nodes, the coupling coefficient is updated, and then the computing instructions and the updated coupling coefficient are sent to the computing nodes until the iteration converges. For each feature, the global probability distribution is calculated based on the probability distribution calculated by the computing nodes containing that feature.
2. The method for large-scale data feature extraction based on mean field as described in claim 1, characterized in that, The steps for updating the coupling coefficients include: for each feature, calculating the global entropy based on the entropy calculated by the computing node containing the feature; calculating the global expectation based on the expectation calculated by the computing node containing the feature; calculating the variance based on the global expectation and the expectation calculated by the computing node containing the feature; and updating the coupling coefficients based on the global entropy and variance through an entropy-variance weighted mechanism.
3. The method for large-scale data feature extraction based on mean field as described in claim 2, characterized in that, The entropy-variance weighting mechanism is as follows: in, Represents the coupling coefficient obtained in the t-th iteration; Representation of features global entropy; Representation of features global entropy; Representation of features Variance between different computation nodes; Representation of features Variance between different computation nodes; It's a hyperparameter.
4. The method for large-scale data feature extraction based on mean field as described in claim 1, characterized in that, The steps for constructing the feature association graph include: calculating the feature mutual information matrix based on the dataset using the mutual information formula; treating each feature in the dataset as a graph node and constructing the feature association graph using the value of the mutual information matrix as the edge weight.
5. The method for large-scale data feature extraction based on mean field as described in claim 1, characterized in that, The condition for the iteration to converge is that the Jensen-Shannon divergence between the global feature field distribution of the current round and the previous round is less than a threshold, and the global feature field distribution is the product of the global probability distributions of all features.
6. The method for large-scale data feature extraction based on mean field as described in claim 5, characterized in that, The threshold is: ;in, As the initial threshold, The attenuation rate, This represents the number of iterations.
7. The method for large-scale data feature extraction based on mean field as described in claim 1, characterized in that, The distributed mean-field equation is: ; in, Indicates the first During the round of iteration, the first Features computed by each computing node The probability distribution; It is a temperature hyperparameter; It is the first Normalization factor for each computing node; Representative characteristics The set of indices of the neighbor features; Representing the Features of all neighbors on each computing node Features The average resultant force produced; partial derivatives Energy function Features The gradient.
8. A method for feature extraction from large-scale data based on mean field, characterized in that, Applied to compute nodes, including: Receive a subset of parameters sent by the server; In response to the computation instructions sent by the parameter server and receiving the coupling coefficients sent by the parameter server, for each feature in the subset, the probability distribution, expectation, and entropy are calculated by combining the coupling coefficients and the distributed mean field equations, and sent to the parameter server so that the parameter server updates the coupling coefficients based on the probability distribution, expectation, and entropy of each feature calculated by all computing nodes until the iteration converges. For each feature, the global probability distribution is calculated based on the probability distribution calculated by the computing nodes containing that feature. The parameter server divides the feature association graph of the dataset into several subgraphs, and divides the dataset into the same number of subsets as the subgraphs. Each subset contains all data samples and only contains the features of the corresponding subgraph.
9. A large-scale data feature extraction system based on mean field, characterized in that, Includes a parameter server and multiple computing nodes; The parameter server is used to acquire the dataset, divide the feature association graph of the dataset into several subgraphs, divide the dataset into the same number of subsets as the subgraphs, and each subset contains all data samples and only contains the features of the corresponding subgraph. The computing node is used to receive a subset sent by the parameter server; respond to the computing instructions sent by the parameter server and receive the coupling coefficient sent by the parameter server; for each feature in the subset, in combination with the coupling coefficient, calculate the probability distribution, expectation and entropy through the distributed mean field equation; The parameter server is also used to update the coupling coefficients based on the probability distribution, expectation, and entropy of each feature calculated by all computing nodes, and then issue calculation instructions and updated coupling coefficients to the computing nodes until the iteration converges. For each feature, the global probability distribution is calculated based on the probability distribution calculated by the computing nodes containing that feature.
10. A computer device comprising a computer-readable storage medium, a processor, and a computer program stored on the computer-readable storage medium and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the method for large-scale data feature extraction based on mean field as described in any one of claims 1-7 or 8.
Citation Information
Patent Citations
Graph structure data node classification method and device in heterogeneous federated environment
CN117171628A
Unstructured data processing method and system
CN119513922A
Content marketing effect evaluation and optimization method and system based on deep learning
CN120450744A
Hybrid load flow calculation method and device based on cross entropy algorithm
CN120613731A
Joint estimation with space-time entropy regularization
US20200051255A1