Multi-view feature selection method and device based on multi-agent reinforcement learning

Through the multi-view feature selection method of multi-agent reinforcement learning, the synergy of multiple agents is used to optimize the feature selection process, and the problem of insufficient utilization of view dependencies and complementarity in the existing technology is solved, and efficient and accurate feature extraction and transparent selection paths are achieved.

CN120451716APending Publication Date: 2025-08-08SOUTHWESTERN UNIV OF FINANCE & ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510488082.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the prior art, the multi-view feature selection method fails to effectively utilize the dependencies of internal features of the view and the complementarity between views, resulting in a decrease in search efficiency of global optimal solutions and lacks transparency and interpretation capabilities.

Method used

A multi-view feature selection model is constructed using a method based on multi-agent reinforcement learning. By selecting features based on shared state and own behavioral strategies of multiple agents, combining Pearson's correlation coefficient, utility function and global-local joint reward mechanism, the feature selection process is optimized.

Benefits of technology

It improves the efficiency and transparency of multi-view feature selection, significantly improves the accuracy and robustness of feature extraction, and can explain the selection path of the optimal subset.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451716A_ABST
    Figure CN120451716A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-view feature selection method and device based on multi-agent reinforcement learning, and relates to the technical field of machine learning. The method comprises the following steps: firstly, acquiring multi-view sample feature data from different sources and preprocessing the multi-view sample feature data; then, a multi-view feature selection model is constructed and trained, the multi-view feature selection model comprises a plurality of intelligent agents, and the intelligent agents are constructed based on different views and used for selecting features in the views to which the intelligent agents belong according to the sharing state in combination with the behavior strategy of the intelligent agents and generating corresponding feature subsets and updated states; and finally, inputting the preprocessed multi-view sample feature data into the multi-view feature selection model for feature selection, and outputting a target multi-view feature selection result. Efficient and accurate selection of multi-view features is achieved through the synergistic effect of multiple agents, each agent independently formulates a strategy while sharing state information, and the accuracy and robustness of feature extraction are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine learning technology, and in particular to a multi-agent based multi-view feature selection method and device. Background Art

[0002] Multi-view data typically refers to different feature representations of the same object obtained from multiple angles or sources. For example, in image processing, the same image may contain pixel-level features, edge features, texture features, and more. This multi-view data is often complementary, correlated, and redundant. How to effectively integrate this information is a core issue that needs to be addressed in multi-view feature selection.

[0003] Existing techniques typically construct a unified feature space by simply concatenating data from each view, then design a reinforcement learning framework for heuristic search within this space. However, this approach ignores the dependencies between features within views and the complementarity and correlation between views, failing to fully exploit the unique structural properties of multi-view data, resulting in reduced efficiency in searching for the global optimal solution. Furthermore, existing multi-view feature selection methods lack effective explanations for feature subset selection and are unable to clearly explain the reasons for selecting specific features. This results in a lack of transparency in the feature selection process, making it difficult to provide a clear basis for subsequent tasks.

[0004] Therefore, in the related technology, there is an urgent need for a method that can improve the efficiency, transparency and credibility of multi-view feature selection. Summary of the Invention

[0005] Based on this, it is necessary to provide a multi-view feature selection method and device based on multi-agent reinforcement learning that can improve the efficiency, transparency and credibility of multi-view feature selection in response to the above technical problems.

[0006] In a first aspect, the present application provides a multi-view feature selection method based on multi-agent reinforcement learning. The method comprises: Obtain multi-view sample feature data from different sources and perform preprocessing; Constructing and training a multi-view feature selection model, wherein the multi-view feature selection model includes multiple agents, each constructed based on different views, configured to select features in the corresponding view based on a shared state combined with its own behavior strategy, and generate corresponding feature subsets and updated states; The preprocessed multi-view sample feature data is input into the multi-view feature selection model for feature selection, and a target multi-view feature selection result is output.

[0007] Optionally, in one embodiment of the present application, constructing and training a multi-view feature selection model includes: The joint state representation function based on the relevant transformation determines the shared state of the agents.

[0008] Optionally, in one embodiment of the present application, determining the agent shared state based on the joint state representation function of the relevant transformation includes: Calculating the Pearson correlation coefficients between the features in the selected view sample feature set, and constructing a correlation matrix based on the Pearson correlation coefficients; Performing weighting and normalization processing on the correlation matrix to obtain a weight matrix; Performing a correlation transformation on the feature subspace based on the weight matrix to obtain a transformation feature matrix; The average of each row of the transformation feature matrix is calculated to determine the agent shared state.

[0009] Optionally, in one embodiment of the present application, constructing and training a multi-view feature selection model further includes: The utility function is used to evaluate the prediction performance and information redundancy of the feature subspace, and the utility function is expressed as:

[0010]

[0011] in, Representing feature subspace Prediction accuracy in downstream tasks, is the dimension of the feature subspace, To adjust the parameters, it is used to balance the impact of prediction accuracy and redundancy. For quantitative features and The mutual information measurement function of the redundant relationship between Representation characteristics and The joint distribution function of and Represents characteristics and The marginal distribution of Representing feature subspace The degree of redundancy.

[0012] Optionally, in one embodiment of the present application, constructing and training a multi-view feature selection model further includes: Based on the global-local joint reward mechanism, the local reward value of the feature subspace explored by each agent and the global reward value of the joint feature space composed of the feature subspaces explored by all agents are calculated.

[0013] Optionally, in one embodiment of the present application, a global reward for view feature selection is allocated based on the principle of more work, more pay.

[0014] In a second aspect, the present application also provides a multi-view feature selection device based on multi-agent reinforcement learning. The device comprises: Feature data acquisition and preprocessing module, used to acquire and preprocess feature data of multi-view samples from different sources; A multi-view feature selection model construction and training module is used to build and train a multi-view feature selection model. The multi-view feature selection model includes multiple agents, which are built based on different views and are used to select features in their respective views based on shared states and their own behavior strategies, and generate corresponding feature subsets and updated states. The multi-view feature selection module is used to input the pre-processed multi-view sample feature data into the multi-view feature selection model for feature selection, and output the target multi-view feature selection result.

[0015] The multi-view feature selection method and apparatus based on multi-agent reinforcement learning first obtains and preprocesses multi-view sample feature data from different sources. Next, a multi-view feature selection model is constructed and trained. The multi-view feature selection model comprises multiple agents, each constructed based on different views, which select features from their respective views based on a shared state and their own behavioral strategies, generating corresponding feature subsets and updated states. Finally, the pre-processed multi-view sample feature data is input into the multi-view feature selection model for feature selection, outputting the target multi-view feature selection results. In other words, by leveraging the synergy of multiple agents, efficient and accurate multi-view feature selection is achieved. Each agent independently formulates its own strategy while sharing state information, significantly improving the accuracy and robustness of feature extraction. Secondly, a joint reward mechanism and an optimized reward distribution scheme are introduced to effectively balance local and global utility, ensuring the preservation of key features while effectively eliminating redundant information. Through an experience replay mechanism and policy optimization using a Deep Q-Network (DQN), the learning efficiency and decision-making ability of the agents are improved, thereby enabling the selection of high-quality feature subsets. In addition, the selection path of the optimal subset can be explained from the trajectory of the agent's interaction with the environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 This is a diagram illustrating an application environment of a multi-view feature selection method based on multi-agent reinforcement learning in one embodiment; Figure 2 1 is a flow chart of a multi-view feature selection method based on multi-agent reinforcement learning in one embodiment; Figure 3 A schematic diagram of a data structure in one embodiment; Figure 4 Schematic diagram of the structure of a multi-view feature selection model in one embodiment; Figure 5 Schematic diagram of the calculation process of the joint state representation function of the relevant transformations in one embodiment; Figure 6 1 is a structural block diagram of a multi-view feature selection device based on multi-agent reinforcement learning in one embodiment; Figure 7 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0017] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0018] The embodiment of the present application provides a multi-view feature selection method based on multi-agent reinforcement learning, which can be applied to Figure 1 In the application environment shown, the terminal communicates with the server through the network. The data storage system can store data that the server needs to process. The data storage system can be integrated on the server or placed on the cloud or other network servers. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can be smart speakers, smart TVs, smart air conditioners, smart car devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server can be implemented as a standalone server or a server cluster consisting of multiple servers.

[0019] In one embodiment, Figure 2 As shown in the figure, a multi-view feature selection method based on multi-agent reinforcement learning is provided. Figure 1 The following steps are used as an example to illustrate the server in the example: S201: Acquire multi-view sample feature data from different sources and perform preprocessing.

[0020] In the embodiment of the present application, first, multi-view sample feature data from different sources is obtained, and each view contains the same number of samples, and each view describes the features of these samples from a different perspective. For example, a single sample may contain sound, fingerprint, and iris features at the same time, and each view provides different information related to the sample. These multi-view sample feature data are pre-processed according to their source and type, and presented in a table form, such as Figure 3 As shown in the figure, for image modalities, image features are extracted based on color, shape, and texture. Specifically, color features are extracted by calculating the RGB color histogram and statistically analyzing the distribution characteristics of each color channel in the image to capture the global color information of the image. Shape features are based on the Hu moment, extracting invariant moment features to describe the image shape and maintain invariance to transformations such as rotation, scaling, and mirroring. Texture features use the Histogram of Oriented Gradients (HOG) method to enhance the depiction of image edges and contours by calculating the distribution of gradient directions within a local region. For acoustic modalities (such as speech or ambient sound signals), the signal is first pre-emphasized to enhance the signal energy in the high-frequency portion. It is then segmented into short-term frames (for example, using a 25ms frame length and a 10ms frame shift), and a window function (such as a Hamming window) is applied to each frame to reduce the impact of spectral leakage. Next, a Fast Fourier Transform (FFT) is calculated to extract feature information for the first 100 frequency points to capture the frequency distribution of the acoustic signal. After the above preprocessing process, the unstructured data of all modalities are converted into structured tabular data, ready to be input into the model for feature selection.

[0021] S203: Constructing and training a multi-view feature selection model. The multi-view feature selection model includes multiple agents. The agents are constructed based on different views and are used to select features in their respective views based on shared states and their own behavior strategies, and generate corresponding feature subsets and updated states. In the embodiment of the present application, for the multi-view feature selection problem, an independent agent is constructed for each view. In the process of interacting with the environment, each agent dynamically adjusts the feature set based on the "exploration-exploitation" strategy, that is, adds or removes a feature from the selected feature set. This setting allows the agent to gradually optimize the feature selection decision to adapt to the feature distribution of different views. And based on all the agents, such as Figure 4 The multi-view feature selection model shown in the figure mainly includes the control and training processes. The control process is responsible for each agent performing feature selection operations based on shared state information and its own behavioral strategy, and generating new feature subsets and states. The training process continuously optimizes each agent's strategy through the experience replay mechanism and the Deep Q-Network (DQN), continuously improving its feature selection decision-making ability.

[0022] The action space design of the agent takes into account the sophistication of feature selection and the complexity of the overall reinforcement learning framework. view, the action space of the agent responsible for this view is ,in Is a view The number of features of action Indicates that the agent maintains the current view The selection state of the feature subspace remains unchanged, while the action (for ) indicates that the agent changes the The selection state of the feature. In the step, the agent The action performed is , then the characteristic subspace Remain unchanged. If the action performed is , then the agent switches to The selection status of a feature: If the feature Previously selected ( ), then cancel the selection, and the updated feature subspace is ; If the feature Not previously selected ( ), then select this feature, then the next feature subspace after the selection is completed is .

[0023] Multiple agents are responsible for different views. Each agent is based on the shared state (i.e., comprehensive representation of feature subsets) and its strategy Select features by action Change the state of the feature in the feature subset. After the selection is completed, the feature set of all views Combine into new feature subsets , and express the function through the state Generate a new state , the state is passed to all agents as shared information. Through the utility function Evaluate the overall utility of a subset of features and use it in the reward function In the reward distribution mechanism, each agent is rewarded. Specifically, each agent's reward consists of two parts: the utility of the local feature subset associated with the view for which the agent is responsible, and the utility of the global feature subset. This reward mechanism not only considers the utility of the local view feature set (i.e., the agent's performance on the view it is responsible for), but also balances the agent's contribution to the utility of the global feature subset. Therefore, based on this reward, the agent can achieve a balance between exploration and exploitation in the view it controls, while also promoting the selection of the optimal feature subset from a global perspective.

[0024] Each agent controls the data generated by the interaction with the environment The data is stored in the experience replay pool. To break the correlation of time series data and improve training stability, the experience replay strategy randomly extracts some historical interaction trajectories. This data is used to update the agent's policy network (DQN) through backpropagation and gradient descent, enabling it to make better choices in subsequent iterations. Throughout the training process, the control module and the training module are alternating. In the control module, each agent performs feature selection operations and interacts with the environment to generate experience data. In the training module, the agent uses the stored experience data through the DQN learning mechanism to continuously optimize its policy. As training progresses, the agent's policy network gradually converges, enabling it to make better decisions during future feature selection processes.

[0025] Specifically, the agent's strategy Updated by DQN to optimize the exploration decision of feature subspace. The core idea of DQN is to use action value function To guide the agent in different states Select the best action , to maximize long-term returns.

[0026] Strategy The design adopts - Greedy strategy, as follows:

[0027] This strategy is to The probability of selecting the currently estimated optimal action to be used is The probability of exploration is 2, which ensures that the agent can not only use existing experience but also discover new effective actions.

[0028] DQN uses a neural network to approximate the action-value function , which is updated through the following loss function:

[0029] in, is the target value calculated based on the Bellman equation, that is, the current reward Add the next step status The estimated value of the future cumulative discounted reward. The objective function network Is a training network The network with the same structure but fixed parameters, every Step and train the network By minimizing this loss function, the agent can gradually update its action-value function, allowing the policy to optimize its decisions based on experience.

[0030] In one embodiment of the present application, constructing and training a multi-view feature selection model includes: The joint state representation function based on the relevant transformation determines the shared state of the agents.

[0031] In one embodiment of the present application, all views in the The characteristic subspace obtained at each moment is considered as a state in a Markov decision process (MDP) where It is The view in The feature set selected in step 1 is Indicates the horizontal concatenation of two matrices with the same number of rows. In order to ensure that the length of the state representation does not change over time, a joint state representation function of related transformations is proposed. Based on the joint state representation function of related transformations, the shared state of the agent is determined. While extracting the feature set information, it ensures that no matter how the size of the feature set changes, the obtained state vector always maintains a fixed length, thereby ensuring the consistency of information transmission and the stability of the feature selection training process. The specific calculation process of the function is as follows: Figure 5 shown.

[0032] Specifically, in one embodiment of the present application, determining the agent shared state based on the joint state representation function of the relevant transformation includes: S301: Calculate the Pearson correlation coefficients between the features in the selected view sample feature set, and construct a correlation matrix based on the Pearson correlation coefficients.

[0033] S303: Perform weighting and normalization processing on the correlation matrix to obtain a weight matrix.

[0034] S305: Performing correlation transformation on the feature subspace based on the weight matrix to obtain a transformation feature matrix.

[0035] S307: Calculate the average of each row of the transformation feature matrix to determine the agent sharing state.

[0036] In one embodiment of the present application, first, the feature set selected for all views is Calculate the correlation between each feature in pairwise and obtain the Pearson correlation coefficient (where is the sample size, for The number of features selected in all views at that moment).

[0037]

[0038] in, , , Indicates the The sample in The value of each feature. According to the above formula, the correlation matrix can be constructed ,in .

[0039] To eliminate the influence of autocorrelation, the matrix The diagonal elements of are set to 0. Then, the elements of each row are normalized to obtain the normalized weight matrix .

[0040] Then use the weight matrix For the characteristic subspace Perform relevant transformations to obtain the transformation feature matrix:

[0041] Transformation feature matrix Contains the information between the features after fusion processing. Dynamic changes, so the matrix needs to be further By averaging each row of .

[0042] The specific form of the joint state representation function is as follows:

[0043] in, is the weight matrix No. List, , the length of the state vector does not change with time. As a state shared by all agents, it can effectively capture global information from different views, ensuring that during the feature selection process, the agent not only considers local information but also makes full use of the complementarity between views, thereby improving the accuracy and efficiency of feature selection.

[0044] In one embodiment of the present application, constructing and training a multi-view feature selection model further includes: The utility function is used to evaluate the prediction performance and information redundancy of the feature subspace, and the utility function is expressed as:

[0045]

[0046] in, Representing feature subspace Prediction accuracy in downstream tasks, is the dimension of the feature subspace, To adjust the parameters, it is used to balance the impact of prediction accuracy and redundancy. For quantitative features and The mutual information measurement function of the redundant relationship between Representation characteristics and The joint distribution function of and Represents characteristics and The marginal distribution of Representing feature subspace The degree of redundancy.

[0047] In one embodiment of the present application, in order to effectively evaluate the quality of the feature subspace selected by each agent, a utility function is designed. , which is used to comprehensively evaluate the prediction performance and information redundancy of the feature subspace. The function is designed to balance the discriminative power of the feature subspace in downstream tasks and the degree of redundancy between features. The evaluation result is fed back to the agent through the reward function to guide it to select a more effective feature combination. Its specific expression is as follows:

[0048]

[0049] in, Representing feature subspace Prediction accuracy in downstream tasks, is the dimension of the feature subspace, To adjust the parameters, it is used to balance the impact of prediction accuracy and redundancy. For quantitative features and The mutual information measurement function of the redundant relationship between Representation characteristics and The joint distribution function of and Represents characteristics and The marginal distribution of Representing feature subspace The degree of redundancy.

[0050] By adjusting the parameters , The function can achieve a dynamic balance between prediction accuracy and redundancy. This balance mechanism further affects the reward function value, thereby optimizing the overall performance of feature selection in the multi-agent system.

[0051] In one embodiment of the present application, constructing and training a multi-view feature selection model further includes: Based on the global-local joint reward mechanism, the local reward value of the feature subspace explored by each agent and the global reward value of the joint feature space composed of the feature subspaces explored by all agents are calculated.

[0052] In one embodiment of the present application, in order to promote information interaction between agents, a global-local joint reward mechanism is designed to optimize the collaboration and competition of multiple agents in multi-view feature selection. Each agent will not only select the best feature subspace based on its own exploration, but also select the best feature subspace based on its own exploration. Obtaining local rewards is also based on the joint feature subspace searched by all agents at time t Get global rewards.

[0053] First, construct local rewards To measure the agent at time t The exploration results in its corresponding view, this score combines the feature set redundancy (based on mutual information calculation) and downstream task performance (such as classification accuracy ACC). Its calculation formula is:

[0054] Second, build global rewards Used to reflect the joint feature space explored by all agents at time t The overall performance is calculated as follows:

[0055] In order to achieve a reasonable distribution of rewards, the following distribution strategy is proposed:

[0056] in, Representing an agent The comprehensive rewards obtained, To adjust the parameters, control the ratio of local rewards to global rewards.

[0057] In one embodiment of the present application, a global reward for view feature selection is allocated based on the principle of more work, more pay.

[0058] In one embodiment of the present application, the final reward of each agent consists of two parts: the score of the feature set selected by itself And global rewards The global reward distribution follows the principle of more work, more pay - if a view's feature selection contribution is large (i.e. its is higher), the view This mechanism ensures that the agent not only optimizes the feature selection of its own view, but also optimizes towards the overall optimal feature set. This not only promotes the optimization of each view, but also improves the overall performance of the multi-view joint feature space, ultimately improving the efficiency and quality of feature selection.

[0059] S205: Inputting the pre-processed multi-view sample feature data into the multi-view feature selection model to perform feature selection, and outputting a target multi-view feature selection result.

[0060] In an embodiment of the present application, preprocessed multi-view sample feature data is input into a multi-view feature selection model, which efficiently explores the feature space to obtain the optimal feature subset. The model then outputs the target multi-view feature selection results, enabling downstream tasks such as speech recognition, sentiment analysis, classification, clustering, or prediction to be performed, effectively improving the performance of these tasks. In specific applications, this model can be used for medical diagnosis and analysis. In disease diagnosis, multi-view features include imaging data (MRI / CT), gene expression data, clinical records, and pathology reports. The selected features can help improve medical diagnostic modeling. Anomaly detection can also be performed, such as using selected multimodal features to identify abnormal transactions in finance. Environmental monitoring or weather forecasting can also be performed. Multimodal data may be composed of multiple sensors or satellite imagery data. The selected multi-view features can be used for weather forecasting, natural disaster prediction, and other applications.

[0061] In the multi-view feature selection method based on multi-agent reinforcement learning, first, multi-view sample feature data from different sources is acquired and preprocessed. Next, a multi-view feature selection model is constructed and trained. The multi-view feature selection model includes multiple agents, each constructed based on different views, which select features in their respective views based on a shared state combined with their own behavioral strategies, generating corresponding feature subsets and updated states. Finally, the preprocessed multi-view sample feature data is input into the multi-view feature selection model for feature selection, outputting the target multi-view feature selection results. In other words, by leveraging the synergy of multiple agents, efficient and accurate multi-view feature selection is achieved. Each agent independently formulates its own strategy while sharing state information, significantly improving the accuracy and robustness of feature extraction. Secondly, a joint reward mechanism and an optimized reward distribution scheme are introduced to effectively balance local and global utility, ensuring the preservation of key features while effectively eliminating redundant information. Through the experience replay mechanism and DQN strategy optimization, the learning efficiency and decision-making ability of the agents are improved, thereby achieving the selection of high-quality feature subsets. In addition, the selection path of the optimal subset can be explained from the trajectory of the agent's interaction with the environment.

[0062] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0063] Based on the same inventive concept, the embodiments of the present application also provide a multi-view feature selection device based on multi-agent reinforcement learning for implementing the multi-view feature selection method based on multi-agent reinforcement learning mentioned above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more embodiments of a multi-view feature selection device based on multi-agent reinforcement learning provided below can be found in the above limitations of a multi-view feature selection method based on multi-agent reinforcement learning, and will not be repeated here.

[0064] In one embodiment, Figure 6 As shown, a multi-view feature selection device 600 based on multi-agent reinforcement learning is provided, comprising: a feature data acquisition and preprocessing module 601, a multi-view feature selection model building and training module 603 and a multi-view feature selection module 605, wherein: The feature data acquisition and preprocessing module 601 is used to acquire feature data of multi-view samples from different sources and perform preprocessing.

[0065] The multi-view feature selection model construction and training module 603 is used to construct and train a multi-view feature selection model. The multi-view feature selection model includes multiple intelligent agents, which are constructed based on different views and are used to select features in the view to which they belong based on the shared state and their own behavior strategies, and generate corresponding feature subsets and updated states.

[0066] The multi-view feature selection module 605 is configured to input the pre-processed multi-view sample feature data into the multi-view feature selection model for feature selection, and output a target multi-view feature selection result.

[0067] In one embodiment of the present application, the multi-view feature selection model building and training module is further used to: The joint state representation function based on the relevant transformation determines the shared state of the agents.

[0068] In one embodiment of the present application, the multi-view feature selection model building and training module is further used to: Calculating the Pearson correlation coefficients between the features in the selected view sample feature set, and constructing a correlation matrix based on the Pearson correlation coefficients; Performing weighting and normalization processing on the correlation matrix to obtain a weight matrix; Performing a correlation transformation on the feature subspace based on the weight matrix to obtain a transformation feature matrix; The average of each row of the transformation feature matrix is calculated to determine the agent shared state.

[0069] In one embodiment of the present application, the multi-view feature selection model building and training module is further used to: The utility function is used to evaluate the prediction performance and information redundancy of the feature subspace, and the utility function is expressed as:

[0070]

[0071] in, Representing feature subspace Prediction accuracy in downstream tasks, is the dimension of the feature subspace, To adjust the parameters, it is used to balance the impact of prediction accuracy and redundancy. For quantitative features and The mutual information measurement function of the redundant relationship between Representation characteristics and The joint distribution function of and Represents characteristics and The marginal distribution of Representing feature subspace The degree of redundancy.

[0072] In one embodiment of the present application, the multi-view feature selection model building and training module is further used to: Based on the global-local joint reward mechanism, the local reward value of the feature subspace explored by each agent and the global reward value of the joint feature space composed of the feature subspaces explored by all agents are calculated.

[0073] In one embodiment of the present application, a global reward for view feature selection is allocated based on the principle of more work, more pay.

[0074] Each module in the aforementioned multi-view feature selection device based on multi-agent reinforcement learning can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0075] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 7 As shown. The computer device includes a processor, memory, a communication interface, a display screen, and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal via wired or wireless communication. The wireless communication can be achieved via Wi-Fi, a mobile cellular network, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a multi-view feature selection method based on multi-agent reinforcement learning. The display screen of the computer device can be a liquid crystal display or an electronic ink display. The input device of the computer device can be a touch layer covering the display screen, or keys, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse.

[0076] Those skilled in the art will understand that Figure 7 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0077] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0078] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0079] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0080] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A multi-view feature selection method based on multi-agent reinforcement learning, characterized in that: The method comprises: Obtain multi-view sample feature data from different sources and perform preprocessing; Constructing and training a multi-view feature selection model, wherein the multi-view feature selection model includes multiple agents, each constructed based on different views, configured to select features in the corresponding view based on a shared state combined with its own behavior strategy, and generate corresponding feature subsets and updated states; The preprocessed multi-view sample feature data is input into the multi-view feature selection model for feature selection, and a target multi-view feature selection result is output.

2. A multi-view feature selection method based on multi-agent reinforcement learning according to claim 1, characterized in that: The constructing and training of a multi-view feature selection model includes: The joint state representation function based on the relevant transformation determines the shared state of the agents.

3. The multi-view feature selection method based on multi-agent reinforcement learning according to claim 2, characterized in that: The determination of the agent shared state by the joint state representation function based on the relevant transformation includes: Calculating the Pearson correlation coefficients between the features in the selected view sample feature set, and constructing a correlation matrix based on the Pearson correlation coefficients; Performing weighting and normalization processing on the correlation matrix to obtain a weight matrix; Performing a correlation transformation on the feature subspace based on the weight matrix to obtain a transformation feature matrix; The average of each row of the transformation feature matrix is calculated to determine the agent shared state.

4. The multi-view feature selection method based on multi-agent reinforcement learning according to claim 1, characterized in that: The constructing and training of the multi-view feature selection model further includes: The utility function is used to evaluate the prediction performance and information redundancy of the feature subspace, and the utility function is expressed as: in, Representing feature subspace Prediction accuracy in downstream tasks, is the dimension of the feature subspace, To adjust the parameters, used to balance the impact of prediction accuracy and redundancy, For quantitative features and The mutual information measurement function of the redundant relationship between Representation characteristics and The joint distribution function of and Represents characteristics and The marginal distribution of Representing feature subspace The degree of redundancy.

5. The multi-view feature selection method based on multi-agent reinforcement learning according to claim 1, characterized in that: The constructing and training of the multi-view feature selection model further includes: Based on the global-local joint reward mechanism, the local reward value of the feature subspace explored by each agent and the global reward value of the joint feature space composed of the feature subspaces explored by all agents are calculated.

6. The multi-view feature selection method based on multi-agent reinforcement learning according to claim 5, characterized in that: Global rewards for view feature selection are distributed based on the principle of more work, more pay.

7. A multi-view feature selection device based on multi-agent reinforcement learning, characterized in that: The device comprises: Feature data acquisition and preprocessing module, used to acquire and preprocess feature data of multi-view samples from different sources; A multi-view feature selection model construction and training module is used to build and train a multi-view feature selection model. The multi-view feature selection model includes multiple agents, which are built based on different views and are used to select features in their respective views based on shared states and their own behavior strategies, and generate corresponding feature subsets and updated states. The multi-view feature selection module is used to input the pre-processed multi-view sample feature data into the multi-view feature selection model for feature selection, and output the target multi-view feature selection result.