Longitudinal federated learning-based principal component analysis method and related device

By generating the same random Gaussian matrix and performing matrix operations in longitudinal federated learning, the contradiction between data privacy protection and dimensionality reduction is resolved. This enables secure data processing and efficient dimensionality reduction across devices, improving model training efficiency and accuracy. It is applicable to fields such as healthcare, finance, intelligent recommendation, and industrial forecasting.

CN120849933APending Publication Date: 2025-10-28BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510840484.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In vertical federated learning, how to effectively integrate the covariance information of multiple data sets with different characteristics to achieve data privacy protection and extract the intrinsic correlation of data of different dimensions of different participants through principal component analysis for data dimensionality reduction.

Method used

By generating the same random Gaussian matrix, the business data matrix is ​​multiplied in each communication device to obtain an intermediate matrix. After concatenation and orthogonalization, the decomposition matrix is ​​extracted, singular value decomposition is performed, a dimension-reduced matrix is ​​generated, and data privacy is ensured through encryption and hash verification, thereby achieving secure transmission and processing of data across devices.

Benefits of technology

While protecting data privacy, it achieves efficient data dimensionality reduction, improves model training efficiency and accuracy, and reduces data transmission and computational pressure, making it suitable for fields such as healthcare, finance, intelligent recommendation, and industrial forecasting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120849933A_ABST
    Figure CN120849933A_ABST
Patent Text Reader

Abstract

A principal component analysis method based on longitudinal federated learning is applied to first communication equipment and comprises the steps that random Gaussian matrixes are generated, and the random Gaussian matrixes are the same in the first communication equipment and second communication equipment; multiplying the service data matrix of the first communication equipment by the random Gaussian matrix to obtain a first intermediate matrix; acquiring a second intermediate matrix sent by a second communication device; splicing the first intermediate matrix and the at least one second intermediate matrix to obtain a spliced matrix; orthogonalizing the spliced matrix to obtain an orthogonal matrix; extracting a first decomposition matrix from the orthogonal matrix based on the position relation of the first intermediate matrix in the splicing matrix; multiplying the transpose of the first decomposition matrix by a service data matrix to obtain a first switching matrix; acquiring a second switching matrix sent by second traffic equipment; fusing the first switching matrix and the second switching matrix to obtain a fusion matrix; and performing accurate singular value decomposition on the fusion matrix to obtain left and right singular dimension reduction matrixes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of privacy computing, specifically to privacy computing based on federated learning, and particularly to a principal component analysis method and related apparatus based on longitudinal federated learning. Background Art

[0002] In today's data-driven intelligent systems, the high dimensionality and distributed nature of data are becoming increasingly prominent, especially in fields such as healthcare, finance, and the Internet of Things. Traditional data processing methods often rely on centralized data collection and analysis, which not only faces the risk of data privacy breaches but may also violate relevant data protection regulations.

[0003] Federated learning, as an emerging distributed machine learning framework, allows multiple data holders to collaboratively train models without sharing the original data, thereby achieving collaborative modeling while protecting data privacy. In horizontal federated learning scenarios, participants possess datasets with the same feature space but different samples, while in vertical federated learning scenarios, participants possess datasets with the same sample space but different features. Summary of the Invention

[0004] This disclosure provides a principal component analysis method and related apparatus based on longitudinal federated learning.

[0005] According to the first aspect, a principal component analysis method based on longitudinal federated learning is provided, characterized in that the method includes:

[0006] Generate a random Gaussian matrix, wherein the random Gaussian matrix is ​​identical in both the first communication device and the second communication device;

[0007] Multiply the random Gaussian matrix by the business data matrix to obtain the intermediate matrix;

[0008] Based on the malicious behavior, the driver is guided to avoid risks.

[0009] Optionally, generating a random Gaussian matrix includes:

[0010] Obtain a random seed, which is the same in both the first communication device and the second communication device;

[0011] A random Gaussian matrix is ​​generated based on a pseudo-random number generator.

[0012] Optionally, the generation of the random Gaussian matrix further includes:

[0013] Obtain the hash value of the random Gaussian matrix;

[0014] Verify that the hash value is the same in both the first communication device and the second communication device.

[0015] Optionally, obtaining the second intermediate matrix sent by the second communication device includes:

[0016] Obtain the encrypted data sent by the second communication device;

[0017] Decrypt the encrypted data to obtain the second intermediate matrix.

[0018] Optionally, the step of concatenating the first intermediate matrix with at least one second intermediate matrix to obtain a concatenated matrix includes:

[0019] Determine the order in which the first intermediate matrix and at least one of the second intermediate matrices are arranged;

[0020] Based on the stated arrangement order, the first intermediate matrix and at least one second intermediate matrix are combined to obtain a spliced ​​matrix.

[0021] Optionally, the step of extracting the first decomposition matrix from the orthogonal matrix based on the positional relationship of the first intermediate matrix in the concatenated matrix includes:

[0022] Based on the arrangement order, obtain the start and end row numbers of the first intermediate matrix in the spliced ​​matrix;

[0023] The first decomposition matrix is ​​extracted from the orthogonal matrix based on the start and end row numbers.

[0024] According to the second aspect, a principal component analysis apparatus based on longitudinal federated learning is provided, characterized in that the apparatus comprises:

[0025] A random Gaussian matrix generation module is used to generate a random Gaussian matrix, wherein the random Gaussian matrix is ​​the same in both the first communication device and the second communication device;

[0026] The first intermediate matrix acquisition module is used to multiply the service data matrix of the first communication device with the random Gaussian matrix to obtain the first intermediate matrix;

[0027] An orthogonal matrix acquisition module is used to acquire a second intermediate matrix sent by the second communication device; concatenate the first intermediate matrix with at least one of the second intermediate matrices to obtain a concatenated matrix; and orthogonalize the concatenated matrix to obtain an orthogonal matrix.

[0028] The first exchange matrix acquisition module is used to extract a first decomposition matrix from the orthogonal matrix based on the positional relationship of the first intermediate matrix in the concatenated matrix; and multiply the transpose of the first decomposition matrix by the business data matrix to obtain the first exchange matrix.

[0029] The fusion matrix acquisition module is used to acquire the second exchange matrix sent by the second transportation equipment; and to add the first exchange matrix and the second exchange matrix to obtain the fusion matrix.

[0030] The singular value decomposition module is used to perform precise singular value decomposition on the fusion matrix to obtain left and right singular dimension-reduced matrices.

[0031] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method as described in any implementation of the first aspect.

[0032] According to a fourth aspect, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing a computer to perform the method described in any implementation of the first aspect.

[0033] According to a fifth aspect, a computer program product is provided, including a computer program that, when executed by a processor, implements the method as described in any implementation of the first aspect.

[0034] The technical solution disclosed herein achieves accurate identification and early warning of malicious behavior through real-time monitoring and in-depth analysis of external vehicle traffic information. It innovatively integrates emotional reassurance and intelligent risk avoidance guidance mechanisms, alleviating driver psychological stress while providing dynamic and scenario-based risk avoidance strategies. This significantly reduces driving risks and operational errors caused by external malicious behavior, effectively improving driving safety and the driving experience.

[0035] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Attached Figure Description

[0036] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0037] Figure 1 This is a flowchart of a principal component analysis method based on longitudinal federated learning provided in an embodiment of this disclosure;

[0038] Figure 2 This is an interactive schematic diagram of a principal component analysis method based on longitudinal federated learning provided in an embodiment of this disclosure;

[0039] Figure 3 This is a schematic diagram of a principal component analysis device based on longitudinal federated learning provided in an embodiment of this disclosure;

[0040] Figure 4 This is a block diagram of an electronic device used to implement the principal component analysis method based on longitudinal federated learning in the embodiments of this disclosure. DETAILED DESCRIPTION

[0041] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0042] In today's privacy-preserving computing field, large-scale business models built upon business data heavily rely on massive amounts of data. Distributed storage and processing models adapt to diverse application scenarios, widely covering areas such as joint diagnostics, financial risk control, intelligent customer service, and industrial forecasting. Faced with the massive volume resulting from the rapid growth of multimodal data, the pressure on data storage, transmission, and computation is increasing dramatically. How to efficiently process data while strictly ensuring data privacy and security has become crucial for optimizing model training. To address this, our proposed method reduces the dimensionality of distributed multimodal data, ensuring data privacy and security while decreasing data complexity. The resulting dimensionality reduction matrix can be directly applied to the distributed training tasks of business models, effectively improving training efficiency and helping models achieve more accurate and efficient predictions and decisions in complex business scenarios.

[0043] Business data is widely distributed across various fields, encompassing both structured and unstructured information. Structured data, such as transaction records and communication behavior data in the financial industry, is stored in a well-organized tabular format, facilitating data retrieval and analysis. Unstructured data, on the other hand, includes patient CT images and genetic data in the medical field, user browsing history texts on e-commerce platforms, and sensor waveform data such as equipment vibration frequencies and temperature curves in industrial settings. This type of data has diverse formats and lacks a unified structure.

[0044] Within the vertical federated learning framework, leveraging techniques such as principal component analysis, feature extraction can be achieved that balances privacy protection and efficient dimensionality reduction for both structured and unstructured business data. This processed business data is then applied to scenarios such as collaborative medical diagnosis, financial risk control, intelligent recommendation, and industrial forecasting. This approach breaks down data barriers between institutions to achieve collaboration while ensuring privacy and security through localized data processing, significantly improving business efficiency and service quality across various sectors.

[0045] This disclosure proposes a principal component analysis method and related apparatus based on longitudinal federated learning.

[0046] Principal component analysis (PCA) is a common data preprocessing method for dimensionality reduction. However, when PCA is used in longitudinal federated learning, it requires centralized processing of data from all participants, which is detrimental to protecting the data privacy of each participant.

[0047] The urgent problem to be solved is how to effectively integrate the covariance information of different feature datasets from multiple parties in vertical federated learning, so as to both protect data privacy and extract the intrinsic correlation of data from different participants in different dimensions through principal component analysis for data dimensionality reduction.

[0048] This method can be integrated into various privacy computing devices, including but not limited to central control systems, computing units, smart terminals, sensor fusion terminals, and other terminal devices with data processing capabilities, as well as server systems such as local servers and vehicle-to-everything (V2X) cloud servers (whether single servers or clustered server systems). This method does not rely on a specific hardware platform or software architecture. During execution, whether it is an independently running terminal device or a terminal-server architecture working collaboratively through the V2X network, the method disclosed herein can be effectively utilized. When the terminal device executes independently, the method can meet offline computing needs without relying on an external network. In scenarios requiring higher processing performance or broader data resource support, the method of this invention also supports communication between the terminal device and the V2X server, leveraging the powerful computing capabilities and abundant data resources of the cloud to jointly complete the method. This method can adapt to different operating systems and platform environments, including various dedicated operating systems, embedded real-time operating systems, vehicle-to-everything (V2X) systems, and various server-side operating systems.

[0049] It is understood that the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0050] It is understood that the various embodiments covered in this disclosure are described and explained in detail from a unique perspective and different dimensions of the overall technical solution. In practical application scenarios, based on diverse needs, the technical solutions involved in different embodiments can be flexibly and freely combined according to specific circumstances. This combination is not a simple patchwork, but rather a comprehensive technical architecture that is more comprehensive, efficient, and tailored to actual business needs through careful integration and adaptation, further expanding the applicability and feasibility of the technical solutions disclosed in this disclosure in different fields and contexts.

[0051] Example 1

[0052] This disclosure has good applicability in a variety of scenarios. For ease of understanding, Figure 1 A flowchart of an embodiment of this disclosure is shown.

[0053] like Figure 1 As shown, in this embodiment, a principal component analysis method 100 based on vertical federated learning is provided, and the specific process is as follows:

[0054] Step 101: Generate a random Gaussian matrix, wherein the random Gaussian matrix is ​​identical in both the first communication device and the second communication device;

[0055] Step 102: Multiply the service data matrix of the first communication device with the random Gaussian matrix to obtain the first intermediate matrix;

[0056] Step 103: Obtain the second intermediate matrix sent by the second communication device; concatenate the first intermediate matrix with at least one of the second intermediate matrices to obtain a concatenated matrix; orthogonalize the concatenated matrix to obtain an orthogonal matrix;

[0057] Step 104: Based on the positional relationship of the first intermediate matrix in the splicing matrix, extract the first decomposition matrix from the orthogonal matrix; multiply the transpose of the first decomposition matrix with the business data matrix to obtain the first exchange matrix;

[0058] Step 105: Obtain the second exchange matrix sent by the second transportation device; add the first exchange matrix and the second exchange matrix to obtain the fusion matrix;

[0059] Step 106: Perform precise singular value decomposition on the fusion matrix to obtain left and right singular dimension reduction matrices.

[0060] In distributed computing scenarios, multiple communicable general-purpose computing devices need to collaborate on data processing. The local business data stored on these devices all require privacy-preserving computation through distributed communication.

[0061] For ease of understanding, Figure 2 A schematic diagram involving two communication devices is shown.

[0062] First, in step 101, all communication devices in the learning group of the vertical federated learning generate the same random Gaussian matrix. The learning group of the vertical federated learning includes a first communication device and at least one second communication device.

[0063] Communication equipment is essentially a general-purpose computing device that integrates communication functions, achieving efficient information transmission and processing through hardware and software collaboration. These devices have built-in core computing components such as a central processing unit and memory, enabling them to analyze, transform, and store data according to pre-set algorithms and protocols. Simultaneously, communication equipment is equipped with communication modules such as antennas and modems, supporting wireless or wired signal transmission and reception, and enabling data transfer between different network nodes.

[0064] The participants in vertical federated learning include a first communication device and at least one second communication device. In fact, there is no essential difference between the first and second communication devices. As a learning mode in which multiple participants work together, each participant can be regarded as the first communication device, while the other participants are considered as the second communication devices.

[0065] All communication devices generate the same random Gaussian matrix Ω with dimension n×l based on the same pre-shared random seed, where n is the original data dimension and l is the target dimension for dimensionality reduction.

[0066] Step 102 involves performing matrix multiplication on the device service data A of the first communication device and the randomly generated Gaussian matrix Ω.

[0067] For the first communication data, the data after this multiplication process is called the first intermediate matrix. For the second communication data, the data after this multiplication process is called the second intermediate data.

[0068] In the first communication device, the service data matrix A1∈R^(m1×n), where R is a real number, n is the original data dimension, and m1 is the number of data samples.

[0069] The business data matrix A1 is multiplied by the Gaussian matrix Ω to obtain the first intermediate matrix C1 = A1·Ω∈R^(m1×l), where R is a real number, m1 is the number of data samples, and l is the dimension of the dimensionality reduction target. The calculation result integrates the characteristics of the business data and equipment parameters of the first communication device.

[0070] Matrix multiplication, as a core computational operation in this disclosure, directly impacts the overall system performance. To address the computational challenges posed by large-scale data, a matrix block parallelization strategy can be employed, dividing the data matrix into multiple independent sub-blocks and implementing parallel computation through multithreading, GPU acceleration, or distributed computing frameworks.

[0071] At least one second communication device performs calculations similar to those in the first communication device. The service data matrix A2∈R^(m2×n), where R is a real number, n is the original data dimension, and m2 is the number of service data samples in the second communication device.

[0072] Multiplying the business data matrix A2 with the Gaussian matrix Ω yields the second intermediate matrix C2 = A2·Ω∈R^(m2×l), where R is a real number, m2 is the number of data samples, and l is the target dimension for dimensionality reduction.

[0073] The first communication device and at least one second communication device process service data A1 and A2 using the same logic to obtain a first intermediate matrix C1 and at least one second intermediate matrix C2, respectively.

[0074] Step 103 is cross-device data splicing.

[0075] To achieve cross-device data splicing, data transfer between devices must first be implemented. The process of the second communication device transferring data to the first communication device can be compared with the general process of distributed computing. For example... Figure 2 As shown, the two communication devices exchanged the first intermediate matrix C1 and at least one second intermediate matrix C2.

[0076] First, the second communication device serializes the data to be transmitted, converting it into a format suitable for network transmission, such as JSON or a byte stream. For data security, encryption can be used, so the encrypted data is transmitted here. To ensure transmission reliability, the TCP protocol can be used, and a retransmission mechanism can be set up to avoid data loss.

[0077] When receiving data, the first communication device first assembles the fragmented data and verifies its integrity, such as through hash verification. If encryption was previously applied, the first communication device uses the corresponding private key to decrypt the encrypted data and obtain the original data.

[0078] After data transmission is completed, the first communication device concatenates a first intermediate matrix C1 with at least one second intermediate matrix C2 to obtain a concatenated matrix E. Taking the concatenation of the first intermediate matrix C1 and a second intermediate matrix C2 as an example, the concatenated matrix E = concat(C1, C2) is obtained.

[0079] The "concat" operation mentioned here refers to concatenating matrices along a specific dimension. Generally, this concatenation is done along the channel dimension, but the specific method depends on the actual application scenario.

[0080] Throughout the system, the first communication device and at least one second communication device need to obtain the same splicing matrix E. To ensure this, each communication device must obtain the intermediate matrices of all other communication devices. Furthermore, both parties must strictly adhere to the same order when splicing the intermediate matrices. For example, if the first communication device places C1 first and then C2, the second communication device must also splice C1 first, then C2. If multiple second intermediate matrices exist, such as C2, C3, etc., all devices must use the same sorting for splicing. This could be based on principles such as generation time or a preset ID order to prevent deviations in the splicing matrix structure due to inconsistent order.

[0081] After constructing the splicing matrix E, the next step is to perform orthogonalization to obtain the orthogonal matrix Q.

[0082] In step 104, the orthogonal matrix Q is decomposed in the same way as the intermediate matrix was previously assembled.

[0083] The decomposition process is designed to be both inverse and consistent with the previous concatenation operation. Specifically, if the original concatenated matrix E is formed by concatenating the first intermediate matrix C1 (dimension m1×l) and the second intermediate matrix C2 (dimension m2×l) column-wise (i.e., m = m1 + m2), then the orthogonal matrix Q is also partitioned according to the same column division rules. The orthogonal matrix Q is divided into Q = [Q1|Q2], where Q1 contains the first m1 columns, Q2 contains the last m2 columns, and both Q1 and Q2 maintain orthogonality. This partitioning method strictly corresponds to the concatenation structure of the original intermediate matrices, ensuring a direct mapping relationship in dimension and data distribution between the decomposed matrix blocks and the unconcatenated matrix.

[0084] In multi-device collaborative scenarios, all communication devices follow the same partitioning rules and transformation strategies. For example, if the first communication device generates an orthogonal matrix Q and decomposes it into Q1 and Q2 according to columns m1 and m2, then the second communication device, after generating its own orthogonal matrix Q, will decompose it according to the exact same column index range, ensuring that the decomposed matrices obtained by each device are completely consistent. This decomposition method based on structural symmetry effectively avoids subsequent calculation deviations caused by inconsistent decomposition methods, providing a reliable guarantee for collaborative processing in distributed systems.

[0085] The first communication device multiplies the transpose of the first decomposition matrix Q1 obtained by decomposition with the service data matrix A1 to obtain the first exchange matrix F1.

[0086] In the first communication device, the service data matrix A1 ∈ R^(m1×n), where R is a real number, n is the original data dimension, and m1 is the number of data samples. The first decomposition matrix Q1 ∈ R^(m1×l), where l is the target dimension for dimensionality reduction. Therefore, the transpose of the first decomposition matrix... Then the first commutation matrix

[0087] At least one second communication device performs calculations similar to those in the first communication device. The second communication device multiplies the transpose of the second decomposition matrix Q2 obtained from the decomposition with the service data matrix A2 to obtain the second exchange matrix F2.

[0088] In the second communication device, the service data matrix A2 ∈ R^(m2×n), where R is a real number, n is the original data dimension, and m2 is the number of data samples. The second decomposition matrix Q2 ∈ R^(m2×l), where l is the target dimension for dimensionality reduction. Therefore, the transpose of the second decomposition matrix... Then the second commutation matrix

[0089] Step 105 is cross-device data fusion.

[0090] To achieve cross-device data fusion, data transfer between devices must first be realized. The process by which a second communication device transmits data to a first communication device can be referenced from the signed cross-device data splicing process. For example... Figure 2 As shown, the two communication devices exchanged the first exchange matrix F1 and at least one second exchange matrix F2.

[0091] After data transmission is completed, the first communication device adds the first exchange matrix F1 to at least one second exchange matrix F2 to obtain the fusion matrix G. Taking the matrix addition of the first exchange matrix F1 and a second exchange matrix F2 as an example, the fusion matrix G = F1 + F2 is obtained.

[0092] Throughout the system, the first communication device and at least one second communication device need to obtain the same fusion matrix G. To ensure this, each communication device must obtain the exchange matrices of all other communication devices. Furthermore, the order in which the two devices add their exchange matrices is not strictly identical. For example, the first communication device obtains the fusion matrix G = F1 + F2. The second communication device can obtain the fusion matrix G = F2 + F1.

[0093] In step 106, after the construction of the fusion matrix G is completed, the next step is to perform singular value decomposition.

[0094] The first communication device performs singular value decomposition (SVD) on the fused matrix G, obtaining left and right singular vector matrices, where the left singular vector matrix U is the final dimensionality-reduced matrix. When processing the fused matrix G, the first communication device can use stochastic singular value decomposition (SVFD) to perform the SVFD operation, thereby obtaining the left and right singular vector matrices. Here, the left singular vector matrix is ​​used as the final dimensionality-reduced matrix. In handling large-scale matrix factorization tasks, traditional SVFD algorithms often encounter efficiency bottlenecks due to high computational complexity and large memory requirements. In contrast, stochastic SVFD uses random projection and subspace approximation strategies to map the original high-dimensional matrix to a low-dimensional space, thus greatly reducing the computational load. This technique samples the original matrix by constructing a random matrix, which can quickly capture the main eigenvectors of the matrix. It improves the efficiency of matrix factorization while maintaining decomposition accuracy.

[0095] At the same time, the second communication device will also generate its own dimensionality-reduced matrix U according to the same singular value decomposition algorithm.

[0096] From steps 101 to 106, through multiple data processing steps and cross-device data exchanges, the data privacy of the business data of each device is guaranteed. This also ensures that the principal component analysis preprocessing step of the vertical federated learning is carried out while maintaining data privacy.

[0097] In a federated learning framework, all participants need to share data to collaboratively train the model. However, existing principal component analysis (PCA) techniques require collecting raw data from all devices for dimensionality reduction. This leads to the following drawbacks when PCA is directly applied to federated learning: firstly, data privacy can be compromised when business data from one communication device is transmitted to other communication devices. This disclosure, through ingenious design, not only achieves the goal of "usable but invisible" data but also allows for secure and efficient data flow between different devices. This enables all participants to collaboratively train a high-quality machine learning model without compromising privacy, effectively balancing the conflict between data sharing needs and privacy protection.

[0098] Although this transformed data originates from the original information, it cannot be reverse-engineered to reconstruct the original data content using conventional methods. Therefore, when communication devices exchange this transformed data, they can achieve the goal of collaborative data analysis without worrying about the privacy of their local original data, providing a secure and reliable foundation for data sharing in distributed computing.

[0099] In this disclosure, the data exchanged between communication devices includes an intermediate matrix C and an exchange matrix F. The communication devices possess the same random Gaussian matrix Ω, a spliced ​​matrix E, and an orthogonal matrix Q. Taking the random Gaussian matrix Ω and the orthogonal matrix Q as examples, it can be concluded through the following discussion that it is impossible to reverse-engineer the other party's data matrix A. For example, the first communication device has service data A1, and the second communication device has service data A2. The first communication device sends a first intermediate matrix C1 to the second communication device, and the second communication device sends a second intermediate matrix C2 to the first communication device.

[0100] When matrix Ω or Q is invertible, then the second communication device can communicate via C1Ω. -1 The original data matrix A1 of the first communication device is obtained. Similarly, the original data matrix A2 of the second communication device can be obtained by the same calculation.

[0101] When matrices Ω and Q are not invertible, i.e., singular matrices det(Ω) = 0 and det(Q) = 0, multiple different A1s satisfy C1 = A1Ω. Therefore, the second communication device holding Ω cannot determine the true A1 of the first communication device.

[0102] This processing method has significant security advantages. It is closely linked to the original data of the first or second communication device, but avoids directly disclosing the original data of the first or second communication device, and also avoids disclosing the original data of the first or second communication device through reversible restoration.

[0103] In a federated learning system consisting of a first communication device and at least one second communication device, after each communication device reduces the dimensionality of its local business data to obtain a dimensionality reduction matrix, it collaboratively trains a large business model through the following steps.

[0104] First, a large-scale business model is trained based on a federated learning framework. Each communication device calculates the model gradient based on the locally reduced dimensionality matrix U and sends the calculated model gradient to other communication devices. Global gradient updates are generated based on federated learning, and the large-scale business model is trained accordingly.

[0105] Secondly, the large-scale business model is incrementally optimized through federated learning. For the already trained large-scale business model, the client fine-tunes the model parameters using a locally generated dimensionality reduction matrix U. For communication devices with large differences in data distribution, their gradient contribution weights are dynamically adjusted, or knowledge distillation is used to transfer global model knowledge to the local model of the communication device, balancing generalization and personalization needs. After each round of training, the communication device can periodically incorporate newly added data increments into the dimensionality reduction matrix to achieve continuous model optimization.

[0106] Finally, the performance of the large-scale business model is evaluated and iterated. Each communication device evaluates the performance of the global model on a distributed test set through cross-validation; if convergence is not achieved, the next round of training is triggered. For multimodal data (such as image + text), the dimensionality reduction matrices of each modality can be fused within a federated framework through a cross-modal attention mechanism. For example, in intelligent customer service scenarios, the voice feature matrix and the text semantic matrix are jointly trained under encrypted conditions to improve the accuracy of intent recognition. The final model is deployed to various clients via edge computing to support real-time business decisions.

[0107] Principal component analysis (PCA) based on vertical federated learning, along with subsequent large-scale business model training and applications, has been widely used in numerous technological fields. The following examples, using large-scale business model training and applications in four domains, illustrate how, in the face of the massive volume brought about by the rapid growth of big data, the pressure on data storage, transmission, and computation has increased dramatically. PCA based on vertical federated learning technology achieves efficient data dimensionality reduction while strictly ensuring data privacy and security, becoming a key factor in optimizing large-scale business model training.

[0108] (1) Medical joint diagnosis

[0109] In collaborative medical diagnostic scenarios, sensitive medical information such as patient CT images and genetic data stored on communication devices across multiple hospitals faces a dual challenge: on the one hand, directly transmitting raw data would severely violate privacy regulations; on the other hand, massive amounts of medical image data would consume significant network bandwidth. This solution innovatively addresses this problem by combining a vertical federated learning framework with principal component analysis technology.

[0110] Through localized dimensionality reduction processing, patient data in each hospital's communication equipment undergoes privacy-preserving dimensionality reduction locally, compressing the original high-dimensional medical data into a low-dimensional feature representation. This process achieves two core advantages: significantly reduced data volume (actual measurements show a reduction of over 85% in the original data volume), while ensuring that the patient's original privacy data remains locally, completely avoiding the network transmission of sensitive medical information.

[0111] In the federated learning architecture design, a layered collaborative mechanism is adopted: edge nodes are responsible for local encoder training and only upload encrypted model parameters; the cloud server centrally aggregates the global model, and homomorphic encryption technology ensures the security and reliability of gradient transmission. To further improve system efficiency, federated knowledge distillation technology is introduced, effectively compressing the number of model parameters (reducing computational resource consumption by 60% in typical scenarios), enabling cross-institutional medical collaborative diagnosis to both protect data privacy and significantly improve the efficiency and performance of the collaborative diagnostic model.

[0112] (2) Financial risk control

[0113] In financial risk control scenarios, sensitive information such as user transaction records, communication behaviors, and consumption data stored by institutions such as banks, telecom operators, and e-commerce companies faces a dual challenge: on the one hand, directly sharing raw data violates regulations such as the Personal Information Protection Law; on the other hand, integrating massive amounts of heterogeneous data (such as bank transaction logs and telecom operator call records) can lead to extremely high computational and communication overhead. This solution innovatively addresses this problem by combining a vertical federated learning framework with feature engineering optimization techniques.

[0114] This publication utilizes principal component analysis (PCA) based on longitudinal federated learning to achieve global feature extraction and dimensionality reduction while protecting the privacy of local business data. On the bank side, financial features such as user transaction frequency and amount fluctuations are extracted; on the operator side, behavioral features such as call duration and cross-regional logins are analyzed; and on the e-commerce side, indicators such as consumer category preferences and return rates are calculated. This process employs privacy set intersection (PSI) technology to securely match user IDs, ensuring that only users from the intersection participate in modeling and that the original data remains within the specified domain. Experimental results show that feature engineering optimization can reduce data dimensionality by more than 70%, while homomorphic encrypted gradient transmission reduces communication overhead by 62%.

[0115] Based on the dimensionality reduction matrix of the aforementioned business data obtained through principal component analysis, a large-scale business model for financial risk control is trained on a federated learning architecture. Institutions with high-quality data (such as banks with complete transaction chains) receive higher weights during model aggregation, while participants with sparse features (such as newly integrated government data) gradually increase their contribution through transfer learning. To further improve real-time performance, a lightweight knowledge distillation technique is introduced to compress the global business model into a lightweight version suitable for edge devices (such as mobile banking apps), reducing fraud detection latency from 100ms to 10ms and the false positive rate by 30%.

[0116] (3) Intelligent Recommendation

[0117] E-commerce platforms extract user browsing history data to identify inquiry intent features such as "noise cancellation needs" and "battery life," and combine this with their add-to-cart behavior to generate local preference vectors. Telecommunications operators analyze past customer service calls to identify implicit needs such as "frequent data plan inquiries" and "concern for call quality." Social media data uncovers user interaction records in headphone-related communities, such as liking "sports headphone recommendations."

[0118] After semantic alignment of the three-party data using word embedding technology, dimensionality reduction is achieved through principal component analysis in a vertical federated learning approach. The compressed feature matrix is ​​then used to train a large model within a federated architecture. The model aggregates feature increments from each platform in real time based on the user's current browsing context. E-commerce platforms, due to active user behavior during promotional periods, prioritize uploading the latest click data for model updates; while telecom operators and social media platforms dynamically adjust their aggregation frequency according to regular cycles.

[0119] The final recommendation system generates a combination solution: headphone product pages prioritize displaying models with noise cancellation and suitable for sports activities, simultaneously pushing discounted "data plan + headphones" combinations, along with genuine user reviews from community groups. Throughout the process, user's original conversations and browsing history are stored locally, participating in collaborative modeling only through encrypted feature vectors. This achieves accurate cross-platform recommendations while ensuring privacy and security.

[0120] (4) Industrial Forecasting

[0121] As the digital transformation of smart factories accelerates, the massive amounts of data generated by equipment operation hold immense value, but also present challenges in data collaboration and privacy protection. Principal Component Analysis (PCA) dimensionality reduction technology based on longitudinal federated learning has become key to solving the dilemma of multi-factory joint training of federated fault diagnosis models.

[0122] Within this framework, each factory relies on a vertical federated learning architecture and utilizes PCA technology to perform deep dimensionality reduction processing on local equipment operation data (such as multi-dimensional sensor data like vibration frequency, temperature curves, and current fluctuations). By retaining key features and removing redundant information, high-value shared feature vectors are generated, enabling cross-factory data collaboration while ensuring that the original data does not leave the factory.

[0123] Subsequently, the factory servers participating in federated learning initiate a collaborative training process, securely aggregating the dimensionality-reduced feature vectors to construct a global fault diagnosis model. This model can accurately predict the remaining useful life (RUL) of equipment, providing a scientific basis for preventative maintenance. Taking the FedRUL method developed by Southwest Jiaotong University as an example, in the scenario of milling cutter life prediction, this technology successfully reduced the prediction error by 19.8%, helping enterprises reduce downtime losses caused by sudden equipment failures, significantly improving the accuracy and timeliness of equipment maintenance, and effectively reducing operation and maintenance costs.

[0124] Example 2

[0125] In this embodiment, the generation of a random Gaussian matrix in method 100 includes:

[0126] Obtain a random seed, which is the same in both the first communication device and the second communication device;

[0127] A random Gaussian matrix is ​​generated based on a pseudo-random number generator.

[0128] The process of generating a random Gaussian matrix further includes:

[0129] Obtain the hash value of the random Gaussian matrix;

[0130] Verify that the hash value is the same in both the first communication device and the second communication device.

[0131] In a distributed communication system, the first and second communication devices need to collaborate to complete matrix operations. To efficiently generate identical random Gaussian matrices and avoid data transmission overhead, this engineering discipline creatively proposes a method for generating random Gaussian matrices based on a shared random seed.

[0132] Taking a first communication device and at least one second communication device as an example, the communication devices share an initial seed value S in advance through a secure channel. The scheme disclosed herein does not require a trusted third party; the communication devices can securely share the initial seed value through a cryptographic protocol.

[0133] Taking three communication devices in federated learning as an example, the initial seed value is securely shared through a distributed cryptographic protocol, as follows.

[0134] Step 1: Negotiate and disclose encryption parameters

[0135] Three devices (devices A, B, and C) pre-agree to use the same secure encryption foundation (e.g., elliptic curve cryptography) and jointly select a set of publicly available mathematical parameters, such as specifying an elliptic curve and a point G on the curve. This set of parameters acts like a "mathematical rule" agreed upon by the three parties, ensuring that subsequent calculations have a unified standard.

[0136] Step 2: Each person generates their private and public keys.

[0137] Device A generates a random number 'a' as its private key, calculates the public key A = a·G using the private key 'a' and the public parameter G, and then sends the public key A to devices B and C.

[0138] Device B generates a random number b as its private key, calculates the public key B = b·G, and sends it to devices A and C.

[0139] Device C generates a random number c as its private key, calculates the public key C = c·G, and sends it to devices A and B.

[0140] The private key is kept by each device, while the public key is transmitted publicly over the network. However, due to the security of the encryption principle, even if an attacker intercepts the public key, they cannot deduce the private key.

[0141] Step 3: Calculate the shared key among the three parties

[0142] After receiving the public key B from device B and the public key C from device C, device A uses its own private key a to calculate: K = a·B·C (where "·" represents the dot product operation on the elliptic curve).

[0143] After receiving the public key A from device A and the public key C from device C, device B uses its own private key b to calculate: K = b·A·C;

[0144] After receiving the public key A from device A and the public key B from device B, device C uses its own private key c to calculate: K = c·A·B.

[0145] Due to the commutative and associative laws of elliptic curve operations, the shared key K calculated by the three parties is exactly the same, just as the three parties each use different ingredients to make the same "cryptographic cake", and the result is consistent.

[0146] Step 4: Derive the initial seed value

[0147] After obtaining the shared key K, the three parties need to convert it into a seed value suitable for generating random numbers. For example, the coordinates of K on an elliptic curve can be extracted as raw material, processed by a secure hash function (such as SHA-256), and custom identification information can be added during hashing to ensure the uniqueness and randomness of the seed. The final hash result is the initial seed value Seed shared by all three parties.

[0148] Step 5: Verify seed consistency

[0149] To avoid transmission or calculation errors, each of the three parties calculates a hash value for the seed (e.g., H(Seed)) and exchanges the hash values ​​through the communication channel. If the hash values ​​of the three devices are completely consistent, it means that seed sharing was successful; if they are inconsistent, the above steps are checked again to rule out errors or possible attacks.

[0150] In this embodiment, the use of a deterministic pseudo-random number generator (PRNG) is a key technology for enabling multiple devices to collaboratively generate the same random Gaussian matrix.

[0151] Deterministic PRNGs, based on the same initial seed, can strictly repeat the generation of the same random number sequence, ensuring that all participating devices obtain completely consistent results when computing independently on their own. In this disclosure, communication devices generate the same Gaussian random number matrix by sharing a seed and running the same PRNG algorithm, thus avoiding the direct transmission of sensitive matrix data and ensuring the accuracy of collaborative computation.

[0152] For example, when the first and second communication devices start the PRNG based on the same seed, the algorithm will generate the same uniformly distributed random numbers in sequence, and then convert them into a Gaussian distribution through methods such as Box-Muller transformation, and finally fill them into a random Gaussian matrix with consistent dimensions.

[0153] Example 3

[0154] In this embodiment, the method 100 further includes: obtaining the second intermediate matrix sent by the second communication device, comprising:

[0155] Obtain the encrypted data sent by the second communication device;

[0156] Decrypt the encrypted data to obtain the second intermediate matrix.

[0157] Taking a scenario where the first communication device is a server and at least one second communication device is a client as an example, the server and client need to collaborate to complete privacy-preserving computations based on federated learning. This involves cross-device interaction of business data such as intermediate matrices and exchange matrices during the training process.

[0158] In a certain calculation, a client in the second communication device calculates an intermediate matrix based on business data. This matrix contains key information after processing the business data. To ensure data security during transmission and prevent information theft or tampering, the client can encrypt the matrix. The client uses the AES-256 encryption algorithm and a pre-negotiated symmetric key with the server to encrypt the generated second intermediate matrix, converting the original matrix data into encrypted data in ciphertext form.

[0159] After encryption is complete, the client sends the encrypted data to the server through the established secure communication channel. Upon receiving the encrypted data from the client, the server first performs a preliminary verification of the data's integrity, such as checking whether the hash value of the data matches the expectation, to ensure that the data has not been corrupted during transmission.

[0160] After successful verification, the server uses the same symmetric key as the client and the AES-256 decryption algorithm to decrypt the encrypted data. During decryption, the server gradually restores the ciphertext data to its original matrix form according to the AES-256 algorithm, ultimately obtaining the second intermediate matrix. Once the server obtains the second intermediate matrix, it can integrate it with its own computed data or data transmitted from other devices to continue the training process of the distributed machine learning model.

[0161] By encrypting and decrypting data exchanged across devices, the security and integrity of data interaction are ensured. The communication equipment employs high-strength encryption algorithms such as AES-256 to encrypt the intermediate matrix, resisting the risks of data eavesdropping and tampering, ensuring that sensitive information is not leaked during transmission. The decryption process, combined with an integrity verification mechanism, can promptly detect transmission errors or malicious attacks, guaranteeing the accuracy of the obtained matrix. This provides a reliable data interaction foundation for multi-device collaborative computing, effectively balancing privacy protection and computational efficiency.

[0162] Example 4

[0163] In this embodiment, in method 100, the step of splicing the first intermediate matrix with at least one second intermediate matrix to obtain a spliced ​​matrix includes:

[0164] Determine the order in which the first intermediate matrix and at least one of the second intermediate matrices are arranged;

[0165] Based on the stated arrangement order, the first intermediate matrix and at least one second intermediate matrix are combined to obtain a spliced ​​matrix.

[0166] The step of extracting the first decomposition matrix from the orthogonal matrix based on the positional relationship of the first intermediate matrix in the concatenated matrix includes:

[0167] Based on the arrangement order, obtain the start and end row numbers of the first intermediate matrix in the spliced ​​matrix;

[0168] The first decomposition matrix is ​​extracted from the orthogonal matrix based on the start and end row numbers.

[0169] In a vertical federated learning scenario, taking a learning group consisting of a first communication device and a second communication device as an example, the first and second communication devices generate identical random Gaussian matrices Ω of dimension n×l based on a pre-shared random seed, where n is the original data dimension and l is the target dimension for dimensionality reduction.

[0170] The service data matrix A1∈R^(m1×n) of the first communication device is divided into multiple sub-blocks using a matrix block parallelization strategy. Utilizing GPU acceleration technology, each sub-block is multiplied in parallel with a Gaussian matrix Ω to obtain the first intermediate matrix C1=A1·Ω∈R^(m1×l). Similarly, the service data matrix A2∈R^(m2×n) of the second communication device is distributed across multiple computing nodes using a distributed computing framework. Each node computes the product of its sub-blocks and Ω in parallel to obtain the second intermediate matrix C2=A2·Ω∈R^(m2×l).

[0171] During the data transmission phase, the second communication device serializes the second intermediate matrix C2 into a byte stream and transmits it to the first communication device. The first communication device receives the second intermediate matrix C2.

[0172] The first communication device concatenates C1 and C2 along the channel dimension to obtain a concatenation matrix E = concat(C1, C2). To ensure that the second communication device also obtains the same concatenation matrix E, the first communication device sends the concatenation order (C1 first, then C2) and related metadata to the second communication device through a secure channel. After receiving C1, the second communication device completes the concatenation in the same order.

[0173] After constructing the concatenated matrix E, the QR decomposition algorithm is used to orthogonalize E, resulting in an orthogonal matrix Q. Subsequently, according to the column partitioning rules used during concatenation, the orthogonal matrix Q is divided into Q = [Q1|Q2], where Q1 contains the first m1 columns and Q2 contains the last m2 columns. During partitioning, the orthogonality of Q1 and Q2 is ensured by verifying that the inner product of their column vectors is zero. This decomposition method ensures that Q1 and Q2 strictly correspond to C1 and C2 before concatenation in terms of dimension and data distribution. Subsequently, during model training, the first and second communication devices can perform calculations based on Q1 and Q2 respectively, guaranteeing the consistency and accuracy of the calculation results.

[0174] Example 5

[0175] like Figure 3As shown, the principal component analysis device 300 based on longitudinal federated learning provided in this embodiment includes:

[0176] A random Gaussian matrix generation module 301 is used to generate a random Gaussian matrix, which is the same in both the first communication device and the second communication device;

[0177] The first intermediate matrix acquisition module 302 is used to multiply the service data matrix of the first communication device with the random Gaussian matrix to obtain the first intermediate matrix;

[0178] The orthogonal matrix acquisition module 303 is used to acquire the second intermediate matrix sent by the second communication device; concatenate the first intermediate matrix with at least one of the second intermediate matrices to obtain a concatenated matrix; and orthogonalize the concatenated matrix to obtain an orthogonal matrix.

[0179] The first exchange matrix acquisition module 304 is used to extract a first decomposition matrix from the orthogonal matrix based on the positional relationship of the first intermediate matrix in the concatenated matrix; and multiply the transpose of the first decomposition matrix by the business data matrix to obtain the first exchange matrix.

[0180] The fusion matrix acquisition module 305 is used to acquire the second exchange matrix sent by the second traffic equipment; and to add the first exchange matrix and the second exchange matrix to obtain the fusion matrix;

[0181] The singular value decomposition module 306 is used to perform precise singular value decomposition on the fusion matrix to obtain left and right singular dimension-reduced matrices.

[0182] In this embodiment, the specific processing of each unit and the resulting technical effects in the principal component analysis device 300 based on vertical federated learning can be referred to the relevant descriptions of each step in the foregoing embodiments, and will not be repeated here.

[0183] Example 6

[0184] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0185] Figure 4A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0186] like Figure 4 As shown, device 400 includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 402 or a computer program loaded from storage unit 408 into random access memory (RAM) 403. RAM 403 may also store various programs and data required for the operation of device 400. The computing unit 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.

[0187] Multiple components in device 400 are connected to I / O interface 405, including: input unit 406, such as keyboard, mouse, etc.; output unit 407, such as various types of monitors, speakers, etc.; storage unit 408, such as disk, optical disk, etc.; and communication unit 409, such as network card, modem, wireless transceiver, etc. Communication unit 409 allows device 400 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0188] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above, such as the principal component analysis method based on longitudinal federated learning. For example, in some embodiments, the principal component analysis method based on longitudinal federated learning can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on device 400 via ROM 402 and / or communication unit 409. When the computer program is loaded into RAM 403 and executed by the computing unit 401, one or more steps of the principal component analysis method based on longitudinal federated learning described above can be performed. Alternatively, in other embodiments, computing unit 401 may be configured by any other suitable means (e.g., by means of firmware) to perform principal component analysis based on longitudinal federated learning.

[0189] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0190] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable information processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0191] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0192] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0193] The systems and technologies described herein can be implemented in computing systems that include back-end components (e.g., as a data server), or computing systems that include middleware components (e.g., an information server), or computing systems that include front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0194] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.

[0195] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0196] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A principal component analysis method based on vertical federated learning, applied to a first communication device, wherein the first communication device and at least one second communication device constitute a learning group for vertical federated learning, characterized in that, The method includes: Generate a random Gaussian matrix, wherein the random Gaussian matrix is ​​identical in both the first communication device and the second communication device; Multiply the service data matrix of the first communication device by the random Gaussian matrix to obtain the first intermediate matrix; Obtain the second intermediate matrix sent by the second communication device; concatenate the first intermediate matrix with at least one of the second intermediate matrices to obtain a concatenated matrix; orthogonalize the concatenated matrix to obtain an orthogonal matrix; Based on the positional relationship of the first intermediate matrix in the concatenated matrix, a first decomposition matrix is ​​extracted from the orthogonal matrix; the transpose of the first decomposition matrix is ​​multiplied by the business data matrix to obtain a first exchange matrix; Obtain the second exchange matrix sent by the second traffic device; add the first exchange matrix and the second exchange matrix to obtain the fusion matrix; Perform precise singular value decomposition on the fusion matrix to obtain left and right singular dimension-reduced matrices.

2. The principal component analysis method based on longitudinal federated learning according to claim 1, characterized in that, The generation of the random Gaussian matrix includes: Obtain a random seed, which is the same in both the first communication device and the second communication device; A random Gaussian matrix is ​​generated based on a pseudo-random number generator.

3. The principal component analysis method based on longitudinal federated learning according to claim 2, characterized in that, The process of generating a random Gaussian matrix further includes: Obtain the hash value of the random Gaussian matrix; Verify that the hash value is the same in both the first communication device and the second communication device.

4. The principal component analysis method based on longitudinal federated learning according to claim 1, characterized in that, The step of obtaining the second intermediate matrix sent by the second communication device includes: Obtain the encrypted data sent by the second communication device; Decrypt the encrypted data to obtain the second intermediate matrix.

5. The principal component analysis method based on longitudinal federated learning according to claim 1, characterized in that, The step of concatenating the first intermediate matrix with at least one second intermediate matrix to obtain a concatenated matrix includes: Determine the order in which the first intermediate matrix and at least one of the second intermediate matrices are arranged; Based on the stated arrangement order, the first intermediate matrix and at least one second intermediate matrix are combined to obtain a spliced ​​matrix.

6. The principal component analysis method based on longitudinal federated learning according to claim 5, characterized in that, The step of extracting the first decomposition matrix from the orthogonal matrix based on the positional relationship of the first intermediate matrix in the concatenated matrix includes: Based on the arrangement order, obtain the start and end row numbers of the first intermediate matrix in the spliced ​​matrix; The first decomposition matrix is ​​extracted from the orthogonal matrix based on the start and end row numbers.

7. A principal component analysis device based on longitudinal federated learning, characterized in that, The device includes: A random Gaussian matrix generation module is used to generate a random Gaussian matrix, wherein the random Gaussian matrix is ​​the same in both the first communication device and the second communication device; The first intermediate matrix acquisition module is used to multiply the service data matrix of the first communication device with the random Gaussian matrix to obtain the first intermediate matrix; An orthogonal matrix acquisition module is used to acquire a second intermediate matrix sent by the second communication device; concatenate the first intermediate matrix with at least one of the second intermediate matrices to obtain a concatenated matrix; and orthogonalize the concatenated matrix to obtain an orthogonal matrix. The first exchange matrix acquisition module is used to extract a first decomposition matrix from the orthogonal matrix based on the positional relationship of the first intermediate matrix in the concatenated matrix; and multiply the transpose of the first decomposition matrix by the business data matrix to obtain the first exchange matrix. The fusion matrix acquisition module is used to acquire the second exchange matrix sent by the second transportation equipment; and to add the first exchange matrix and the second exchange matrix to obtain the fusion matrix. The singular value decomposition module is used to perform precise singular value decomposition on the fusion matrix to obtain left and right singular dimension-reduced matrices.

8. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.

10. A computer program product comprising a computer program that, when executed by a processor, implements the method of any one of claims 1-6.