A system and method for securely transferring data amongst a fleet of motor vehicles and a network for the same

WO2026201852A1PCT designated stage Publication Date: 2026-10-01CONTINENTAL AUTOMOTIVE TECHNOLOGIES GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2026/058052
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-26
Filing Date
2026-03-23
Publication Date
2026-10-01

Smart Images

  • Figure EP2026058052_01102026_PF_FP_ABST
    Figure EP2026058052_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Invention relates to a method for securely transferring data amongst motor vehicles, the method comprising: executing, by way of a global autoencoder, an aggregation process for aggregating a first output of labelled dataset generated by a machine learning model using a first set of unlabelled vehicular sensing data and a second output of labelled dataset generated by the machine learning model using a second set of unlabelled vehicular sensing data; correlating by way of a co-attention mechanism, the first output of labelled dataset with the second output of labelled dataset for identifying at least one hidden representation between the first output of labelled dataset and the second output of labelled dataset; training, by way of a shared classifier, classification of hidden representation using the that least one hidden representation identified; and synchronising, by way of a server, a set of trained parameters with the global autoencoder, the shared classifier and each motor vehicle in the fleet. A system and an edge network comprising the same is also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] 202407982

[0002] 1

[0003] A SYSTEM AND METHOD FOR SECURELY TRANSFERRING DATA AMONGST A FLEET OF MOTOR VEHICLES AND A NETWORK FOR THE SAME

[0004] TECHNICAL FIELD

[0005] This disclosure relates to a multi-modal federated learning framework to train a model with a various data of modality.

[0006] BACKGROUND

[0007] Automotive industry is increasingly moving towards implementation of connected vehicles on the road. With the rapid development of Internet of Vehicles (loV), the computing power of the connected motor vehicle hardware requires constant improvement, to catch up with the power consumption needs.

[0008] In addition, vehicle owners are also concern over data privacy issues and cybersecurity attacks. At present, different types of computational processing units onboard motor vehicles implement strict encryption to ensure privacy and security of each processing unit. However, such solutions present problem of lack of flexibility in terms of sharing of data amongst the processing units due to model heterogeneity. Existing solutions to solve the aforesaid issue focusses on addressing data heterogeneity where the model heterogeneity exists in application which is still an impediment for collaborative training tasks amongst involving different data forms, for example an intrusion detection task onboard motor vehicle.

[0009] Other objects, features and characteristics, as well as the methods of operation and the functions of the related elements of the structure, the combination of parts and economics of manufacture will become more apparent upon consideration of the following detailed description and appended claims with reference to the accompanying drawings, all of which form a part of this specification. It should be understood that the detailed description and specific examples, while indicating the202407982

[0010] 2

[0011] non-limiting embodiments of the disclosure, are intended for purposes of illustration only and are not intended to limit the scope of the disclosure.

[0012] SUMMARY

[0013] A purpose of this disclosure is to ameliorate some of the problem(s) as discussed by providing the subject-matter of the independent claims.

[0014] Further purposes of this disclosure are set out in the accompanying dependent claims.

[0015] In a first aspect of this disclosure, a method for securely transferring data amongst motor vehicles is provided. The method comprises (a) executing by way of a global autoencoder, an aggregation process for aggregating a first output of labelled dataset generated by a machine learning model using a first set of unlabelled vehicular sensing data and a second output of labelled dataset generated by the machine learning model using a second set of unlabelled vehicular sensing data; (b) correlating, by way of a co-attention mechanism, the first output of labelled dataset with the second output of labelled dataset for identifying at least one hidden representation between the first output of labelled dataset and the second output of labelled dataset; (c) training, by way of a shared classifier, classification of hidden representation using the at least one hidden representation identified; and (d) synchronising, by way of a server, a set of trained parameters with the global autoencoder, the shared classifier and each motor vehicle in the fleet.

[0016] Advantageously, the aforesaid sequence or method labels data collected by a machine learning model located in client devices, i.e. onboard motor vehicles and transmit the unlabelled dataset collected to a server, of which the server may label all the dataset received or label a certain number of auxiliary datasets. In addition, the correlation of at least two sets of labelled datasets by a co-attention mechanism aids identification of at least one hidden variable or hidden representation between the sets of labelled datasets correlated. The hidden representation identified by the202407982

[0017] 3

[0018] co-attention mechanism is classified by a shared classifier and the classification of hidden representation may be use for training the shared classifier, strengthening the multi-modal federal learning process. As a final sequence, the server will in turn synchronise the global model and the shared classifier with all the motor vehicles in the fleet, to complete one round of communication, such that a set of trained parameters generated from the aforesaid sequence of steps are uniformly distributed.

[0019] In some embodiment, the sequence of (b) correlating the first output labelled dataset and the second output labelled dataset comprises (b)(i) encoding, by way of the co-attention mechanism, the first output of labelled dataset and the second output of labelled dataset with the at least one hidden representation identified.

[0020] In some embodiment, the sequence of (b) correlating the first output of the labelled dataset with the second output of labelled dataset for identifying at least one hidden representation comprises (b)(ii) identifying pairwise datapoint from the first output labelled dataset with the second output of labelled dataset by way of an affinity matrix.

[0021] In some embodiment, the method includes correlating by way of a deep canonical correlation analysis executed by the machine learning model, covariance of at least one multimodality datapoint between the first set of unlabelled vehicular sensing data collected and the second set of unlabelled vehicular sensing data collected, with the first output of labelled dataset and the second output of labelled dataset generated by the machine learning model, for identifying at least one maximally correlated projection.

[0022] In some embodiment, the method includes a sequence of (e) correlating covariance of at least one multimodality datapoint for identifying at least one maximally correlated projection comprises (e)(i) maintaining, by way of an integrated deep canonical correlation analysis model, an individual autoencoder for each of the at least one multimodality datapoint, for maximising the canonical correlation between the hidden representations from two different multimodality datapoints.202407982

[0023] 4

[0024] In some embodiment, the method comprises (e) correlating covariance of at least one multimodality datapoint for identifying at least one maximally correlated projection comprises: (e)(ii) executing, by way of a separated deep canonically correlated autoencoder, a forward propagation process comprising: exchanging, by way of an encrypted transmission channel, a first multimodality datapoint containing a first hidden representation identified from the first output of labelled dataset with a second multimodality datapoint containing a second hidden representation identified from the second output of labelled dataset, thereby obtaining a corresponding encoder for each of the first hidden representation and the second hidden representation.

[0025] In some embodiment, the method comprises: correlating, by way of a deep canonical correlation analysis executed by the machine learning model, covariance of at least one multimodality datapoint between the first set of unlabelled vehicular sensing data collected and the second set of unlabelled vehicular sensing data collected, with the first output of labelled dataset and the second output of labelled dataset generated by the machine learning model, for identifying at least one maximally correlated projection.

[0026] In some embodiment, the method comprises (e) correlating covariance of at least one multimodality datapoint for identifying at least one maximally correlated projection comprises: (e)(i) maintaining, by way of an integrated deep canonical correlation analysis model, an individual autoencoder for each of the at least one multimodality datapoint, for maximising the canonical correlation between the hidden representations from two different multimodality datapoints.

[0027] In some embodiment, the method comprises (e) correlating covariance of at least one multimodality datapoint for identifying at least one maximally correlated projection comprises: (e)(ii) executing, by way of a separated deep canonically correlated autoencoder, a forward propagation process comprising: exchanging, by way of an encrypted transmission channel, a first multimodality datapoint containing a first hidden representation identified from the first output of labelled dataset with a202407982

[0028] 5

[0029] second multimodality datapoint containing a second hidden representation identified from the second output of labelled dataset, thereby obtaining a corresponding encoder for each of the first hidden representation and the second hidden representation.

[0030] In some embodiment, the method comprises: initiating, by way of the machine learning model, an unsupervised vertical federated learning model using the first set of unlabelled vehicular sensing data and the second set of unlabelled vehicular sensing data collected by the machine learning model.

[0031] In some embodiment, the method comprises generating, by way of an adaptive federated aggregation algorithm, a similarity score for measuring a difference between a set of local parameters stored in the at least one machine learning model and a set of global parameters stored in the global autoencoder.

[0032] In some embodiment, the method comprises: the deep canonical correlation analysis is selected from a group consisting of an integrated deep canonical correlation analysis; a separate deep canonical correlation analysis framework, and combination thereof.

[0033] In a second aspect of this disclosure, a system for secure transfer of data amongst motor vehicles is provided. The system comprises: (a) a global autoencoder operable to execute an aggregated process to aggregate a first output of labelled dataset using a first set of unlabelled vehicular sensing data received, and a second output of labelled dataset using a second set of unlabelled vehicular sensing data received by the server, of which the first set of unlabelled vehicular sensing data and the second set of unlabelled vehicular sensing data are collected by the machine learning model; (b) a co-attention mechanism operable to correlate the first output of labelled dataset with the second output of labelled dataset to identify at least one hidden representation between the first output of labelled dataset and the second output of labelled dataset; (c) a shared classifier operable to classify the at least one hidden representation identified by the co-attention mechanism; and (d) a server202407982

[0034] 6

[0035] operable to synchronise a set of trained parameters with the global autoencoder, the shared classifier and each motor vehicle in a fleet.

[0036] Advantageously, the aforesaid configuration provides an edge computing solution, of which a machine learning model located in client devices, i.e. onboard motor vehicles, collects unlabelled dataset from one or more vehicular sensing devices and transmit the unlabelled dataset collected to a server. In turn, the server may label all the dataset received or label a certain number of auxiliary datasets, such that processing of the unlabelled dataset are not done within the motor vehicle. In addition, the correlation of at least two sets of labelled datasets by a co-attention mechanism aids identification of at least one hidden variable or hidden representation between the sets of labelled datasets correlated. The hidden representation identified by the co-attention mechanism is classified by a shared classifier and the classification of hidden representation may be use for training the shared classifier, strengthening the multi-modal federal learning process. Since the processing of dataset are done on a vehicle of the framework instead of onboard the motor vehicle, which requires power consumption. The edge computer system includes a server, of which the server will in turn synchronise the global model and the shared classifier with all the motor vehicles in the fleet, to complete one round of communication, such that a set of trained parameters generated are uniformly distributed. More advantageously, the aforesaid system yields a multimodal hybrid federated learning model to aggregate the features extracted from different units without sharing the raw data, thereby ensuring security, privacy, and performance.

[0037] In some embodiment, the global autoencoder is operable to encode an aggregated output data generated from the aggregated process as an input global data, the input global data for initiate a supervised federated learning model by the global autoencoder.

[0038] In some embodiment, the machine learning model is operable to initiate an unsupervised vertical federated learning model using the first set of unlabelled vehicular sensing data and the second set of unlabelled vehicular sensing data collected by the machine learning model.202407982

[0039] 7

[0040] In some embodiment, the machine learning model is operable to execute a deep canonical correlation analysis algorithm to correlate covariance of at least one modality datapoint between the first set of unlabelled vehicular sensing data collected and the second set of unlabelled vehicular sensing data collected, with the first output of labelled dataset and the second output of labelled dataset generated by the machine learning model, to identify at least one maximally correlated projection.

[0041] In a third aspect of this disclosure, an edge network for secure transfer of data amongst motor vehicles over the air is provided. The edge network comprises: a global autoencoder and a co-attention mechanism; a fleet of motor vehicles, each motor vehicle in the fleet comprising a machine learning model and a plurality of vehicular subsystems; and a wireless communication network in communication with the server and each motor vehicle in the fleet, the wireless communication network operable to receive and transmit data between the server and each motor vehicle in the fleet, (a) the global autoencoder is operable to execute an aggregated process to aggregate a first output of labelled dataset using a first set of unlabelled vehicular sensing data received by the server, and a second output of labelled dataset using a second set of unlabelled vehicular sensing data received by the server, of which the first set of unlabelled vehicular sensing data and the second set of unlabelled vehicular sensing data are collected by the machine learning model; (b) the co-attention mechanism is operable to correlate the first output of labelled dataset with the second output of labelled dataset to identify at least one hidden representation between the first output of labelled dataset and the second output of labelled dataset, (c) wherein the at least one hidden representation identified by the co-attention mechanism is classified by a shared classifier; and (d) a server is operable to synchronise a set of trained parameters with the global autoencoder, the shared classifier and each motor vehicle in the fleet.

[0042] Advantageously, the aforesaid edge network provides a solution for synchronising a set of trained parameters amongst a server, a fleet of motor vehicles in an aggregated manner, thereby ensuring the data are distributed amongst nodes and202407982

[0043] 8

[0044] devices of the edge network. Further thereto, since the unlabelled dataset collected by the machine learning model are transferred to global autoencoder, and the server label all the dataset received or label a certain number of auxiliary datasets, data processing of the unlabelled dataset are not done within the motor vehicle. In addition, the correlation of at least two sets of labelled datasets by a co-attention mechanism aids identification of at least one hidden variable or hidden representation between the sets of labelled datasets correlated. The hidden representation identified by the co-attention mechanism is classified by a shared classifier and the classification of hidden representation may be use for training the shared classifier, strengthening the multi-modal federal learning process. The edge computer system includes a server, of which the server will in turn synchronise the global model and the shared classifier with all the motor vehicles in the fleet, to complete one round of communication, such that a set of trained parameters generated are uniformly distributed. More advantageously, the aforesaid system yields a multimodal hybrid federated learning model to aggregate the features extracted from different units without sharing the raw data, thereby ensuring security, privacy, and performance.

[0045] BRIEF DESCRIPTION OF DRAWINGS

[0046] The present disclosure will become more fully understood from the detailed description and the accompanying drawings, wherein:

[0047] FIG. 1 shows a flowchart for a method for securely transferring data amongst motor vehicles in accordance with an embodiment.

[0048] FIG. 2 shows a flowchart for correlating labelled datasets for identifying variant or hidden representations in accordance with an embodiment.

[0049] FIG. 3 shows a schematic of a system for securely transferring data amongst motor vehicles in accordance with an embodiment.202407982

[0050] 9

[0051] FIG. 4a shows a deep canonically correlated autoencoder in accordance with a preferred embodiment.

[0052] FIG. 4b shows a proposed separated deep canonically correlated autoencoder in accordance with a preferred embodiment.

[0053] DETAILED DESCRIPTION OF EMBODIMENTS

[0054] It should be understood that like reference numerals identify corresponding or similar elements throughout the several drawings. It should be understood that although a particular component arrangement is disclosed and illustrated in these exemplary embodiments, other arrangements could also benefit from the teachings of this disclosure.

[0055] Hereinafter, the term “first”, “second”, “third” and the like used in the context of this disclosure may refer to modification of different elements in accordance with various exemplary embodiments, but not limited thereto. The expressions may be used to distinguish one element from another element, regardless of sequence of importance. By way of an example, “a first output data” and “a second output data” may indicate two different output data generated by a processing unit. On a similar note, “a first hidden representation” may be referred to as the “second hidden representation” and vice versa without departing from the scope of this disclosure.

[0056] Method 100

[0057] FIG. 1 shows a flowchart for a method 100 for securely transferring data amongst motor vehicles in accordance with an embodiment. At step 102, a global autoencoder executes an aggregated process. The aggregated process generates a first output of labelled dataset using a first set of unlabelled vehicular sensing data collected by a machine learning model onboard at least one motor vehicle in a fleet, of which the first set of unlabelled vehicular sensing data is transmitted to a server, comprising the global autoencoder. Similarly, the aggregated process generates a second output of labelled dataset using a second set of unlabelled vehicular sensing202407982

[0058] 10

[0059] data collected by the machine learning model, the second set of unlabelled vehicular sensing data sent from the machine learning model to the server side of the architecture. In some embodiment as shown in FIG. 3, the vehicular sensing data may be collected from different sources of vehicular sensing devices onboard each of the motor vehicle 302, 302’, 302”, 302”’, for example an image device 306, 306’ and a sensor for road detection, such as LiDAR 308, 308’. In some embodiment, the sensing device may be from a single, integrated vehicular system such as a surround view system 304.

[0060] In a next step at 104, a co-attention mechanism correlates the first output of labelled dataset generated by the global autoencoder with the second output of labelled dataset generated by the global autoencoder, to identify at least one hidden variable or hidden representation between the first output of labelled dataset and the second output of labelled dataset. As with the global autoencoder, the co-attention mechanism may be part of mechanism in the server. The co-attention mechanism function to identify and amplify the relevant latent or hidden representations’ features, to beneficially obtain a more fine-grained task-driven correlation distribution.

[0061] At step 106, a shared classifier classifies hidden representation in the first output of labelled dataset or the second output of labelled dataset using the at least one hidden representation identified from the previous step 104. In all embodiments, the classified hidden representation(s) is used as input data for training the shared classifier.

[0062] At step 108, the server synchronises a set of trained parameters yielded from all of the above sequence with the global autoencoder, the shared classifier and each motor vehicle in the fleet to complete one cycle of communication.

[0063] In some embodiment, the step of correlating the first output labelled dataset and the second output labelled from step 104 includes encoding 110 the first output of labelled dataset and the second output of labelled dataset with the at least one202407982

[0064] 11

[0065] hidden representation identified. The correlation is carried out by the co-attention mechanism 310, as shown in FIG. 3.

[0066] In some embodiment, the step of correlating 104 the first output of labelled dataset with the second output of labelled dataset for identifying at least one hidden representation from step 104 includes identifying 112 pairwise datapoint from the first output labelled dataset with the second output of labelled dataset by way of an affinity matrix.

[0067] In all of the embodiments, the method includes a step of correlating covariance of at least one multimodality datapoint between the first set of unlabelled vehicular sensing data collected and the second set of unlabelled vehicular sensing data collected, with the first output of labelled dataset and the second output of labelled dataset generated by the machine learning model, for identifying at least one maximally correlated projection, using a deep canonical correlation analysis. The deep canonical correlation analysis is executable by the machine learning model located onboard each of the motor vehicle in the fleet.

[0068] In some embodiments, the step of correlating covariance of at least correlating covariance of at least one multimodality datapoint for identifying at least one maximally correlated projection includes maintaining an individual autoencoder for each of the at least one multimodality datapoint, for maximising the canonical correlation between the hidden representations from two different multimodality datapoints. This may be executable by an integrated deep canonical correlation analysis model.

[0069] In some embodiment, the step of correlating covariance of at least one multimodality datapoint for identifying at least one maximally correlated projection includes executing a forward propagation process using a separated deep canonically correlated autoencoder. The forward propagation process includes exchanging a first multimodality datapoint containing a first hidden representation identified from the first output of labelled dataset with a second multimodality datapoint containing a second hidden representation identified from the second output of labelled dataset,202407982

[0070] 12

[0071] thereby obtaining a corresponding encoder for each of the first hidden representation and the second hidden representation. In such embodiment, the exchange of multimodality datapoint may be through an encrypted transmission channel 402, as shown in FIG. 4b. Consequently, data may be transmitted or transferred in a secure manner.

[0072] In some embodiment, the method includes a step of encoding an aggregated global output data generated from the aggregated process as an input global data for initiating a supervised federal learning model by the global autoencoder.

[0073] In some embodiment, the method includes a step of initiating an unsupervised vertical federated learning model using the first set of unlabelled vehicular sensing data and the second set of unlabelled vehicular sensing data collected by the machine learning model. This is executable by the machine learning model located onboard each of the motor vehicles in the fleet.

[0074] In some embodiment, the method includes a step of generating a similarity score for measuring a difference between a set of local parameters stored in the at least one machine learning model and a set of global parameters stored in the global autoencoder.

[0075] In some embodiments, the deep canonical correlation analysis is an integrated deep canonical correlation analysis framework. In some embodiments, the deep canonical correlation analysis is a separate deep canonical correlation analysis framework.

[0076] System 300

[0077] FIG. 3 shows a schematic of a system 300 for securely transferring data amongst motor vehicles in accordance with an embodiment. The system 300 is operable to securely transfer data amongst a fleet of motor vehicles 302, 302’, 302”, 302”’. As shown in FIG. 3, a global autoencoder 314 is initialised on a server side of the framework as a global model. The server may randomly select motor vehicles 302,202407982

[0078] 13

[0079] 302’, 302”, 302”’ in the fleet for communication. This may be done wirelessly, through transmission of signals via a wireless network A vehicle side synchronises the global autoencoder 314 with a machine learning model and train the local encoder using unlabelled dataset collected from the motor vehicle. In some embodiment, the collection of unlabelled dataset may be from two different types of vehicles sensing devices, such as image unit 306 and LIDAR unit 308. In some embodiment, the collection of unlabelled dataset may be from a single vehicular system 304 having at least two type of sensing devices.

[0080] The system 300 includes a deep canonical correlated autoencoder. In some embodiment, the deep canonical correlated autoencoder may an integrated deep canonical correlation autoencoder. In some embodiment, the deep canonical correlated autoencoder may be a separated deep canonically correlated autoencoder. Beneficially, a separated deep canonically correlated autoencoder allows the separated processing units to exchange encoded hidden representations significantly smaller than the original data to maximum the correlation objective. This approach not only greatly reduces communication costs between the processing units but also extracts correlated hidden representations from the two modalities of data. Data collected from the machine learning models are sent to the server for aggregation. An adaptive aggregation method that is described herein achieves objective of accelerating and stabilizing the convergence of the autoencoders. After aggregating the local models into the global model on the server side, the server encodes the labelled dataset using the global model, and the encoded latent representations are used as inputs to train a shared classifier. To obtain a more fine-grained task-driven correlation distribution, it is the work of this invention to incorporate a co-attention mechanism 310 to identify and amplify the latent representations’ features relevant to the current task. Finally, the server synchronizes the global model and the shared classifier 312 with all the vehicles 302, 302’, 302”, 302’”, completing one round of communication.

[0081] The system 300 includes a global autoencoder 314, a co-attention mechanism 310, a shared classifier 312 and a server (cloud). The global autoencoder 314 is operable to execute an aggregated process to aggregate a first output of labelled dataset202407982

[0082] 14

[0083] using a first set of unlabelled vehicular sensing data received, and a second output of labelled dataset using a second set of unlabelled vehicular sensing data received, of which the first set of unlabelled vehicular sensing data and the second set of unlabelled vehicular sensing data are collected by the machine learning model. The co-attention mechanism 310 is operable to correlate the first output of labelled dataset with the second output of labelled dataset to identify at least one hidden representation between the first output of labelled dataset and the second output of labelled dataset. The shared classifier 312 is operable to classify the at least one hidden representation identified by the co-attention mechanism 310. The server is operable to synchronise a set of trained parameters with the global autoencoder 314, the shared classifier 312 and each motor vehicle 302, 302’, 302”, 302”’ in a fleet, to complete one round of communication. The set of trained parameters comprises one or more hidden variant or hidden representatives.

[0084] Beneficially, all the learning process and training processes are executed at different nodes and synchronised in throughout the network. Since intensive machine learning training and processes are not executed onboard the fleet of motor vehicles, power consumption for dataset training does not affect each of the motor vehicles in the fleet. Further, since the trained parameters are distributed and synchronised, all motor vehicles in the fleet can benefit from the different data collected from different motor vehicles, thereby improving accuracy and robustness of the set of parameters generated.

[0085] In some embodiment, an integrated deep canonical correlation analysis model may be applicable to execute correlation of covariance for identification of at least one maximally correlated projection. In some embodiment, a separated deep canonically correlated autoencoder may be implemented, to the deep canonical correlation analysis.

[0086] Communication-efficient training for separated deep canonical correlated autoencoder202407982

[0087] 15

[0088] For clarity and brevity, possible exemplary embodiments of canonical correlated autoencoder are as explained below.

[0089] 1) Canonically Correlated Analysis (CCA):

[0090] A CCA algorithm comprises a set of sequence to identify a least one maximally correlated projection (XTwx, YTwy)

[0091] (wx*, wy*) = arg max (wxTSXYwy) / √((wxTSXXwx)(wyTSYYwy))

[0092]

[0093] (1)

[0094] In this embodiment, let X ∈ Rd X nY ∈ Rd X n,, and (SXX, SYY) are covariance of each modality, SXYrepresents a cross-covariance.

[0095] To ensure the validity of following computation as the invariant objectives to achieve regularization, the following computation are used Sxx+ rxI and SY Y+ ryI are applicable, to constrain the projection for determining a unit variance.

[0096] — argmax: ^ SA-y(r;;

[0097] A W'l ='<’3

[0098]

[0099] (2)

[0100] In the above embodiment, wx, wyare representative of the matrices of top k projection vectors which are uncorrelated,

[0101]

[0102] < 0 jsrepresentative of i < j. The objective may be denoted using the following formulation:

[0103] max tr(WxTSXYWy) s.t. WxTSXXWx= WyTSYYWy= I

[0104]

[0105] (3)202407982

[0106] 16

[0107] using singular value decomposition to express the solution. T is defined as T = S1x~x1 / / 2SYV Y S’ Y-Y1 / 2then the objective is equivalent to the sum of the top k singular values of T. Let Uk, Vkbe representative of left- and right- singular matrices, that is:

[0108] max tr(UkTTVk) s.t. WxTSXXWx= WyTSYYWy= I

[0109]

[0110] (4)

[0111] The optimum value of (3) is obtained or generated, corresponding to the matrices, (

[0112]

[0113] (Wx, Wy) = SXX-1 / 2Uk, SYY-1 / 2Vk.

[0114] 2) Deep Canonically Correlated Analysis (DCCA)

[0115] A DCCA algorithm is an extension of CCA with deep neural network (DNN) framework, for extracting nonlinear features. In this embodiment using DCCA, the computation for extracting the nonlinear features may hidden representation

[0116]

[0117] extracted by DNNs, where dx', dy' denote the dimension of representation. DCCA identifies a maximum of correlation between Hxand Hy. Similar computation method as CCA, T is defined by correlating covariance of each modality (SXX, SYY, and the cross-covariance SXYmay be represented as:

[0118] SXX= (1 / N) HxHxTSYY= (1 / N) HyHyTSXY= (1 / N) HxHyT

[0119]

[0120] (5)202407982

[0121] 17

[0122] Then the objective is equivalent to find the sum of the top k singular values of T as represented by (3) as discussed above.

[0123] 3) Separated Deep Canonically Correlated Autoencoder

[0124] In some embodiment, the deep canonical correlation analysis model may be a separated deep canonically correlated autoencoder. For brevity, a conventional type of deep canonically correlated autoencoder (DCCAE) proposed in the following paper:

[0125] • Wang, R. Arora, K. Livescu, and J. A. Bilmes, “On deep multi-view representation learning," in Proceedings of the 32nd International Conference on Machine Learning, F. R. Bach and D. M. Blei, Eds., vol. 37, 2015, pp. 1083-1092.

[0126] In this known technique, DCCAE is to keep an individual autoencoder for each modality and seeks to maximize the canonical correlation between the hidden representations from two independent modalities, thereby allowing the encoded features to retain the information while learning information from a mutual correlation space within the framework.

[0127] FIG. 4a shows an exemplary structure of a deep canonically correlated autoencoder, where encoders fx, fyextract hidden representations from unlabelled dataset while decoder gx, gyreconstruct or generate the original input with hidden representations. To measure an error between the original input and the generated output, let Lx, Lybe the loss function, and let a be the hyperparameter to balance the reconstruction objectives to maximize the correlation. The loss function of correlation Lzand the objective of DCCAE may be defined as follows:

[0128]

[0129] (6)202407982

[0130] arg min (α(Lx+ Ly) + Lz)

[0131]

[0132] (7)

[0133] It shall be known to a skilled practitioner, a direct implementation of the DCCAE framework to motor vehicles is impossible due to limitation of hardware architecture of processing units for modality separation in motor vehicles. The existing training of DCCAE requires the transmission of raw data between different processing units, which leads to significant communication operational expenses and limitation on the utilization of only a single unit’s computational capacity.

[0134] Consequently, to address the aforesaid discussed issue with implementation of DCCAE directly to motor vehicles, in some embodiment as disclosed herein, a gradient-based training process of DCCAE may be applicable to ease searching for parameters. The objective function for each modality may be represented as follows:

[0135] For modality X: arg min α Lx+ LzFor modality Y: arg min α Ly+ Lz

[0136] Accordingly, it is the work of the present invention to propose a separated deep canonically correlated autoencoder. In some embodiment, the deep canonical correlation analysis may be executed by a separate deep canonically correlated autoencoder, such as one as shown in FIG. 4b. This is implemented by separating the autoencoders for two different modalities, X and Y. In a forward propagation process executable by the separated deep canonically correlated autoencoder, a first processing unit for modality X and a second processing unit for modality Y exchange latent representations obtained from each of the first processing unit and the second processing unit. The latent representations yield through the forward propagation process is significantly smaller in size as compared to the original data.202407982

[0137] Using the latent representations, the loss function of correlation, Lzcan be calculated. In some embodiment, calculation of the gradients in backward propagation in the autoencoder of modality X, Hyis assumed as a constant. Therefore, when computing the gradients of Lzin backward propagation, only Hyis required. Likewise, Lxdoes not affect the model parameter updates of the modality Y autoencoder during backward propagation. In some embodiment, the DCCAE can be trained in a separated manner, and after a number of iterations, the autoencoders can be transmitted to the server individually for subsequent aggregation and assembly.

[0138] Federated Aggregation Algorithm

[0139] In some embodiment, adaptive local model aggregation may be executed using a similarity score for measuring a difference between a local model and a global model. The similarity score may be represented by the follow formula:

[0140]

[0141] (9)

[0142] where n denotes the amount of data and a is a hyperparameter. |

[0143]

[0144] ||wlocalk- wglobk|| refers to L2 Norm, also known as Euclidean norm, of model parameters, for use in optimization of loss function. To accelerate the convergence and mitigate biased data from misleading the global model, an adaptive update may be implemented to using model-wise similarity score as discussed above:

[0145] K

[0146] ...glob glob,nk ( ( k ^,slob\, rk i.,glob\\wt+iwt + ZJVVT+I - FJ “ St+lAwo ~Wt J J k=l

[0147]

[0148] (10)

[0149] where / 3 is the hyperparameter controlling the size of step back. It is important to note that the step back is applied to the model from round t, not the updated model202407982

[0150] 20

[0151] from round t + 1. This is because the step size for local model updates is controlled by the learning rate. A direct application of step back to the updated model affects the information from parameter updates, thus making it difficult for the model to converge. By stepping back the model from the previous round, which has already been aggregated, the updated model may yield a similarity score closer to a guided model without affecting the parameter updates for the new round. An exemplary embodiment referred to as Adaptive Modal-specific AE Aggregation algorithm is shown below:

[0152] Algorithm 1 Adaptive Modal-specific AE Aggregation

[0153] Require: uf: parameters of autoencoder on client k at round t; rnx, my: parameters of modality X and F; / 3 hyperparameter of adaptive aggregation;

[0154] 1: initialize

[0155] 2: for each round t = 1, 2,... do

[0156] 3: Ct•— RandomSelect(m) Random select m clients 4: for each client k G Ctin parallel do

[0157] 5: if not wtkthen:

[0158] 6:

[0159] 7: end if

[0160] 8: it't M ClientUpdateAEfA’.

[0161]

[0162] 9 end for

[0163] 10: {εt+1} = SimilarityScore({wt+1k}k=1K, wglob)

[0164] 11: n u AdaAggregate({+xKU- ^h’b) 12: end for

[0165] 13: wx← wglob∧ mx

[0166] 14:

[0167]

[0168] wy← wglob∧ my

[0169] 15: return wglob, wx, wy

[0170] In various embodiments, the aggregated uploaded models, (3 controls the constraint on the model updates. A larger / 3 results in the updates being closer to the guided model. In various embodiments, when β = 0, the algorithm is equivalent to federated averaging (FedAvg).

[0171] Hidden Representation Learning with Co-Attention Mechanism

[0172] To leverage multimodal knowledge in supervised learning, an encoded labelled dataset for each modality X, Y, is used for hiding representation Hx, Hy. The affinity matrix M is calculated by the following formula:202407982

[0173] 21

[0174] M = tanh H^WafHy)

[0175] (11)

[0176] Where

[0177]

[0178] " is a trainable parameter. Consider M as a feature to learn and the attention maps are computed using the following:

[0179] / / , tank 11, • ( n, 4 / ,. ).w ) / / . tank (it, / / 7* HI dl, ] }!1) n sotlmax ( tr1, / / , I < / „ softmax ( a1II,, }

[0180]

[0181] (12)

[0182] Here, ax, aydenotes the normalized attention weights and affinity matrix M, MT* _ H ’. - JI ',, '. J‘ ’ ‘lt -. K-. “J ___ employs co-attention across two modalities.1- <> are the trainable weight parameters where h denotes the dimension of attention space. With the attention weights, the features may be sum up with attention score to attain the weighted representation i.e.,

[0183]

[0184] (13)

[0185] Based on the weighted representations of the correlated modalities, the fused feature is calculated by addition and used for training the shared classifier.

[0186] Thus, it can be seen, the aforesaid system and method can be implemented in an edge network in communication with a wireless communication network for receiving and transmitting data between the sever and each motor vehicle in a fleet, to yields a multimodal hybrid federated learning model to aggregate the features202407982

[0187] 22

[0188] extracted from different units without sharing the raw data, thereby ensuring security, privacy, and performance.

[0189] The foregoing description shall be interpreted as illustrative and not be limited thereto. One of ordinary skill in the art would understand that certain modifications may come within the scope of this disclosure. Although the different non-limiting embodiments are illustrated as having specific components or steps, the embodiments of this disclosure are not limited to those combinations. Some of the components or features from any of the non-limiting embodiments may be used in combination with features or components from any of the other non-limiting embodiments. For these reasons, the appended claims should be studied to determine the true scope and content of this disclosure.

Claims

20240798223Patent claims1. A method (100) for securely transferring data amongst motor vehicles, the method comprising;(a) executing (102), by way of a global autoencoder, an aggregation process for aggregating a first output of labelled dataset generated by a machine learning model using a first set of unlabelled vehicular sensing data and a second output of labelled dataset generated by the machine learning model using a second set of unlabelled vehicular sensing data;(b) correlating (104), by way of a co-attention mechanism, the first output of labelled dataset with the second output of labelled dataset for identifying at least one hidden representation between the first output of labelled dataset and the second output of labelled dataset;(c) training (106), by way of a shared classifier, classification of hidden representation using the that least one hidden representation identified; and(d) synchronising (108), by way of a server, a set of trained parameters with the global autoencoder, the shared classifier and each motor vehicle in the fleet, wherein the set of trained parameters comprises one or more hidden representations.

2. The method (100) according to claim 1, characterised by that (b) correlating (104) the first output labelled dataset and the second output labelled dataset comprises:(b)(i) encoding (110), by way of the co-attention mechanism, the first output of labelled dataset and the second output of labelled dataset with the at least one hidden representation identified.

3. The method (100) according to claim 1, characterised by that (b) correlating (104) the first output of labelled dataset with the second output of labelled dataset for identifying at least one hidden representation comprises:(b)(ii) identifying (112) pairwise datapoint from the first output labelled dataset with the second output of labelled dataset by way of an affinity matrix.202407982244. The method (100) according to claims 1-3, characterised by that the method (100) comprises:correlating, by way of a deep canonical correlation analysis executed by the machine learning model, covariance of at least one multimodality datapoint between the first set of unlabelled vehicular sensing data collected and the second set of unlabelled vehicular sensing data collected, with the first output of labelled dataset and the second output of labelled dataset generated by the machine learning model, for identifying at least one maximally correlated projection.

5. The method (100) according to claim 4, characterised by that(e) correlating covariance of at least one multimodality datapoint for identifying at least one maximally correlated projection comprises:(e)(i) maintaining, by way of an integrated deep canonical correlation analysis model, an individual autoencoder for each of the at least one multimodality datapoint, for maximising the canonical correlation between the hidden representations from two different multimodality datapoints.

6. The method (100) according to claim 4, characterised by that(e) correlating covariance of at least one multimodality datapoint for identifying at least one maximally correlated projection comprises:(e)(ii) executing, by way of a separated deep canonically correlated autoencoder, a forward propagation process comprising:exchanging, byway of an encrypted transmission channel (402), a first multimodality datapoint containing a first hidden representation identified from the first output of labelled dataset with a second multimodality datapoint containing a second hidden representation identified from the second output of labelled dataset, thereby obtaining a corresponding encoder for each of the first hidden representation and the second hidden representation.

7. The method (100) according to any of the preceding claims, characterised by that the method comprises:20240798225encoding, by way of the global autoencoder, an aggregated global output data generated from the aggregated process as an input global data for initiating a supervised federal learning model by the global autoencoder.

8. The method (100) according to any of the preceding claims, characterised by that the method (100) comprises:initiating, by way of the machine learning model, an unsupervised vertical federated learning model using the first set of unlabelled vehicular sensing data and the second set of unlabelled vehicular sensing data collected by the machine learning model.

9. The method (100) according to any of the preceding claims, characterised by that the method (100) comprise:generating, by way of an adaptive federated aggregation algorithm, a similarity score for measuring a difference between a set of local parameters stored in the at least one machine learning model and a set of global parameters stored in the global autoencoder.

10. The method (100) according to any of the preceding claims, characterised by that the deep canonical correlation analysis is selected from a group consisting of:an integrated deep canonical correlation analysis framework; a separate deep canonical correlation analysis framework, and combination thereof.

11. A system (300) for secure transfer of data amongst motor vehicles, the system (300) comprising:(a) a global autoencoder (314) operable to execute an aggregated process to aggregate a first output of labelled dataset using a first set of unlabelled vehicular sensing data received, and a second output of labelled dataset using a second set of unlabelled vehicular sensing data (306, 306’, 308, 308”) received by the server, of which the first set of unlabelled vehicular20240798226sensing data and the second set of unlabelled vehicular sensing data are collected by the machine learning model;(b) a co-attention mechanism (310) operable to correlate the first output of labelled dataset with the second output of labelled dataset to identify at least one hidden representation between the first output of labelled dataset and the second output of labelled dataset;(c) a shared classifier (312) operable to classify the at least one hidden representation identified by the co-attention mechanism;and(d) a server operable to synchronise a set of trained parameters with the global autoencoder, the shared classifier (312) and each motor vehicle (302, 302’, 302”, 302”’) in a fleet, wherein the set of trained parameters comprises one or more hidden representations.

12. The system (300) according to claim 10, characterised in that the global autoencoder (314) is operable to encode an aggregated output data generated from the aggregated process as an input global data, the input global data for initiate a supervised federated learning model by the global autoencoder.

13. The system (300) according to claim 10, characterised in that the machine learning model is operable to initiate an unsupervised vertical federated learning model using the first set of unlabelled vehicular sensing data and the second set of unlabelled vehicular sensing data collected by the machine learning model.

14. The system (300) according to claim 12, characterised in that the machine learning model is operable to execute a deep canonical correlation analysis algorithm to correlate covariance of at least one modality datapoint between the first set of unlabelled vehicular sensing data collected and the second set of unlabelled vehicular sensing data collected, with the first output of labelled dataset and the second output of labelled dataset generated by the machine learning model, to identify at least one maximally correlated projection.2024079822715. An edge network for secure transfer of data amongst motor vehicles over the air, the edge network comprising:a global autoencoder (314) and a co-attention mechanism (310); a fleet of motor vehicles (302, 302’, 302”, 302”’), each motor vehicle (302, 302’, 302”, 302’”), in the fleet comprising a machine learning model and a plurality of vehicular subsystems;anda wireless communication network in communication with the server and each motor vehicle (302, 302’, 302”, 302’”) in the fleet, the wireless communication network operable to receive and transmit data between the server and each motor vehicle in the fleet, characterised in that(a) the global autoencoder (314) is operable to execute an aggregated process to aggregate a first output of labelled dataset using a first set of unlabelled vehicular sensing data received by the server, and a second output of labelled dataset using a second set of unlabelled vehicular sensing data received by the server, of which the first set of unlabelled vehicular sensing data and the second set of unlabelled vehicular sensing data are collected by the machine learning model;(b) the co-attention mechanism (310) is operable to correlate the first output of labelled dataset with the second output of labelled dataset to identify at least one hidden representation between the first output of labelled dataset and the second output of labelled dataset,(c) wherein the at least one hidden representation identified by the co-attention mechanism (310) is classified by a shared classifier (312);and(d) a server is operable to synchronise a set of trained parameters with the global autoencoder (314), the shared classifier (312) and each motor vehicle (302, 302’, 302”, 302’”) in the fleet, wherein the set of trained parameters comprises one or more hidden representations.