Method and system for training neural network model through cross-island federated learning

By training identical copies of neural network models on local servers in different geographical locations through a cross-island federated learning system, and utilizing synchronous learning and gradient aggregation techniques, the problems of low training efficiency and accuracy in cross-border distributed systems are solved, achieving efficient and secure data utilization and improving the performance of vehicle driver assistance systems.

CN120898210APending Publication Date: 2025-11-04HARMAN INT IND INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380096793.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-04-24
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

In existing cross-border distributed systems, federated learning methods are affected by bandwidth limitations and differences in data distribution across different geographical locations, resulting in decreased training efficiency and model accuracy. This makes it difficult to effectively utilize cross-border data to train neural network models for vehicle driver monitoring systems and occupant monitoring systems.

Method used

A cross-island federated learning system is adopted, which trains identical copies of a neural network model on local servers in different geographical locations. By utilizing synchronous learning and gradient aggregation techniques, data exchange is reduced, improving training efficiency and accuracy. Specifically, this involves training multiple mini-batch datasets on each server, accumulating and aggregating gradients, and adjusting batch size and learning rate to optimize the model.

Benefits of technology

While ensuring data privacy and security, the training efficiency and accuracy of neural network models have been improved, bandwidth requirements have been reduced, and the system has been adapted to the differences in data distribution across different geographical locations, thereby enhancing the performance of vehicle driver assistance systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120898210A_ABST
    Figure CN120898210A_ABST
Patent Text Reader

Abstract

A method for improving the efficiency of training a neural network using synchronous learning within a federated learning system is provided. In one example, training a neural network model includes performing a weighted aggregation of a first set of cumulative gradients of a plurality of model parameters over a first plurality of batches of a first training data set and a second set of cumulative gradients of a second identical neural network model trained over a second plurality of batches of a second training data set; updating parameters of the neural network model based on the aggregation gradient; and in response to a decrease in accuracy of the neural network model, adjusting one or more of a batch size of the first plurality of batches, a first aggregate weighting coefficient of the first set of cumulative gradients, and a learning rate used during training.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to systems and methods for improving neural network model performance using cross-silo federated learning. BACKGROUND

[0002] Standards governing data privacy and security protect data privacy and ensure the safety of users’ data. In some cases, these standards govern the creation of data, the use of data, the storage of data, and the transfer of data between foreign countries. Since vehicle driver monitoring system (DMS) / occupant monitoring system (OMS) features are based on monitoring operators and occupants within a vehicle cabin, the data collected can be inherently personal. As such, these standards can hinder global businesses that rely on training data collected across multiple countries to develop machine learning (ML) / deep learning (DL) models for DMS / OMS features in vehicles. To ensure compliance with data privacy and security standards, federated learning (FL) can allow for training of ML / DL models among decentralized edge devices and / or servers that store local data samples without exchanging the data samples between them.

[0003] However, current methods for implementing FL in a cross-country distributed system are hindered by cross-country network communication bandwidth limitations, as bandwidth limitations can reduce training speed, and different data distributions of local data in different geographic locations can result in different convergence speeds and different model generalization forces during model training. Thus, due to the shortcomings of existing FL techniques, the training efficiency of models and model accuracy can be reduced. SUMMARY

[0004] The present disclosure addresses at least one of the above problems, in part, by a method for an advanced driver assistance system (ADAS) of a vehicle, the method comprising: adjusting an operation of the vehicle based on vehicle occupant images captured by an in-cabin monitoring system of the vehicle, wherein the in-cabin monitoring system relies on a neural network model trained using synchronous learning within a cross-silo federated learning (FL) system, and training the neural network model includes performing a weighted aggregation of a first set of accumulated gradients of a plurality of parameters of the neural network model over a first plurality of batches of a first training data set and a second set of accumulated gradients of a second same neural network model trained over a second plurality of batches of a second training data set; updating the parameters of the neural network model based on the aggregated gradients; and in response to an accuracy of the neural network model decreasing over a predetermined number of iterations, adjusting one or more of a batch size of the first plurality of batches, a first aggregation weighting factor assigned to the first set of accumulated gradients, and a learning rate. The in-cabin monitoring system can be a DMS or an OMS of the vehicle.

[0005] Accordingly, the proposed cross-island FL system advantageously makes global adjustments to the training parameters of different replicas of the same neural network model that are individually trained in a decentralized manner on different servers, based on gradients aggregated across the different replicas. In this way, multiple different replicas of a neural network can benefit from training data acquired in a cross-country distributed system, where data privacy and security standards prevent the direct exchange of data stored in servers located at different geographical locations, referred to herein as local servers. Each server can be located at a particular geographical location, and can generate and store data from that particular location. By accumulating gradients locally and then exchanging gradients between different servers, rather than sending training data to a centralized global neural network model, the amount of bandwidth used and the amount of time required to exchange data during training can be reduced, thereby improving training efficiency, while ensuring that personal data is maintained within a geographical region.

[0006] Other systems, methods, features, and advantages of the present disclosure will be or become apparent to one with skill in the art upon examination of the following drawings and detailed description. It is intended that all such additional systems, methods, features, and advantages be included within this description, be within the scope of the present disclosure, and be protected by the accompanying claims.

[0007] Unless otherwise defined, all technical and / or scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art in the field of the implementations described herein. Although methods and materials similar or equivalent to those described herein can be used in the practice or testing of implementations described herein, exemplary methods and / or materials are described below. In case of conflict, the patent specification, including definitions, will control. In addition, the materials, methods, and examples are illustrative only and are not intended to be limiting.

[0008] Implementation of the methods and / or systems described herein can involve performing or completing selected tasks manually, automatically, or a combination thereof. Moreover, according to actual instrument and equipment and methods of those skilled in the related art, several selected tasks could be implemented, by hardware, software, or firmware or combinations thereof, using an operating system.

[0009] For example, hardware for performing selected tasks according to embodiments described herein could be implemented as a chip or a circuit. As software, selected tasks according to embodiments described herein could be implemented as a plurality of software instructions being executed by a computer using any suitable operating system. In an exemplary embodiment, one or more tasks according to exemplary embodiments as described herein are performed by a data processor, such as a computing platform for executing a plurality of instructions. Optionally, the data processor includes a volatile memory for storing instructions and / or data and / or a non-volatile storage, for example, a magnetic hard disk and / or removable media. Optionally, a network connection is provided as well. A display and / or a user input device such as a keyboard or mouse are optionally provided as well. BRIEF DESCRIPTION OF DRAWINGS

[0010] Some embodiments are described herein, by way of example only, with reference to the accompanying drawings. With specific reference now to the drawings in detail, it is stressed that the particulars shown are by way of example and for purposes of illustrative discussion of embodiments described herein. In this regard, the description taken with the drawings makes apparent to those skilled in the art how embodiments described herein can be practiced.

[0011] In the drawings:

[0012] Figure 1 A vehicle including an ADAS according to one or more embodiments of the disclosure is shown systematically;

[0013] Figure 2 An example FL system for storing cross-country distributed vehicle data in a plurality of local servers according to one or more embodiments of the disclosure is shown;

[0014] Figure 3A A first example system for updating a local neural network model at a local server in a cross-country distributed system according to one or more embodiments of the disclosure is shown;

[0015] Figure 3B A second example system for updating a local neural network model at a local server in a cross-country distributed system according to one or more embodiments of the disclosure is shown;

[0016] Figure 4 A block diagram of an exemplary embodiment of a node of a FL system configured to train a neural network model according to embodiments of the disclosure is shown;

[0017] Figure 5 A method for training a neural network model using synchronous learning according to one or more embodiments of the disclosure is shown;

[0018] Figure 6 A method for accumulating, at a local server, multiple sets of gradients for respective multiple mini-batches of training data is shown in accordance with one or more embodiments of the present disclosure; and

[0019] Figure 7 A method for implementing ADAS intervention using a trained local neural network model is shown in accordance with one or more embodiments of the present disclosure. DETAILED DESCRIPTION

[0020] Methods are provided for improving the efficiency of training neural networks using synchronous learning within a federated learning (FL) system, where training data stored at multiple different servers is available for training a neural network without copying, moving, exchanging, or sharing the training data. For example, the servers can be located in different geographic locations, and the training data can include personal data that is protected by different laws applicable to the different geographic locations.

[0021] In one embodiment, the personal data includes image or facial image data of vehicle drivers in different locations, and the neural network is used to detect facial features and / or expressions of the drivers. For example, a driver monitoring system (DMS) or occupant monitoring system (OMS) of a vehicle can capture facial features and / or expressions of a driver, and the neural network can be trained to identify the driver, and / or predict a state (e.g., fatigue, alertness, anxiety, etc.) of the driver. However, it can be appreciated that the systems and methods described herein can be applied to other types of distributed data, and the neural network models described herein can be trained to perform other actions based on different types of data.

[0022] In the proposed cross-island FL system, during training, the same copy of the neural network is trained in a synchronized manner at different local servers (e.g., corresponding to different geographic locations). Each local neural network model of each local server can be trained by dividing a larger batch of training data into multiple smaller mini-batches. For example, the larger batch of training data can include 1280 training pairs, and the smaller mini-batches can include 64 training pairs. Each mini-batch can have a fixed batch size. Each mini-batch of training data is loaded into the respective local neural network model one by one, and a set of gradients of multiple network parameters computed by forward and backward propagation on each mini-batch of training data are collected and stored (e.g., accumulated).

[0023] Each local server can send multiple sets of accumulated gradients to other local servers. Each local server can aggregate the accumulated gradients received from other local servers and use the aggregated gradients to update the local neural network model stored in the associated local server by updating the parameters of the local neural network model using the aggregated gradients. Since the aggregated gradients are the same at each local server, the parameters of each local model are updated in the same way, so that at the end of each training iteration, the parameters of each local model will be the same.

[0024] In response to the decreasing model accuracy on the local validation dataset after a predetermined number of iterations, the batch size of the training data can be reduced by a predetermined amount across all local servers. The global learning rate can be adjusted based on the new batch size, and the aggregation weighting coefficients for specific local servers can be increased. By globally adjusting the parameters of each individual neural network model in response to the decreasing model accuracy, training efficiency and model accuracy can be improved. By training local neural network models using multiple mini-batches of training data on local servers, the computational requirements of the local servers can be reduced without increasing the implementation time of FL, since the time-constrained step occurs during the aggregation and accumulation of gradients.

[0025] After training the local neural network model, the trained neural network model at the local server can be stored in the memory of multiple vehicles located in the same geographic area as the local server, where training data used to train the local neural network model is collected. For example, the DMS and / or OMS can use this model to detect facial features of the vehicle operator and / or occupants, such as loading a driver profile. For example, the vehicle's DMS and / or OMS can be used by the vehicle's Advanced Driver Assistance Systems (ADAS), and ADAS can use the trained neural network model to determine the driver's emotional state based on the detected facial expressions. ADAS can then adjust the vehicle's operating parameters based on the operator's emotional state.

[0026] The following description relates to a system and method for training a local neural network model using synchronous learning at a local server within a multi-local server communication coupled to an FL system. The trained neural network model can be integrated into vehicles (such as...) Figure 1 The ADAS system of the vehicle shown. Figure 2 An example of an FL system that trains a neural network model on multinational distributed vehicle data is shown. Figure 3A and Figure 3B This illustrates a system in which the local neural network model is updated at an associated local server. Figure 4 An example node of an FL system configured to update a local neural network model is shown. It can be based on... Figure 5and Figure 6 The method described in the trains a local neural network model. The trained local neural network model can be used to determine a Figure 7 The method described in the enables an ADAS intervention.

[0027] Turning now to the drawings, Figure 1 An example vehicle 100 is schematically illustrated. The vehicle 100 includes an instrument panel 102, a driver seat 104, a first passenger seat 106, a second passenger seat 108, and a third passenger seat 110. In other examples, the vehicle 100 can include more or fewer passenger seats. The driver seat 104 and the first passenger seat 106 are located in a front portion of the vehicle, near the instrument panel 102, and can therefore be referred to as front seats. The second passenger seat 108 and the third passenger seat 110 are located in a rear portion of the vehicle and can be referred to as rear (or back) seats.

[0028] Additionally, the vehicle 100 includes a plurality of integrated speakers 114, which can be arranged around a periphery of the vehicle 100. In some embodiments, the integrated speakers 114 are electronically coupled to an electronic control system of the vehicle, such as the computing system 120, by a wired connection. In other embodiments, the integrated speakers 114 can be in wireless communication with the computing system 120. As an example, an occupant of the vehicle 100, such as a driver passenger, can select an audio file through the user interface 116, and the selected audio file can be projected through the integrated speakers 114. In some examples, an audio alert can be generated by the computing system 120 and can also be projected by the integrated speakers 114, such as will be explained in detail herein.

[0029] The vehicle 100 includes a steering wheel 112 and a steering column 122 through which a driver can input steering instructions for the vehicle 100. The vehicle 100 also includes one or more cameras 118. The camera 118 can be one of a plurality of cameras. In Figure 1 In the illustrated embodiment, the camera 118 is positioned to the side of the driver seat 104, which can facilitate monitoring the driver from the side. However, in other examples, the camera 118 can be located in other locations in the vehicle, such as on the steering column 122 and directly in front of the driver seat 104. Additionally, the camera 118 can be located to the side of the passenger seats 106, 108, and 110, or directly in front of the passenger seats 106, 108, and 110.

[0030] Additionally, as some examples, in Figure 1In the illustrated embodiment, the camera 119 can be positioned externally at the rear end of the vehicle 100, which can help monitor the position of the vehicle 100 in the lane and / or monitor the position of the vehicle 100 relative to other vehicles and / or the surrounding environment. In other examples, the camera 119 can be positioned in other locations in the vehicle, such as externally at the front end of the vehicle 100 or at the sides of the vehicle 100.

[0031] The camera 118 can include one or more optical (e.g., visible light) cameras, one or more infrared (IR) cameras, or a combination of optical and IR cameras with one or more perspectives. In some examples, the camera 118 can have an interior perspective as well as an exterior perspective. In some examples, the camera 118 can include more than one lens and more than one image sensor. For example, the camera 118 can include a first lens that directs light to a first visible light image sensor (e.g., a charge-coupled device or a metal-oxide-semiconductor) and a second lens that directs light to a second thermal imaging sensor (e.g., a focal plane array), thereby enabling the camera 118 to collect light of different wavelength ranges to produce both visible and thermal images. In some examples, the camera 118 can also include a depth camera and / or sensor, such as a time-of-flight camera or a LiDAR sensor.

[0032] In some examples, the camera 118 can be a digital camera configured to acquire a series of images (e.g., frames) at a programmable frequency (e.g., frame rate) and can be electronically and / or communicatively coupled to the computing system 120. Further, the camera 118 can output acquired images to the computing system 120 in real-time, such that the computing system 120 and / or a computer network can process these images in real-time. As used herein, the term “real-time” denotes a process that occurs instantaneously and without intentional delay. For example, “real-time” can refer to a response time of less than or equal to about 1 second. In some examples, “real-time” can refer to simultaneous or substantially simultaneous processing, detection, or recognition. Further, in some examples, the camera 118 can be calibrated with respect to a world coordinate system (e.g., world space x, y, z). In other examples, the camera 118 can acquire images to determine the state, attributes, and posture of a driver or passenger (e.g., occupant), such as will be explained in detail herein.

[0033] The vehicle 100 can also include a driver seat sensor 124 coupled to or within the driver seat 104 and a passenger seat sensor 126 coupled to or within the first passenger seat 106. The rear seats can also include seat sensors, such as a passenger seat sensor 128 coupled to the second passenger seat 108 and a passenger seat sensor 130 coupled to the third passenger seat 110. The driver seat sensor 124 and the passenger seat sensor 126 can each include one or more sensors, such as a weight sensor, a pressure sensor, and one or more seat position sensors, that output measurement signals to the computing system 120. For example, the computing system 120 can use the output of the weight sensor or the pressure sensor to determine whether the respective seat is occupied, and if so, the weight of the person occupying the seat. As another example, the computing system 120 can use the output of the one or more seat position sensors to determine one or more of: a seat height, a longitudinal position relative to the dashboard 102 and the rear seats, and an angle (e.g., recline) of a seat back of the corresponding seat.

[0034] In some examples, the vehicle 100 also includes a driver seat motor 134 coupled to or positioned within the driver seat 104 and a passenger seat motor 138 coupled to or positioned within the first passenger seat 106. Although not shown, in some embodiments, the rear seats can also include seat motors. The driver seat motor 134 can be used to adjust seat positions, including seat height, longitudinal seat position, and angle of a seat back of the driver seat 104, and can include adjustment inputs 136. For example, the adjustment inputs 136 can include one or more toggle switches, buttons, and switches. The passenger seat motor 138 can be used to adjust seat positions, including seat height, longitudinal seat position, and angle of a seat back of the first passenger seat 106, and can include adjustment inputs 140. The adjustment inputs 140 can include one or more toggle switches, buttons, and switches. Although not shown, in some embodiments, the rear seats can be adjusted in a similar manner.

[0035] The computing system 120 can receive inputs and output information to the user interface 116. The user interface 116 can be included in, for example, a digital cockpit, and can include a display and one or more input devices. The one or more input devices can include one or more touchscreens, knobs, dials, hard buttons, and soft buttons for receiving user inputs from vehicle occupants.

[0036] The computing system 120 includes a processor 142 configured to execute machine readable instructions stored in a memory 144. The processor 142 can be single or multi-core, and programs executed by the processor 142 can be configured for parallel or distributed processing. In some embodiments, the processor 142 is a microcontroller. The processor 142 can optionally include components distributed across two or more devices, which can be remotely located and / or configured for coordinated processing. In some embodiments, one or more aspects of the processor 142 can be virtualized and executed by a remotely accessible network computing device configured as a cloud computing configuration. For example, the computing system 120 can be communicatively coupled with the wireless network 132 through a transceiver 146, and the computing system 120 can communicate with network computing devices through the wireless network 132.

[0037] The computing system 120 can include a DMS 147 that can monitor a driver of the vehicle 100. For example, a camera 118 can be located on a front of the vehicle 100 (e.g., on the dashboard 102) or on a side of the vehicle 100 proximate to the driver seat 104, and can be positioned to view a face of the driver. The DMS can detect facial features of the driver. In some embodiments, the DMS 147 can be used to retrieve a driver profile of the driver based on the facial features, which can be used to customize a position of the driver seat 104, the steering wheel 112, and / or other components or software of the vehicle 100.

[0038] The computing system 120 can include an OMS 148 that can monitor one or more passengers of the vehicle 100. For example, a camera 118 can be located on a side of the vehicle 100 proximate to one or more of the passenger seats 106, 108, and 110, and can be positioned to view a face of a passenger of the vehicle.

[0039] The computing system 120 can include an ADAS 149 that can provide assistance to the driver based at least in part on the DMS 147. For example, the ADAS 149 can receive facial expression data from the DMS 147, and the ADAS 149 can process the facial expression data to provide assistance to the driver. For example, the ADAS 149 can process the facial expression data to determine whether the driver appears to be fatigued or stressed. In response to detecting a fatigue or stress condition of the driver, the ADAS 149 can alert the driver, or play music, or perform a different action to address the fatigue or stress condition of the driver.

[0040] As described in greater detail herein, trained neural network models can be integrated into the DMS 147, OMS 148, or ADAS 149, which can facilitate detection of facial features or expressions. For example, a trained model can utilize sensor and / or camera data from the camera 118 to determine whether ADAS intervention is required. The trained neural network model can be trained from driving data of multiple drivers. Further, the trained neural network model can be trained using synchronous learning and / or federated learning (FL). Reference is made to the following Figure 5 and Figure 6 Training of the neural network model is described in greater detail.

[0041] Additionally or alternatively, the computing system 120 can communicate directly with network computing devices through a short-range communication protocol such as Bluetooth®. In some embodiments, the computing system 120 can include other electronic components capable of performing processing functions, such as a digital signal processor, a field programmable gate array (FPGA), or a graphics board. In some embodiments, the processor 142 can include multiple electronic components capable of performing processing functions. For example, the processor 142 can include two or more electronic components selected from a plurality of possible electronic components, including a central processor, a digital signal processor, a field programmable gate array, and a graphics board. In still further embodiments, the processor 142 can be configured as a graphics processing unit (GPU), including parallel computing architecture and parallel processing capabilities.

[0042] Further, the memory 144 can include any non-transitory, tangible computer- readable medium having stored therein programming instructions. As used herein, the term “tangible computer-readable medium” is expressly defined to include any type of computer- readable storage. The example methods described herein can be implemented using encoded instructions (e.g., computer-readable instructions) stored on non-transitory computer- readable media such as flash memory, read-only memory (ROM), random-access memory (RAM), a cache or any other storage medium, in which the information is stored for any duration (e.g., for extended periods of time, permanently, for brief instances, for temporarily buffering, and / or for caching of the information).

[0043] The computer memory of the computer-readable storage medium referred to herein can include volatile and non-volatile or removable and non-removable media used to store electronically formatted information, such as computer-readable program instructions or modules of computer-readable program instructions, data, and the like, which can be independent or part of a computing device. Examples of computer memory can include any other medium that can be used to store information in the desired electronic format and that is accessible by at least a portion of one or more processors or computing devices. In various embodiments, the memory 144 can include an SD memory card, an internal and / or external hard drive, a USB memory device, or similar modular memory.

[0044] Further, in some examples, the computing system 120 can include a plurality of subsystems or modules tasked with performing specific functions related to performing image acquisition and analysis. As used herein, the term "system," "unit," or "module" can include a hardware and / or software system that operates to perform one or more functions. For example, a module, unit, or system can include a computer processor, controller, or other logic-based device that performs operations based on instructions stored on a tangible and non-transitory computer readable storage medium, such as a computer memory. Alternatively, a module, unit, or system can include a hard-wired device that performs operations based on hard-wired logic of the device. The various modules or units shown in the figures can represent hardware, software, or a combination thereof that operates based on software or hard-wired instructions to perform the operations.

[0045] Figure 2 A cross-national distributed system 200 is shown for utilizing networked vehicle data and resources from multiple local servers from different geographic regions, which have different data protection laws, to prevent sharing of personal data between different geographic regions. Each local server can be a cloud network. Each local server can be considered a node of a FL system, as described in more detail below.

[0046] Each local server (e.g., node) can include resources (e.g., memory, processors) that can be allocated to update a local neural network model and store and execute instructions to update the local neural network model based on vehicle data collected from multiple vehicles in different geographic locations on the local server (e.g., as shown in FIG. 3 and Figure 4 without sharing the vehicle data.

[0047] Each local server can store data that can be used to initialize and train one or more local neural network models. The cross-country distributed system 200 includes a first local server 212 for a first region 202, a second local server 214 for a second region 204, a third local server 216 for a third region 206, and a fourth local server 218 for a fourth region 208. However, it can be appreciated that other embodiments of the cross-country distributed system 200 can include fewer or additional regions and corresponding local servers without departing from the scope of the present disclosure.

[0048] The first local server 212, the second local server 214, the third local server 216, and the fourth local server 218 can each be one of a plurality of local servers. Each of the plurality of local servers can be communicatively coupled to a local database that stores vehicle data, such as image data or video data, collected within a particular geographic location. The first region 202, the second region 204, the third region 206, and the fourth region 208 can correspond to particular geographic locations. As one example, the first region 202 can be associated with the United States, the second region 204 can be associated with India, the third region 206 can be associated with China, and the fourth region 208 can be associated with the United Kingdom.

[0049] In this manner, the first local server 212 can receive vehicle data from a plurality of vehicles 222 located in the first region 202, the second local server 214 can receive vehicle data from a plurality of vehicles 224 located in the second region 204, the third local server 216 can receive vehicle data from a plurality of vehicles 226 located in the third region 206, and the fourth local server 218 can receive data from a plurality of vehicles 228 located in the fourth region 208. The vehicle data collected at a particular geographic location can be used as training data for a local neural network model trained at the local server of that particular geographic location. The plurality of local servers can be communicatively coupled, thereby enabling each local server to transfer data, including but not limited to, sets of accumulated gradients, batch sizes of training data, and other data, to other local servers to train the plurality of local neural network models using synchronous learning.

[0050] While multiple local servers can be communicatively coupled, the local servers can be configured according to a set of instructions to prevent certain data from being exchanged between the multiple servers. For example, in one embodiment, vehicle data of a particular geographic location can not be accessible to a different local server located in a different geographic location. Specifically, a first local server 212 of a first region 202 (e.g., located in the United States) can not directly receive vehicle data from a plurality of vehicles 224 located in a second region 204 (e.g., located in India). Rather, the plurality of local servers can indirectly receive vehicle data from different geographic locations in the form of multiple sets of accumulated gradients computed at each local server.

[0051] Figure 3A A cross-island FL system 300 based on a cross-country distributed system is shown, which can be the cross-country distributed system described above with respect to Figure 2 The cross-country distributed system includes a plurality of vehicles 222 located in a first region 202, a plurality of vehicles 224 located in a second region 204, and a plurality of vehicles 226 located in a third region 206. Vehicle data of the plurality of vehicles 222, 224, and 226 can be stored in a first local database 302, a second local database 304, and a third local database 306, respectively. It can be appreciated that the cross-country distributed system can include additional or fewer regions and corresponding plurality of vehicles and local databases without departing from the scope of the present disclosure. The cross-island FL system 300 includes a plurality of local servers, such as a first local server 212, a second local server 214, and a third local server 216, which can be located within the first region 202, the second region 204, and the third region 206, respectively. The first local server 212, the second local server 214, and the third local server 216 can receive batch training data from the local databases 302, 304, and 306, respectively. To train the local neural network model, a first training data set 303 can be generated from vehicle data from the plurality of vehicles 222 and stored in the first local database 302; a second training data set 305 can be generated from vehicle data from the plurality of vehicles 224 and stored in the second local database 304; and a third training data set 307 can be generated from vehicle data from the plurality of vehicles 226 and stored in the third local database 306.

[0052] Each batch of training data received at each local server can train one of a plurality of local neural network models. For example, the first training data set 303 can be used to train a first local neural network model 312; the second training data set 305 can be used to train a second local neural network model 314; and the third training data set 307 can be used to train a third local neural network model 316. The first, second, and third local neural network models 312, 314, and 316 can be trained simultaneously using synchronous learning, as described in greater detail below.

[0053] During the computation phase 320 of training, during each iteration of training the first, second, and third local neural network models 312, 314, and 316, a first set of accumulated gradients computed during the backpropagation phase of training can be stored at the first local server 212, a second set of accumulated gradients computed during the backpropagation phase of training can be stored at the second local server 214, and a third set of accumulated gradients computed during the backpropagation phase of training can be stored at the third local server 216. The accumulation of gradients is described in greater detail below with respect to Figure 5 The accumulation of gradients is described in greater detail below with respect to

[0054] Further, during the communication phase 321 of training, during each iteration of training the first, second, and third local neural network models 312, 314, and 316, the first set of accumulated gradients from the first local server 212 can be transmitted to the second and third local servers 214 and 216; the second set of accumulated gradients from the second local server 214 can be transmitted to the first and third local servers 212 and 216; and the third set of accumulated gradients from the third local server 216 can be transmitted to the first and second local servers 212 and 214.

[0055] The training of the first, second, and third local neural network models 312, 314, and 316 can include aggregating the first, second, and third sets of accumulated gradients at each of the first, second, and third local servers 212, 214, and 216. At the end of each iteration of training the first, second, and third local neural network models 312, 314, and 316, the parameters of the first, second, and third local neural network models 312, 314, and 316 can be updated based on the same set of aggregated gradients. The aggregation of gradients is described in greater detail below with respect to Figure 5 The aggregation of gradients is described in greater detail below with respect to

[0056] Figure 3BThe cross-isl and FL system 300 is shown during a validation phase in which the accuracy of the plurality of local neural model network models is validated. Validating the accuracy of the first local neural network model 312, the second local neural network model 314, and the third local neural network model 316 includes testing a first performance of the first local neural network model 312 based on a local validation dataset 352 collected at the first local server 212, testing a first performance of the second local neural network model 314 based on a local validation dataset 354 collected at the second local server 214, and testing a first performance of the third local neural network model 316 based on a local validation dataset 356 collected at the third local server 216.

[0057] The local validation dataset 352 can include data from vehicles 222 of the first region 202; the local validation dataset 354 can include data from vehicles 224 of the second region 204; and the local validation dataset 356 can include data from vehicles 226 of the third region 206, such that validation data is not shared between different regions. Thus, the performance of each local neural network is evaluated based on local data. In other words, each neural network model is trained on a local server associated with a certain geographic region, such that a first training dataset from a first geographic region does not include data from a second geographic region, and a second training dataset does not include data from the first geographic region. Additionally, the training of each neural network model on the respective local server is performed without receiving or exchanging training data between geographic regions.

[0058] It can be appreciated that the examples provided are exemplary and do not limit the scope of the present disclosure. The examples provided can include the cross-isl and FL system 300 with additional or fewer local servers in the plurality of local servers. As one example, the cross-isl and FL system 300 can include a fourth local server in which a fourth neural network model is trained using synchronous learning.

[0059] Reference is made to Figure 4 , a local server 400 of a cross-isl and FL system, such as the cross-isl and FL system 300 of Figure 3A and Figure 3B The local server 400 includes a node 402 of a cross-isl and FL system for training a neural network on vehicle data of a cross-country distributed system, such as the cross-country distributed system 200 of Figure 2

[0060] ​The node 402 includes one or more processors 404 configured to execute machine readable instructions stored in a non-transitory memory 406. For example, the processor 404 can be any suitable processor, processing unit, or microprocessor. The processor can be a multi-processor system and thus can include one or more additional processors that are identical or similar to each other and communicatively coupled by an interconnection bus. The processor 404 can be single- or multi-core and programs executing thereon can be configured for parallel processing or distributed processing. In some embodiments, the processor 404 can optionally include components distributed among two or more devices, which can be remotely located and / or configured for coordinated processing. In some embodiments, one or more aspects of the processor 404 can be virtualized and executed by remotely accessible networked computing devices configured as a cloud computing configuration.

[0061] The non-transitory memory 406 can include one or more data storage structures, such as optical, magnetic, or solid-state memory devices, for storing programs and routines executed by the local server’s processor to implement the various functions disclosed herein. The memory can include any desired type of volatile and / or non-volatile memory, such as, for example, static random access memory (SRAM), dynamic random access memory (DRAM), flash memory, read-only memory (ROM), and the like.

[0062] The non-transitory memory 406 can include a cumulative gradient module 408, an aggregated gradient module 410, a neural network module 412, a training module 414, and a communication module 416. In some embodiments, the processor 406 can include components disposed at two or more devices, which can be remotely located and / or configured for coordinated processing. In some embodiments, one or more aspects of the non-transitory processor 406 can include remotely accessible networked computing devices configured as a cloud computing configuration.

[0063] The neural network module 412 can include one or more trained and / or untrained neural networks, and can also include various data or metadata related to the one or more neural networks stored therein. The neural networks stored in the neural network module 412 can be trained using synchronous learning as part of a FL learning system, as described in greater detail herein. The neural network module 412 can include a plurality of local neural network model parameters, batch sizes, weighting coefficients, global learning rates, and the like. The training module 414 can include instructions for training one or more of the neural networks stored in the neural network module 412. In particular, the training module 414 can include instructions that, when executed by the processor 404, cause the node 402 to perform one or more steps of the method 500 for training a neural network model to detect facial features of a vehicle driver using data collected from vehicles in different geographic regions participating in a FL learning system, as discussed in greater detail below with reference to Figure 5 are discussed in greater detail.

[0064] In various embodiments, the facial feature detection performed by the trained neural network model can be used in a DMS / OMS system of a vehicle. For example, the DMS / OMS system can rely on the trained neural network model to determine a driver profile of a vehicle driver, which is discussed in greater detail below with reference to Figure 7 In some embodiments, the training module 414 includes instructions for implementing one or more gradient descent algorithms, applying one or more loss functions and / or training routines for adjusting parameters of one or more neural networks.

[0065] The accumulated gradient module 408 can include a plurality of sets of accumulated gradients computed during training of a neural network model. The aggregated gradient module 410 can include an aggregation of a plurality of sets of accumulated gradients received from other neural network models trained on other servers in other regions of a cross-country distributed system.

[0066] The node 402 can include a communication module 416. The communication module 416 can facilitate transmission of electronic data (e.g., accumulated gradient data, batch size data, learning rates, etc.) used or generated during training of a neural network model to other local servers having other neural network models in other geographic locations of a cross-country distributed system.

[0067] Communication by the communication module 416 can be implemented using one or more protocols. The communication module can be a wired interface (e.g., a data bus, a universal serial bus (USB) connection, etc.) and / or a wireless interface (e.g., radio frequency, infrared, near field communication (NFC), etc.). For example, the communication module can communicate over a wired local area network (LAN), a wireless LAN, a wide area network (WAN), etc. using any past, present, or future communication protocol (e.g., BLUETOOTH™, USB 2.0, USB 3.0, etc.).

[0068] The node 402 can be operatively / communicatively coupled to a user input device 432 and a display device 434. The user input device 432 can include one or more of a touchscreen, a keyboard, a mouse, a trackpad, a motion-sensing camera, or other devices configured to enable a user to interact with and manipulate data within the node 402. The display device 434 can include one or more display devices that utilize virtually any type of technology. In some embodiments, the display device 434 can include a computer monitor. The display device 434 can be combined with the processor 404, the non-transitory memory 406, and / or the user input device 432 in a shared housing, or can be a peripheral display device, and can include a monitor, a touchscreen, a projector, or other display devices known in the art that can enable a user to interact with various data stored in the non-transitory memory 406, such as adjusting batch size and weighted coefficients of a set of accumulated gradients.

[0069] It can be appreciated that Figure 4 The illustrated node 402 is for illustration and not limitation. Another suitable node of a FL system can include more, less, or different components.

[0070] Turning now to Figure 5 , an example method 500 for training a neural network model using synchronous learning within a FL system is shown, such that a neural network model can be trained on data stored in different regions of a cross-national distributed system (e.g., the cross-national distributed system 200 of Figure 2 and / or the cross-national distributed system 300 of Figure 3A ). In the embodiments described with reference to Figure 5 , Figure 6 and Figure 7 , the cross-national distributed system is a vehicle system, where vehicle data is collected at each region of the vehicle system and is not shared with other regions. The vehicle data includes image data acquired from one or more cameras (e.g., the cameras 118 of Figure 1 ) installed in each vehicle of the vehicle system, where connection data is assumed to be available.

[0071] Connectivity data availability can include communicating with a connected fleet of vehicles or multiple databases and / or servers using V2V networks, V2I networks, cloud networks, etc. Image data is collected at each vehicle and transmitted to a local server in the area (e.g., local servers 212, 214, 216, and 218). Image data collected within each area can be processed at the local server of the area, where each local server acts as a node in the FL system, such as node 402 of FIG. 4. Figure 4 Instructions for implementing at least a portion of method 500 can be stored in non-transitory memory 406 and executed by processor 404 of node 402. Figure 4

[0072] In the embodiments described in Figure 5 , Figure 6 and Figure 7 , image data is used to train copies of a neural network model to detect facial features and / or expressions of drivers or passengers of vehicles. The trained neural network model can be installed on a plurality of vehicles. For example, the trained neural network model can be incorporated into a DMS / OMS of a vehicle, and detected facial features can be made available to an ADAS system of the vehicle, as described in more detail below with reference to Figure 7 .

[0073] Image data can be protected by privacy laws that prevent image data from being transmitted outside of the region. However, image data from all of the different regions can be used to train multiple copies of a neural network model stored in different regions without transmitting image data. It can be appreciated that the use of vehicle data is for illustration, and that method 500 can be applied to other types of training data collected at other cross-country distributed systems without departing from the scope of the present disclosure.

[0074] During synchronous learning, a copy of the neural network model is installed at each node of the FL system (e.g., at each local server). Each copy of the neural network model is trained on local data, but gradients computed during each training iteration at each copy of the neural network model can be shared between models. The gradients can be used to adjust training parameters of each copy of the neural network model.

[0075] ​At 502, the method includes performing forward and backward propagation using multiple mini-batches of training data on the local copy of the neural network model and accumulating (e.g., storing) the gradients from each mini-batch. Each mini-batch can have a fixed batch size. Unlike training the local copy of the neural network model (also referred to herein as the local neural network model) using larger batches of training data, using multiple smaller mini-batches of training data can reduce the computational load on the local server when training the local neural network model, while improving training efficiency and improving the accuracy of the local neural network model. In accordance with the described method, training on multiple mini-batches of training data can include accumulating multiple sets of gradients for multiple parameters. Figure 6 The described method, training on multiple mini-batches of training data can include accumulating multiple sets of gradients for multiple parameters.

[0076] In one embodiment, training the local copy of the neural network model includes inputting pairs of training data into the neural network, where each training pair includes an input image of a face of a vehicle driver and a ground truth classification of the driver. For example, the ground truth classification can be an identification of the driver, which can be used to load a driving profile of the driver at an inference stage. In other examples, the ground truth classification can be a level of fatigue of the driver, or a known emotional state of the driver, or a different characteristic of the driver.

[0077] During forward propagation, the input image data can be keyed into the input layer of the neural network. In various embodiments, the neural network can be a convolutional neural network (CNN) having various hidden layers. For example, each pixel of the input image can be input into an input node of the neural network. The image data can be multiplied by a parameter (e.g., a weight) of the neural network at each input node to generate an output, which can then be input into multiple nodes of a first hidden layer of the various hidden layers. The output of each hidden layer can be input into a next hidden layer until a final output (e.g., a classification) of the neural network is generated at an output layer of the neural network.

[0078] During backward propagation, the final output generated at the output layer is compared to the ground truth classification using one or more loss functions, and the computed loss is propagated backward through the various layers of the neural network to determine a gradient at each node of each layer of the neural network. The gradients can be accumulated and stored for use in updating the parameters at a later stage, as described below.

[0079] At 504, the method includes sending the accumulated gradients to other local neural network models at other nodes of the FL system. As described herein with respect to the described systems and methods, multiple local servers can be communicatively coupled, and when executed, the training modules stored at the local servers can be configured to train the local copy of the neural network model using the accumulated gradients from other local servers. For example, the training module stored at the local server can be configured to receive the accumulated gradients from other local servers and use the accumulated gradients to update the parameters of the local copy of the neural network model. Figure 4The instructions in the training module 414) can cause the processor to transmit the plurality of sets of accumulated gradients stored in the non-transitory memory to other local servers in the plurality of local servers.

[0080] At 506, the method includes receiving accumulated gradients from other local neural network models at other nodes of the FL system. In other words, the plurality of sets of accumulated gradients used during training can be exchanged among the plurality of local servers.

[0081] At 508, the method includes aggregating the accumulated gradients of each local neural network model received. The instructions stored in the training module can be executed by the processor to cause the processor to aggregate the plurality of sets of gradients received from the plurality of local servers. The gradient aggregation based on the plurality of sets of accumulated gradients of the plurality of local neural network models can be performed according to the following equation:

[0082]

[0083] wherein is the aggregated gradient, is the accumulated gradient, is a weighting coefficient of the accumulated gradient from the mth local server. The aggregated gradient can be stored in at least one address in the non-transitory memory of the plurality of local servers, such as Figure 4 the aggregated gradient module 410. The weighting coefficient can be an aggregation weighting coefficient that indicates a factor by which the accumulated gradient of a given neural network model can be multiplied to improve training efficiency.

[0084] For example, a cross-country distributed system can include two local servers, where the training can include aggregating a first set of accumulated gradients of a first local neural network model with a second set of accumulated gradients of a second neural network model received from the second server, the second neural network model being the same as the first neural network model (e.g., a copy of the same neural network), the second neural network being trained on a second plurality of mini-batches of a second training data set on the second server. Aggregating the first set of accumulated gradients with the second set of accumulated gradients can include summing a product of a first weighting coefficient and the first set of accumulated gradients with a product of a second weighting coefficient and the second set of accumulated gradients, the second set of accumulated gradients being received from the second server.

[0085] In one implementation, the product of the first weighting coefficient and the first set of accumulated gradients can be computed locally. In another implementation, for example, the first weighting coefficient can be sent from the first server and received by the second server, while the second weighting coefficient can be sent from the second server and received by the first server. After all the accumulated gradients are sent by the multiple servers, the product of the first weighting coefficient and the first set of accumulated gradients and the product of the second weighting coefficient and the second set of accumulated gradients can be computed globally. As one example, the first weighting coefficient can be 0.5, and the second weighting coefficient can be 0.5.

[0086] In some implementations, the first training data set of the first local neural network model can be smaller than the second training data set of the second local neural network model. A smaller first weighting coefficient can be selected to be applied to the accumulated gradients from the first training data set during the aggregation, and a larger second weighting coefficient can be selected to be applied to the accumulated gradients from the second training data set during the aggregation. It can be appreciated that the weighting coefficients can be automatically adjusted (e.g., without human input) based on validation results. The smaller value of the first weighting coefficient can be selected to prevent catastrophic interference. That is, the smaller value of the first weighting coefficient can be selected to prevent the other local neural network models from forgetting the first training data set, thereby ensuring that the training of the other local neural network models is based on the first training data set. In this way, the local network models do not forget how to complete previously learned tasks when encountering new tasks. In addition, the larger value of the second weighting coefficient is selected to improve the convergence speed of the second training data set of the second local neural network model.

[0087] At 510, the method includes updating the parameters (e.g., weights) of the local neural network model based on the aggregated gradients. The processor can execute the instructions stored in the training module to cause the processor to update the parameters based on the aggregated gradients stored in the non-transitory memory. The parameters can be updated according to the following formula:

[0088]

[0089] wherein is a learning rate generated by a learning rate scheduler that is scaled based on the batch size, is the aggregated gradient, and is a parameter of the local neural network model. There can be multiple parameters, each having its own respective aggregated gradient based on its own respective multiple sets of accumulated gradients. The instructions can also include storing the updated parameters in the non-transitory memory of each of the multiple local servers.

[0090] For example, in a cross-country distributed system including two local servers, the training can include updating parameters of a first local neural network model based on aggregated gradients at the first local server. Updating the parameters of the first local neural network can include updating each of the parameters at the first local server using Equation 2 based on the aggregated gradients associated with the parameters and a learning rate. Similarly, parameters of a second local neural network model can be updated based on aggregated gradients at the second local server, and updating the parameters of the second local neural network model can include updating each of the parameters at the second local server using the same equation based on the aggregated gradients associated with the parameters and a learning rate. For both the first and second local neural networks, the set of aggregated gradients, batch size, and learning rate can be the same, whereby the same set of updated parameters can be generated for both the first and second local neural networks for a subsequent training iteration.

[0091] At 512, the method includes checking accuracy of the local neural network model by a local validation dataset after a predetermined number of training iterations. In other words, each local neural network model of the plurality of local servers in the overall FL system can evaluate accuracy of the local neural network model by a local validation set after a predetermined number of iterations, as described above with respect to Figure 3B The local validation set can include validation pairs similar to the training pairs of the respective training dataset. For example, each training pair of the local validation set can include an image taken by a camera of a vehicle in the respective geographic region and a respective true classification.

[0092] At 514, the method includes determining whether the model accuracy on the local validation dataset is decreasing in a number of iterations of the local neural network model using the local validation dataset. For example, the model accuracy can be checked on the local validation set every three cycles. If the error rate of the local neural network model is constantly increasing with an increase in the number of iterations, then it can be determined that the accuracy is decreasing. If it is determined at 514 that the accuracy of the local neural network model is decreasing, then the method proceeds to 516.

[0093] At 516, due to the decrease in model accuracy, the method includes determining a new batch size, new weighting coefficients for a set of accumulated gradients of the local neural network model, and a new global learning rate based on the new batch size. In other words, to improve the accuracy of the local neural network model and balance the training speed, the batch size can be adjusted by a predetermined amount in the local server. In some embodiments, the batch size can be automatically (e.g., without human input) reduced by a predetermined amount for the local neural network model. The learning rate can be adjusted based on the new batch size using the learning rate scheduler described above. Different learning rate schedulers can be used, including LineaLR, Constant LR, Exponential, and CosineAnnealingLR. The frequency of operation of the learning rate scheduler can be every cycle, every n iterations, etc.

[0094] During the training initialization, the weighting coefficients for each local server can be set. The values of the weighting coefficients can be the same or different. As an example, the weighting coefficients can be one divided by the number of local servers. Each local server can have a target accuracy for the local neural network model. When estimating the local neural model accuracy using the local validation dataset during training, after the local neural network model converges at a particular local server, the model accuracy is compared to the target model accuracy for that particular local server. If the model accuracy is not within the threshold accuracy, the weighting coefficient for the particular local server can be decreased to improve the model accuracy and improve the rate of convergence during training.

[0095] Since the local neural network models are trained using synchronous training, all of the local neural network models have the same parameters, and the local neural network models are also global models. The global model can continually learn and achieve the target accuracy on all of the local validation datasets. The accuracy between different local validation datasets can be compared. The accuracy between a model trained based on global data and a model trained based on local data can be compared. The model trained based on global data can have equal or higher accuracy.

[0096] In some embodiments, the adjustments to the batch size, the weighting coefficients, and the global learning rate can be performed according to a predetermined set of instructions stored in the training module. In other embodiments, manual adjustments to the batch size, the weighting coefficients, and the global learning rate can be performed through a user input device communicatively coupled to the user input device, as shown in Figure 4 .

[0097] At 517, the method includes synchronizing the batch size with other local neural network models in the FL system. The adjusted batch size of each local server is sent to other local servers, and each local server receives the adjusted batch size from other local servers. The smallest adjusted batch size from all of the adjusted batch sizes of the local servers is selected as the new batch size for each local server. For example, for a cross-country distributed system that includes at least two local servers, the batch size of the first training dataset at the first local server is adjusted by a predetermined amount that can be different from the predetermined amount by which the batch size of the second training dataset at the second local server is adjusted. The adjusted batch size of both the first training dataset and the second training dataset at the first local server can be the smallest of the adjusted batch size of the first training dataset and the adjusted batch size of the second training dataset.

[0098] At 518, the method includes applying the adjusted batch size, the new weighting factor, and the updated learning rate to subsequent training iterations, and the method returns to 502 to continue training the local neural network model.

[0099] Returning to 514, if it is determined that the accuracy of the local neural network model has not decreased, i.e., the local neural network is continuing to increase, the method proceeds to 520. At 520, the method includes determining whether a termination condition of the local neural network has been reached. In some embodiments, the termination condition can be a maximum number of epochs or iterations of model training. The maximum number of epochs or iterations can be set prior to training. In other embodiments, the termination condition can include reaching a threshold accuracy. The threshold accuracy can depend on the type of training data used. For example, if the training data includes images of a driver’s face and associated driver identification, the threshold accuracy can be a 95% correct driver identification rate. If the model accuracy on a particular local server is continually increasing (e.g., during the first validation and the second validation), the weighting factor on that particular local server can be reduced by a predetermined amount until the local neural network model converges on that particular local server.

[0100] If the termination condition has not been reached, the method returns to 502, and the method includes continuing to train the local neural network model, as described above. Alternatively, if it is determined at 520 that the termination condition has been reached, the method proceeds to 522.

[0101] At 522, the method includes deploying the trained local neural network in the region corresponding to the node and the local server. The deployment of the trained local neural network will be described in more detail below with reference to Figure 7 The deployment of the trained local neural network is described in more detail. The method 500 ends.

[0102] It can be appreciated that training of different local neural network models can have different convergence rates in accumulating multiple sets of gradients and aggregating multiple sets of accumulated gradients. Since the topology of the cross-country distributed system is defined before the training starts, the time issue of non-uniform distribution of training data among multiple local servers can be addressed by configuring the system according to the instruction set to wait for the completion of the computation of multiple sets of accumulated gradients in each local server before aggregating the accumulated gradients of each iteration. In this way, each local neural network model can be trained on multiple mini-batches at different rates on the corresponding local server without affecting the aggregation of accumulated gradients and the premature update of multiple local neural network models.

[0103] Now turning to Figure 6 , an exemplary method 600 for collecting accumulated gradients during training of a neural network model at a node of a cross-island FL system, such as the cross-island FL system 300 of Figure 3A and Figure 3B In various embodiments, the local server can be a cloud-based server, such as the local server 400 of Figure 4 Instructions for implementing the method 600 can be stored on and executed by the cloud-based server. In one example, the method 600 is performed as part of the method 500 described above. Figure 5

[0104] At 602, the method includes obtaining a larger batch of training data. Instructions stored in a training module (e.g., the training module 414 of Figure 4 and executed by a processor can cause the processor to enable the local server to receive training data by accessing a local database, where the larger batch of training data is stored on the local server. The local database can be for a particular geographic location, such as, for example, China. The larger batch of training data can include vehicle data, such as image or video data, for the particular geographic location.

[0105] In some examples, the image data or video data can include images or videos of a vehicle cabin, where the subject of the images or videos are vehicle operators and / or occupants. In other examples, the image data or video data can include images or videos of the external environment surrounding the vehicle.

[0106] ​As an example, for a cross-country distributed system that includes at least two local servers, the larger batch of training data can include a first set of training data on a first local server and a second set of training data on a second local server. The first set of training data is stored in a first local database communicatively coupled to the first server and includes training data collected at a first geographic location, and the second set of training data is stored in a second local database communicatively coupled to the second server and includes training data collected at a second geographic location. The first set of training data includes image data and / or video data from a first plurality of vehicles located in the first geographic location, and the second set of training data includes image data and / or video data from a second plurality of vehicles located in the second geographic location. Instructions executable in each of the first and second local servers can cause the first local server to have access to the first set of training data and the second local server to have access to the second set of training data.

[0107] It can also be noted that the training of the first neural network model at the first local server and the training of the second neural network model at the second local server are performed without exchanging training data between the first and second servers. In this way, since the first and second servers are configured to not exchange training data between the first and second servers, the second server does not have access to the training data collected and stored in the first local database at the first geographic location, and the first server does not have access to the training data collected and stored in the second local database at the second geographic location.

[0108] At 604, the method 600 includes dividing the larger batch into smaller mini-batches. Instructions stored in the training module and executed by the processor can cause the processor to divide the larger batch of training data into a plurality of mini-batches of training data, where the size of the mini-batches is fixed during training. As noted above Figure 5 The local server stores information about the local neural network model in the memory of the local server (e.g., in the training module 414), such as the batch size or the adjusted batch size. In this way, the instructions divide the larger batch of training data into a plurality of mini-batches of training data based on the fixed batch size.

[0109] At 606, the method includes loading one mini-batch of localized training data into the local neural network model. Instructions stored in the training module and executed by the processor can cause the processor to load one mini-batch of training data into the local neural network model at the local server.

[0110] At 608, the method includes collecting a plurality of sets of gradients for a respective plurality of parameters of the local neural network model. The instructions stored in the training module and executed by the processor can cause the processor to train the local neural network model by performing forward propagation and backpropagation using the mini-batches of training data, as described above with reference to FIG. 4. Backpropagation of the local neural network model at the local server computes a plurality of sets of gradients for the respective plurality of parameters of the local neural network model. The gradients can be collected according to the following equation: Figure 5

[0111]

[0112] wherein is the training data in a mini-batch at each local server, is a loss function, is an index of the local server, and is a corresponding gradient.

[0113] At 610, the method includes computing cumulative gradients for the local neural network model. The plurality of sets of cumulative gradients can be computed by accumulating the plurality of sets of gradients computed by the local servers. Gradient accumulation can be performed according to the following equation:

[0114]

[0115] wherein is a cumulative gradient, is a corresponding gradient, and m refers to the mth local server. In this way, accumulating a set of gradients can include performing a summation of the plurality of sets of gradients that have been computed. In other words, a set of gradients computed during a backpropagation phase when training the local neural network model using a mini-batch of training data can be summed with other sets of gradients computed during a backpropagation phase when training the local neural network model using other mini-batches of training data. The plurality of sets of gradients for the respective plurality of parameters can be stored in an address in the non-transitory memory of the local server, such as the cumulative gradient module 408 of FIG. 4. Figure 4

[0116] At 612, the method 600 includes determining whether there is additional mini-batch training data to train the local neural network model. The instructions stored in the training module and executed by the processor can cause the processor to determine a total amount of mini-batches in the plurality of mini-batches with a batch size of the mini-batches. In this way, the instructions can be able to determine a number of mini-batches of localized training data that have been trained and a number of mini-batches of localized training data that have not been trained.

[0117] ​​In response to there being additional mini-batch training data to train the local neural network model, the method 600 includes loading one mini-batch of localized training data into the local neural network model at 606. As such, the method 600 continues to collect multiple sets of gradients for the respective multiple parameters and performs gradient accumulation to obtain a set of accumulated gradients for each local neural network model of the multiple local servers until there is no remaining mini-batch of localized training data. In response to there being no additional mini-batch training data to train the local neural network model, the method 600 then ends.

[0118] As shown in FIG. 7, Figure 7 an example method 700 for determining whether to adjust operation of a vehicle based on output of a trained neural network model for a vehicle’s ADAS is shown. The trained neural network model can be a trained version of a local neural network model trained in accordance with the methods described in Figure 5 and Figure 6 . The vehicle can operate within a geographic location of a cross-country distributed vehicle system, such as the cross-country distributed system 200 in Figure 2 . As the neural network model was trained using the cross-island FL system described herein (e.g., the cross-island FL system 300), the same neural network sharing the same parameters can be used to perform inference in multiple geographic regions of the cross-country distributed vehicle system independently of the neural network models in other geographic regions without relying on a network. Instructions for implementing the method 700 can be stored in a non-transitory memory and executed by a processor of a vehicle computing system, such as the computing system 120 of Figure 1 and / or the vehicle’s ADAS, such as the ADAS 149 of Figure 1 .

[0119] The method 700 begins at 702, where the method includes estimating and / or measuring vehicle operating conditions. The vehicle operating conditions can be estimated based on one or more outputs of various sensors of the vehicle, such as oil temperature sensors, engine or wheel speed sensors, torque sensors, etc. The vehicle operating conditions can include engine speed and load, vehicle speed, transmission oil temperature, exhaust gas flow, air mass flow, coolant temperature, coolant flow, engine oil pressure (e.g., oil gallery pressure), operating mode of one or more intake and / or exhaust valves, electric motor speed, battery charge, engine torque output, wheel torque, etc. Estimating and / or measuring the vehicle operating conditions can include determining whether the vehicle is being driven by the engine or the electric motor.

[0120] Estimating and measuring the vehicle operating conditions can also include capturing images of the vehicle’s surroundings by one or more cameras and / or sensors positioned on or in the vehicle, such as the cameras 145 of Figure 1The camera 118 of the vehicle determines the proximity of other vehicles in the vehicle’s lane or other lanes to the vehicle. The vehicle operating conditions can include road conditions (e.g., whether the road is paved or unpaved, wet, icy, etc.), weather conditions (e.g., whether it is raining or snowing), and / or other external conditions of the vehicle (e.g., whether it is daytime or nighttime, etc.).

[0121] At 704, the method includes receiving image data of a driver of the vehicle from a DMS of the vehicle. In various embodiments, the image data can be received from a camera installed in a cabin of the vehicle, such as the camera 118 of the vehicle. For example, the DMS can capture an image of the driver when the driver is looking at a forward road segment, including traffic and / or other vehicles in front of the vehicle. The DMS can capture an image of the driver looking out of one or more side windows of the vehicle and / or looking at a rearview mirror. The DMS can capture an image of the driver interacting with other occupants of the vehicle. Figure 1

[0122] At 706, the method includes processing the image data to determine whether to adjust operation of the vehicle based on the image data. In the depicted embodiment, the ADAS can adjust operation of the vehicle in response to a predicted alertness level of the driver falling below a threshold alertness level. For example, the alertness level of the driver can be predicted based on a facial expression of the driver, signs of fatigue in the driver’s posture, such as slouching or hunching, a period of time in which the driver is not moving, detected movement or slippage of the driver’s eyelids, etc.

[0123] At 708, processing the image data to determine whether to adjust operation of the vehicle includes using a trained neural network to predict the alertness level. The image data can be input into the trained neural network, and the neural network can output a predicted driver alertness. For example, the predicted alertness can be a value between 0.0 and 1.0, where 1.0 indicates a highest alertness level and 0.0 indicates a lowest alertness level. As described above with reference to FIGS. 1-3, the trained neural network can be trained on facial images that include images of drivers with different alertness levels collected from various geographic regions without sharing the facial images across the geographic regions. Figure 5 Figure 6 As described above with reference to FIGS. 1-3, the trained neural network can be trained on facial images that include images of drivers with different alertness levels collected from various geographic regions without sharing the facial images across the geographic regions.

[0124] In other embodiments, the ADAS can adjust operation of the vehicle in response to different predicted states of the driver, recognition of the driver, or different characteristics of the driver’s behavior. It can be appreciated that the examples and embodiments described herein are for illustrative purposes only, and that the trained neural network model can be used by the ADAS or DMS to adjust operation of the vehicle in other ways without departing from the scope of the present disclosure.

[0125] ​​At 710, the method includes determining whether the predicted alertness level output by the trained neural network model is below a threshold alertness level. For example, the threshold alertness level can be.5, and if the predicted alertness level is.7, then it is determined that the driver’s alertness is above the threshold alertness level (e.g., the alertness is sufficient to drive without ADAS intervention). Alternatively, if the predicted alertness level is.4, then it is determined that the driver’s alertness is below the threshold alertness level (e.g., the alertness is not sufficient to drive without ADAS intervention).

[0126] If it is determined at 710 that the predicted alertness level output by the trained neural network model is not below the threshold alertness level, then the method proceeds to 712. At 712, the method includes continuing to operate the vehicle without ADAS system intervention, and the method ends. Alternatively, if it is determined at 710 that the predicted alertness level is below the threshold alertness level, then the method proceeds to 714.

[0127] At 714, the method includes adjusting operation of the vehicle based on the predicted alertness level. In other words, the ADAS can perform one or more actions to increase the driver’s alertness level. For example, the ADAS can notify the driver through an audible sound, warning the driver that they appear to be fatigued, or the ADAS can play an audio file with music or other content that can increase the driver’s alertness level. In some embodiments, the ADAS can adjust the acceleration or speed of the vehicle, or perform a different action or operation of the vehicle, in response to the alertness level being below the threshold alertness level.

[0128] In this way, the trained neural network model can be integrated within an ADAS of a vehicle, which can use the model to determine whether and how to intervene during driving based on a driver condition detected or predicted from facial features or expressions of the driver, which are captured through a DMS system of the vehicle. The neural network model can be trained using the cross-island FL system for synchronous learning described herein. By using the proposed cross-island FL system, the accuracy of the trained neural network model can be improved with image data from a larger, more diverse population of drivers (rather than data obtained from image data collected within a single geographic region), while respecting personal data privacy protections that can differ across different geographic regions. The method 700 then ends.

[0129] Neural network models can be trained using synchronous learning within a federated learning (FL) system. The FL system can include a plurality of local servers (e.g., nodes), where for each local server in the FL system, a local neural network model is trained on the respective local server using vehicle data received from a geographic location where the local server is located. A batch of vehicle data (e.g., training data) can be received at the local servers, where the training data is divided into a plurality of mini-batches of training data. In this way, bandwidth limitations can be reduced as a single mini-batch can be trained on the local servers at a time. For each local server, the local neural network model can continue to load a single mini-batch of training data and accumulate a plurality of gradients for a plurality of parameters during a backpropagation phase until there are no remaining mini-batches of training data. The accumulated gradients can be sent to the plurality of local servers for aggregation. Each local server can then update the local neural network model using the aggregated gradients received from the plurality of local servers.

[0130] In response to an accuracy of the local neural network model of the local server decreasing over a predetermined number of iterations when validating the local neural network model using the local validation set, the batch size can be reduced and the training of the local neural network model on the local server can continue. Adjusting the batch size and continuing the training of the local neural network model on the local server can also include adjusting at least one of the weighting coefficient and the global learning rate before continuing the training of the first neural network. In particular, adjusting the batch size, the global learning rate based on the adjusted batch size, and the weighting coefficient can also include reducing the batch size by a predetermined amount, adjusting the global learning rate based on the reduced batch size, and increasing the weighting coefficient by a predetermined amount. By training the local neural network model in this way, vehicle data from a particular geographic location is not directly exchanged, improving the training efficiency and accuracy of the local neural network model while protecting user privacy and maintaining data security.

[0131] As an example of how the cross-island FL system can be used in practice, the cross-country distributed system can include two local servers. When the accuracy of the plurality of local neural network models continues to increase over a predetermined number of iterations during validation using the plurality of local validation data sets, a first training data set from a first local database can be collected on a first local server to train a first local neural network model on the first training data set at the first local server using synchronous learning. A second training data set from a second local database can be collected on a second local server to train a second local neural network model on the second training data set at the second local server using synchronous learning.

[0132] During each iteration of training the first and second local neural networks, a first set of accumulated gradients computed during a backpropagation phase of the training can be stored at the first local server and a second set of accumulated gradients computed during the backpropagation phase of the training can be stored at the second local server. Further, during each iteration of training the first and second local neural networks, the first set of accumulated gradients from the first local server can be transmitted to the second local server and the second set of accumulated gradients from the second local server can be transmitted to the first local server.

[0133] Training the first and second local neural network models can include aggregating the first and second sets of accumulated gradients at each of the first and second local servers. Each iteration of training the first and second local neural network models can terminate when updating parameters of the first and second neural network models based on the aggregated gradients.

[0134] In response to detecting that at least one of the accuracy of the first local neural network model training and the accuracy of the second local neural network model training is decreasing consistently over a predetermined number of iterations, the training includes updating a batch size of the first and second training data sets and updating a first weight applied to the first set of accumulated gradients and a second weight applied to the second set of accumulated gradients when aggregating the first and second sets of gradients.

[0135] In particular, the training includes reducing a first batch size of the first training data set by a first predetermined amount and reducing a second batch size of the second training data set by a second predetermined amount. In some embodiments, the first predetermined amount can be different than the second predetermined amount. As such, the adjusted batch size of both the first and second training data sets can be the minimum of the adjusted batch size of the first training data set and the adjusted batch size of the second training data set. In response to the plurality of local servers receiving the adjusted batch size, the training can continue.

[0136] During the synchronized training across the islanded FL system, the technical effects of adjusting the batch size, learning rate, and aggregation weight coefficients of the neural network model based on the accuracy of the plurality of replicas of the neural network model distributed across different geographic regions is that the training efficiency of the neural network model can be improved, resulting in a more accurate prediction model based on the personal data while reducing bandwidth limitations and protecting the personal data.

[0137] The present disclosure also provides support for a method of an advanced driver assistance system (ADAS) for a vehicle, the method comprising: adjusting operation of the vehicle based on vehicle occupant images captured by an in-cabin monitoring system of the vehicle, wherein the in-cabin monitoring system relies on a neural network model trained using synchronous learning across a federated learning (FL) system of islands, and training the neural network model includes performing a weighted aggregation of a first set of accumulated gradients of a plurality of parameters of the neural network model over a first plurality of batches of a first training dataset and a second set of accumulated gradients of a second identical neural network model trained over a second plurality of batches of a second training dataset; updating the parameters of the neural network model based on the aggregated gradients; and responsive to an accuracy of the neural network model decreasing over a predetermined number of iterations, adjusting one or more of a batch size of the first plurality of batches, a first aggregation weighting factor assigned to the first set of accumulated gradients, and a learning rate.

[0138] In a first example of the method, the in-cabin monitoring system is one of a driver monitoring system (DMS) and an occupant monitoring system (OMS) of the vehicle. In a second example of the method (optionally including the first example), the first training dataset is collected from a first geographic region, the neural network model is trained on a first local server within the first geographic region, the second training dataset is collected from a second geographic region, and the second neural network model is trained on a second local server within the second geographic region. In a third example of the method (optionally including one or both of the first and second examples), the training of the neural network model on the first local server and the training of the second neural network model on the second local server are performed without exchanging training data between the first local server and the second local server.

[0139] In a fourth example of the method (optionally including one or more or each of the first through third examples), the first training dataset does not include data from the second geographic region, and the second training dataset does not include data from the first geographic region. In a fifth example of the method (optionally including one or more or each of the first through fourth examples), performing the weighted aggregation further includes performing a summation of a product of a first weighting factor and the first set of accumulated gradients and a product of a second weighting factor and the second set of accumulated gradients, the second set of accumulated gradients being received from the second local server. In a sixth example of the method (optionally including one or more or each of the first through fifth examples), the second weighting factor is received from the second local server.

[0140] In a seventh example of the method (optionally including one or more or each of the first through sixth examples), the method further includes sending at least one of the adjusted batch size, the adjusted aggregation weighting factor, and the adjusted learning rate to the second local server for training the second neural network model. In an eighth example of the method (optionally including one or more or each of the first through seventh examples), the accuracy of the neural network model is determined using a validation dataset collected from a first geographic region. In a ninth example of the method (optionally including one or more or each of the first through eighth examples), adjusting one or more of the batch size of the first plurality of batches, the first aggregation weighting factor assigned to the first set of accumulated gradients, and the learning rate further includes at least one of reducing the batch size by a first predetermined amount, adjusting the learning rate based on the reduced batch size, and increasing the aggregation weighting factor by a second predetermined amount. In a tenth example of the method (optionally including one or more or each of the first through ninth examples), adjusting the operation of the vehicle based on the images of the occupants of the vehicle captured by the in-cabin monitoring system further includes one of adjusting the operation of the vehicle based on the identification of the vehicle operator and adjusting the operation of the vehicle based on the predicted state of the vehicle operator.

[0141] The present disclosure also provides support for an advanced driver assistance system (ADAS) of a vehicle, the ADAS including a user input device, a display device, a processor, and a non-transitory memory storing instructions that, when executed, cause the processor to adjust the operation of the vehicle based on an output of a neural network model trained on images of a driver of the vehicle captured by a driver monitoring system (DMS) of the vehicle using synchronous learning within a federated learning (FL) system, wherein the in-cabin monitoring system relies on a neural network model trained using synchronous learning within a cross-island federated learning (FL) system, wherein training the neural network model includes performing a weighted aggregation of a first set of accumulated gradients of a plurality of parameters of the neural network model on a first plurality of batches of a first training dataset and a second set of accumulated gradients of a second identical neural network model trained on a second plurality of batches of a second training dataset, updating the parameters of the neural network model based on the aggregated gradients, and in response to an accuracy of the neural network model decreasing for a predetermined number of iterations, adjusting one or more of a batch size of the first plurality of batches, a first aggregation weighting factor assigned to the first set of accumulated gradients, and a learning rate.

[0142] In a first example of the system, the first training dataset is collected from a first geographic region, the neural network model is trained on a first local server within the first geographic region, the second training dataset is collected from a second geographic region, the second neural network model is trained on a second local server within the second geographic region, the first training dataset does not include data from the second geographic region, and the second training dataset does not include data from the first geographic region, and the training of the second neural network model on the second local server is performed without exchanging training data between the first local server and the second local server. In a second example of the system (optionally including the first example), a validation dataset collected from the first geographic region is used to determine an accuracy of the neural network model.

[0143] In a third example of the system (optionally including one or both of the first and second examples), adjusting one or more of a batch size of the first plurality of batches, a first aggregation weighting factor assigned to the first set of accumulated gradients, and a learning rate further includes at least one of: reducing the batch size, adjusting the learning rate based on the reduced batch size, and increasing the aggregation weighting factor. In a fourth example of the system (optionally including one or more or each of the first through third examples), the batch size is reduced by a predetermined amount, and the aggregation weighting factor is increased by an amount based on the accuracy of the neural network model. In a fifth example of the system (optionally including one or more or each of the first through fourth examples), adjusting operation of the vehicle based on the image of the occupant of the vehicle captured by the DMS further includes one of: adjusting operation of the vehicle based on an identification of the driver, and adjusting operation of the vehicle based on a predicted state of the driver.

[0144] The present disclosure also provides support for a method comprising: collecting, at a first server, a first set of images of drivers of a first set of vehicles at a first geographic location; collecting, at a second server, a second set of images of drivers of a second set of vehicles at a second geographic location different from the first geographic location; training, at the first server, a first copy of a neural network model on the first set of images; training, at the second server, a second copy of the neural network model on the second set of images; during each iteration of the training of the first copy of the neural network model and the training of the second copy of the neural network model, using synchronous learning: storing, at the first server, a first set of accumulated gradients computed during a backpropagation phase of the training; storing, at the second server, a second set of accumulated gradients computed during the backpropagation phase of the training; transmitting the first set of accumulated gradients from the first server to the second server; transmitting the second set of accumulated gradients from the second server to the first server; aggregating, at both the first server and the second server, the first set of accumulated gradients and the second set of accumulated gradients; updating parameters of the first copy of the neural network model and the second copy of the neural network model based on the aggregated gradients; in response to a decrease in accuracy of the first copy of the neural network model on a first validation dataset collected at the first server from the first geographic location, or a decrease in accuracy of the second copy of the neural network model on a second validation dataset collected at the second server from the second geographic location: reducing a batch size of the first training dataset and the second training dataset, updating a first weighting coefficient applied to the first set of accumulated gradients and a second weighting coefficient applied to the second set of accumulated gradients when aggregating the first set of accumulated gradients and the second set of accumulated gradients, updating a learning rate during the training of the first copy of the neural network model on the first server and the training of the second copy of the neural network model on the second server based on the batch size; and in response to the first copy of the neural network model reaching a threshold accuracy, using the trained first copy of the neural network model in a first vehicle at the first geographic location to adjust operation of the first vehicle based on images of a first driver of the first vehicle captured by a first driver monitoring system (DMS); and in response to the second copy of the neural network model reaching a threshold accuracy, using the trained second copy of the neural network model in a second vehicle at the second geographic location to adjust operation of the second vehicle based on images of a second driver of the second vehicle captured by a second DMS.

[0145] In a first example of the method, storing, at the first server, a first set of accumulated gradients computed during a backpropagation phase of the training and storing, at the second server, a second set of accumulated gradients computed during the backpropagation phase of the training further comprises: splitting the first training dataset into a first plurality of mini-batches of a fixed batch size; splitting the second training dataset into a second plurality of mini-batches of the fixed batch size; performing forward and backward propagation to compute a first plurality of gradients of the plurality of parameters for each mini-batch of the first plurality of mini-batches; performing forward and backward propagation to compute a second plurality of gradients of the plurality of parameters for each mini-batch of the second plurality of mini-batches; accumulating the first set of gradients by performing summation of the first plurality of gradients on the first server; and accumulating the second set of gradients by performing summation of the second plurality of gradients on the second server. In a second example of the method (optionally including the first example), updating the first weighting coefficient applied to the first set of accumulated gradients and the second weighting coefficient applied to the second set of accumulated gradients further comprises: increasing the first weighting coefficient in response to a decrease in accuracy of the first copy of the neural network model on a first validation dataset; and increasing the second weighting coefficient in response to a decrease in accuracy of the second copy of the neural network model on a second set of validation data.

[0146] A “system,” “unit,” or “module” can include or represent hardware and associated instructions (for example, software stored on a tangible and non-transitory computer- readable storage medium) that perform one or more of the operations described herein. The hardware can include an electronic circuit that includes and / or is connected to one or more logic-based devices, such as microprocessors, processors, controllers, or the like. The devices can be off-the-shelf devices that are suitably programmed or instructed to perform the operations described herein.

[0147] The descriptions of various embodiments described herein have been presented for the purpose of illustration, but not limitation, to the best of the Applicant’s knowledge. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technology found in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

[0148] It is expected that during the life of this patent many relevant systems, methods and computer programs will be developed and the scope of the terms ML model, DL model, neural network and vehicle operation data is intended to include all such new technologies a priori.

[0149] The terms "comprises", "comprising", "includes", "including", "having" and their conjugates mean "including but not limited to". This term encompasses the terms "consisting of" and "consisting essentially of".

[0150] As used herein, the singular forms "a", "an" and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a compound" or "at least one compound" can include a plurality of compounds, including mixtures thereof.

[0151] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.

[0152] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any implementation described as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.

[0153] It should be appreciated that certain features of the implementations described herein, which are, for clarity, described in the context of separate implementations, can also be provided in combination in a single implementation. Conversely, various features of the described implementations, which are, for brevity, described in the context of a single implementation, can also be provided separately or in any suitable subcombination or as suitable in any other implementation of the described implementations. Certain features described in the context of various implementations are not to be interpreted as essential features of those implementations, unless explicitly stated otherwise.

[0154] Although the present application has been described in connection with certain specific embodiments, it will be understood that it is capable of further modifications. This application is intended to cover any alternatives, modifications, and equivalents within the spirit and scope of the claims appended hereto, in accordance with the statutes.

[0155] It is the intent of the applicant(s) that all publications, patents, and patent applications referred to in this specification be incorporated by reference herein in their entirety for all purposes. In the event of inconsistencies between the disclosure of the present specification and the disclosures of the publications, patents, and patent applications incorporated herein by reference, the disclosure of the present specification shall prevail. Additionally, any reference to claimed subject matter shall not be construed as an admission of prior art status. To the extent that section headings are used, they could not be construed as necessarily limiting. Additionally, any priority document of this application is hereby incorporated by reference in its entirety.

Claims

1. A method for an advanced driver assistance system (ADAS) of a vehicle, the method comprising: adjusting operation of the vehicle based on images of an occupant of the vehicle captured by an in-cabin monitoring system of the vehicle, wherein the in-cabin monitoring system relies on a neural network model trained using synchronous learning across a federated learning (FL) system, and training the neural network model comprises: performing a weighted aggregation of a first set of accumulated gradients of a plurality of parameters of the neural network model over a first plurality of batches of a first training dataset and a second set of accumulated gradients of a second same neural network model trained over a second plurality of batches of a second training dataset; updating parameters of the neural network model based on the aggregated gradients; and in response to an accuracy of the neural network model decreasing over a predetermined number of iterations, adjusting one or more of a batch size of the first plurality of batches, a first aggregation weighting factor assigned to the first set of accumulated gradients, and a learning rate.

2. The method of claim 1, wherein the in-cabin monitoring system is one of a driver monitoring system (DMS) and an occupant monitoring system (OMS) of the vehicle.

3. The method of any one of claims 1 or 2, wherein: the first training dataset is collected from a first geographic region, the neural network model is trained on a first local server within the first geographic region; the second training dataset is collected from a second geographic region, a second neural network model is trained on a second local server within the second geographic region.

4. The method of claim 3, wherein the training of the neural network model on the first local server and the training of the second neural network model on the second local server are performed without exchanging training data between the first local server and the second local server.

5. The method of any one of claims 3 or 4, wherein the first training dataset does not include data from the second geographic region, and the second training dataset does not include data from the first geographic region.

6. The method of any one of claims 3 to 5, wherein performing the weighted aggregation further comprises performing a summation of a product of a first weighting factor and the first set of accumulated gradients and a product of a second weighting factor and the second set of accumulated gradients, the second set of accumulated gradients being received from the second local server.

7. The method of claim 6, wherein the second weighting factor is received from the second local server.

8. The method of any one of claims 3 to 7, further comprising sending at least one of an adjusted batch size, an adjusted aggregation weighting factor, and an adjusted learning rate to the second local server for training the second neural network model.

9. The method of any one of claims 3 to 8, wherein the accuracy of the neural network model is determined using a validation dataset collected from the first geographic region.

10. The method of any one of claims 1 to 9, wherein adjusting one or more of the batch size of the first plurality of batches, the first aggregated weighting factor assigned to the first set of accumulated gradients, and the learning rate further comprises at least one of: reducing the batch size by a first predetermined amount; adjusting the learning rate based on the reduced batch size; and increasing the aggregated weighting factor by a second predetermined amount.

11. The method of any one of claims 2 to 9, wherein adjusting the operation of the vehicle based on the images of the occupants of the vehicle captured by the in-cab monitoring system further comprises one of: adjusting the operation of the vehicle based on an identification of a vehicle operator; and adjusting the operation of the vehicle based on a predicted state of the vehicle operator.

12. An advanced driver assistance system (ADAS) of a vehicle, the ADAS comprising: a user input device; a display device; a processor; and a non-transitory memory storing instructions that, when executed, cause the processor to: adjust an operation of the vehicle based on an output of a first neural network model trained on images of a driver of the vehicle captured by a driver monitoring system (DMS) of the vehicle using synchronous learning within a federated learning (FL) system, wherein training the first neural network model comprises: performing a weighted aggregation of a first set of accumulated gradients of a plurality of parameters of the first neural network model over a first plurality of batches of a first training dataset and a second set of accumulated gradients of a second identical neural network model trained over a second plurality of batches of a second training dataset; updating parameters of the first neural network model based on the aggregated gradients; and in response to an accuracy of the first neural network model decreasing over a predetermined number of iterations, adjusting one or more of a batch size of the first plurality of batches, a first aggregated weighting factor assigned to the first set of accumulated gradients, and a learning rate.

13. The system of claim 12, wherein: the first training dataset is collected from a first geographic region, the first neural network model is trained on a first local server within the first geographic region; the second training dataset is collected from a second geographic region, the second neural network model is trained on a second local server within the second geographic region; the first training dataset does not include data from the second geographic region, and the second training dataset does not include data from the first geographic region; and the training of the second neural network model on the second local server is performed without exchanging training data between the first local server and the second local server.

14. The system of claim 13, wherein the accuracy of the first neural network model is determined using a first validation dataset collected from the first geographic region, and the accuracy of the second neural network model is determined using a second validation dataset collected from the second geographic location. ​ 15. The system of any one of claims 12 to 14, wherein adjusting one or more of the batch size of the first plurality of batches, the first aggregation weighting factor assigned to the first set of accumulated gradients, and the learning rate further comprises at least one of: reducing the batch size; adjusting the learning rate based on the reduced batch size; and increasing the aggregation weighting factor.

16. The system of claim 15, wherein the batch size is reduced by a predetermined amount, and the aggregation weighting factor is increased by an amount based on the accuracy of the first neural network model.

17. The system of any one of claims 12 to 16, wherein adjusting the operation of the vehicle based on the images of the occupants of the vehicle captured by the DMS further comprises one of: adjusting the operation of the vehicle based on an identification of the driver; adjusting the operation of the vehicle based on a predicted state of the driver.

18. A method comprising: collecting, on a first server, a first set of images of drivers of a first set of vehicles at a first geographic location; collecting, on a second server, a second set of images of drivers of a second set of vehicles at a second geographic location different from the first geographic location; training, using simultaneous learning, a first copy of a neural network model on the first set of images at the first server and a second copy of the neural network model on the second set of images on the second server; at each iteration of the training of the first copy of the neural network model and the training of the second copy of the neural network model: storing, at the first server, a first set of accumulated gradients computed during a backpropagation phase of the training; storing, at the second server, a first set of accumulated gradients computed during the backpropagation phase of the training; transmitting the first set of accumulated gradients from the first server to the second server; transmitting the second set of accumulated gradients from the second server to the first server; aggregating the first set of accumulated gradients and the second set of accumulated gradients at both the first server and the second server; updating parameters of the first copy of the neural network model and the second copy of the neural network model based on the aggregated gradients; in response to a decrease in accuracy of the first copy of the neural network model on a first validation dataset collected at the first server from the first geographic location, or a decrease in accuracy of the second copy of the neural network model on a second validation dataset collected at the second server from the second geographic location: reducing a batch size of the first training dataset and the second training dataset; updating a first weighting factor applied to the first set of accumulated gradients and a second weighting factor applied to the second set of accumulated gradients when aggregating the first set of accumulated gradients and the second set of accumulated gradients; updating a learning rate during training of the first copy of the neural network model on the first server and the second copy of the neural network model on the second server based on the batch size; and in response to the first copy of the neural network model reaching a threshold accuracy, adjusting operation of a first vehicle in the first geographic location using the trained first copy of the neural network model in the first vehicle based on images of a first driver of the first vehicle captured by a first driver monitoring system (DMS); and in response to the second copy of the neural network model reaching the threshold accuracy, adjusting operation of a second vehicle in the second geographic location using the trained second copy of the neural network model in the second vehicle based on images of a second driver of the second vehicle captured by a second DMS.

19. The method of claim 18, wherein storing the first set of accumulated gradients computed during the backpropagation phase of the training at the first server and the second set of accumulated gradients computed during the backpropagation phase of the training at the second server further comprises: splitting the first training dataset into a first plurality of mini-batches of a fixed batch size; splitting the second training dataset into a second plurality of mini-batches of the fixed batch size; performing forward and backward propagation to compute a first plurality of gradients of a plurality of parameters for each mini-batch of the first plurality of mini-batches; performing forward and backward propagation to compute a second plurality of gradients of the plurality of parameters for each mini-batch of the second plurality of mini-batches; accumulating the first set of gradients by performing summation of the first plurality of gradients on the first server; and accumulating the second set of gradients by performing summation of the second plurality of gradients on the second server.

20. The method of claim 18 or 19, wherein updating the first weighting coefficient applied to the first set of accumulated gradients and the second weighting coefficient applied to the second set of accumulated gradients further comprises: in response to the accuracy of the first copy of the neural network model on the first validation dataset decreasing, increasing the first weighting coefficient; and in response to the accuracy of the second copy of the neural network model on the second validation dataset decreasing, increasing the second weighting coefficient.

Citation Information

Cited By

  • Neural network model training system and training method

    CN122264009A