New method for communicating between AI / ML-enabled clients during joint learning
By using the control message format to encapsulate and transmit messages during the federated learning process, the inefficient message delivery problem between the device and the server is solved, and efficient synchronization and update of the AI/ML model is achieved.
Patent Information
- Application Number
- CN202480004246.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-14
- Filing Date
- 2024-05-15
- Publication Date
- 2025-05-13
AI Technical Summary
The lack of effective communication mechanisms in the prior art to manage messaging between devices and servers during the joint learning process, resulting in inefficient collaboration between networks and devices.
Provided is an apparatus and method to encapsulate and transmit joint learning messages using the control message format by encapsulating code and control code, including the identifier, size, type and body of the message, for controlling synchronization, qualification, model evaluation, model update and error handling in the AI/ML joint learning process.
It realizes efficient messaging and control in the AI/ML joint learning process, improves the collaboration efficiency between devices and servers, and ensures the smooth progress of the model synchronization and update process.
Smart Images

Figure CN119998820A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 466,604 filed on May 15, 2023, U.S. Provisional Application No. 63 / 466,609 filed on May 15, 2023, and U.S. Application No. 18 / 663,634 filed on May 14, 2024, the disclosures of which are incorporated herein by reference in their entirety. Technical Field
[0002] The present disclosure provides a method for communicating between devices having AI / ML capabilities through a service to achieve federated learning between the devices. Background Art
[0003] In federated learning, each device uses its local data and possibly a portion of the data provided by the server to improve its AI / ML model, and then communicates its improvements to the server, and thus to the other devices.
[0004] Artificial Intelligence (AI) and Machine Learning (ML) have been developed in recent years and have found many applications. The application of AI / ML in 5G networks is a new topic. Recently, 3GPP SA4 launched a research project on AI / ML for media, which will produce a technical report on this topic.
[0005] The goal of SA4's "Artificial Intelligence (AI) and Machine Learning (ML) for Media" is to identify media service architectures and related service flows, model operation configurations, data components including available data formats, and data flows. The results of the research project are saved in a permanent document (PD).
[0006] Although PD includes different collaboration scenarios between the network and the devices, including federated learning, PD does not discuss the communication between the network and the devices during the federated learning, and thus, there is such a defect in computer technology.
[0007] Therefore, for any of these reasons, there is a need for technical solutions to these problems arising in computer technology. Summary of the invention
[0008] According to one aspect of the present disclosure, a device and similarly a method and a computer-readable medium are provided, the device comprising: at least one memory configured to store computer program code; and at least one processor configured to access the computer program code and operate as instructed by the computer program code, the computer program code comprising: encapsulation code configured to cause at least one processor to encapsulate a message in one or more joint learning messages by controlling a message format, the control message format comprising a plurality of fields, the plurality of fields respectively indicating one of the following: an identifier of a message, a size of a message, a type of a message, and a body of a message; and a control code configured to The invention is configured to cause at least one processor to control artificial intelligence / machine learning (AI / ML) joint learning based on a message encapsulated by a control message format, wherein the AI / ML joint learning includes a server that controls multiple separate devices to implement a joint part of the AI / ML joint learning, and reports results of implementing the joint part to the server from each of the multiple separate devices, respectively, and wherein a body of the message indicates at least one of: AI / ML joint learning synchronization between the multiple separate devices, device qualification of the AI / ML joint learning, model evaluation of the AI / ML joint learning, model update of the AI / ML joint learning, and an error of the AI / ML joint learning.
[0009] AI / ML joint learning may include multiple rounds of joint learning, each round of joint learning being a back-and-forth iteration between a server and multiple individual devices, and AI / ML joint learning may be implemented in parallel at multiple individual devices.
[0010] The body of the message may indicate AI / ML federated learning synchronization between multiple devices, and that in one round of multiple rounds of federated learning, federated learning will start simultaneously at each of the multiple individual devices.
[0011] The body of the message may indicate device eligibility for AI / ML federated learning and one or more criteria for device eligibility.
[0012] The one or more criteria for device qualification for AI / ML federated learning may be any of operating system, processor speed, available memory, available image library, number of images, geographic location, language setting.
[0013] The AI / ML joint learning model evaluation may instruct multiple separate devices to implement the evaluation of the AI / ML joint learning model.
[0014] Model evaluation for AI / ML federated learning may instruct multiple separate devices to implement evaluation of models separate from AI / ML federated learning.
[0015] The AI / ML joint learning model update may instruct multiple individual devices to update the parameters of the AI / ML joint learning model.
[0016] The model update may be one of: a first instruction from a server to the plurality of individual devices; and a second instruction from at least one of the plurality of individual devices to the server. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and other features and aspects of embodiments of the present disclosure will become apparent from the following description taken in conjunction with the accompanying drawings, in which:
[0018] Figure 1 is a simplified diagram according to an embodiment;
[0019] Figure 2 is a simplified diagram according to an embodiment;
[0020] Figure 3 is a simplified diagram according to an embodiment;
[0021] Figure 4 is a simplified diagram according to an embodiment;
[0022] Figure 5 is a simplified diagram according to an embodiment;
[0023] Figure 6 is a simplified diagram according to an embodiment; and
[0024] Figure 7 is a simplified diagram according to an implementation. DETAILED DESCRIPTION
[0025] The following detailed description of example embodiments refers to the accompanying drawings.The same reference numbers in different drawings may identify the same or similar elements.
[0026] Figure 1 A simplified block diagram of a communication system 100 according to an embodiment of the present disclosure is shown. The communication system 100 may include at least two terminals 102 and 103 interconnected via a network 105. For unidirectional transmission of data, a first terminal 103 may encode video data at a local location for transmission to another terminal 102 via the network 105. The second terminal 102 may receive the encoded video data of the other terminal from the network 105, decode the encoded data and display the recovered video data. Unidirectional data transmission may be common in media service applications and the like.
[0027] Figure 1A second pair of terminals 101 and 104 are shown, which are arranged to support two-way transmission of encoded video, such as may occur during a video conference. For two-way transmission of data, each terminal 101 and 104 can encode video data captured at a local location for transmission to the other terminal via a network 105. Each terminal 101 and 104 can also receive encoded video data sent by the other terminal, can decode the encoded data, and can display the recovered video data at a local display device.
[0028] exist Figure 1 In the embodiment of the present disclosure, terminals 101, 102, 103 and 104 can be shown as servers, personal computers and smart phones, but the principles of the present disclosure are not limited thereto. The embodiments of the present disclosure are applicable to laptop computers, tablet computers, media players and / or special video conferencing equipment. Network 105 represents any number of networks that transmit coded video data between terminals 101, 102, 103 and 104, including, for example, wired and / or wireless communication networks. Communication network 105 can exchange data in circuit switching channels and / or packet switching channels. Representative networks include telecommunication networks, local area networks, wide area networks and / or the Internet. For the purpose of this discussion, unless otherwise specified below, the architecture and topology of network 105 may be insignificant for the operation of the present disclosure.
[0029] Figure 2 The placement of a video encoder and a video decoder in a streaming environment is shown as an example of an application for the disclosed subject matter. The disclosed subject matter can be equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CD (Compact Disc, CD), DVD (Digital Versatile Disc, DVD), memory sticks, etc.
[0030] The streaming system may include a capture subsystem 203, which may include a video source 201, such as a digital camera, that creates, for example, an uncompressed video sample stream 213. The sample stream 213 may be emphasized as a high amount of data when compared to an encoded video bitstream, and may be processed by an encoder 202 coupled to the camera 201. The encoder 202 may include hardware, software, or a combination thereof to implement or implement aspects of the disclosed subject matter as described in more detail below. The encoded video bitstream 204, which may be emphasized as a lower amount of data when compared to the sample stream, may be stored on a streaming server 205 for future use. One or more streaming clients 212 and 207 may access the streaming server 205 to retrieve a copy 208 and a copy 206 of the encoded video bitstream 204. The client 212 may include a video decoder 211 that decodes the incoming copy of the encoded video bitstream 208 and creates an outgoing video sample stream 210 that may be rendered on a display 209 or other rendering device (not shown). In some streaming systems, the video bitstreams 204, 206, and 208 may be encoded according to certain video encoding / compression standards. Examples of these standards are mentioned above and further described herein.
[0031] Figure 3 may be a functional block diagram of a video decoder 300 according to an embodiment of the present invention.
[0032] Receiver 302 may receive one or more coded video sequences to be decoded by decoder 300; in the same or another embodiment, one coded video sequence is received at a time, wherein the decoding of each coded video sequence is independent of the other coded video sequences. The coded video sequence may be received from channel 301, which may be a hardware / software link to a storage device storing coded video data. Receiver 302 may receive coded video data as well as other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to their respective consuming entities (not depicted). Receiver 302 may separate the coded video sequence from the other data. To prevent network jitter, a buffer memory 303 may be coupled between receiver 302 and entropy decoder / parser 304 (hereinafter referred to as "parser"). When receiver 302 is receiving data from a storage / forward device with sufficient bandwidth and controllability or from an isochronous network, buffer 303 may not be needed or may be small. For use on a best-effort packet network such as the Internet, buffer 303 may be needed, which may be relatively large and may advantageously have an adaptive size.
[0033] The video decoder 300 may include a parser 304 to reconstruct symbols 313 from an entropy coded video sequence. The categories of these symbols include: information for managing the operation of the decoder 300; and information potentially used to control a rendering device such as a display 312, which is not part of the decoder but can be coupled to the decoder. The control information for the rendering device can be in the form of auxiliary enhancement information (SEI (Supplementary Enhancement Information, SEI) message) or a video availability information parameter set fragment (not depicted). The parser 304 can parse / entropy decode the received coded video sequence. The encoding of the coded video sequence can be performed according to a video coding technique or standard, and can follow principles known to those skilled in the art, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, etc. The parser 304 can extract a subgroup parameter set from at least one subgroup of the subgroups of pixels in the video decoder in the coded video sequence based on at least one parameter corresponding to the group. Subgroups may include Group Of Picture (GOP), pictures, tiles, slices, macroblocks, Coding Units (CU), blocks, Transform Units (TU), Prediction Units (PU), etc. The entropy decoder / parser may also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.
[0034] The parser 304 may perform an entropy decoding / parsing operation on the video sequence received from the buffer 303 to create the symbol 313. The parser 304 may receive the encoded data and selectively decode a specific symbol 313. In addition, the parser 304 may determine whether to provide the specific symbol 313 to the motion compensation prediction unit 306, the scaler / inverse transform unit 305, the intra prediction unit 307, or the loop filter 311.
[0035] Depending on the type of the coded video picture or its portion (e.g., inter-picture and intra-picture, inter-block and intra-block) and other factors, the reconstruction of the symbol 313 may involve multiple different units. Which units are involved and the way in which the units are involved can be controlled by subgroup control information parsed from the coded video sequence by the parser 304. For the sake of brevity, such subgroup control information flow between the parser 304 and the following multiple units is not depicted.
[0036] In addition to the functional blocks already mentioned, decoder 300 can be conceptually subdivided into a plurality of functional units as described below. In a practical implementation operating under commercial constraints, many of these units interact closely with each other and can be at least partially integrated with each other. However, for the purpose of describing the disclosed subject matter, it is appropriate to conceptually subdivide into the following functional units.
[0037] The first unit is the sealer / inverse transform unit 305. The sealer / inverse transform unit 305 receives the quantized transform coefficients as symbols 313 from the parser 304, as well as control information, including which transform to use, block size, quantization factor, quantization scaling matrix, etc. The sealer / inverse transform unit 305 may output blocks including sample values, which may be input into the aggregator 310.
[0038] In some cases, the output samples of the sealer / inverse transform 305 may belong to an intra-coded block; that is, a block that does not use predictive information from a previously reconstructed picture, but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit 307. In some cases, the intra-picture prediction unit 307 uses surrounding already reconstructed information extracted from the current (partially reconstructed) picture 309 to generate a block of the same size and shape as the block under reconstruction. In some cases, the aggregator 310 adds the prediction information already generated by the intra-prediction unit 307 to the output sample information as provided by the sealer / inverse transform unit 305 on a per-sample basis.
[0039] In other cases, the output samples of the scaler / inverse transform unit 305 may belong to an inter-coded and possibly motion compensated block. In such a case, the motion compensated prediction unit 306 may access the reference picture memory 308 to extract samples for prediction. After the extracted samples are motion compensated according to the symbol 313 belonging to the block, these samples may be added by the aggregator 310 to the output of the scaler / inverse transform unit (in this case referred to as residual samples or residual signal) to generate output sample information. The address within the reference picture memory from which the motion compensation unit extracts the predicted samples may be controlled by a motion vector, which is available to the motion compensation unit in the form of a symbol 313, which may have, for example, an X component, a Y component, and a reference picture component. Motion compensation may also include interpolation of sample values extracted from the reference picture memory when using sub-sample accurate motion vectors, motion vector prediction mechanisms, etc.
[0040] The output samples of the aggregator 310 may be subjected to various loop filtering techniques in a loop filter unit 311. The video compression techniques may include in-loop filter techniques controlled by parameters included in the coded video bitstream and available to the loop filter unit 311 as symbols 313 from the parser 304, but the video compression techniques may also be responsive to meta-information obtained during decoding of a previous (in decoding order) portion of a coded picture or coded video sequence, and to previously reconstructed and loop filtered sample values.
[0041] The output of the loop filter unit 311 may be a sample stream, which may be output to the rendering device 312 and stored in the reference picture memory 557 for future inter picture prediction.
[0042] Certain coded pictures, once fully reconstructed, may be used as reference pictures for future prediction. Once a coded picture is fully reconstructed and the coded picture has been identified as a reference picture (e.g., by parser 304), current reference picture 309 may become part of reference picture buffer 308, and new current picture memory may be reallocated before starting reconstruction of subsequent coded pictures.
[0043] The video decoder 300 may perform decoding operations according to a predetermined video compression technique that may be recorded in a standard such as ITU-T (International Telecommunication Union-Telecommunication Standardization Sector, ITU-T) H.265 Recommendation. As specified in a video compression technology document or standard and in particular in a profile therein, the coded video sequence may conform to the syntax specified by the video compression technology or standard used in the sense that the coded video sequence follows the syntax of the video compression technology or standard. For compliance, the complexity of the coded video sequence is also required to be within the range defined by the hierarchy of the video compression technology or standard. In some cases, the hierarchy limits the maximum picture size, the maximum frame rate, the maximum reconstruction sampling rate (measured in, for example, megasamples per second), the maximum reference picture size, etc. In some cases, the limits set by the hierarchy may be further defined by a hypothetical reference decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.
[0044] In an embodiment, the receiver 302 may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence. The additional data may be used by the video decoder 300 to properly decode the data and / or more accurately reconstruct the original video data. The additional data may be in the form of, for example, temporal, spatial or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0045] Figure 4 It may be a functional block diagram of a video encoder 400 according to an embodiment of the present disclosure.
[0046] Encoder 400 may receive video samples from a video source 401 (which is not part of the encoder), which may capture video images to be encoded by encoder 400 .
[0047] The video source 401 may provide a source video sequence in the form of a digital video sample stream to be encoded by the encoder (303), which may have any suitable bit depth (e.g., 8 bits, 10 bits, 12 bits, ...), any color space (e.g., BT.601 Y CrCB, RGB, ...), and any suitable sampling structure (e.g., YCrCb 4:2:0, Y CrCb 4:4:4). In a media service system, the video source 401 may be a storage device that stores previously prepared videos. In a video conferencing system, the video source 401 may be a camera that captures local image information as a video sequence. The video data may be provided as a plurality of separate pictures that are given motion when viewed sequentially. The pictures themselves may be organized as a spatial pixel array, wherein each pixel may include one or more samples, depending on the sampling structure, color space, etc. used. The relationship between pixels and samples may be readily understood by those skilled in the art. The following description focuses on samples.
[0048] According to an embodiment, the encoder 400 can encode and compress the pictures of the source video sequence into a coded video sequence 410 in real time or according to any other time constraints required by the application. It is a function of the controller 402 to implement the appropriate encoding speed. The controller controls other functional units as described below and is functionally coupled to these units. For clarity, the coupling is not depicted. The parameters set by the controller may include: rate control related parameters (picture skipping, quantizer, lambda value of rate distortion optimization technology, ...), picture size, picture group (GOP) layout, maximum motion vector search range, etc. Those skilled in the art can easily identify other functions of the controller 402, because these functions may belong to the video encoder 400 optimized for a specific system design.
[0049] Some video encoders operate in a manner that is readily recognizable to those skilled in the art as a "codec loop". As an overly simplified description, the codec loop may include: an encoding portion of an encoder 402 (hereinafter "source encoder") (responsible for creating symbols based on the input picture to be encoded and the reference picture); and a (local) decoder 406 embedded in the encoder 400, which reconstructs the symbols to create sample data that the (remote) decoder will also create (because in the video compression techniques considered in the disclosed subject matter, any compression between the symbols and the encoded video bitstream is lossless). The reconstructed sample stream is input to a reference picture memory 405. Since decoding the symbol stream produces bit-accurate results that are independent of the decoder location (local or remote), the reference picture buffer contents are also bit-accurate between the local encoder and the remote encoder. In other words, the reference picture samples "seen" by the prediction portion of the encoder are exactly the same sample values that the decoder will "see" when using prediction during decoding. The basic principles of this reference picture synchronization (and the resulting drift if synchronization cannot be maintained, for example due to channel errors) are well known to those skilled in the art.
[0050] The operation of the "local" decoder 406 can be combined with the above Figure 3 The operation of the "remote" decoder 300 described in detail is the same. However, reference is also briefly made to Figure 4 Since the symbols are available and the entropy encoder 408 and the parser 304 can losslessly encode / decode the symbols into a coded video sequence, the entropy decoding part of the decoder 300 including the channel 301, the receiver 302, the buffer 303 and the parser 304 may not be fully implemented in the local decoder 406.
[0051] It can be observed at this point that any decoder technology other than parsing / entropy decoding present in a decoder must also necessarily be present in a corresponding encoder in the form of substantially the same functionality. Since encoder technology is inverse to the fully described decoder technology, the description of encoder technology can be simplified. A more detailed description is only required and provided in certain places below.
[0052] As part of the operation of source encoder 403, source encoder 403 may perform motion compensated predictive encoding, which predictively encodes an input frame with reference to one or more previously encoded frames from a video sequence designated as "reference frames." In this manner, encoding engine 407 encodes the differences between pixel blocks of an input frame and pixel blocks of a reference frame, which may be selected as a prediction reference for the input frame.
[0053] The local video decoder 406 can decode the encoded video data of the frame that can be designated as the reference frame based on the symbol created by the source encoder 403. The operation of the encoding engine 407 can advantageously be a lossy process. When the encoded video data can be decoded at the video decoder ( Figure 4 When decoded at a remote video decoder (not shown), the reconstructed video sequence may typically be a copy of the source video sequence with some errors. The local video decoder 406 replicates the decoding process that may be performed on the reference frame by the video decoder, and may cause the reconstructed reference frame to be stored in the reference picture cache 405. In this manner, the encoder 400 may locally store a copy of the reconstructed reference frame that has common content (without transmission errors) with the reconstructed reference frame to be obtained by the remote video decoder.
[0054] The predictor 404 may perform a prediction search for the encoding engine 407. That is, for a new frame to be encoded, the predictor 404 may search the reference picture memory 405 for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc. that may be used as appropriate prediction references for the new picture. The predictor 404 may operate on a pixel block by pixel block basis to find an appropriate prediction reference. In some cases, as determined by the search results obtained by the predictor 404, the input picture may have prediction references taken from multiple reference pictures stored in the reference picture memory 405.
[0055] The controller 402 may manage encoding operations of the video encoder 403 , including, for example, setting parameters and subgroup parameters for encoding video data.
[0056] The outputs of all the above-mentioned functional units may be subjected to entropy encoding in the entropy encoder 408. The entropy encoder converts the symbols as generated by the various functional units into a coded video sequence by losslessly compressing them according to techniques known to those skilled in the art, such as Huffman encoding, variable length encoding, arithmetic coding, etc.
[0057] Transmitter 409 may buffer the encoded video sequence created by entropy encoder 408 in preparation for transmission via communication channel 411, which may be a hardware / software link to a storage device where the encoded video data will be stored. Transmitter 409 may merge the encoded video data from video encoder 403 with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0058] Controller 402 may manage the operation of encoder 400. During encoding, controller 405 may assign a certain encoding picture type to each encoded picture, which may affect the encoding techniques that may be applied to the corresponding picture. For example, a picture may generally be assigned one of the following frame types:
[0059] An intra picture (I picture) may be a picture that can be encoded and decoded without using any other frame in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, independent decoder refresh pictures. Those skilled in the art are aware of those variations of I pictures and their corresponding applications and features.
[0060] A predictive picture (P picture) may be a picture that can be encoded and decoded using intra prediction or inter prediction to predict sample values of each block using at most one motion vector and a reference index.
[0061] Bi-directional predictive pictures (B pictures), which can be pictures that can be encoded and decoded using intra-prediction or inter-prediction that uses up to two motion vectors and reference indices to predict sample values for each block. Similarly, multi-predictive pictures can use more than two reference pictures and associated metadata for reconstruction of a single block.
[0062] The source picture may typically be spatially subdivided into a number of blocks of samples (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples, respectively), and encoded on a block-by-block basis. The blocks may be predictively encoded with reference to other (already encoded) blocks, which are determined by the coding allocation applied to the corresponding picture of the block. For example, blocks of an I picture may be non-predictively encoded, or may be predictively encoded (spatial prediction or intra-frame prediction) with reference to already encoded blocks of the same picture. Pixel blocks of a P picture may be non-predictively encoded via spatial prediction or via temporal prediction with reference to one previously encoded reference picture. Blocks of a B picture may be non-predictively encoded via spatial prediction or via temporal prediction with reference to one or two previously encoded reference pictures.
[0063] The video encoder 400 may perform encoding operations according to a predetermined video encoding technique or standard, such as ITU-T H.265 Recommendation. In the operation of the video encoder 400, the video encoder 400 may perform various compression operations, including predictive encoding operations that exploit temporal redundancy and spatial redundancy in an input video sequence. Thus, the encoded video data may conform to the syntax specified by the video encoding technique or standard being used.
[0064] In an embodiment, the transmitter 409 may transmit additional data along with the encoded video. The source encoder 403 may include such data as part of the encoded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, auxiliary enhancement information (SEI) messages, visual usability information (VUI) parameter set fragments, etc.
[0065] As 3GPP SA4 explores AI / ML for media, the embodiments herein provide a messaging framework for federated learning. For example, see Figure 5 , which shows an example 500 of a federated learning architecture according to an implementation herein based on the AI / ML research of 3GPP SA4.
[0066] like Figure 5 As shown, network 502 can provide the AI model to device 501. This can also be done in Figure 6 As can be seen at step S601 in example 600 of FIG. 5 . Device 501 can perform local training on the model and provide the results back to network 501. This can also be done in Figure 6 As can be seen at steps S602, S602', and S602" in Example 600, each of steps S602, S602', and S602" represents a corresponding federated device in the federated devices, one of which is device 501. The service can use the training results to update the AI model, and optionally provide the updated model to the device and other devices. This can be considered an example of another iteration of Example 600, and can be seen at step S601.
[0067] Embodiments herein provide a messaging format for communication between a central service and devices, such as device 501, located on a network 502 to exchange messages during a federated learning process. The control messages have a common envelope that identifies the source / destination, the specific model in the federated learning, the federated learning state at the destination, the quality / accuracy of the locally updated model, and other information during the federated learning process.
[0068] The structure of the joint learning message is shown in Table 1: Table 1 - High-level structure of a federated learning message type Require Encapsulation Mandatory Messages_number Mandatory Message_body 1 Mandatory Message_body 2 Optional … … Message_body N Optional
[0069] According to an embodiment and according to Table 1, a message consists of an envelope and one or more message bodies. The envelope provides general information about the message. The number of messages defines the number of message bodies in the message. Each message body defines a specific message.
[0070] In addition, the joint learning message encapsulation includes the following parameters, as shown in Table 2: Table 2 - Federated Learning Message Encapsulation
[0071] According to an embodiment, and as shown in Table 2 above: 1. Source_id: A unique id that identifies the source of a message. The source can be a device or node participating in federated learning, or a network service / server that coordinates federated learning. 2. Destination or group id (Destination_id): A unique id used to identify the destination or destination group of the message. The destination can be a device or node participating in federated learning, or a network service / server coordinating federated learning. If not present, the message is for all devices / nodes participating in federated learning. 3. Model_id: A unique id that identifies the model in federated learning, so that the recipient of the message applies the message to a specific model. 4. Messages_number defines the number of message bodies in the message. A message can contain one or more message bodies.
[0072] According to an embodiment, when a set of devices, such as one of the devices 501, participate in federated learning, any of those devices can send and receive messages to a service server, such as a service server of the network 502. Each message should include the above-mentioned encapsulation. The encapsulation enables identification of the source, destination, and model that is the subject of the message.
[0073] The joint learning message body includes the following parameters, as shown in Table 3: Table 3 – Federated Learning Message Body
[0074] According to an embodiment, and as shown in Table 3 above: 1.message_ids should be assigned in increasing order so that the order of published messages can be processed in the order they were published when needed. 2.message_size defines the length of the entire message body in bytes. 3.message_type defines the type of the message. The body format of the message is defined by the type of the message. A message can have multiple message types. In this case, the message type is followed by the corresponding body. 4. The rest of the message body format and semantics are defined by the message type.
[0075] Referring to Tables 1 to 3, the design defines a message framework for control messages during joint learning, which has the following advantages achieved by the embodiments of this article: 1. Defines a common encapsulation and body structure that enables one or more message bodies to be embedded in one transmission, thereby achieving efficient transmission of multiple messages. 2. The source and destination or destination group are identified. Therefore, the receiver can always find out the source of the message from the message, regardless of the transport protocol. 3. Topic models are identified so that multiple federated learning sessions can work in parallel. 4. Messages are uniquely identified. Therefore, duplicate messages can be detected and only one message is processed. In addition, the order of published messages from one source can be identified, even if the messages are delivered at different times. 5. Since the length of each message is defined, parsing can skip the message body if necessary. 6. The message type enables the detection of the type of the message body. Thus, the receiver can parse the message and identify the various message bodies and process them accordingly.
[0076] Thus, embodiments herein provide a control message format for exchanging messages between a device and a network and between devices during federated learning, wherein the messages include an encapsulation and one or more message bodies, wherein the message encapsulation defines a source, one or a set of destinations, and a sender, and a topic model under training, and one or more message bodies, wherein each message body indicates a message identifier to detect duplicate messages, a body size, and a message type that defines the remainder of the message body syntax and semantics.
[0077] Additionally, AI / ML for media according to example embodiments is also provided, and a messaging framework, particularly during federated learning, and specific messages for the above exchanges are provided herein.
[0078] Therefore, control messaging during federated learning is provided herein, and the innovation defines a messaging format for communication between a central service and devices located on a network to exchange messages during a federated learning process. The control messages have a common encapsulation that identifies the source / destination, the specific model in the federated learning, the federated learning state at the destination, the quality / accuracy of the locally updated model, and other information during the federated learning process.
[0079] Regarding the definition of "synchronization message" according to the embodiments herein, the synchronization message is used to ensure that all devices start the training process at the same time and at the same speed. For example, the server can send a synchronization message to all UEs to start a new round of training.
[0080] The behavior of the synchronization message is from a network server (such as the network server of network 502) to a device (such as device 501). The server sends a synchronization message to all UEs to start a new round of training at the same time. The message contains the round number and may also contain a timestamp indicating when the training round should start. This is in Figure 6 It is implemented at S601.
[0081] Regarding the syntax of the synchronization message, Table 4 is provided: Table 4 – Synchronous message body Cardinality: 0=not allowed, 1=only once, 0..1=at most once, 0..N=zero or more times, and 1..N=one or more times.
[0082] See Table 4: 1. Round_number indicates the training round in the model training. Therefore, the value 4 means that this synchronization message is for the 4th round of training. Depending on the implementation, it can be Figure 6 The 4th iteration of example 600. 2. Start_time indicates the start time of training. This value is optional. 3. Duration indicates the expected duration of the training. This value simply shows an indication of the expected time to complete a training round.
[0083] According to an embodiment, a "device eligibility message" is also provided. For example, embodiments herein use a device eligibility message to define criteria for selecting devices that will participate in the training process. For example, a server (such as a server of network 502) can send a device eligibility message to all devices belonging to an application-defined group, including device 501.
[0084] The behavior of device qualification messages is from a network server (such as a network server of network 502) to a device (such as device 501). The server sends device qualification messages to select devices that meet certain criteria defined by the application. Based on the number of criteria met, the application assigns a group id to the device. For example, the criteria can contain information about the device's operating system, processor speed, available memory, available image library (number of images, ...), the device's geographic location, language settings, and other properties. This is in Figure 6 It is implemented at S601.
[0085] Regarding the syntax of the synchronization message, Table 5 is provided: Table 5 – Syntax (Device Qualification Message) Cardinality: 0=not allowed, 1=only once, 0..1=at most once, 0..N=zero or more times, and 1..N=one or more times.
[0086] In Table 5 above: 1. Group_id is used to assign a new id to the device that meets the eligibility criteria for this message. If the device is eligible, the device uses this value as one of its group ids, and from now on, the device reacts to messages with the same group id. 2. Application_group_id is assigned by the application on the device, and if the value is equal to the value of this field, the device is eligible. 3. The Hardware, Location, and Language parameters define the device's hardware, location, and language qualification criteria, respectively. 4. Data_library_id defines the database that the eligible device should have.
[0087] Note that if more than one eligibility field is present, then according to an exemplary embodiment, a device should meet all criteria to become eligible.
[0088] Depending on the implementation, a "model evaluation message" is also provided. For example, the implementation herein uses a model evaluation message to evaluate the performance of the global model of each device and make decisions about the training process. After running the learning phase, the device sends a model evaluation message to the server that measures the accuracy of the model. The server can then decide whether to continue another round of training or stop. Alternatively, the server can use this message to request the device to evaluate the newly downloaded global model. As with the above features, the server and device here can also be the server and device 501 of the network device 502, respectively.
[0089] According to an embodiment, the behavior of the device qualification message is from a network server (such as a network server of network 502) to a device (such as device 501). The message contains the metrics used for evaluation. This Figure 6 It is implemented at S601.
[0090] According to an embodiment, the device qualification message is also sent from a device (such as device 501) to a network server (such as network 502). The message contains the metrics used for evaluation. Figure 6 It is implemented at any one of S602, S602', and S602".
[0091] Regarding the syntax of the device qualification message, Table 6 is provided: Table 6 - Syntax (Model Evaluation Message) Cardinality: 0=not allowed, 1=only once, 0..1=at most once, 0..N=zero or more times, and 1..N=one or more times.
[0092] In Table 6 above: 1. Round_number shows the round after which the evaluation is performed. 2. Metric_number shows the number of metrics included in the message body. 3. A metric is one or more of the name-value pairs showing the name of the metric (Name) and the corresponding value (Value) obtained in the evaluation.
[0093] According to the implementation mode, a "model update message" is also provided. For example, the implementation mode in this article uses the model update message to update the model parameters on the device after each round of training. For example, the server can send a model update message to all devices to update the global model with new model parameters. The model update message is also used to update the global model on the server with the new parameters updated by local training on the device. Like the above features, the server and device here can also be the server and device 501 of the network device 502, respectively, and can be implemented at S601, S602, S602', S602", respectively.
[0094] According to an embodiment, the behavior of the device qualification message is from a network server (such as the network server of network 502) to a device (such as device 501). The server sends a model update message to all devices to update the AI / ML model with new model parameters. The message contains the model id of the AI / ML model to be updated, the updated model parameters that the UE will use to train the model in the next round, and the new model id when the parameters are updated. According to an embodiment, this Figure 6 It is implemented at S601.
[0095] According to an embodiment, the behavior of device qualification messages is also from a device (such as device 501) to a network server (such as a network server of network 502). After running training locally, each device can send a model update message with updated parameters to the server. Together with the received model evaluation message, the server can decide whether the global model needs to be updated. The model update message then contains only the model id of the AI / ML model used for local training and the updated parameters. This is in Figure 6 It is implemented at any one of S602, S602', and S602".
[0096] Regarding the syntax of the device qualification message, Table 7 is provided: Table 7 - Syntax (Model Update Message) Cardinality: 0=not allowed, 1=only once, 0..1=at most once, 0..N=zero or more times, and 1..N=one or more times.
[0097] In Table 7 above: 1. Parameters contains the new model value vector. 2. New_model_id is the id of the new model when the server sends the model to one or more devices.
[0098] According to an embodiment, a "failure reporting message" is also provided. Error messages are used to handle unexpected errors or exceptions that may occur during the training process. For example, the server can send error messages to all devices to handle device failures or network interruptions.
[0099] According to an embodiment, the behavior of the device qualification message is from a network server (such as the network server of network 502) to a device (such as device 501). The server sends a request to all devices to report a device failure or network outage. For example, if a device fails to send its model parameters back to the server, the device should notify the server so that the device is removed from the training process. According to an embodiment, this is Figure 6 It is implemented at S601.
[0100] According to an embodiment, the behavior of device qualification messages is also from a device (such as device 501) to a network server (such as a network server of network 502). If a failure occurs, the device sends a failure message to the server. This is Figure 6 It is implemented at any one of S602, S602', and S602".
[0101] Regarding the syntax of the device qualification message, Table 8 is provided: Table 8 - Syntax (Fault Report Message) Cardinality: 0=not allowed, 1=only once, 0..1=at most once, 0..N=zero or more times, and 1..N=one or more times.
[0102] In Table 8 above, Message describes the cause of the failure.
[0103] Therefore, according to the embodiments herein, a set of control messages with a common header to be used during AI / ML federated learning is also provided, wherein the message id, type and size are defined, wherein various types are developed, including a) qualifications for defining which devices are eligible to participate, b) synchronizing model evaluation by defining the time when a specific round of evaluation needs to start, c) requesting evaluation of a newly downloaded model using local data, d) reporting evaluation for each request, or reporting evaluation at the end of a specific round of training, e) sending an updated model, and finally f) requesting reporting of any failures and reporting failures.
[0104] The above techniques may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media, or implemented by one or more specially configured hardware processors. Figure 7A computer system 700 suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0105] Computer software may be encoded using any suitable machine code or computer language, which may be subjected to mechanisms such as assembly, compilation, and linking to create code comprising instructions that may be directly executed by a computer central processing unit (CPU), graphics processing unit (GPU), etc., or executed through interpretation, microcode execution, etc.
[0106] The instructions may be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, IoT devices, and the like.
[0107] Figure 7 The components for computer system 700 shown in the example are exemplary in nature and are not intended to suggest any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Neither should the configuration of components be interpreted as having any dependency or requirement related to any one component or combination of components shown in the exemplary embodiment of computer system 700.
[0108] Computer system 700 may include certain human interface input devices. Such human interface input devices may be responsive to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, data glove movements), audio input (e.g., voice, tapping), visual input (e.g., gestures), olfactory input (not depicted). Human interface devices may also be used to capture certain media that are not necessarily directly related to conscious input by humans, such as audio (e.g., voice, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).
[0109] Input human interface devices may include one or more of the following (only one of each is depicted): keyboard 701 , mouse 702 , trackpad 703 , touch screen 710 , joystick 705 , microphone 706 , scanner 708 , camera 707 .
[0110] The computer system 700 may also include certain human-computer interface output devices. Such human-computer interface output devices may stimulate one or more human user senses through, for example, tactile output, sound, light, and smell / taste. Such human-computer interface output devices may include: tactile output devices (e.g., tactile feedback through touch screen 710, or joystick 705, but there may also be tactile feedback devices that are not used as input devices); audio output devices (e.g., speakers 709, headphones (not depicted)); visual output devices (e.g., screen 710, including CRT (Cathode Ray Tube, CRT) screens, LCD (Liquid Crystal Display, LCD) screens, plasma screens, OLED (Organic Light Emitting Diode, OLED) screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which may be able to output two-dimensional visual outputs or more than three-dimensional outputs through methods such as stereoscopic image output; virtual reality glasses (not depicted); holographic displays and cigarette cans (not depicted)); and printers (not depicted).
[0111] Computer system 700 may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM (Read Only Memory, ROM) / RW 720 with media such as CD / DVD 711, thumb drive 722, removable hard disk drive or solid state drive 723, traditional magnetic media such as magnetic tapes and floppy disks (not depicted), dedicated ROM / ASIC (Application Specific Integrated Circuit, ASIC) / PLD based (Programmable Logic Device, PLD) devices such as security dongles (not depicted).
[0112] Those skilled in the art should also understand that the term "computer-readable media" used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transient signals.
[0113] The computer system 700 may also include an interface 799 to one or more communication networks 798. The network 798 may be, for example, wireless, wired, optical. The network 798 may also be local, wide, metropolitan, vehicle and industrial, real-time, delay-tolerant, etc. Examples of the network 798 include: local area networks such as Ethernet; wireless LAN (Local Area Network, LAN); cellular networks including GSM (Global System for Mobile Communications, GSM), 3G (the Third Generation, 3G), 4G (the Fourth Generation, 4G), 5G (the Fifth Generation, 5G), LTE (Long Term Evolution, LTE), etc.; TV wired or wireless wide area digital networks including cable TV, satellite TV and terrestrial broadcast TV; including CANBus (Controller Area Network Bus, CANBus) vehicle and industrial networks, etc. Some networks 798 usually require external network interface adapters attached to certain universal data ports or peripheral buses (750 and 751) (such as, for example, the USB (Universal Serial Bus, USB) port of the computer system 700); other networks are usually integrated into the core of the computer system 700 by attaching to the system bus described below (for example, to the Ethernet interface in a PC computer system or to the cellular network interface in a smart phone computer system). Using any of these networks 798, the computer system 700 can communicate with other entities. Such communication can be one-way receive-only (for example, broadcast TV), one-way send-only (for example, CANbus to certain CANbus devices), or two-way, for example, to other computer systems using local area digital networks or wide area digital networks. As described above, certain protocols and protocol stacks can be used on each of these networks and network interfaces.
[0114] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core 740 of the computer system 700 .
[0115] The core 740 may include one or more central processing units (CPUs) 741, graphics processing units (GPUs) 742, graphics adapters 717, dedicated programmable processing units in the form of field programmable gate areas (FPGAs) 743, hardware accelerators 744 for certain tasks, etc. These devices, along with read-only memory (ROM) 745, random access memory 746, internal mass storage devices 747 such as internal non-user accessible hard drives, SSDs (Solid State Drives, SSDs), etc., may be connected via a system bus 748. In some computer systems, the system bus 748 may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached to the core's system bus 748 directly or via a peripheral bus 749. The architecture of the peripheral bus includes PCI (Peripheral Component Interconnect, PCI), USB, etc.
[0116] The CPU 741, GPU 742, FPGA 743, and accelerator 744 may execute certain instructions, which may constitute the above-mentioned computer code in combination. The computer code may be stored in ROM 745 or RAM (Random Access Memory, RAM) 746. Transient data may also be stored in RAM 746, while permanent data may be stored in, for example, an internal mass storage device 747. Fast storage and retrieval of any of the memory devices may be achieved by using a cache memory, which may be closely associated with one or more CPUs 741, GPUs 742, mass storage devices 747, ROM 745, RAM 746, and the like.
[0117] The computer readable medium may have computer codes for performing various computer-implemented operations. The media and computer codes may be those specially designed and constructed for the purposes of the present disclosure, or they may be of a type well known and available to those skilled in the art of computer software.
[0118] As an example and not limitation, a computer system having architecture 700 and in particular core 740 can provide functionality due to the execution of software implemented in one or more tangible computer-readable media by a processor (including a CPU, GPU, FPGA, accelerator, etc.). Such a computer-readable medium can be a medium associated with a user-accessible mass storage device as introduced above and certain storage devices of the core 740 with non-transient properties such as a core internal mass storage device 747 or ROM 745. Software implementing various embodiments of the present disclosure can be stored in such a device and executed by the core 740. Depending on specific needs, the computer-readable medium may include one or more memory devices or chips. The software can enable the core 740 and in particular the processor therein (including a CPU, GPU, FPGA, etc.) to perform specific processing or specific parts of specific processing described herein, including defining data structures stored in RAM 746 and modifying such data structures according to software-defined processing. Additionally or as an alternative, the computer system may provide functionality due to logic hard-wired or otherwise implemented in a circuit (e.g., accelerator 744), which may replace software or operate with software to perform a specific process or a specific portion of a specific process described herein. Where appropriate, reference to software may include logic and vice versa. Where appropriate, reference to a computer-readable medium may include a circuit (such as an integrated circuit (IC)) storing software for execution, a circuit implementing logic for execution, or both. The present disclosure includes any suitable combination of hardware and software.
[0119] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various substitute equivalents that fall within the scope of the present disclosure. It will therefore be appreciated that those skilled in the art will be able to design a variety of systems and methods that, although not explicitly shown or described herein, implement the principles of the present disclosure and are therefore within the spirit and scope of the present disclosure.
Claims
1. A method for artificial intelligence / machine learning (AI / ML) joint learning, the method comprising: encapsulating a message in one or more federated learning messages by a control message format, the control message format comprising a plurality of fields, the plurality of fields respectively indicating one of: an identifier of the message, a size of the message, a type of the message, and a body of the message; and controlling the AI / ML joint learning based on the message encapsulated by the control message format, wherein the AI / ML joint learning comprises a server, the server controls a plurality of separate devices to implement a joint part of the AI / ML joint learning, and reports results of implementing the joint part to the server from each of the plurality of separate devices, respectively, and Wherein, the body of the message indicates at least one of the following: AI / ML federated learning synchronization between the multiple individual devices, device qualification of the AI / ML federated learning, model evaluation of the AI / ML federated learning, model update of the AI / ML federated learning, and errors of the AI / ML federated learning.
2. The method according to claim 1, in, The AI / ML joint learning includes multiple rounds of joint learning, each round of joint learning is a back-and-forth iteration between the server and the multiple individual devices, and Wherein, the AI / ML joint learning is implemented in parallel at the multiple separate devices.
3. The method according to claim 2, wherein: The body of the message indicates AI / ML federated learning synchronization between the multiple devices, and in one round of the multiple rounds of federated learning, the federated learning will start simultaneously at each of the multiple individual devices.
4. The method according to claim 1, wherein: The body of the message indicates device eligibility for the AI / ML federated learning and one or more criteria for device eligibility.
5. The method according to claim 4, wherein: The one or more criteria for device qualification for the AI / ML joint learning is any of operating system, processor speed, available memory, available image library, number of images, geographic location, language setting.
6. The method according to claim 1, wherein: The AI / ML joint learning model evaluation instructs the multiple individual devices to implement the evaluation of the AI / ML joint learning model.
7. The method according to claim 1, wherein: The model evaluation of the AI / ML joint learning instructs the multiple separate devices to implement evaluation of the model separate from the AI / ML joint learning.
8. The method according to claim 1, wherein: The AI / ML joint learning model update instructs the multiple individual devices to update parameters of the AI / ML joint learning model.
9. The method according to claim 1, wherein: The model updates are instructions from the server to the plurality of individual devices.
10. The method according to claim 1, wherein: The model update is an instruction from at least one of the plurality of individual devices to the server.
11. An apparatus comprising: at least one memory configured to store computer program code; At least one processor is configured to access the computer program code and operate according to the instructions of the computer program code, the computer program code comprising: Packaging code configured to cause the at least one processor to package a message in one or more federated learning messages through a control message format, the control message format comprising a plurality of fields, the plurality of fields respectively indicating one of: an identifier of the message, a size of the message, a type of the message, and a body of the message; and control code configured to cause the at least one processor to control artificial intelligence / machine learning (AI / ML) joint learning based on the message encapsulated by the control message format, wherein the AI / ML joint learning comprises a server, the server controls a plurality of separate devices to implement a joint part of the AI / ML joint learning, and reports results of implementing the joint part to the server from each of the plurality of separate devices, respectively, and Wherein, the body of the message indicates at least one of the following: AI / ML federated learning synchronization between the multiple individual devices, device qualification of the AI / ML federated learning, model evaluation of the AI / ML federated learning, model update of the AI / ML federated learning, and errors of the AI / ML federated learning.
12. The device according to claim 11, in, The AI / ML joint learning includes multiple rounds of joint learning, each round of joint learning is a back-and-forth iteration between the server and the multiple individual devices, and Wherein, the AI / ML joint learning is implemented in parallel at the multiple separate devices.
13. The device according to claim 12, wherein: The body of the message indicates AI / ML federated learning synchronization between the multiple devices, and in one round of the multiple rounds of federated learning, the federated learning will start simultaneously at each of the multiple individual devices.
14. The device according to claim 11, wherein: The body of the message indicates device eligibility for the AI / ML federated learning and one or more criteria for device eligibility.
15. The device according to claim 14, wherein: The one or more criteria for device qualification for the AI / ML joint learning is any of operating system, processor speed, available memory, available image library, number of images, geographic location, language setting.
16. The device according to claim 11, wherein The AI / ML joint learning model evaluation instructs the multiple individual devices to implement the evaluation of the AI / ML joint learning model.
17. The device according to claim 11, wherein: The model evaluation of the AI / ML joint learning instructs the multiple separate devices to implement evaluation of the model separate from the AI / ML joint learning.
18. The device according to claim 11, wherein The AI / ML joint learning model update instructs the multiple individual devices to update parameters of the AI / ML joint learning model.
19. The device according to claim 11, wherein: The model update is one of: a first instruction from the server to the plurality of individual devices; and a second instruction from at least one of the plurality of individual devices to the server.
20. A non-transitory computer-readable medium storing a program, the program causing a computer to: encapsulating a message in one or more federated learning messages by a control message format, the control message format comprising a plurality of fields, the plurality of fields respectively indicating one of: an identifier of the message, a size of the message, a type of the message, and a body of the message; and controlling artificial intelligence / machine learning (AI / ML) joint learning based on the message encapsulated by the control message format, in, The AI / ML joint learning includes a server that controls a plurality of separate devices to implement a joint part of the AI / ML joint learning, and reports results of implementing the joint part to the server from each of the plurality of separate devices, respectively, and Wherein, the body of the message indicates at least one of the following: AI / ML federated learning synchronization between the multiple individual devices, device qualification of the AI / ML federated learning, model evaluation of the AI / ML federated learning, model update of the AI / ML federated learning, and errors of the AI / ML federated learning.