Data communication method, device, equipment, medium and product

By coordinating scheduling at the cloud and edge, and utilizing a dynamic reversible encoder-decoder and a dynamic weight determination module, dynamic fusion and bidirectional conversion of multimodal data are achieved. This solves the problems of feature drift and resource efficiency of multimodal data in dynamic environments, and reduces latency and memory consumption.

CN120980084APending Publication Date: 2025-11-18CHINA MOBILE (XIONGAN) ICT CO LTD +4
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511435797.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing multimodal data fusion methods are difficult to adapt to dynamic environments, resulting in feature drift and semantic mismatch between modalities. They are also resource inefficient, with traditional encoding and decoding architectures having high memory consumption and being unable to achieve bidirectional conversion of arbitrary modalities.

Method used

By acquiring task instructions from the cloud and distributing them to the edge, a dynamic sharding migration strategy is used to acquire multimodal data. This data is then fused and encoded using a dynamic reversible encoder/decoder and a dynamic weight determination module, enabling cross-modal interaction and bidirectional conversion.

Benefits of technology

It solves the problem of fragmented multimodal data, realizes bidirectional conversion of arbitrary modes, reduces end-to-end latency, and optimizes resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980084A_ABST
    Figure CN120980084A_ABST
Patent Text Reader

Abstract

The invention discloses a data communication method and device, equipment, a medium and a product, and the method comprises the steps: obtaining a to-be-processed task instruction through a cloud end, and distributing the to-be-processed task instruction to an edge end through a dynamic fragmentation migration strategy; obtaining a current channel state and multi-modal data corresponding to the task instruction to be processed through the edge end; according to the current channel state, the multi-mode data and a dynamic reversible codec, fused coded data is determined and transmitted back to the cloud, and the dynamic reversible codec comprises a bidirectional coding and decoding module and a dynamic weight determination module; and receiving the fused coded data through the cloud, processing and feeding back. Through edge and cloud collaborative scheduling, the dynamic weight is generated through the edge end according to the current channel state, the channel state change is responded in real time, and the fusion coding data is generated through combination of the multi-modal data and the dynamic weight, so that the problem of splitting of the multi-modal data is solved, bidirectional conversion of any modal is realized, and the end-to-end delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a data communication method, apparatus, device, medium, and product. Background Technology

[0002] Multimodal tasks typically involve processing data from different sources (such as text, images, and speech) and integrating this data to obtain richer semantic information. For example, the transmission of multimodal data is crucial in applications such as security monitoring, environmental perception of autonomous vehicles, and intelligent medical diagnosis.

[0003] Multimodal fusion achieves cross-modal interaction through feature-level concatenation (such as direct concatenation of text and image vectors) or static attention mechanisms (such as dual cross-attention modules DCA); reversible neural networks (such as RevNet) achieve lossless transmission of single-modal data (such as images) through reversible blocks with a fixed number of layers.

[0004] However, this multimodal fusion is a static fusion mechanism, which is difficult to adapt to dynamic environments, leading to feature drift and semantic mismatch between modalities, and it cannot achieve bidirectional conversion between arbitrary modalities. Reversible networks have significant limitations. Although existing methods (such as HiNet and Steg-cINN) achieve lossless transmission in single-modal tasks, they are not adapted to multimodal scenarios. For example, RevNet's fixed feature decomposition method limits the flexibility of cross-modal interaction. The problem of low resource efficiency is prominent. Traditional encoding and decoding architectures (such as LSTM and Transformer) need to cache all intermediate states, resulting in high memory consumption and difficulty in scaling to large-scale tasks. Segmentation methods that rely on external models (such as SAM) have high computational costs. Summary of the Invention

[0005] This invention provides a data communication method, apparatus, device, medium, and product to achieve dynamic fusion and bidirectional conversion of multimodal data.

[0006] According to one aspect of the present invention, a data communication method is provided, comprising:

[0007] The task instructions to be processed are obtained from the cloud and distributed to the edge end through a dynamic sharding migration strategy.

[0008] The current channel state and the multimodal data corresponding to the task instruction to be processed are obtained through the edge terminal.

[0009] Based on the current channel state, the multimodal data, and the dynamic reversible codec, the fused coded data is determined and transmitted back to the cloud. The dynamic reversible codec includes a bidirectional encoding / decoding module and a dynamic weight determination module.

[0010] The cloud receives the fused encoded data, processes it, and then provides feedback.

[0011] According to a second aspect of the present invention, a data communication device is provided, comprising:

[0012] The task allocation module is used to obtain the task instructions to be processed through the cloud and allocate the task instructions to the edge end through a dynamic sharding migration strategy.

[0013] The data acquisition module is used to acquire the current channel state and the multimodal data corresponding to the task instruction to be processed through the edge terminal;

[0014] The data backhaul module is used to determine the fused coded data and backhaul it to the cloud based on the current channel state, the multimodal data and the reversible transformation layer. The reversible transformation layer performs bidirectional propagation symmetric calculation through a hidden state buffer.

[0015] The data feedback module is used to receive the fused encoded data through the cloud, process it, and then provide feedback.

[0016] According to a third aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data communication method described in any embodiment of the present invention.

[0020] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the data communication method according to any embodiment of the present invention.

[0021] According to a fifth aspect of the present invention, embodiments of the present invention also provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the data communication method of any embodiment of the present invention.

[0022] The technical solution of this invention obtains the task instructions to be processed from the cloud and distributes them to the edge end through a dynamic fragmentation migration strategy. The edge end obtains the current channel state and the multimodal data corresponding to the task instructions. Based on the current channel state, multimodal data, and a dynamic reversible codec, fused coded data is determined and transmitted back to the cloud. The dynamic reversible codec includes a bidirectional encoding / decoding module and a dynamic weight determination module. The cloud receives the fused coded data, processes it, and provides feedback. Through collaborative scheduling between the edge and cloud, the edge end generates dynamic weights according to the current channel state, responds to channel state changes in real time, and generates fused coded data by combining multimodal data with dynamic weights. This solves the problem of fragmented multimodal data, achieves bidirectional conversion of arbitrary modes, and reduces end-to-end latency.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of a data communication method provided according to Embodiment 1 of the present invention;

[0026] Figure 2 This is a flowchart of a data communication method provided according to Embodiment 2 of the present invention;

[0027] Figure 3 This is a schematic diagram of the structure of a data communication device according to Embodiment 3 of the present invention;

[0028] Figure 4 This is a schematic diagram of the structure of an electronic device that implements an embodiment of the present invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] Example 1

[0032] Figure 1 This is a flowchart illustrating a data communication method provided in Embodiment 1 of the present invention. This embodiment is applicable to the encoding and decoding of multimodal data. The method can be executed by a data communication device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0033] S110: Obtain the task instructions to be processed through the cloud, and distribute the task instructions to the edge end through a dynamic sharding migration strategy.

[0034] In this embodiment, the cloud can be understood as a centralized resource pool centered on cloud computing and cloud storage. The edge can be a network edge node (such as an edge server and smart device) close to the data source, responsible for localized real-time data processing and low-latency response. Together, they form an integrated "cloud-edge-device" architecture to support the efficient operation of digital scenarios. The dynamic sharding migration strategy can be understood as a migration mechanism that dynamically adjusts data or service shards between nodes in real time. By intelligently sensing changes in system load, resource utilization, and business needs, data shards or service modules are dynamically scheduled between different nodes (such as cloud servers and edge nodes) to achieve load balancing, resource optimization, and service quality improvement. The task instructions to be processed can be understood as task calculation instructions that need to be completed, such as image diagnosis tasks in the medical field (such as image → text report generation, pathological description → image reconstruction), multi-sensor data fusion tasks in the industrial Internet of Things (such as the transmission tasks of sensor signals from different key devices under different channel conditions), and video streaming tasks.

[0035] Specifically, user-uploaded task instructions can be obtained through the cloud, and computationally intensive task instructions can be offloaded to the edge through a dynamic sharding migration strategy using reinforcement learning (state space definition, reward function design, and action space modeling, etc.).

[0036] S120. Obtain the current channel status and multimodal data corresponding to the task instructions to be processed through the edge terminal.

[0037] In this embodiment, the current channel state can be understood as a key indicator reflecting the quality of the current communication link, such as signal-to-noise ratio (SNR) and time delay. Multimodal data can be understood as a collection containing various different types of data, such as text, images, and audio.

[0038] Specifically, the current channel status and multimodal data corresponding to the task instructions to be processed are obtained through the edge device, such as by collecting multimodal data through connected sensors or obtaining multimodal data to be processed from the task instructions.

[0039] S130. Based on the current channel state, multimodal data, and dynamic reversible codec, determine the fused coded data and transmit it back to the cloud. The dynamic reversible codec includes a bidirectional encoding / decoding module and a dynamic weight determination module.

[0040] In this embodiment, the dynamic reversible codec can be understood as a tool for multimodal data encoding and decoding in conjunction with the current channel state. The fused coded data can be understood as a fused feature vector formed by mapping features from different modalities to a shared low-dimensional feature space. The bidirectional encoding / decoding module can be understood as coded data that supports forward conversion and reverse recovery of data, allowing lossless conversion between the original data and the encoded data. The dynamic weight determination module can be understood as a module that determines the weights corresponding to different modal types of data in real time. For example, it can model a dynamic function for a lightweight LSTM model, such as α(t) = f(SNR, delay), with the input being the current channel state time series and the output being normalized weight coefficients.

[0041] Specifically, the edge device can analyze the current channel state through a dynamic weight determination module to determine the weights of different modal types, encode multimodal data through a bidirectional encoding and decoding module to determine multimodal features, and dynamically fuse multimodal features through weights to obtain fused encoded data and transmit it back to the cloud.

[0042] For example, the weight allocation strategy of the dynamic weight module can automatically increase the weight of the speech modality in low signal-to-noise ratio (SNR < 10dB) scenarios, suppress the image transmission load, and achieve cross-modal feature ratio optimization.

[0043] S140: Receive fused encoded data via the cloud, process it, and provide feedback.

[0044] Specifically, the system receives fused encoded data from various edge devices via the cloud, processes it, and then feeds it back to the user.

[0045] The technical solution of this invention obtains the task instructions to be processed from the cloud and distributes them to the edge end through a dynamic fragmentation migration strategy. The edge end obtains the current channel state and the multimodal data corresponding to the task instructions. Based on the current channel state, multimodal data, and a dynamic reversible codec, fused coded data is determined and transmitted back to the cloud. The dynamic reversible codec includes a bidirectional encoding / decoding module and a dynamic weight determination module. The cloud receives the fused coded data, processes it, and provides feedback. Through collaborative scheduling between the edge and cloud, the edge end generates dynamic weights according to the current channel state, responds to channel state changes in real time, and generates fused coded data by combining multimodal data with dynamic weights. This solves the problem of fragmented multimodal data, achieves bidirectional conversion of arbitrary modes, and reduces end-to-end latency.

[0046] Furthermore, based on the above embodiments, the steps of determining the fused coded data and transmitting it back to the cloud according to the current channel state, multimodal data, and dynamic reversible codec can be refined as follows:

[0047] Based on the current channel state and the dynamic weight determination module, the dynamic weight coefficients for different modal types are determined; the multimodal data is encoded by the bidirectional encoding and decoding module based on encoders of different modal types to determine the multimodal encoded data; based on the multimodal encoded data, the reversible transformation layer of the bidirectional encoding and decoding module, and the dynamic weight coefficients, the fused encoded data is determined and transmitted back to the cloud.

[0048] In this embodiment, modality type is used to distinguish different data in multimodal data, such as text, image, and speech. Dynamic weighting coefficients can be understood as coefficients used to adjust the fusion ratio of different modality types. An encoder can be understood as a module that converts input data into a specific form of coded signal.

[0049] Specifically, the edge device can input the current channel state into the dynamic weight determination module to obtain dynamic weight coefficients for different modal types. The bidirectional encoding / decoding module encodes the multimodal data based on encoders of different modal types to determine the encoded multimodal data. The edge device can then determine the fused encoded data and transmit it back to the cloud based on the multimodal encoded data, the reversible transform layer of the bidirectional encoding / decoding module, and the dynamic weight coefficients.

[0050] For example, a text encoder can use Transformer encoding, and the resulting data can be encoded using V. textThis indicates that the corresponding dynamic weight coefficient is α1(t). Image encoders can use CNN+VIT encoding, and the obtained data can be processed through V... img This indicates that the corresponding dynamic weight coefficient is α2(t). Speech encoders can use LSTM encoding, and the obtained data can be processed through V. audio This indicates that the corresponding dynamic weight coefficient is α3(t), then the fused encoded data V fused It can be determined using the following formula: .

[0051] Based on the above embodiments, the steps of determining the fused encoded data and transmitting it back to the cloud according to the multimodal encoded data, the reversible transform layer of the bidirectional encoding / decoding module, and the dynamic weighting coefficients can be refined, including:

[0052] Multimodal encoded data is input into a reversible transformation layer to obtain multimodal features. The reversible transformation layer performs bidirectional propagation symmetric calculations through a hidden state cache. The multimodal features are fused using dynamic weight coefficients to obtain fused encoded data, which is then transmitted back to the cloud.

[0053] In this embodiment, the reversible transform layer can be understood as a module that transforms the input data and can recover the original input through inverse transform. The hidden state cache can be understood as a mechanism for temporarily storing and reusing the model's hidden states in sequence data processing. Specifically, the edge can input multimodal encoded data into the reversible transform layer, and through bidirectional interaction, obtain multimodal features. The reversible transform layer performs bidirectional propagation symmetric computation through the hidden state cache. The edge can multiply the dynamic weight coefficients with the multimodal features under the corresponding modality, and add the product results from different modalities to obtain fused encoded data.

[0054] Based on the above embodiments, the step of inputting multimodal encoded data into a reversible transform layer to obtain multimodal features can be refined as follows:

[0055] The encoded data of each modality type in the multimodal encoded data is used as input features. The input features are segmented into temporal features and spatial features through an invertible transform layer. The spatial features are processed according to the convolution residual function in the invertible transform layer to determine the first feature. The temporal features are processed according to the temporal modeling function in the invertible transform layer to determine the second feature. The multimodal features are determined based on the first feature and the second feature.

[0056] In this embodiment, temporal features can be understood as focusing on the characteristics of data changing over time. Spatial features can be understood as the distribution, location, and relationships of data in physical or abstract space. The convolutional residual function is used to extract local spatial features. The first feature can be understood as the feature after fusing local spatial features. The temporal modeling LSTM function can be understood as being used to capture long-range dependencies. The second feature can be understood as the feature after fusing long-range dependencies.

[0057] Specifically, at the edge, the encoded data for each modality in the multimodal encoded data can be used as input features. A reversible transform layer is used to segment the input features into temporal and spatial features. Local spatial features are extracted from the spatial features using the convolutional residual function in the reversible transform layer and fused with the spatial features to determine the first feature. Long-distance dependencies in the temporal features are captured using the temporal modeling function in the reversible transform layer and fused with the temporal features to determine the second feature. The first and second features under different modality types are fused to form the features for that modality type, and multimodal features are determined based on the features under different modality types.

[0058] For example, the forward encoding formula for the reversible transform layer can be: .

[0059] Where x1 and x2 are the two parts after input feature segmentation (such as the temporal and spatial features of a video frame sequence), F(·) is the convolution residual function, y1 is the first feature, G(·) is the temporal modeling function, and y2 is the second feature.

[0060] The formula for the inverse transformation layer during reverse recovery is: .

[0061] During reverse recovery, lossless feature recovery is achieved through hidden state caching.

[0062] The technical solution of this invention dynamically generates fusion weight coefficients based on the current channel state, optimizes the interaction ratio between modes, and fuses multimodal data under different mode types through a bidirectional encoding and decoding module combined with dynamic weight coefficients, thus solving the problem of multimodal data fragmentation. Through the constructed bidirectional reversible transformation layer, it supports lossless transmission of heterogeneous modes and reduces memory consumption, realizes bidirectional conversion of arbitrary modes, and reduces end-to-end latency.

[0063] As a first optional embodiment of this embodiment, after determining the fused coded data based on the multimodal coded data, the reversible transform layer, and the dynamic weight coefficients, the method further includes:

[0064] The dynamic weight determination module is optimized based on fused encoded data using KL divergence and contrastive learning loss.

[0065] Specifically, after determining the fused encoded data based on multimodal encoded data, reversible transform layers, and dynamic weight coefficients, the fused encoded data is used as input. The dynamic weight determination module is then optimized using a loss function constructed with KL divergence and contrastive learning loss. KL divergence constrains the similarity between the fused features and the single-modal distribution, avoiding semantic shift. Contrastive learning loss forces cross-modal feature alignment by constructing positive and negative sample pairs (such as matched text-image pairs and random noise pairs). For example, in image diagnosis tasks, this ensures that the latent space distance between pathological description text and CT image regions is less than that of unrelated pairs.

[0066] In the first optional embodiment of this example, a unified latent space mapping matrix is ​​constructed by adversarial latent space alignment (KL divergence + contrastive learning loss), which forces the consistency of the distribution of text, image and speech features, solves the semantic gap between modalities, and enhances the robustness of the dynamic weight module in noisy environments.

[0067] For example, a specific example can be used to demonstrate the architecture of a dynamic reversible encoder. Figure 2 This is an example diagram of a dynamic reversible encoder architecture in a data communication method provided in Embodiment 1 of the present invention, as shown below. Figure 2 As shown, the system mainly consists of two parts: a bidirectional encoding / decoding module and a dynamic weight determination module. The bidirectional encoding / decoding module includes a text encoder (Transformer), an image encoder (CNN+ViT hybrid structure), and a speech encoder (LSTM), which are connected through a reversible transformation layer. After segmenting the input text, the text encoder extracts the semantic vector V_text through a multi-head self-attention mechanism. For example, in a medical diagnosis scenario, it transforms the patient's chief complaint into structured semantic features. The image encoder adopts a hierarchical processing strategy, first extracting local texture features (such as the edges of lesion areas in CT images) through CNN, and then encoding global spatial correlations through ViT to generate V_img. The speech encoder processes MFCC acoustic features based on an LSTM network and extracts temporal context information V_audio. The above multimodal features interact bidirectionally through the reversible transformation layer and are fused with the generated dynamic weight coefficients to obtain fused encoded data. Finally, the dynamic weight determination module is jointly optimized through KL divergence and contrastive learning loss.

[0068] As a second optional embodiment of this first embodiment, based on the above embodiment, it further includes:

[0069] The reversible transformation layer at the edge is used to reconstruct the unified modal data corresponding to the task instruction to be processed, and the reconstructed data is obtained. The reversible transformation layer is then fixed into FPGA logic units in the form of parameters.

[0070] Specifically, the edge layer can be embedded as a reversible transformation layer within the FPGA logic unit to reconstruct data of the same modality. This involves using the formula from the reverse recovery example above to perform reverse recovery and obtain the reconstructed data. This design enables the system to process video streams up to 10 minutes long (30fps) with memory consumption only 1 / 8 that of traditional LSTM. Furthermore, the reverse process does not require storing intermediate activation values, reducing memory complexity from O(LT) to O(L) (where L is the sequence length and T is the time step). It supports end-to-end reversible processing of video streams (such as 30fps sequences) with a reconstruction error of ||x - z||. 2 < 10 −6 (Measured values) are for improving real-time performance in edge scenarios.

[0071] Example 2

[0072] Figure 3 This is a schematic diagram of a data communication device provided in Embodiment 2 of the present invention. Figure 3 As shown, the device includes:

[0073] Task allocation module 31 is used to obtain the task instructions to be processed through the cloud and allocate the task instructions to the edge end through a dynamic sharding migration strategy.

[0074] Data acquisition module 32 is used to acquire the current channel state and the multimodal data corresponding to the task instruction to be processed through the edge terminal;

[0075] Data backhaul module 33 is used to determine fused coded data and backhaul it to the cloud based on the current channel state, the multimodal data and the reversible transformation layer. The reversible transformation layer performs bidirectional propagation symmetric calculation through hidden state buffer.

[0076] The data feedback module 34 is used to receive the fused encoded data through the cloud, process it, and provide feedback.

[0077] The technical solution of this invention obtains the task instructions to be processed from the cloud and distributes them to the edge end through a dynamic fragmentation migration strategy. The edge end obtains the current channel state and the multimodal data corresponding to the task instructions. Based on the current channel state, multimodal data, and a dynamic reversible codec, fused coded data is determined and transmitted back to the cloud. The dynamic reversible codec includes a bidirectional encoding / decoding module and a dynamic weight determination module. The cloud receives the fused coded data, processes it, and provides feedback. Through collaborative scheduling between the edge and cloud, the edge end generates dynamic weights according to the current channel state, responds to channel state changes in real time, and generates fused coded data by combining multimodal data with dynamic weights. This solves the problem of fragmented multimodal data, achieves bidirectional conversion of arbitrary modes, and reduces end-to-end latency.

[0078] Furthermore, the data return module includes:

[0079] The first determining unit is used to determine the dynamic weight coefficients under different mode types based on the current channel state and the dynamic weight determining module.

[0080] The second determining unit is used to encode the multimodal data based on encoders of different modal types using a bidirectional encoding / decoding module to determine the multimodal encoded data.

[0081] The third determining unit is used to determine the fused encoded data and transmit it back to the cloud based on the multimodal encoded data, the reversible transformation layer of the bidirectional encoding and decoding module, and the dynamic weighting coefficients.

[0082] The third determining unit includes:

[0083] The first determining subunit is used to input the multimodal encoded data into the reversible transform layer to obtain multimodal features. The reversible transform layer performs bidirectional propagation symmetric calculation through a hidden state buffer.

[0084] The second determining subunit is used to fuse the multimodal features through the dynamic weight coefficients to obtain fused encoded data and transmit it back to the cloud.

[0085] Specifically, the first determining subunit is used for:

[0086] The encoded data of each modality type in the multimodal encoded data is used as input features, and the input features are segmented into temporal features and spatial features through an invertible transformation layer;

[0087] The spatial features are processed according to the convolution residual function in the reversible transform layer to determine the first feature;

[0088] The temporal features are processed according to the temporal modeling function in the reversible transformation layer to determine the second feature;

[0089] Based on each of the first features and each of the second features, multimodal features are determined.

[0090] Optionally, the device may also include:

[0091] The module optimization module is used to optimize the dynamic weight determination module based on the fused coding data after determining the fused coding data according to the multimodal coding data, the reversible transformation layer and the dynamic weight coefficients, using KL divergence and contrastive learning loss.

[0092] Optionally, the device may also include:

[0093] The data reconstruction module is used to reconstruct the unified modal data corresponding to the task instruction to be processed through the reversible transformation layer at the edge, and to obtain the reconstructed data. The reversible transformation layer is fixed as an FPGA logic unit in the form of parameters.

[0094] The data communication device provided in the embodiments of the present invention can execute the data communication method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0095] Example 3

[0096] Figure 4 A schematic diagram of an electronic device 40 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0097] like Figure 4 As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded from storage unit 48 into the RAM 43. The RAM 43 may also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.

[0098] Multiple components in electronic device 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0099] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as data communication methods.

[0100] In some embodiments, the data communication method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the data communication method described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to perform the data communication method by any other suitable means (e.g., by means of firmware).

[0101] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0102] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0103] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0104] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0105] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0106] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0107] In one embodiment, the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the data communication method of any embodiment of the present invention.

[0108] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof. Programming languages ​​include object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0109] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0110] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data communication method, characterized in that, include: The task instructions to be processed are obtained from the cloud and distributed to the edge end through a dynamic sharding migration strategy. The current channel state and the multimodal data corresponding to the task instruction to be processed are obtained through the edge terminal. Based on the current channel state, the multimodal data, and the dynamic reversible codec, the fused coded data is determined and transmitted back to the cloud. The dynamic reversible codec includes a bidirectional encoding / decoding module and a dynamic weight determination module. The cloud receives the fused encoded data, processes it, and then provides feedback.

2. The method according to claim 1, characterized in that, The step of determining fused coded data and transmitting it back to the cloud based on the current channel state, the multimodal data, and the dynamic reversible codec includes: Based on the current channel state and the dynamic weight determination module, determine the dynamic weight coefficients for different mode types; The multimodal data is encoded using a bidirectional encoding / decoding module based on encoders of different modal types to determine the multimodal encoded data; Based on the multimodal encoded data, the reversible transformation layer of the bidirectional encoding and decoding module, and the dynamic weight coefficients, the fused encoded data is determined and transmitted back to the cloud.

3. The method according to claim 2, characterized in that, The step of determining the fused encoded data and transmitting it back to the cloud based on the multimodal encoded data, the reversible transform layer of the bidirectional encoding / decoding module, and the dynamic weighting coefficients includes: The multimodal encoded data is input into the reversible transform layer to obtain multimodal features. The reversible transform layer performs bidirectional propagation symmetric computation through a hidden state buffer. The multimodal features are fused using the dynamic weighting coefficients to obtain fused encoded data, which is then transmitted back to the cloud.

4. The method according to claim 3, characterized in that, The step of inputting the multimodal encoded data into a reversible transform layer to obtain multimodal features includes: The encoded data of each modality type in the multimodal encoded data is used as input features, and the input features are segmented into temporal features and spatial features through an invertible transformation layer; The spatial features are processed according to the convolution residual function in the reversible transform layer to determine the first feature; The temporal features are processed according to the temporal modeling function in the reversible transformation layer to determine the second feature; Based on each of the first features and each of the second features, multimodal features are determined.

5. The method according to claim 2, characterized in that, After determining the fused coded data based on the multimodal coded data, the reversible transform layer, and the dynamic weight coefficients, the method further includes: The dynamic weight determination module is optimized based on the fused encoded data using KL divergence and contrastive learning loss.

6. The method according to claim 1, characterized in that, Also includes: The reversible transformation layer at the edge is used to reconstruct the unified modal data corresponding to the task instruction to be processed, thereby obtaining reconstructed data. The reversible transformation layer is then fixed into FPGA logic units in the form of parameters.

7. A data communication device, characterized in that, include: The task allocation module is used to obtain the task instructions to be processed through the cloud and allocate the task instructions to the edge end through a dynamic sharding migration strategy. The data acquisition module is used to acquire the current channel state and the multimodal data corresponding to the task instruction to be processed through the edge terminal; The data backhaul module is used to determine the fused coded data and backhaul it to the cloud based on the current channel state, the multimodal data and the reversible transformation layer. The reversible transformation layer performs bidirectional propagation symmetric calculation through a hidden state buffer. The data feedback module is used to receive the fused encoded data through the cloud, process it, and then provide feedback.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data communication method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the data communication method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the data communication method according to any one of claims 1-6.