Data transmission method, communication apparatus, and communication system
By selecting the most matching compression scheme based on the distribution characteristics of the AI model data blocks and performing special processing for important and stable data, the problem of large resource overhead for data transmission of AI model is solved, and efficient resource utilization and information retention of important data is achieved.
Patent Information
- Application Number
- PCT/CN2023/084824
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-03-29
- Publication Date
- 2025-06-19
AI Technical Summary
In the field of AI, the interaction between ultra-large-scale AI model data leads to huge overhead in transmission resources, which has become a key issue that needs to be solved urgently.
By determining the compression scheme of each data block based on the data block distribution characteristics of the AI model, the most matching compression of different data blocks is achieved, and no compression is performed for data blocks with high importance, and no transmission is performed for data blocks with low changes, reducing resource usage.
It effectively reduces the resources required to transmit AI model data, improves transmission efficiency and resource utilization, and ensures that the amount of information of important data is not lost.
Smart Images

Figure CN2023084824_19062025_PF_FP_ABST
Abstract
Description
Data transmission method, communication device and communication system Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a data transmission method, a communication device, and a communication system. Background Art
[0002] Depending on the application scenario of artificial intelligence (AI) technology, the scale of AI models can be very large. For large-scale AI models, the number of AI model data can reach millions or even tens of millions. When AI model data needs to be exchanged between different devices, it will cause huge transmission resource overhead.
[0003] With transmission resources becoming increasingly scarce, how to transmit data for ultra-large-scale AI models has become a key issue that needs to be urgently addressed in the AI field.
[0004] Summary of the Invention
[0005] The present application provides a data transmission method, a communication device, and a communication system for reducing the resources required to transmit AI model data.
[0006] In a first aspect, an embodiment of the present application provides a data transmission method, which is performed by a first device, which may be a terminal device, a network device, or other device, or a device, module, etc. used in conjunction with the aforementioned device. The method comprises: determining compressed data based on N data blocks corresponding to an AI model, wherein the compression scheme corresponding to each of the N data blocks is related to the distribution characteristics of the data blocks, the compressed data including M compressed blocks, and the M compressed blocks corresponding one-to-one to the M data blocks of the N data blocks; wherein N is an integer greater than 1 and M is an integer not greater than N; and outputting the compressed data.
[0007] In one implementation, outputting the compressed data includes sending the compressed data to a second device.
[0008] The above scheme determines the compression scheme corresponding to each data block based on the distribution characteristics of the N data blocks corresponding to the AI model. It can select the most matching compression scheme for different data blocks instead of selecting the same compression scheme for all data blocks. This helps to improve the compression effect of the AI model and save the transmission resources used when transmitting the compressed data blocks.
[0009] In a possible implementation method, the M compressed blocks include at least one first data block, the first data block is a data block among the N data blocks, and the importance of the first data block is greater than a first threshold.
[0010] The above solution does not compress the data blocks with higher importance, but transmits them directly, which can ensure that the data with higher importance will not lose information due to compression.
[0011] In a possible implementation method, the N data blocks include at least one second data block, and the second data block does not correspond to the M compressed blocks.
[0012] In one possible implementation method, the second data block is weight data and the degree of change of the second data block is lower than a second threshold; or, the second data block is gradient data and the probability that the second data block is in a stable interval is greater than a third threshold.
[0013] In the above solution, the first device does not transmit weight data with a change degree lower than the second threshold and gradient data with a probability of being in a stable range greater than the third threshold to the second device, which can reduce resource occupation and improve resource utilization.
[0014] In one possible implementation method, the method further includes: sending compression parameters to the second device; wherein the compression parameters are used to indicate one or more of the following information: model parameters of the AI model, the compression granularity of each of the N data blocks, or the compression scheme corresponding to each of the N data blocks; the model parameters include at least one of the following: the type of the AI model, the number of layers of the AI model, the number of nodes in each layer of the AI model, or the function of each layer in the AI model.
[0015] The above solution helps the second device to correctly and quickly determine the AI model by sending compression parameters to the second device.
[0016] In one possible implementation method, the compression granularity of each of the N data blocks is a partial node in a layer of the AI model, a layer, or multiple layers.
[0017] The above solution can flexibly map the data that needs to be transmitted in the AI model into N data blocks, which helps improve transmission efficiency.
[0018] In one possible implementation method, the M data blocks include at least one third data block, the third data block includes one or more sub-data blocks, and the difference between distribution information of any two sub-data blocks among the multiple sub-data blocks is less than a fourth threshold.
[0019] The above solution can accurately determine the data blocks, help improve the compression effect of the AI model, and thus save the transmission resources used when transmitting the compressed data blocks.
[0020] In a second aspect, an embodiment of the present application provides a data transmission method, which is performed by a second device, which may be a terminal device, a network device, or other device, or a device, module, etc. used in conjunction with the aforementioned device. The method comprises: receiving compressed data from a first device, wherein the compressed data comprises M compressed blocks, wherein the M compressed blocks correspond one-to-one to M data blocks of N data blocks corresponding to an artificial intelligence (AI) model, and wherein the compression scheme corresponding to each of the N data blocks is related to the distribution characteristics of the data blocks; wherein N is an integer greater than 1, and M is an integer not greater than N; and determining the AI model based on the compressed data.
[0021] The above scheme determines the compression scheme corresponding to each data block based on the distribution characteristics of the N data blocks corresponding to the AI model. It can select the most matching compression scheme for different data blocks instead of selecting the same compression scheme for all data blocks. This helps to improve the compression effect of the AI model and save the transmission resources used when transmitting the compressed data blocks.
[0022] In a possible implementation method, the M compressed blocks include at least one first data block, the first data block is a data block among the N data blocks, and the importance of the first data block is greater than a first threshold.
[0023] The above solution does not compress the data blocks with higher importance, but transmits them directly, which can ensure that the data with higher importance will not lose information due to compression.
[0024] In a possible implementation method, the N data blocks include at least one second data block, and the second data block does not correspond to the M compressed blocks.
[0025] In one possible implementation method, the second data block is weight data and the degree of change of the second data block is lower than a second threshold; or, the second data block is gradient data and the probability that the second data block is in a stable interval is greater than a third threshold.
[0026] In the above solution, the first device does not transmit weight data with a change degree lower than the second threshold and gradient data with a probability of being in a stable range greater than the third threshold to the second device, which can reduce resource occupation and improve resource utilization.
[0027] In a possible implementation method, the method also includes: receiving compression parameters from the first device; determining the AI model based on the compressed data includes: determining the AI model based on the compressed data and the compression parameters; wherein the compression parameters are used to indicate one or more of the following information: model parameters of the AI model, the compression granularity of each data block in the N data blocks, or the compression scheme corresponding to each data block in the N data blocks; the model parameters include at least one of the following: the type of the AI model, the number of layers of the AI model, the number of nodes in each layer of the AI model, or the function of each layer in the AI model.
[0028] The above solution helps the second device to accurately and quickly determine the AI model by receiving compression parameters from the first device.
[0029] In one possible implementation method, the compression granularity of each of the N data blocks is a partial node in a layer of the AI model, a layer, or multiple layers.
[0030] The above solution can flexibly map the data to be transmitted in the AI model into M data blocks, which helps improve transmission efficiency.
[0031] In one possible implementation method, the M data blocks include at least one third data block, the third data block includes one or more sub-data blocks, and the difference between distribution information of any two sub-data blocks among the multiple sub-data blocks is less than a fourth threshold.
[0032] The above solution can accurately determine the data blocks, help improve the compression effect of the AI model, and thus save the transmission resources used when transmitting the compressed data blocks.
[0033] In a third aspect, embodiments of the present application provide a communication device. The device has the function of implementing any of the implementation methods of aspects 1 to 2 above. The function can be implemented by hardware or by hardware executing corresponding software implementations. The hardware or software includes one or more modules corresponding to the above functions.
[0034] In a fourth aspect, an embodiment of the present application provides a communication device, comprising a unit or means for executing each step of any implementation method in the above-mentioned first to second aspects.
[0035] In a fifth aspect, an embodiment of the present application provides a communication device comprising at least one processor, wherein the processor is configured to execute any implementation method in the above-mentioned first to second aspects by at least one of the following: running computer instructions or programs, or logic circuits.
[0036] In one possible implementation, the processor is coupled to a memory that stores the aforementioned computer instructions or programs; alternatively, the memory stores a configuration file for a logic circuit. Alternatively, the memory may be located within the device, i.e., the device includes the memory; alternatively, the processor and memory are integrated together; alternatively, the memory may be located external to the device.
[0037] In a possible implementation, the communication device further includes: an interface circuit, the interface circuit being used to input and / or output signals. The processor is used to communicate with other devices through the interface circuit.
[0038] In a possible implementation, the communication device is a chip.
[0039] In a sixth aspect, an embodiment of the present application further provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are run by a communication device, any implementation method in the above-mentioned first to second aspects is executed.
[0040] In the seventh aspect, an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when run on a communication device, enables any implementation method in the above-mentioned first to second aspects to be executed.
[0041] In an eighth aspect, an embodiment of the present application further provides a chip system, comprising: a processor for executing any implementation method of the above-mentioned first to second aspects.
[0042] In a ninth aspect, an embodiment of the present application further provides a communication system, comprising a first device for executing any implementation method of the first aspect, and a second device for executing any implementation method of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] FIG1( a ) is a schematic diagram of a possible, non-limiting system applicable to the embodiments of the application;
[0044] FIG1( b ) is a schematic diagram of the architecture of a communication system used in an embodiment of the present application;
[0045] FIG2 is a flow chart of a data transmission method provided in an embodiment of the present application;
[0046] FIG3( a ) is an example diagram of the association between data blocks and compressed blocks provided in an embodiment of the present application;
[0047] FIG3( b ) is an example diagram of a quantization method provided in an embodiment of the present application;
[0048] FIG4 is an example diagram of an AI model provided in an embodiment of the present application;
[0049] FIG5 is an example diagram of a method for determining a stable layer according to an embodiment of the present application;
[0050] FIG6 is an example diagram of transmitted data provided in an embodiment of the present application;
[0051] FIG7 is a schematic structural diagram of a communication device provided in an embodiment of the present application;
[0052] FIG8 is a schematic diagram of the structure of a communication device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] [Corrected 08.05.2025 according to Rule 91] To facilitate understanding of the embodiments of the present application, the application scenarios used in the present application are described using the communication system architecture shown in Figure 1(a) as an example. Figure 1(a) is a possible, non-limiting system schematic diagram. As shown in Figure 1(a), the communication system 1000 includes a radio access network (RAN) 100 and a core network (CN) 200. The RAN 100 includes at least one network device (such as 110a and 110b in Figure 1(a), collectively referred to as 110) and at least one terminal device (such as 120a-120j in Figure 1(a), collectively referred to as 120). The RAN 100 may also include other RAN nodes, such as wireless relay devices and / or wireless backhaul devices (not shown in Figure 1(a)). The terminal device 120 is connected to the network device 110 via a wireless connection. The network device 110 is connected to the core network 200 via a wireless or wired connection. The core network device in the core network 200 and the network device 110 in the RAN 100 may be different physical devices, or may be the same physical device that integrates core network logical functions and radio access network logical functions.
[0054] The RAN 100 may be a cellular system related to the Third Generation Partnership Project (3GPP), such as a fourth generation (4G) or fifth generation (5G) mobile communication system, or an evolved system after 5G (such as a sixth generation (6G) mobile communication system). The RAN 100 may also be an open access network (O-RAN or ORAN), a cloud radio access network (CRAN), or a wireless fidelity (WiFi) system. The RAN 100 may also be a communication system that integrates two or more of the above systems.
[0055] [Corrected 08.05.2025 according to Rule 91] The apparatus provided in the embodiments of the present application can be applied to the network device 110 or to the terminal device 120. It will be understood that FIG1(a) only illustrates one possible communication system architecture to which the embodiments of the present application can be applied. In other possible scenarios, the communication system architecture may also include other devices.
[0056] Figure 1(b) is a schematic diagram of the architecture of a communication system used in an embodiment of the present application. The communication system includes a first device and a second device.
[0057] In one implementation method, the first device is a network device or a module for a network device, and the second device is a terminal device or a module for a terminal device. The first device and the second device communicate with each other via an air interface.
[0058] In another implementation method, the first device is a terminal device or a module for a terminal device, and the second device is a network device or a module for a network device. The first device and the second device communicate with each other via an air interface.
[0059] In another implementation method, the first device is a network device or a module for a network device, and the second device is a network device or a module for a network device. The first device and the second device communicate with each other via an air interface or a wired manner.
[0060] In another implementation method, the first device is a terminal device or a module for a terminal device, and the second device is a terminal device or a module for a terminal device. The first device and the second device communicate with each other via an air interface.
[0061] Of course, the first device and the second device in the embodiment of the present application can also be other types of devices. For example, the first device can also be a cloud device or a cloud server and the second device can be a cloud device or a cloud server. This application does not limit this.
[0062] In the implementation of this application, a terminal device is a device with wireless transceiver capabilities, and may specifically refer to user equipment (UE), access terminal, subscriber unit, user station, mobile station, remote station, remote terminal, mobile device, user terminal, wireless communication device, user agent, or user device. The terminal device can be deployed on land, including indoors or outdoors, handheld or vehicle-mounted; can also be deployed on water (such as ships); and can also be deployed in the air (such as aircraft, balloons, and satellites). The terminal device can be a cellular phone, a mobile phone, a tablet computer (pad), a wireless data card, a wireless modem, a satellite terminal, a vehicle (e.g., a car, a bicycle, an electric car, an airplane, a ship, a train, a high-speed rail, etc.) onboard equipment, a robotic arm, a workshop equipment, a wearable device (e.g., a smart watch, a smart bracelet, a pedometer, etc.), a drone, a robot, a smart point of sale (POS) machine, a customer-premises equipment (CPE), a computer with a wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a terminal device in industrial control, a terminal device in self-driving, a terminal device in remote medical care, a terminal device in a smart grid, a terminal in transportation safety, a terminal device in a smart city, a terminal in a smart home (e.g., a refrigerator, a television, an air conditioner, an electric meter, and other smart home devices). The terminal device can also be other devices with terminal functions. The embodiments of this application do not limit the device form factor of the terminal. The device used to implement the functions of the terminal device can be the terminal device; it can also be a device that supports the terminal device to implement the functions, such as a chip system. The device can be installed in the terminal device or used in conjunction with the terminal device. In the embodiments of this application, the chip system can be composed of chips or include chips and other discrete devices.
[0063] In the implementation of this application, the network device is a device with wireless transceiver functions, which is used to communicate with the terminal device or other network devices; it can also be a device that can access the terminal device to the wireless network, such as a radio access network (RAN) device or node. The network devices in the embodiments of the present application may include various forms of base stations, such as: base stations, evolved NodeBs (eNodeBs), next generation NodeBs (gNBs), macro base stations, micro base stations (also known as small stations), relay stations, access points, devices that implement base station functions in communication systems evolved after the fifth generation (5G) technology, access points (APs) in wireless local area networks (WLAN) systems, integrated access and backhaul (IAB) nodes, transmission points (TRPs), transmitting points (TPs), mobile switching centers, and devices that perform base station functions in device-to-device (D2D), vehicle-to-everything (V2X), and machine-to-machine (M2M) communications, etc., and may also include network devices in non-terrestrial network (NTN) communication systems, that is, they can be deployed on high-altitude platforms or satellites. In some possible scenarios, different network devices implement part of the functions of the base station respectively. For example, the network device can be a centralized unit (CU), a distributed unit (DU), a CU-control plane (CP), a CU-user plane (UP), or a radio unit (RU). The CU and DU can be set separately, or they can be included in the same network element, such as a baseband unit (BBU). The RU can be included in a radio frequency device or radio frequency unit, such as a remote radio unit (RRU), an active antenna unit (AAU), or a remote radio head (RRH).It is understood that the network device may be a CU node, a DU node, or a device including a CU node and a DU node. In addition, the CU may be classified as a network device in the access network RAN, or may be classified as a network device in the core network CN, without limitation herein.
[0064] In different systems, CU (or CU-CP and CU-UP), DU or RU may also have different names, but those skilled in the art can understand their meanings. For example, in an open RAN (open RAN, ORAN) system, CU may also be called O-CU (open CU), DU may also be called O-DU, CU-CP may also be called O-CU-CP, CU-UP may also be called O-CU-UP, and RU may also be called O-RU. For the convenience of description, this application uses CU, CU-CP, CU-UP, DU and RU as examples for description. Any unit of CU (or CU-CP, CU-UP), DU and RU in this application can be implemented by a software module, a hardware module, or a combination of a software module and a hardware module.
[0065] In the embodiments of the present application, the form of the network device is not limited. The device used to implement the function of the network device can be a network device; it can also be a device that can support the network device to implement the function, such as a chip system. The device can be installed in the network device or used in conjunction with the network device.
[0066] Throughout the evolution of communication systems, high throughput and a large number of connections have always been core challenges for wireless communication networks. To address these challenges, 5G communications have proposed applications such as enhanced mobile broadband (eMBB), ultra-reliable and low-latency communication (URLLC), and massive machine-type communication (mMTC) as technical goals. The 6G communication system, which will evolve after 5G, will inevitably evolve towards higher throughput, lower latency, higher reliability, a larger number of connections, and greater spectrum utilization.
[0067] With the continuous development of the three major driving forces of AI / machine learning (ML), namely computing power, algorithms, and data-related technologies, AI / ML has important application potential in many areas, including modeling and learning in complex unknown environments, channel prediction, intelligent signal generation and processing, network status tracking and intelligent scheduling, and network optimization and deployment.
[0068] The implementation of AI technology may rely on the interaction between multiple devices. For larger-scale AI models, when the computing power and storage capacity vary greatly between devices, the training and reasoning of the AI model may be located in different devices. For example, compared with terminal devices, network devices with stronger computing and storage capabilities can be used to train the AI model. After the model training is completed, the network device can distribute the AI model to the terminal device, and the terminal device uses the received AI model to implement the reasoning of the AI model. In the reasoning stage, when the collection of sample data is related to changes in the environment, when the external environment changes, the collected sample data also changes accordingly, which may cause the current sample data to not match the previously received AI model. At this time, the network device needs to send the updated AI model to the terminal device.
[0069] The AI model training process can also be completed by multiple devices working together. For example, in distributed learning, multiple terminal devices independently train the AI model using local sample data and send the model weights or gradients during training to the network device. The network device aggregates the weights or gradients received from multiple models and sends the aggregated results to multiple terminal devices. The above process is repeated until the model converges.
[0070] Depending on the application scenario of AI technology, the scale of AI models can be extremely large, with data volumes reaching millions or even tens of millions. Therefore, when AI model data needs to be exchanged between different devices, it incurs significant transmission resource overhead. This is especially true in scenarios like distributed learning, where multiple devices participate. Therefore, with increasingly scarce transmission resources, how to transmit data for ultra-large-scale AI models has become a critical issue that needs to be addressed in the AI field.
[0071] FIG2 is a flow chart of a data transmission method provided in an embodiment of the present application. The method includes the following steps:
[0072] In step 201 , the first device determines compressed data based on N data blocks corresponding to the AI model.
[0073] An AI model includes one or more layers, and the data block of the AI model can be divided into N data blocks, where N is an integer greater than 1. A data block can correspond to some nodes in a layer of the AI model, that is, some nodes in a layer are divided into one data block. Alternatively, a data block can correspond to a layer in the AI model, that is, all nodes in a layer are divided into one data block. Alternatively, a data block can correspond to multiple layers in the AI model, that is, all nodes in multiple layers are divided into one data block. Alternatively, a data block can correspond to some nodes in multiple layers in the AI model, that is, some nodes in multiple layers are divided into one data block.
[0074] The compression scheme corresponding to each of the N data blocks is related to the distribution characteristics of the data block. The determined compressed data includes M compressed blocks, and the M compressed blocks correspond one-to-one to M data blocks in the N data blocks, where M is an integer not greater than N. The compression scheme can also be understood as a compression algorithm.
[0075] In step 202, the first device sends compressed data to the second device, and the second device receives the compressed data accordingly.
[0076] In one implementation method, N=M, which means that after the data of the AI model is divided into N data blocks, each of the N data blocks is compressed into a compressed block to obtain M compressed blocks, and then the M compressed blocks are sent to the second device.
[0077] In another implementation method, N>M>0, which means that after the data of the AI model is divided into N data blocks, M data blocks in the N data blocks are respectively compressed into a compressed block to obtain M compressed blocks, and then the M compressed blocks are sent to the second device. The other NM data blocks in the N data blocks are not transmitted to the second device after compression, nor are they directly transmitted to the second device without compression. This method is applicable to part of the data in the trained AI model that has little or no change compared to the corresponding part of the data in the historical AI model. The first device may not directly send the part of the data with little or no change to the second device, nor may it compress the part of the data with little or no change and send it to the second device. The second device can directly use the data corresponding to the historical AI model.
[0078] In another implementation, M = 0, indicating that none of the AI model data is transmitted to the second device before compression, nor is it transmitted directly to the second device without compression. This method is suitable for situations where the trained AI model has little or no change compared to the historical AI model. The first device does not need to send the trained AI model data directly to the second device or send it to the second device after compression; the second device can directly use the historical AI model.
[0079] For the first two implementation methods described above, a data block is compressed into a compressed block, and the compression ratio used can be equal to 1 or a value greater than 0 and less than 1. When the compression ratio corresponding to a data block is equal to 1, it means that the data block is not compressed, and therefore the compressed block corresponding to the data block is the data block itself, or it can be understood that the compressed block and the data block contain the same data. When the compression ratio corresponding to a data block is greater than 0 and less than 1, the data block is compressed to obtain a compressed block, which contains less data than the data block corresponding to the compressed block.
[0080] In one implementation method, the present application divides the data blocks corresponding to the AI model into at least three types of data blocks, which are respectively referred to as first data blocks, second data blocks, and third data blocks. The N data blocks corresponding to the above-mentioned AI model of the present application may include only one type of data block, two types of data blocks, or three types of data blocks.
[0081] Among them, the first data block refers to a data block whose importance is greater than the first threshold. The importance can be called importance. The first data block is also called an important data block. For the first data block, the first device does not perform compression processing, or it is understood that the compression rate is 1, so the compressed block corresponding to each first data block is the first data block itself. For example, when the above-mentioned N data blocks include N1 first data blocks, N1 is a positive integer, then the M compressed blocks determined by the first device include at least N1 compressed blocks, and the N1 compressed blocks are the N1 first data blocks. Or it is understood that the M compressed blocks determined by the first device include at least the N1 first data blocks.
[0082] In one implementation method, the importance of the first data block can be determined according to the following method: During the training process of the AI model, in the early stages of training, the parameters of the entire AI model have not yet reached a stable state. At this time, the data closer to the input layer is more important than the data closer to the output layer. As the number of training iterations increases and the AI model approaches convergence, the data closer to the output layer is more important than the data closer to the input layer. Taking Figure 4 as an example, when the training is close to convergence, the data of the fully connected layer, which is the last layer, is more important. Specifically, the training stages can be divided according to the accuracy of the AI model. For example, the accuracy less than the first accuracy threshold is the early stage of training, the accuracy greater than the second accuracy threshold is the late stage of training, and the accuracy between the first and second accuracy thresholds is the mid-stage of training. Assume that importance or importance is represented by a number between 0 and 1, with 1 indicating the highest importance and 0 indicating the lowest importance. Sort the N data blocks in the order of forward propagation. For the early stage of training, the importance of the N data blocks is N / N, (N-1) / N, (N-2) / N, ..., 1 / N. For the late stage of training, the importance of the N data blocks is 1 / N, 2 / N, ..., (N-1) / N, N / N. For the middle stage of training, the importance of the N data blocks is x / N, where the first accuracy threshold, the second accuracy threshold, and x are pre-configured parameters. According to the above definition of importance, first determine the current training stage based on the accuracy of the AI model, that is, the early, middle or late stage of training. Then, determine the importance of the N data blocks according to the importance determination method corresponding to the training stage, where the data block whose importance is greater than the first threshold is the aforementioned first data block. The first threshold is also a pre-configured parameter and is a number greater than 0 and less than 1.
[0083] Furthermore, the importance of data can be determined based on the function of each layer in the AI model. For example, the pooling layer performs dimensionality reduction on the data. To ensure that the dimensionality reduction operation does not cause loss of information content, the size of the pooling layer's receptive domain is sized to the data with high importance. Specifically, the importance of the data in the pooling layer can be set to 1, ensuring that the importance of the pooling layer is greater than a first threshold.
[0084] The second data block refers to a data block with less variation. The second data block is also called a stable data block. For example, when the second data block is weight data, the degree of variation of the second data block is lower than the second threshold. For another example, when the second data block is gradient data, the probability that the second data block is in a stable interval is greater than the third threshold. For the second data block, the first device does not transmit the second data block to the second device, nor does it compress the second data block and transmit it to the second device. Or it can be understood that the second data block does not correspond to the M compressed blocks transmitted above. For example, when the above-mentioned N data blocks include N2 second data blocks, and N2 is a positive integer, the M compressed blocks determined by the first device do not include the N2 second data blocks, nor do they include the compressed data corresponding to the N2 second data blocks. For the second data block, it can also be understood that the compression rate corresponding to the second data block is equal to 0.
[0085] The third data block is a data block other than the first data block and the second data block. The third data block is a data block that needs to be truly compressed, wherein truly compressed means that the compression ratio of the third data block is greater than 0 and less than 1. The third data block is also called a compressed data block. With respect to the third data block, the first device needs to compress the third data block and transmit it to the second device, and the compression ratio is greater than 0 and less than 1. For example, when the N data blocks include N3 third data blocks, and N3 is a positive integer, the M compressed blocks determined by the first device include at least N3 compressed blocks, and the N3 compressed blocks correspond one-to-one to the N3 third data blocks, and the data of each compressed block is less than the third data block corresponding to the compressed block.
[0086] Wherein, each third data block includes one or more sub-data blocks. When a third data block includes multiple sub-data blocks, the difference between the distribution information of any two sub-data blocks in the multiple sub-data blocks is less than the fourth threshold value, and each sub-data block corresponds to a partial node of a layer, a partial node of multiple layers, a layer or multiple layers of the AI model. Alternatively, when N data blocks include multiple third sub-data blocks, the difference between the distribution information of any two third data blocks in the multiple third data blocks is greater than the fifth threshold value. In one implementation method, when comparing the difference between the distribution information of two sub-data blocks, or comparing the difference between the distribution information between two third data blocks, the probability density function (PDF) or cumulative distribution function (CDF) of the distribution information of the sub-data blocks and the third data blocks can be determined first, and then the difference between the distribution information of the two sub-data blocks or the distribution information between the two third data blocks can be compared based on the PDF or CDF.
[0087] FIG3(a) is an example diagram of the association between data blocks and compressed blocks provided in an embodiment of the present application. The N data blocks include N1 first data blocks, N2 second data blocks, and N3 third data blocks, where N1 + N3 equals M. The N3 third data blocks are compressed into N3 compressed blocks, and the N1 first data blocks are not compressed and are directly used as N1 compressed blocks, resulting in M compressed blocks. The N2 second data blocks are not compressed and are not directly used as compressed blocks, so the M compressed blocks do not include the second data blocks, nor do they include the compressed second data blocks.
[0088] Step 203: The first device sends a compression parameter to the second device, and the second device receives the compression parameter accordingly.
[0089] Step 203 is optional. When executing step 203, the order of step 203 and step 202 is not limited. Alternatively, step 203 and step 202 can be combined into one step, for example, the first device sends a first message to the second device, the first message including compression parameters and compressed data.
[0090] In one implementation method, when compression parameters and compressed data are transmitted simultaneously, they can be transmitted on the same block of time-frequency resources, where the position of the resources carrying the compression parameters in the entire block of transmission resources can be determined by a predefined method or a direct indication method, and the remaining resources can be used to transmit compressed data.
[0091] In another implementation method, when the compression parameter and the compressed data are transmitted on different resources, the compression parameter can be carried alone in a control message for transmission.
[0092] The compression parameter is used to indicate one or more of the following: a model parameter of the AI model, a compression granularity corresponding to each of the N data blocks, or a compression scheme corresponding to each of the N data blocks.
[0093] The model parameters of the AI model include at least one of the following:
[0094] 1) Types of AI Models
[0095] Types of AI models include but are not limited to: deep neural network (dense neutral network, DNN), convolutional neural network (convolution neutral network, CNN), regression neural network (regression neutral network, RNN), and generation against network (generation against network, GAN).
[0096] 2) Number of layers in the AI model
[0097] 3) Number of nodes per layer in the AI model
[0098] 4) Function of each layer in the AI model
[0099] The functions of the AI model layers include but are not limited to: convolution layer, pooling layer, full connection layer (FC), impulse function layer, batch normalization layer (BN).
[0100] The compression granularity of each of the N data blocks may be a partial node in a layer, or a layer, or multiple layers.
[0101] The compression scheme may include one or more of the following: sparse algorithm, quantization, matrix decomposition, full compression, and no compression.
[0102] In one implementation method, a mapping relationship between the compression scheme index and the compression scheme and the compression scheme parameters can be pre-configured on the first device and the second device, or the mapping relationship can be pre-defined through a protocol. After determining the compression scheme and compression scheme parameters corresponding to each data block, the first device determines the index of the compression scheme based on the mapping relationship and carries the index of the compression scheme in the compression parameters. The second device obtains the index of the compression scheme from the compression parameters and determines the corresponding compression scheme and compression scheme parameters from the mapping relationship based on the index. This method can improve indication efficiency and reduce resource overhead by indicating the compression scheme through an index.
[0103] Table 1-1 exemplarily shows the mapping relationship between the compression scheme index and the compression scheme and the parameters of the compression scheme.
[0104] Table 1-1
[0105] Among them, the sparse algorithms include top sparsification and contour sparsification. Top sparsification sorts the original data by absolute value, ultimately transmitting only the top r% of the data with the largest absolute value. The parameter to be specified is the compression rate r%. Contour sparsification uses a transformation matrix or sequence to transform the original data into compressed data. The parameter to be specified is the transformation matrix or sequence.
[0106] Quantization is the process of converting original data or data compressed by other algorithms into bits. The parameters that need to be indicated are the number of quantization bits, that is, the number of bits into which each original data is converted, and the mapping relationship of the quantization bits. Figure 3(b) is an example diagram of the quantization method. In this example, the data in the range of {-1, 1} is mapped to 0, and the corresponding quantization bit is 00; the data in the range of {1, 3} is mapped to 2, and the corresponding quantization bit is 01; the data in the range of {3, 5} is mapped to 4, and the corresponding quantization bit is 10; the data in the range of {5, 7} is mapped to 6, and the corresponding quantization bit is 11. Table 2 is an example of the mapping relationship. The second device can determine the data corresponding to the received bit based on the mapping relationship. If the received bit is 00, the corresponding data is determined to be 0; if the received bit is 01, the corresponding data is determined to be 2; if the received bit is 10, the corresponding data is determined to be 4; if the received bit is 11, the corresponding data is determined to be 6.
[0107] Table 2
[0108] Index 9 in Table 1-1 corresponds to a complete compression scheme, that is, a compression rate of 0, compressing the data to the minimum (that is, there is no data after compression), corresponding to the aforementioned second data block, that is, the compression rate of the second data block is 0, so no content of the second data block is actually transmitted.
[0109] Index 10 in Table 1-1 corresponds to the no-compression scheme, that is, the compression rate is 1, and the data is not compressed. The transmitted compressed block pair is the data block itself, or it can be understood that the compressed block and the data block contain the same data, corresponding to the aforementioned first data block.
[0110] The above Table 1-1 can also be split into two tables and stored on the first device and the second device. For example, Table 1-1 can be split into the following Table 1-2 and Table 1-3.
[0111] Table 1-2
[0112] Table 1-3
[0113] The above Tables 1-3 may also be split into multiple tables. For example, each compression scheme and the parameters of the compression scheme corresponding to the compression scheme may constitute a separate table.
[0114] In another implementation method, a mapping relationship between the compression scheme index and the compression scheme parameters can be preconfigured on the first device and the second device or predefined through a protocol. After determining the compression scheme and compression scheme parameters corresponding to each data block, the first device determines the compression scheme index based on the mapping relationship and carries the compression scheme index in the compression parameters. The second device obtains the compression scheme index from the compression parameters and determines the corresponding compression scheme parameters from the mapping relationship based on the index. The compression scheme parameters correspond to a compression scheme. This method can improve indication efficiency and reduce resource overhead by indicating the compression scheme through an index.
[0115] Tables 1-4 exemplarily provide the mapping relationship between the index of the compression scheme and the parameters of the compression scheme.
[0116] Table 1-4
[0117] In another implementation method, a first mapping relationship between the index of the compression scheme and the compression scheme, and a second mapping relationship between the index of the compression scheme and the parameters of the compression scheme can also be pre-configured on the first device and the second device or pre-defined through a protocol. After determining the compression scheme and the parameters of the compression scheme corresponding to each data block, the first device determines the index of the compression scheme based on the mapping relationship and carries the index of the compression scheme in the compression parameters. The second device obtains the index of the compression scheme from the compression parameters and determines the corresponding compression scheme from the first mapping relationship based on the index, and determines the corresponding compression parameters from the second mapping relationship. This method can improve indication efficiency and reduce resource overhead by indicating the compression scheme and the parameters of the compression scheme through indexes.
[0118] Table 1-5 exemplarily shows a first mapping relationship between the index of the compression scheme and the compression scheme. Table 1-6 exemplarily shows a second mapping relationship between the index of the compression scheme and the parameters of the compression scheme.
[0119] Table 1-5
[0120] Table 1-6
[0121] In step 204 , the second device determines an AI model based on the compressed data.
[0122] Wherein, when the above step 203 is executed, the step 204 is: the second device determines the AI model according to the compressed data and the compression parameters.
[0123] The above scheme determines the compression scheme corresponding to each data block based on the distribution characteristics of the N data blocks corresponding to the AI model. It can select the most matching compression scheme for different data blocks instead of selecting the same compression scheme for all data blocks. This helps to improve the compression effect of the AI model and save the transmission resources used when transmitting the compressed data blocks.
[0124] The following is an explanation with reference to specific examples. Figure 4 is an example diagram of an AI model provided in an embodiment of the present application. It should be noted that the AI model of the present application is not limited to the specific AI model shown in Figure 4. Referring to Figure 4, the AI model includes 13 layers, which are as follows:
[0125] Layer 1: Convolutional layer, the convolution kernel size is 3*3, and the number of convolution kernels is 64;
[0126] Layer 2: Convolutional layer, the convolution kernel size is 3*3, and the number of convolution kernels is 64;
[0127] Layer 3: The size of the reception field of the pooling layer is 2*2;
[0128] Layer 4: Convolution layer, the convolution kernel size is 3*3, and the number of convolution kernels is 128;
[0129] Layer 5: Convolution layer, the convolution kernel size is 3*3, and the number of convolution kernels is 128;
[0130] Layer 6: The size of the receptive field of the pooling layer is 2*2;
[0131] Layer 7: Convolution layer, the convolution kernel size is 3*3, and the number of convolution kernels is 256;
[0132] Layer 8: Convolution layer, the convolution kernel size is 3*3, and the number of convolution kernels is 256;
[0133] Layer 9: The size of the receptive field of the pooling layer is 2*2;
[0134] Layer 10: Convolution layer, the convolution kernel size is 3*3, and the number of convolution kernels is 512;
[0135] Layer 11: Convolutional layer, the convolution kernel size is 3*3, and the number of convolution kernels is 512;
[0136] Layer 12: The size of the receptive field of the pooling layer is 2*2;
[0137] Layer 13: The output dimension of the fully connected layer is 2*2.
[0138] In one implementation method, this application divides the multiple layers of the AI model into different types of layers. Specifically, the multiple layers of the AI model are divided into fixed layers, stable layers, and compressed layers. These are described below.
[0139] 1. Fixed layer
[0140] The fixed layer consists of one or more highly important layers of the AI model. To ensure that the highly important data does not lose information due to compression, this application does not compress the data in the fixed layer and transmits it directly. The compression rate of the fixed layer data can also be considered to be 1. This method can ensure that the highly important data does not lose information due to compression.
[0141] The data of the fixed layer can be divided into one or more data blocks, which are the aforementioned first data blocks. Taking Figure 4 as an example, the 3rd, 6th, 9th, 12th, and 13th layers of the AI model are determined as fixed layers, and the data of the fixed layer is divided into one data block as a whole, namely data block 1, which is a specific example of the aforementioned first data block. Of course, the fixed layer can also be divided into multiple first data blocks in other division methods, which is not limited in this application.
[0142] 2. Stability layer
[0143] During the training of the AI model, the model's weights and weight gradients will be continuously iterated and updated. When the data of certain layers changes relatively slowly, they can be identified as stable layers. For example, when the first device interacts with the weights of the AI model, if the data gap between a certain layer's weights before and after the update is small, there is no need to interact with the weights of that layer, that is, the compression rate is considered to be zero, and the second device can continue to use the weights of that layer received previously. For another example, when interacting with the gradient of the AI model, if the data of a certain layer's gradient is very small in a certain iteration, it can be considered that the weight corresponding to the gradient changes very slowly, and there is no need to interact with the gradient information of that layer, that is, the compression rate is considered to be zero, and the second device can set the gradient corresponding to that layer to zero. This method does not transmit the data of the stable layer, which can reduce resource usage and improve resource utilization.
[0144] Figure 5 is an example diagram of a method for determining a stable layer provided by an embodiment of the present application. Taking the AI model data as gradient data as an example, when the probability that the data of a certain layer is within the stable interval is greater than a predefined stability threshold, the layer is determined to be a stable layer. In Figure 5, the stable interval is [-0.0025, 0.0025]. This method can accurately determine the stable layer, thereby improving resource utilization.
[0145] The data of the stable layer can be divided into one or more data blocks, which are the aforementioned second data blocks. Taking Figure 4 as an example, the 10th and 11th layers of the AI model are determined to be the stable layer, and the data of the stable layer is divided into one data block as a whole, namely data block 2, which is a specific example of the aforementioned second data block. Of course, the stable layer can also be divided into multiple second data blocks in other division methods, which is not limited in this application.
[0146] 3. Compression layer
[0147] The compression layer is a layer other than the fixed layer and the stable layer. The data in the compression layer has a compression ratio greater than 0 and less than 1. The data in the compression layer can be divided into one or more data blocks, which are the aforementioned third data blocks. Each third data block is compressed using an independent compression scheme. The compression granularity of each third data block can be any of the following:
[0148] 1) Layer: Each compression layer is treated as a compression block (i.e., a third data block). Each compression layer is treated as a compression block (i.e., a third data block) and uses an independent compression scheme.
[0149] 2) Layer Group: In AI models, multiple compression layers with the same structure often appear consecutively. Layers with the same structure can be grouped together into a layer group. Each layer group is treated as a compression block (i.e., a third data block) and uses an independent compression scheme. Layer groups can be formed in the following ways:
[0150] Method 1: Adjacent layers with the same structure form a layer group. When the structure of the AI model is determined, the layer group is also determined. This method does not require indicating the layer index.
[0151] Method 2: Adjacent layers with the same structure can be further divided into multiple layer groups through the split layer index. When identifying layer groups, the first-level layer group can be obtained through the adjacent layers with the same structure, and then the final layer group can be obtained according to the split layer index. This method requires the indication of the split layer index.
[0152] For example, the 5th to 10th layers of an AI model have the same structure. If it is determined that the 5th to 6th layers are divided into one layer group, the 7th to 8th layers are divided into one layer group, and the 9th to 10th layers are divided into one layer group. Assuming that the index of the last layer of each layer group is used as the split layer index, the split layer indexes of the three layer groups are 6 and 8. Therefore, the first-level layer group includes layers 5-10, and the split layer indexes are 6 and 8, indicating that layers 5 and 6 of layers 5-10 are divided into one layer group, layers 7 and 8 of layers 5-10 are divided into one layer group, and layers 9 and 10 of layers 5-10 are divided into one layer group.
[0153] Method three: flexible layer group division, directly indicating the layer index included in each layer group to the second device, and optionally indicating the number of layer groups.
[0154] 3) Data Groups: When the number of nodes in a layer exceeds a preset node number threshold, the data in that layer can be divided into multiple data groups. Alternatively, when the number of nodes in a layer is less than a preset node number threshold, the data from multiple layers can be combined into a single data group. Each data group is treated as a compressed block (i.e., the third data block) and uses an independent compression scheme.
[0155] In one implementation method, the compression ratio corresponding to the third data block can be determined by the following method: first determine the initial compression ratio: δ = L s / N, where L s is the amount of data outside the stable interval in the third data block, N is the total amount of data in the third data block, and L s The unit of N can be bits. The size of the stable interval here can be the same as the size of the stable interval corresponding to the second data block, or it can be different, which is not limited here. Then, based on the modulation and coding scheme (MCS) and the resources used to transmit the third data block, the maximum number of bits L allowed to be carried by the third data block is determined. MCS , if L MCS <L s , then reduce δ until L is satisfied MCS >L s .
[0156] For example, the compression rate can be further updated as the training progresses. The updated criteria can be based on the evaluation of the AI model performance, such as the degree of change in accuracy and loss over a period of time. The specific criteria can be A loss >L loss and / or A accuracy <L accuracy , where A accuracy and A loss is the accuracy and loss of the AI model in the current training phase, L accuracy and L loss The accuracy and loss of the AI model during the previous training phase. If the AI model's accuracy decreases or its loss increases, it indicates that the current compression ratio does not match the existing AI model and requires an update. This method can improve the performance of the AI model by appropriately increasing the compression ratio. This method improves the efficiency of the AI module and the spectral efficiency of transmission resources by timely adjusting the compression ratio.
[0157] For example, L lossIt is the loss of the AI model in the previous round of training, or the loss of the AI model in the previous N rounds of training (N is greater than 1), or the average loss of the AI model in the most recent M rounds of training (M is greater than 1), or the loss of the AI model in the training stage at a previous point in time, or the average loss of the AI model in the training stage over a previous period of time. This application is not limited to this.
[0158] For example, L accuracy It may be the accuracy of the AI model in the previous round of training, or the accuracy of the AI model in the previous K rounds of training (K is greater than 1), or the average accuracy of the AI model in the most recent L rounds of training (L is greater than 1), or the accuracy of the AI model in the training stage at a previous point in time, or the average accuracy of the AI model in the training stage over a previous period of time. This application does not limit this.
[0159] Taking Figure 4 as an example, layers 1, 2, 4, 5, 7, and 8 of the AI model are determined to be compression layers. Since layers 1 and 2 have the same structure, they are determined to be one data block, namely, data block 3. Since layers 4 and 5 have the same structure, they are determined to be one data block, namely, data block 4. Since layers 7 and 8 have the same structure, they are determined to be one data block, namely, data block 5. Data blocks 3, 4, and 5 are all the third data blocks mentioned above.
[0160] In the example of Figure 4, the first device determines the five data blocks corresponding to the AI model, namely data blocks 1 to 5, where data block 1 corresponds to the data of the fixed layer, data block 2 corresponds to the data of the stable layer, and data blocks 3 to 5 correspond to the data of the compressed layer. The first device determines four compressed blocks based on the five data blocks, namely compressed blocks 1 to 4, where compressed block 1 is the same as data block 1, compressed block 2 is the data block compressed from data block 3, compressed block 3 is the data block compressed from data block 4, and compressed block 4 is the data block compressed from data block 5.
[0161] The first device then sends compression parameters and compressed data to the second device. Figure 6 is an example diagram of the transmitted data provided in an embodiment of the present application. The transmitted data includes compression parameters and compressed data, and the compressed data includes compressed blocks 1 to 4. The compression parameters include model parameters of the AI model, the index of the fixed layer of the AI model, the index of the stable layer of the AI model, the compression scheme corresponding to each data block, and other information.
[0162] After receiving the compression parameters and compressed data of the AI model, the second device first identifies the parameters of the AI model and determines the structure of the AI model based on the compression parameters, and then fills the AI model according to the remaining compression parameters to obtain the AI model shown in Figure 4.
[0163] The second device determines that the stable layers are layers 10 and 11 based on the index of the stable layer of the AI model in the compression parameter. If the received compressed data is gradient data, the model parameters of the stable layer are set to zero. If the received compressed data is weight data, the data of the corresponding stable layer received last time is continued to be used.
[0164] The second device extracts the fixed layer from the compressed data based on the index of the fixed layer of the AI model and the number of nodes in the fixed layer in the model parameters of the AI model. In Figure 4, the fixed layer contains a total of 5 layers of the AI model. Assuming that the number of nodes in each layer is 100, each data in each layer is transmitted using 32 bits, and the fixed layer data is located at the front end of the compressed data (compressed block 1 as shown in Figure 6), then starting from the starting position of the compressed data, 32*5*100 bits are extracted and filled into the corresponding position according to the index of the compressed layer.
[0165] The second device extracts the remaining compressed data according to the compression scheme. Assuming that the compression scheme adopted is top compression (Top K), that is, the original data is compressed according to its absolute value, and the original data is sorted according to its absolute value. In the end, only the first r% of the data with the largest absolute value needs to be transmitted, where r% is the compression rate. Assuming that the dimensions of the original data of compressed blocks 2 to 4 in Figure 6 are 1000, 2000, and 3000 respectively, and the compression rates are 1%, 2%, and 3% respectively. And the quantization bits of each data in each compressed block are 2, 2, and 4, then the second device needs to extract 1000*1%*2, 2000*2%*2, and 3000*3%*4 bits from the compressed data respectively and fill them into the corresponding positions.
[0166] Through the above method, the second device can decompress the compressed data according to the compression parameters to obtain the data of the AI model.
[0167] In the above example, the second device identifies the stable layer, the fixed layer, and the compressed layer through the index of the stable layer and the index of the fixed layer in the compression parameter. In another implementation method, the compression parameters may also carry the compression granularity corresponding to data blocks 1 to 5, and the compression scheme corresponding to data blocks 1 to 5, respectively. The compression granularity corresponding to each data block indicates the specific layer of the AI model corresponding to the data block, and then the second device determines the processing method for the data block according to the compression scheme corresponding to the data block. Taking data block 1 as an example, the compression granularity corresponding to data block 1 indicates that the layers of the AI model corresponding to data block 1 include layers 3, 6, 9, 12, and 13, and the compression rate indicated by the compression scheme is 1. Therefore, after the second device obtains the compressed block corresponding to the received data block 1, it does not need to perform a decompression operation, that is, the compressed block received is the data block 1 itself. The second device then fills data block 1 into layers 3, 6, 9, 12, and 13 of the AI model. Similarly, taking data block 2 as an example, the compression granularity corresponding to data block 2 indicates that the layers of the AI model corresponding to data block 2 include layer 10 and layer 11, and the compression rate indicated by the compression scheme is 0. If the received compressed data is gradient data, the model parameters of layer 10 and layer 11 are set to zero. If the received compressed data is weight data, the last received data continues to be used in layer 10 and layer 11.
[0168] It should be noted that this application is illustrated by using the above-mentioned data transmission method to transmit data of an AI model as an example. In actual applications, the above-mentioned data transmission method can also be used to transmit data in other scenarios. The other scenarios here can be, for example, any type of scenario, and this application does not limit it.
[0169] It is understandable that in order to implement the functions in the above embodiments, the first device or the second device includes hardware structures and / or software modules corresponding to the execution of each function. It should be readily apparent to those skilled in the art that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.
[0170] Figures 7 and 8 are schematic diagrams of possible communication devices provided in embodiments of the present application. These communication devices can be used to implement the functions of the first device or the second device in the above-mentioned method embodiments, thereby also achieving the beneficial effects of the above-mentioned method embodiments. In the embodiments of the present application, the communication device can be the first device or the second device shown in Figure 1(b).
[0171] The communication device 700 shown in Figure 7 includes a processing unit 710 and a transceiver unit 720. The communication device 700 is used to implement the functions of the first device or the second device in the above method embodiment.
[0172] When the communication device 700 is used to implement the function of the first device in the above method embodiment, the processing unit 710 is used to determine the compressed data based on N data blocks corresponding to the artificial intelligence AI model, the compression scheme corresponding to each of the N data blocks is related to the distribution characteristics of the data blocks, and the compressed data includes M compressed blocks, and the M compressed blocks correspond one-to-one to M data blocks in the N data blocks; wherein N is an integer greater than 1, and M is an integer not greater than N; the transceiver unit 720 is used to send the compressed data to the second device.
[0173] In a possible implementation method, the M compressed blocks include at least one first data block, the first data block is a data block among the N data blocks, and the importance of the first data block is greater than a first threshold.
[0174] In a possible implementation method, the N data blocks include at least one second data block, and the second data block does not correspond to the M compressed blocks.
[0175] In one possible implementation method, the second data block is weight data and the degree of change of the second data block is lower than a second threshold; or, the second data block is gradient data and the probability that the second data block is in a stable interval is greater than a third threshold.
[0176] In one possible implementation method, the transceiver unit 720 is further used to send compression parameters to the second device; wherein the compression parameters are used to indicate one or more of the following information: model parameters of the AI model, the compression granularity of each of the N data blocks, or the compression scheme corresponding to each of the N data blocks; the model parameters include at least one of the following: the type of the AI model, the number of layers of the AI model, the number of nodes in each layer of the AI model, or the function of each layer in the AI model.
[0177] In one possible implementation method, the compression granularity of each of the N data blocks is a partial node in a layer of the AI model, a layer, or multiple layers.
[0178] In one possible implementation method, the M data blocks include at least one third data block, the third data block includes one or more sub-data blocks, and the difference between distribution information of any two sub-data blocks among the multiple sub-data blocks is less than a fourth threshold.
[0179] When the communication device 700 is used to implement the function of the first device in the above method embodiment, the transceiver unit 720 is used to receive compressed data from the first device, and the compressed data includes M compressed blocks, and the M compressed blocks correspond one-to-one to M data blocks of N data blocks corresponding to the artificial intelligence AI model, and the compression scheme corresponding to each data block of the N data blocks is related to the distribution characteristics of the data blocks; wherein N is an integer greater than 1, and M is an integer not greater than N; the processing unit 710 is used to determine the AI model based on the compressed data.
[0180] In a possible implementation method, the M compressed blocks include at least one first data block, the first data block is a data block among the N data blocks, and the importance of the first data block is greater than a first threshold.
[0181] In a possible implementation method, the N data blocks include at least one second data block, and the second data block does not correspond to the M compressed blocks.
[0182] In one possible implementation method, the second data block is weight data and the degree of change of the second data block is lower than a second threshold; or, the second data block is gradient data and the probability that the second data block is in a stable interval is greater than a third threshold.
[0183] In one possible implementation method, the transceiver unit 720 is further used to receive compression parameters from the first device; determining the AI model based on the compressed data includes: determining the AI model based on the compressed data and the compression parameters; wherein the compression parameters are used to indicate one or more of the following information: model parameters of the AI model, the compression granularity of each data block in the N data blocks, or the compression scheme corresponding to each data block in the N data blocks; the model parameters include at least one of the following: the type of the AI model, the number of layers of the AI model, the number of nodes in each layer of the AI model, or the function of each layer in the AI model.
[0184] In one possible implementation method, the compression granularity of each of the N data blocks is a partial node in a layer of the AI model, a layer, or multiple layers.
[0185] In one possible implementation method, the M data blocks include at least one third data block, the third data block includes one or more sub-data blocks, and the difference between distribution information of any two sub-data blocks among the multiple sub-data blocks is less than a fourth threshold.
[0186] For a more detailed description of the processing unit 710 and the transceiver unit 720, reference can be made to the relevant description in the above method embodiment, which will not be repeated here.
[0187] The communication device 800 shown in Figure 8 includes a processor 810 and an interface circuit 820. The processor 810 and the interface circuit 820 are coupled to each other. It is understood that the interface circuit 820 can be a transceiver or an input / output interface. Optionally, the communication device 800 may also include a memory 830 for storing instructions executed by the processor 810, or storing input data required by the processor 810 to execute instructions, or storing data generated after the processor 810 executes instructions; or when the processor is a logic circuit (device), storing a configuration file of the logic circuit (device).
[0188] When the communication device 800 is used to implement the above method embodiment, the processor 810 is used to implement the functions of the above processing unit 710 , and the interface circuit 820 is used to implement the functions of the above processing unit 710 and the transceiver unit 720 .
[0189] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.
[0190] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, an erasable programmable read-only memory, an electrically erasable programmable read-only memory, a register, a hard disk, a mobile hard disk, a compact disc read-only memory (CD-ROM) or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in the first device or the second device. Of course, the processor and the storage medium can also be present in the first device or the second device as discrete components.
[0191] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. A computer program refers to a set of instructions that instruct an electronic computer or other device with message processing capabilities to perform each step of the action, usually written in a certain programming language and running on a certain target architecture. When the computer program or instruction is loaded and executed on a computer, the process or function described in the embodiment of the present application is executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a terminal device or other programmable device. The computer program or instruction can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instruction can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; an optical medium, such as a digital video disk; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.
[0192] In the various embodiments of the present application, unless otherwise specified or there is a logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced by each other. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0193] In this application, "at least one" means one or more, and "more" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. In the text description of this application, the character " / " generally indicates that the previous and next related objects are in an "or" relationship; in the formulas of this application, the character " / " indicates that the previous and next related objects are in a "division" relationship.
[0194] It is understood that the various numbers used in the embodiments of this application are merely for ease of description and are not intended to limit the scope of the embodiments of this application. The order of the sequence numbers of the above-mentioned processes does not necessarily imply a specific order of execution; the order of execution of the processes should be determined by their functions and inherent logic.
Claims
1. A data transmission method, characterized in that, Applied to a first device, the method includes: Determining compressed data according to N data blocks corresponding to an artificial intelligence (AI) model, where the compression scheme corresponding to each of the N data blocks is related to the distribution characteristics of the data block, the compressed data includes M compressed blocks, and the M compressed blocks correspond one-to-one to M of the N data blocks; where N is an integer greater than 1, and M is an integer not greater than N; Sending the compressed data to a second device.
2. The method according to claim 1, characterized in that, At least one first data block is included in the M compressed blocks, and the first data block is a data block among the N data blocks.
3. The method according to claim 2, characterized in that, The importance level of the first data block is greater than a first threshold.
4. The method according to any one of claims 1 to 3, characterized in that, At least one second data block is included in the N data blocks, and the second data block does not correspond to the M compressed blocks.
5. The method according to claim 4, characterized in that, The second data block is weight data and the degree of change of the second data block is lower than a second threshold; or, The second data block is gradient data and the probability that the second data block is in a stable interval is greater than a third threshold.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: sending compression parameters to the second device; where the compression parameters are used to indicate one or more of the following information: the model parameters of the AI model, the compression granularity of each of the N data blocks, or the compression scheme corresponding to each of the N data blocks; the model parameters include at least one of the following: the type of the AI model, the number of layers of the AI model, the number of nodes in each layer of the AI model, or the function of each layer in the AI model.
7. The method according to claim 6, characterized in that, The compression granularity of each of the N data blocks is part of the nodes, one layer, or multiple layers in a layer of the AI model.
8. The method according to any one of claims 1 to 7, characterized in that, At least one third data block is included in the M data blocks, and the third data block includes one or more sub-data blocks, and the difference degree between the distribution information of any two sub-data blocks among the multiple sub-data blocks is less than a fourth threshold.
9. A data transmission method, characterized in that, Applied to a second device, the method includes: Receiving compressed data from a first device, the compressed data includes M compressed blocks, the M compressed blocks correspond one-to-one to M of the N data blocks corresponding to an artificial intelligence (AI) model, and the compression scheme corresponding to each of the N data blocks is related to the distribution characteristics of the data block; where N is an integer greater than 1, and M is an integer not greater than N; Determining the AI model according to the compressed data.
10. The method according to claim 9, characterized in that, At least one first data block is included in the M compressed blocks, and the first data block is a data block among the N data blocks.
11. The method according to claim 10, characterized in that, The importance level of the first data block is greater than a first threshold.
12. The method according to any one of claims 9 to 11, characterized in that, At least one second data block is included in the N data blocks, and the second data block does not correspond to the M compressed blocks.
13. The method according to claim 12, characterized in that, The second data block is weight data and the degree of change of the second data block is lower than a second threshold; or, The second data block is gradient data and the probability that the second data block is in a stable interval is greater than a third threshold.
14. The method according to any one of claims 9 to 13, characterized in that, The method further includes: Receiving compression parameters from the first device; The determining the AI model according to the compressed data includes: Determine the AI model according to the compressed data and the compression parameters; Among them, the compression parameters are used to indicate one or more of the following information: the model parameters of the AI model, the compression granularity of each of the N data blocks, or the compression scheme corresponding to each of the N data blocks; the model parameters include at least one of the following: the type of the AI model, the number of layers of the AI model, the number of nodes in each layer of the AI model, or the function of each layer in the AI model.
15. The method according to claim 14, characterized in that, The compression granularity of each of the N data blocks is a partial node, one layer, or multiple layers in a layer of the AI model.
16. The method according to any one of claims 9 to 15, characterized in that, At least one third data block is included in the M data blocks, the third data block includes one or more sub-data blocks, and the difference degree between the distribution information of any two sub-data blocks among the multiple sub-data blocks is less than a fourth threshold.
17. A communication device, characterized in that, Including: A processing unit, configured to determine compressed data according to N data blocks corresponding to an artificial intelligence (AI) model, the compression scheme corresponding to each of the N data blocks is related to the distribution characteristics of the data blocks, and the compressed data includes M compressed blocks, and the M compressed blocks correspond one-to-one to M data blocks among the N data blocks; where N is an integer greater than 1, and M is an integer not greater than N; A transceiver unit, configured to send the compressed data to a second device.
18. The device according to claim 17, characterized in that, At least one first data block is included in the M compressed blocks, and the first data block is a data block among the N data blocks.
19. The device according to claim 18, characterized in that, The importance degree of the first data block is greater than a first threshold.
20. The device according to any one of claims 17 to 19, characterized in that, At least one second data block is included in the N data blocks, and the second data block does not correspond to the M compressed blocks.
21. The device according to claim 20, characterized in that, The second data block is weight data and the change degree of the second data block is lower than a second threshold; or The second data block is gradient data and the probability that the second data block is in a stable interval is greater than a third threshold.
22. The device according to any one of claims 17 to 21, characterized in that, The transceiver unit is further configured to send compression parameters to the second device; Among them, the compression parameters are used to indicate one or more of the following information: the model parameters of the AI model, the compression granularity of each of the N data blocks, or the compression scheme corresponding to each of the N data blocks; the model parameters include at least one of the following: the type of the AI model, the number of layers of the AI model, the number of nodes in each layer of the AI model, or the function of each layer in the AI model.
23. The device according to claim 22, characterized in that, The compression granularity of each of the N data blocks is a partial node, one layer, or multiple layers in a layer of the AI model.
24. The device according to any one of claims 17 to 23, characterized in that, At least one third data block is included in the M data blocks, the third data block includes one or more sub-data blocks, and the difference degree between the distribution information of any two sub-data blocks among the multiple sub-data blocks is less than a fourth threshold.
25. A communication device, characterized in that, Including: A transceiver unit, configured to receive compressed data from a first device, where the compressed data includes M compressed blocks, and the M compressed blocks correspond one-to-one to M of the N data blocks corresponding to an artificial intelligence (AI) model. The compression scheme corresponding to each of the N data blocks is related to the distribution characteristics of the data block. Here, N is an integer greater than 1, and M is an integer not greater than N. A processing unit, configured to determine the AI model according to the compressed data.
26. The device according to claim 25, characterized in that, At least one first data block is included in the M compressed blocks, and the first data block is a data block among the N data blocks.
27. The device according to claim 26, characterized in that, The importance level of the first data block is greater than a first threshold.
28. The device according to any one of claims 25 to 27, characterized in that, At least one second data block is included in the N data blocks, and the second data block does not correspond to the M compressed blocks.
29. The device according to claim 28, characterized in that, The second data block is weight data and the degree of change of the second data block is lower than a second threshold; or The second data block is gradient data and the probability that the second data block is located in a stable interval is greater than a third threshold.
30. The device according to any one of claims 25 to 29, characterized in that, The transceiver unit is further configured to receive compression parameters from the first device. The processing unit, configured to determine the AI model according to the compressed data, specifically includes: being configured to determine the AI model according to the compressed data and the compression parameters; where the compression parameters are used to indicate one or more of the following information: the model parameters of the AI model, the compression granularity of each of the N data blocks, or the compression scheme corresponding to each of the N data blocks. The model parameters include at least one of the following: the type of the AI model, the number of layers of the AI model, the number of nodes in each layer of the AI model, or the function of each layer in the AI model.
31. The device according to claim 30, wherein The compression granularity of each of the N data blocks is a partial node, one layer, or multiple layers in one layer of the AI model.
32. The device according to any one of claims 25 to 31, wherein At least one third data block is included in the M data blocks, and the third data block includes one or more sub-data blocks, and the difference degree between the distribution information of any two of the multiple sub-data blocks is less than a fourth threshold.
33. A communication device, wherein including: At least one processor, configured to perform at least one of the following: running computer instructions or programs, or logical circuits, such that the communication device executes the data transmission method according to any one of claims 1 to 8, or executes the data transmission method according to any one of claims 9 to 16.
34. A communication system, wherein including a first device configured to execute the method according to any one of claims 1 to 8, and a second device configured to execute the method according to any one of claims 9 to 16.
35. A computer-readable storage medium, wherein A computer program or instruction is stored in the storage medium, and when the computer program or instruction is executed by the communication device, the method according to any one of claims 1 to 8 is implemented, or the method according to any one of claims 9 to 16 is implemented.
36. A computer program product, wherein The computer program product includes a computer program or instruction, and when the computer program or instruction runs on the processor, the processor is caused to execute the method according to any one of claims 1 to 8, or execute the method according to any one of claims 9 to 16.