Text generation method and device based on matrix multiplication of shared index and terminal equipment

By using a matrix multiplication method based on shared index in the LLM model, floating-point format data is converted into an integer data set for calculation, which solves the problem of high hardware overhead for custom data types and realizes low-power and high-efficiency text generation.

CN120012913APending Publication Date: 2025-05-16NANJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411888199.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In the LLM model, the hardware overhead of custom data types is still high and the versatility is poor, which limits its application in cost-sensitive inference scenarios.

Method used

The matrix multiplication method based on shared index is adopted to convert the text data and model parameters of the input LLM model into a floating-point format data set, and the data is exponentially aligned through the preset shared index algorithm to generate an integer data set for matrix multiplication, thereby reducing calculation power consumption.

Benefits of technology

While maintaining high computational accuracy, it greatly reduces the computing power consumption caused by the original floating-point operation, further reduces the inference delay of the LLM model, and improves the efficiency of text generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012913A_ABST
    Figure CN120012913A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of computers, and provides a text generation method and device based on matrix multiplication of a shared index and terminal equipment, and the method comprises the steps: obtaining input LLM model text data and LLM model parameters, converting the text data and the model parameters into a first floating-point format data set, and storing the first floating-point format data set into a second floating-point format data set; performing index alignment on each data in the first floating point format data set according to a preset shared index algorithm to generate a second floating point format data set, performing matrix multiplication operation on the first matrix and the second matrix to generate a third matrix, inputting the third matrix to a self-attention mechanism layer to output attention weighted representation, and outputting attention weighted representation. And inputting the attention weighted representation into a feedforward neural network to output high-level feature representation, and finally inputting the high-level feature representation into a decoder to output a text corresponding to the text data. According to the method, the calculation power consumption of original floating point calculation is greatly reduced while the high calculation precision is maintained, the reasoning delay of the LLM model is further reduced, and the text generation efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer technology, and in particular, relates to a matrix multiplication method, apparatus and terminal device based on a shared index. Background Art

[0002] In the current rapid development of deep learning and artificial intelligence applications, the large model size and high computational complexity have become key factors restricting their widespread application, especially in latency-sensitive cloud services and resource-constrained edge devices. The slowdown of Moore's Law has made the arithmetic density of computing hardware a key constraint on large-scale inference performance. In order to overcome these performance and energy consumption challenges, researchers have not only focused on streamlining the neural network structure, but also explored the reduction of the bit width of weights and activation functions as a solution path. In this context, a series of innovative data formats have emerged, aiming to facilitate the implementation of low-precision inference through efficient data representation and processing mechanisms. Although fixed-point data types are favored for their low hardware overhead, their limited dynamic range and reliance on manual calibration limit their practicality in large-scale applications.

[0003] To address this problem, the industry has begun to turn to custom data types in order to reduce hardware overhead while maintaining high accuracy. For example, NVIDIA launched the TF32 data type for its A100 GPU. These new data types significantly reduce the accuracy loss during the execution of DNN models by providing a wide dynamic range. However, the hardware overhead of these custom data types is still high, limiting their application in cost-sensitive reasoning scenarios. Summary of the invention

[0004] The embodiments of the present application provide a text generation method, apparatus, terminal device and storage medium based on matrix multiplication of shared exponents, which can solve the problem that the hardware overhead of customized data types in the existing LLM model is still high and the versatility is poor.

[0005] In the first aspect, an embodiment of the present application provides a text generation method based on matrix multiplication of shared exponents, including: obtaining text data input into an LLM model and model parameters of the LLM model; converting the text data and the model parameters into a first floating-point format data set; performing exponential alignment on each data in the first floating-point format data set according to a preset shared exponent algorithm to generate a second floating-point format data set; performing matrix multiplication operations on a first matrix and a second matrix to generate a third matrix, wherein the first matrix and the second matrix are any two matrices in the second floating-point format data set that undergo matrix multiplication operations; inputting the third matrix into a self-attention mechanism layer to output an attention weighted representation; inputting the attention weighted representation into a feedforward neural network to output a high-level feature representation; inputting the high-level feature representation into a decoder to output the text corresponding to the text data.

[0006] In a possible implementation manner of the first aspect, the step of performing index alignment on each data in the first floating-point format data set according to a preset shared index algorithm to generate the second floating-point format data set includes:

[0007] Traverse the first floating point format data set and determine the maximum exponent value e max ;

[0008] Calculate the mantissa shift coefficient:

[0009] shift=e max -e i

[0010] Among them, shift represents the mantissa shift coefficient, e i Represents the exponent part of each data in the first floating point format data set;

[0011] According to the mantissa shift coefficient shift, the mantissa part m of each data in the first floating point format data set is shifted. i Shift right to align the exponent of each data in the first floating-point format data set:

[0012] m i =m i >>shift;

[0013] The first floating-point format data set after the exponent is aligned is determined as the second floating-point format data set.

[0014] Optionally, in another possible implementation manner of the first aspect, performing a matrix multiplication operation on the first matrix and the second matrix to generate a third matrix includes:

[0015] Perform an integer matrix multiplication operation on the mantissa part of the first matrix and the mantissa part of the second matrix to generate the mantissa part of the third matrix:

[0016] I c =I a *I b

[0017] Among them, I a ,I b ,I c are respectively the mantissa part of the first matrix, the mantissa part of the second matrix and the mantissa part of the third matrix;

[0018] Add the exponential part of the first matrix to the exponential part of the second matrix to generate the exponential part of the third matrix:

[0019] exp c =exp a +exp b

[0020] Among them, exp c Represents the exponential part of the third matrix, exp a Represents the exponential part of the first matrix, exp b represents the exponential part of the second matrix;

[0021] Combine the mantissa part of the third matrix with the exponent part of the third matrix to generate the third matrix:

[0022] C=I c *2 expC

[0023] Wherein, C represents the third matrix.

[0024] In the second aspect, an embodiment of the present application provides a text generation device based on matrix multiplication of shared exponents, including: an acquisition module, used to acquire text data input into an LLM model and model parameters of the LLM model; a conversion module, used to convert the text data and the model parameters into a first floating-point format data set; a first generation module, used to perform index alignment on each data in the first floating-point format data set according to a preset shared exponent algorithm to generate a second floating-point format data set; a second generation module, used to perform matrix multiplication operations on the first matrix and the second matrix to generate a third matrix, the first matrix and the second matrix being any two matrices in the second floating-point format data set that undergo matrix multiplication operations; a first output module, used to input the third matrix into a self-attention mechanism layer and output an attention weighted representation; a second output module, used to input the attention weighted representation into a feedforward neural network and output a high-level feature representation; a third output module, used to input the high-level feature representation into a decoder and output the text corresponding to the text data.

[0025] In a third aspect, an embodiment of the present application provides a terminal device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the text generation method based on matrix multiplication of shared exponents as described above is implemented.

[0026] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the text generation method based on matrix multiplication of shared exponents as described above is implemented.

[0027] In the technical solution of the present application, the text data and the model parameters of the LLM model are first obtained as input to the LLM model, and then the text data and the model parameters are converted into a first floating-point format data set, and then the index of each data in the first floating-point format data set is aligned according to the preset shared index algorithm to generate a second floating-point format data set, and then the first matrix and the second matrix are matrix multiplication operations are performed to generate a third matrix, the first matrix and the second matrix are any two matrices in the second floating-point format data set that are matrix multiplied, and then the third matrix is ​​input into the self-attention mechanism layer, and the attention weighted representation is output, and then the attention weighted representation is input into the feedforward neural network, and the high-level feature representation is output, and finally the high-level feature representation is input into the decoder, and the text corresponding to the text data is output. The present application shares a set of floating-point data as the same index, and then converts the floating-point format data into integer processing and designs a matrix multiplication and convolution algorithm based on a shared index. While maintaining high calculation accuracy, the algorithm greatly reduces the computational power consumption caused by the original floating-point operation, further reduces the reasoning delay of the LLM model, and improves the efficiency of text generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0029] Figure 1 It is a flowchart of a text generation method based on matrix multiplication of shared exponents provided in one embodiment of the present application;

[0030] Figure 2 is a flow chart of a sharing index algorithm provided in an embodiment of the present application;

[0031] Figure 3 It is a structural schematic diagram of a text generation device based on matrix multiplication of shared exponents provided in one embodiment of the present application;

[0032] Figure 4 It is a schematic diagram of the structure of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0033] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0034] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0035] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0036] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.

[0037] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0038] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0039] The text generation method, apparatus, terminal device and storage medium based on matrix multiplication of shared exponents provided in the present application are described in detail below with reference to the accompanying drawings.

[0040] Figure 1 A flow chart of a text generation method based on matrix multiplication of shared exponents provided in an embodiment of the present application is shown.

[0041] like Figure 1 As shown, the text generation method based on matrix multiplication of shared exponents includes the following steps:

[0042] Step 101, obtaining text data input into the LLM model and model parameters of the LLM model;

[0043] Step 102, converting the text data and the model parameters into a first floating point format data set;

[0044] Step 103, performing index alignment on each data in the first floating point format data set according to a preset shared index algorithm to generate a second floating point format data set;

[0045] Step 104, performing a matrix multiplication operation on the first matrix and the second matrix to generate a third matrix, where the first matrix and the second matrix are any two matrices in the second floating-point format data set that are subjected to a matrix multiplication operation;

[0046] Step 105, input the third matrix into the self-attention mechanism layer, and output the attention weighted representation;

[0047] Step 106, input the attention weighted representation into a feedforward neural network, and output a high-level feature representation;

[0048] Step 107, input the high-level feature representation into the decoder, and output the text corresponding to the text data.

[0049] Furthermore, in an embodiment of the present application, the above step 103 includes:

[0050] Step 1031, traverse the first floating point format data set to determine the maximum exponent value e max ;

[0051] Step 1032, calculate the mantissa shift coefficient:

[0052] shift=e max -e i

[0053] Among them, shift represents the mantissa shift coefficient, e i Represents the exponent part of each data in the first floating point format data set;

[0054] Step 1033: shift the mantissa part m of each data in the first floating point format data set according to the mantissa shift coefficient shift. i Shift right to align the exponent of each data in the first floating-point format data set:

[0055] m i =m i >>shift;

[0056] Step 1034: determine the first floating-point format data set after exponent alignment as the second floating-point format data set.

[0057] Furthermore, in an embodiment of the present application, the above step 104 includes:

[0058] Step 1041, perform an integer matrix multiplication operation on the mantissa part of the first matrix and the mantissa part of the second matrix to generate the mantissa part of the third matrix:

[0059] I c =I a *I b

[0060] Among them, I a ,I b ,I c are respectively the mantissa part of the first matrix, the mantissa part of the second matrix and the mantissa part of the third matrix;

[0061] Step 1042, add the exponential part of the first matrix and the exponential part of the second matrix to generate the exponential part of the third matrix:

[0062] exp c =exp a +exp b

[0063] Among them, exp c Represents the exponential part of the third matrix, exp a Represents the exponential part of the first matrix, exp b represents the exponential part of the second matrix;

[0064] Step 1043, combining the mantissa part of the third matrix and the exponent part of the third matrix to generate a third matrix:

[0065] C=I c *2 expC

[0066] Wherein, C represents the third matrix.

[0067] The text generation method based on matrix multiplication of shared index provided in the present application first obtains the text data of the input LLM model and the model parameters of the LLM model, then converts the text data and the model parameters into a first floating point format data set, and then performs index alignment on each data in the first floating point format data set according to the preset shared index algorithm to generate a second floating point format data set, and then performs matrix multiplication operation on the first matrix and the second matrix to generate a third matrix, the first matrix and the second matrix are any two matrices in the second floating point format data set that perform matrix multiplication operation, and then inputs the third matrix into the self-attention mechanism layer, outputs the attention weighted representation, and then inputs the attention weighted representation into the feedforward neural network, outputs the high-level feature representation, and finally inputs the high-level feature representation into the decoder, and outputs the text corresponding to the text data. The present application shares a set of floating point data as the same index, and then converts the floating point format data into integer processing and designs a matrix multiplication and convolution algorithm based on shared index. While maintaining high calculation accuracy, the algorithm greatly reduces the calculation power consumption caused by the original floating point operation, further reduces the reasoning delay of the LLM model, and improves the efficiency of text generation.

[0068] Figure 2 A schematic diagram of the flow of a sharing index algorithm in an embodiment of the present application is shown.

[0069] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0070] Corresponding to the text generation method based on matrix multiplication of shared exponents in the above embodiment, Figure 3 A structural block diagram of a text generation device based on matrix multiplication of shared exponents provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0071] Reference Figure 3 , the device 300 comprises:

[0072] An acquisition module 301 is used to acquire text data input into the LLM model and model parameters of the LLM model;

[0073] A conversion module 302, used for converting text data and model parameters into a first floating point format data set;

[0074] A first generating module 303 is used to perform index alignment on each data in the first floating point format data set according to a preset shared index algorithm to generate a second floating point format data set;

[0075] A second generating module 304 is used to perform a matrix multiplication operation on the first matrix and the second matrix to generate a third matrix, where the first matrix and the second matrix are any two matrices in the second floating-point format data set that are subjected to a matrix multiplication operation;

[0076] A first output module 305 is used to input the third matrix into the self-attention mechanism layer and output the attention weighted representation;

[0077] A second output module 306, for inputting the attention weighted representation into a feedforward neural network and outputting a high-level feature representation;

[0078] The third output module 307 is used to input the high-level feature representation into the decoder and output the text corresponding to the text data.

[0079] In actual use, the text generation device based on matrix multiplication of shared indexes provided in the embodiment of the present application can be configured in any terminal device to execute the aforementioned text generation method based on matrix multiplication of shared indexes.

[0080] The text generation device based on matrix multiplication of shared index provided by the present application first obtains the text data of the input LLM model and the model parameters of the LLM model, then converts the text data and the model parameters into a first floating point format data set, and then performs index alignment on each data in the first floating point format data set according to the preset shared index algorithm to generate a second floating point format data set, and then performs matrix multiplication operation on the first matrix and the second matrix to generate a third matrix, the first matrix and the second matrix are any two matrices in the second floating point format data set that perform matrix multiplication operation, and then inputs the third matrix into the self-attention mechanism layer, outputs the attention weighted representation, and then inputs the attention weighted representation into the feedforward neural network, outputs the high-level feature representation, and finally inputs the high-level feature representation into the decoder, and outputs the text corresponding to the text data. The present application shares a set of floating point data as the same index, and then converts the floating point format data into integer processing and designs a matrix multiplication and convolution algorithm based on shared index. While maintaining high calculation accuracy, the algorithm greatly reduces the calculation power consumption caused by the original floating point operation, further reduces the reasoning delay of the LLM model, and improves the efficiency of text generation.

[0081] In a possible implementation of the embodiment of the present application, the first generating module 303 includes:

[0082] The first determining unit is used to traverse the first floating point format data set and determine the maximum exponent value e max ;

[0083] Computational unit, used to calculate the mantissa shift coefficient:

[0084] shift=e max -ei

[0085] Among them, shift represents the mantissa shift coefficient, e i Represents the exponent part of each data in the first floating point format data set;

[0086] A moving unit is used to shift the mantissa part m of each data in the first floating point format data set according to the mantissa shift coefficient shift i Shift right to align the exponent of each data in the first floating-point format data set:

[0087] m i =m i >>shift;

[0088] The second determining unit is configured to determine the first floating-point format data set after exponent alignment as the second floating-point format data set.

[0089] Furthermore, in another possible implementation of the embodiment of the present application, the second generating module 304 includes:

[0090] The first generating unit is used to perform an integer matrix multiplication operation on the mantissa part of the first matrix and the mantissa part of the second matrix to generate the mantissa part of the third matrix:

[0091] I c =I a *I b

[0092] Among them, I a ,I b ,I c are respectively the mantissa part of the first matrix, the mantissa part of the second matrix and the mantissa part of the third matrix;

[0093] The second generating unit is used to add the exponential part of the first matrix and the exponential part of the second matrix to generate the exponential part of the third matrix:

[0094] exp c =exp a +exp b

[0095] Among them, exp c Represents the exponential part of the third matrix, exp a Represents the exponential part of the first matrix, exp b represents the exponential part of the second matrix;

[0096] The third generating unit is used to combine the mantissa part of the third matrix and the exponent part of the third matrix to generate a third matrix:

[0097] C=Ic *2 expC

[0098] Wherein, C represents the third matrix.

[0099] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0100] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0101] In order to implement the above embodiments, the present application also proposes a terminal device.

[0102] Figure 4 A schematic diagram of the structure of a terminal device according to an embodiment of the present application.

[0103] like Figure 4 As shown, the terminal device 200 includes: a memory 210 and at least one processor 220, a bus 230 connecting different components (including the memory 210 and the processor 220), the memory 210 stores a computer program, and when the processor 220 executes the program, the text generation method based on matrix multiplication of shared exponents described in the embodiment of the present application is implemented.

[0104] Bus 230 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor or a local bus using any of a variety of bus architectures. For example, these architectures include but are not limited to Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus and Peripheral Component Interconnect (PCI) bus.

[0105] The terminal device 200 typically includes a variety of electronic device readable media, which can be any available media that can be accessed by the terminal device 200, including volatile and non-volatile media, removable and non-removable media.

[0106] The memory 210 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 240 and / or cache memory 250. The terminal device 200 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 260 may be used to read and write non-removable, non-volatile magnetic media ( Figure 4 not shown, usually called a "hard drive"). Although Figure 4 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical medium) may be provided. In these cases, each drive may be connected to bus 230 via one or more data medium interfaces. Memory 210 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present application.

[0107] A program / utility 280 having a set (at least one) of program modules 270 may be stored, for example, in the memory 210, such program modules 270 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The program modules 270 generally perform the functions and / or methods of the embodiments described herein.

[0108] The terminal device 200 can also communicate with one or more external devices 290 (e.g., keyboard, pointing device, display 291, etc.), and can also communicate with one or more devices that enable users to interact with the terminal device 200, and / or communicate with any device that enables the terminal device 200 to communicate with one or more other computing devices (e.g., network card, modem, etc.). This communication can be carried out through an input / output (I / O) interface 292. In addition, the terminal device 200 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN) and / or public network, such as the Internet) through a network adapter 293. As shown in the figure, the network adapter 293 communicates with other modules of the terminal device 200 through a bus 230. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the terminal device 200, including but not limited to: microcode, device driver, redundant processing unit, external disk drive array, RAID system, tape drive, and data backup storage system, etc.

[0109] The processor 220 executes various functional applications and data processing by running the programs stored in the memory 210 .

[0110] It should be noted that the implementation process and technical principles of the terminal device of this embodiment refer to the aforementioned explanation of the text generation method based on matrix multiplication of shared exponents in the embodiment of the present application, and will not be repeated here.

[0111] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0112] An embodiment of the present application provides a computer program product. When the computer program product is run on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0113] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the camera / terminal device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0114] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0115] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0116] In the embodiments provided in the present application, it should be understood that the disclosed devices / terminal equipment and methods can be implemented in other ways. For example, the device / terminal equipment embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0117] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0118] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A text generation method based on matrix multiplication of shared exponents, characterized in that: include: Obtain text data input into the LLM model and model parameters of the LLM model; Converting the text data and the model parameters into a first floating point format data set; Performing index alignment on each data in the first floating-point format data set according to a preset shared index algorithm to generate a second floating-point format data set; Performing a matrix multiplication operation on a first matrix and a second matrix to generate a third matrix, wherein the first matrix and the second matrix are any two matrices in the second floating-point format data set that are subjected to a matrix multiplication operation; Input the third matrix into the self-attention mechanism layer and output the attention weighted representation; Inputting the attention weighted representation into a feedforward neural network and outputting a high-level feature representation; The high-level feature representation is input into a decoder, and the text corresponding to the text data is output.

2. The method according to claim 1, characterized in that The step of performing index alignment on each data in the first floating point format data set according to a preset shared index algorithm to generate a second floating point format data set includes: Traverse the first floating point format data set and determine the maximum exponent value e max ; Calculate the mantissa shift coefficient: shift=e max -e i Among them, shift represents the mantissa shift coefficient, e i Representing the exponent part of each data in the first floating point format data set; According to the mantissa shift coefficient shift, the mantissa part m of each data in the first floating point format data set is shifted. i Right shift is performed to achieve exponent alignment of each data in the first floating point format data set: m i =m i >>shift; The first floating-point format data set after exponent alignment is determined as the second floating-point format data set.

3. The method according to claim 2, characterized in that The performing a matrix multiplication operation on the first matrix and the second matrix to generate a third matrix includes: An integer matrix multiplication operation is performed on the mantissa part of the first matrix and the mantissa part of the second matrix to generate the mantissa part of the third matrix: I c =I a *I b Among them, I a ,I b ,I c are respectively the mantissa part of the first matrix, the mantissa part of the second matrix and the mantissa part of the third matrix; The exponential part of the first matrix and the exponential part of the second matrix are added to generate the exponential part of the third matrix: exp c =exp a +exp b Among them, exp c represents the exponential part of the third matrix, exp a represents the exponential part of the first matrix, exp b represents the exponential part of the second matrix; The mantissa part of the third matrix and the exponent part of the third matrix are combined to generate the third matrix: C=I c *2 expC Wherein, C represents the third matrix.

4. A text generation device based on matrix multiplication of shared exponents, characterized in that: include: An acquisition module, used for acquiring text data input into the LLM model and model parameters of the LLM model; A conversion module, used for converting the text data and the model parameters into a first floating point format data set; A first generating module, configured to perform exponential alignment on each data in the first floating point format data set according to a preset shared exponent algorithm to generate a second floating point format data set; A second generating module is used to perform a matrix multiplication operation on the first matrix and the second matrix to generate a third matrix, wherein the first matrix and the second matrix are any two matrices in the second floating-point format data set that are subjected to a matrix multiplication operation; A first output module, used to input the third matrix into a self-attention mechanism layer and output an attention weighted representation; A second output module, configured to input the attention weighted representation into a feedforward neural network and output a high-level feature representation; The third output module is used to input the high-level feature representation into a decoder and output the text corresponding to the text data.

5. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 3 is implemented.

6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 3 is implemented.