Confusion degree calculation method, device and equipment and computer readable storage medium
By deploying large language models in parallel on multiple graphics processors and calculating loss values and confusion in parallel, the problem of insufficient video memory is solved, and efficient evaluation of large parameter scale and long sequence length models is achieved.
Patent Information
- Application Number
- CN202510044985.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-06
AI Technical Summary
When the parameter scale of large language models is large and the sequence length is long, the perplexity (PPL) calculation consumes too much video memory, resulting in graphics processors with smaller video memory being unable to meet the computing needs of PPL.
The model to be evaluated is deployed in parallel on multiple graphics processors, and the model input is converted into symbol sequences through multiple graphics processors for forward calculation, obtaining the logical value and calculating the loss value. Finally, the loss value calculated by multiple graphics processors is aggregated on one graphics processor to calculate the confusion.
By amortizing the memory requirements, multiple graphics processors compute the confusion in parallel, solving the problem of insufficient memory for a single graphics processor, and supporting model evaluation of larger parameter scale and longer sequence length.
Smart Images

Figure CN119938332A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a perplexity calculation method, device, equipment, and computer-readable storage medium. Background Art
[0002] After the large language model training is completed, the effect of the large language model needs to be evaluated, and perplexity (PPL) evaluation is one of the methods. The lower the PPL, the better the effect.
[0003] When the parameters of a large language model are large and the sequence length is long, the calculation of PPL consumes a lot of video memory, so a graphics processing unit (GPU) with a small video memory cannot meet the calculation requirements of PPL. Summary of the invention
[0004] In order to solve the above technical problems, the present disclosure provides a perplexity calculation method, device, equipment and computer-readable storage medium to solve the problem that PPL cannot be calculated due to the small video memory of a single graphics processor.
[0005] In a first aspect, an embodiment of the present disclosure provides a perplexity calculation method, comprising:
[0006] Convert model input to a sequence of symbols;
[0007] The symbol sequence is sent to multiple graphics processors, and the models to be evaluated are deployed in parallel on the multiple graphics processors. The graphics processors are used to perform forward calculations on the symbol sequence through the models to be evaluated deployed on the graphics processors to obtain a logic value, and calculate a loss value based on the logic value and the symbol sequence;
[0008] Aggregating the loss values respectively calculated by the multiple graphics processors into one graphics processor to obtain an aggregated loss value;
[0009] According to the aggregated loss value, the perplexity of the model to be evaluated is calculated.
[0010] In some embodiments, the model input is textual information;
[0011] Convert the model input to a sequence of symbols, including:
[0012] dividing the text information into a plurality of text units;
[0013] The symbol identifier corresponding to each text unit in the preset vocabulary is determined, and the symbol identifiers corresponding to the multiple text units respectively constitute the symbol sequence.
[0014] In some embodiments, the graphics processor is used to perform forward calculation on the symbol sequence through the model to be evaluated deployed on the graphics processor, including:
[0015] The graphics processor is used to perform forward calculation on the symbol sequence through the model weights corresponding to the model to be evaluated deployed on the graphics processor.
[0016] In some embodiments, calculating the loss value according to the logic value and the symbol sequence includes:
[0017] A loss value is obtained by performing cross entropy calculation based on the logic value and the symbol sequence.
[0018] In some embodiments, aggregating the loss values respectively calculated by the multiple graphics processors into one graphics processor to obtain the aggregated loss value includes:
[0019] Aggregating the loss values respectively calculated by the multiple graphics processors to a target graphics processor among the multiple graphics processors to obtain an aggregated loss value;
[0020] Calculating the perplexity of the model to be evaluated according to the aggregated loss value, including:
[0021] The target graphics processor calculates the perplexity of the model to be evaluated according to the aggregated loss value.
[0022] In a second aspect, an embodiment of the present disclosure provides a perplexity calculation device, comprising:
[0023] A conversion module, which converts the model input into a sequence of symbols;
[0024] A sending module, used for sending the symbol sequence to multiple graphics processors, on which the models to be evaluated are deployed in parallel, and the graphics processors are used for performing forward calculation on the symbol sequence through the models to be evaluated deployed on the graphics processors to obtain a logic value, and calculating a loss value based on the logic value and the symbol sequence;
[0025] An aggregation module, used to aggregate the loss values respectively calculated by the multiple graphics processors into one graphics processor to obtain an aggregated loss value;
[0026] A calculation module is used to calculate the perplexity of the model to be evaluated according to the aggregated loss value.
[0027] In some embodiments, the model input is textual information;
[0028] When the conversion module converts the model input into a symbol sequence, it is specifically used to:
[0029] dividing the text information into a plurality of text units;
[0030] The symbol identifier corresponding to each text unit in the preset vocabulary is determined, and the symbol identifiers corresponding to the multiple text units respectively constitute the symbol sequence.
[0031] In some embodiments, the graphics processor is used to perform forward calculation on the symbol sequence through the model to be evaluated deployed on the graphics processor, including:
[0032] The graphics processor is used to perform forward calculation on the symbol sequence through the model weights corresponding to the model to be evaluated deployed on the graphics processor.
[0033] In some embodiments, when the graphics processor calculates the loss value according to the logic value and the symbol sequence, it is specifically used to:
[0034] A loss value is obtained by performing cross entropy calculation based on the logic value and the symbol sequence.
[0035] In some embodiments, the aggregation module aggregates the loss values respectively calculated by the multiple graphics processors into one graphics processor to obtain the aggregated loss value, specifically used to: aggregate the loss values respectively calculated by the multiple graphics processors into a target graphics processor among the multiple graphics processors to obtain the aggregated loss value;
[0036] The target graphics processor is also used to calculate the perplexity of the model to be evaluated based on the aggregated loss value.
[0037] In a third aspect, an embodiment of the present disclosure provides an electronic device, including:
[0038] Memory;
[0039] Processor; and
[0040] Computer programs;
[0041] The computer program is stored in the memory and is configured to be executed by the processor to implement the method as described in the first aspect.
[0042] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the method described in the first aspect.
[0043] In a fifth aspect, an embodiment of the present disclosure further provides a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the data acquisition method as described above is implemented.
[0044] The perplexity calculation method, apparatus, device and computer-readable storage medium provided by the embodiments of the present disclosure deploy the model to be evaluated on multiple graphics processors in parallel, so that the multiple graphics processors respectively perform forward calculations on the symbol sequence converted from the model input through the model to be evaluated to obtain a logical value, and calculate a loss value based on the logical value and the symbol sequence, that is, the multiple graphics processors calculate the loss value in parallel, and finally aggregate the loss values calculated by the multiple graphics processors to one graphics processor to obtain the aggregated loss value, and calculate the perplexity of the model to be evaluated based on the aggregated loss value, that is, the video memory consumed for calculating the perplexity can be distributed among the multiple graphics processors, so as to achieve the purpose of calculating the perplexity in parallel by the multiple graphics processors, thereby solving the problem that the PPL cannot be calculated due to the small video memory of a single graphics processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0046] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0047] Figure 1 A flow chart of a method for calculating perplexity provided in an embodiment of the present disclosure;
[0048] Figure 2 A schematic diagram of an application scenario provided by an embodiment of the present disclosure;
[0049] Figure 3 A flow chart of a method for calculating perplexity provided by another embodiment of the present disclosure;
[0050] Figure 4 A flow chart of a method for calculating perplexity provided by another embodiment of the present disclosure;
[0051] Figure 5 A schematic diagram of the structure of a perplexity calculation device provided in an embodiment of the present disclosure;
[0052] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0053] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.
[0054] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0055] After the training of the large language model is completed, it is necessary to evaluate the effect of the large language model, and perplexity (PPL) evaluation is one of the ways. The lower the PPL, the better the effect. When the parameter scale of the large language model is large and the sequence length is long, the calculation of PPL consumes a lot of video memory, so the graphics processing unit (GPU) with smaller video memory cannot meet the calculation requirements of PPL. To solve this problem, the embodiment of the present disclosure provides a perplexity calculation method, which is introduced in conjunction with a specific embodiment below.
[0056] Figure 1 The flowchart of the perplexity calculation method provided in the embodiment of the present disclosure is shown in FIG. Figure 2 The application scenario shown includes a terminal device 21 and a server 22. The terminal device 21 may be a smart phone, a PDA, a tablet computer, a wearable device with a display screen, a desktop computer, a laptop computer, an all-in-one machine, a smart home device, etc. It is understandable that the perplexity calculation method provided in the embodiment of the present disclosure may also be applied in other scenarios.
[0057] Below Figure 1 The method for calculating the perplexity shown in FIG. 1 is introduced, and the specific steps of the method are as follows:
[0058] S101, converting the model input into a symbol sequence.
[0059] For example, a trained large language model is deployed in parallel on multiple GPUs, which can be set on the same server or on different servers. Figure 2The first graphics processor 23 and the second graphics processor 24 shown can be set in the server 22, or the first graphics processor 23 and the second graphics processor 24 can be set in different servers respectively. The embodiment of the present disclosure takes the first graphics processor 23 and the second graphics processor 24 as an example in which the server 22 is set. The large language model is the model to be evaluated. When evaluating the large language model, the terminal device 21 sends the text information as a model input to the server 22, and a tokenizer is deployed on the server 22. The server 22 converts the model input (input) into multiple symbol identifiers (token id) through the tokenizer. The multiple symbol identifiers constitute a symbol sequence, which can be a one-dimensional vector.
[0060] S102, sending the symbol sequence to multiple graphics processors, the model to be evaluated is deployed in parallel on the multiple graphics processors, the graphics processor is used to perform forward calculation on the symbol sequence through the model to be evaluated deployed on the graphics processor to obtain a logical value, and calculate the loss value based on the logical value and the symbol sequence.
[0061] Specifically, the server 22 sends the symbol sequence to multiple graphics processors, such as the first graphics processor 23 and the second graphics processor 24. Since the model to be evaluated is deployed in parallel on the first graphics processor 23 and the second graphics processor 24, the first graphics processor 23 can perform forward calculation on the symbol sequence through the model to be evaluated deployed thereon to obtain a logical value, which is recorded as logits0. Further, the first graphics processor 23 can also calculate a loss value based on logits0 and the symbol sequence, and the loss value is recorded as loss0. The second graphics processor 24 performs forward calculation on the symbol sequence through the model to be evaluated deployed thereon to obtain a logical value, which is recorded as logits1. Further, the second graphics processor 24 can also calculate a loss value based on logits1 and the symbol sequence, and the loss value is recorded as loss1.
[0062] S103: Aggregate the loss values respectively calculated by the multiple graphics processors into one graphics processor to obtain an aggregated loss value.
[0063] Specifically, loss0 calculated by the first graphics processor 23 and loss1 calculated by the second graphics processor 24 are aggregated to one graphics processor, for example, to the first graphics processor 23. Specifically, the second graphics processor 24 sends the loss1 calculated by it to the first graphics processor 23, so that the first graphics processor 23 performs an aggregation operation on loss0 and loss1 to obtain an aggregated loss value, which is recorded as loss, that is, the aggregated loss value is the final loss value loss.
[0064] S104. Calculate the perplexity of the model to be evaluated according to the aggregated loss value.
[0065] For example, after the first graphics processor 23 calculates the final loss value loss, it calculates the perplexity of the model to be evaluated, that is, the PPL value, according to the loss.
[0066] The disclosed embodiment deploys the model to be evaluated on multiple graphics processors in parallel, so that the multiple graphics processors respectively perform forward calculations on the symbol sequence converted from the model input through the model to be evaluated to obtain a logical value, and calculate a loss value based on the logical value and the symbol sequence, that is, the multiple graphics processors calculate the loss value in parallel, and finally aggregate the loss values calculated by the multiple graphics processors to one graphics processor to obtain the aggregated loss value, and calculate the perplexity of the model to be evaluated based on the aggregated loss value, that is, the video memory consumed for calculating the perplexity can be distributed among the multiple graphics processors, so as to achieve the purpose of using multiple graphics processors to calculate the perplexity in parallel, thereby solving the problem that the PPL cannot be calculated due to the small video memory of a single graphics processor.
[0067] Optionally, the model input is text information; the model input is converted into a symbol sequence, including Figure 3 The following steps are shown:
[0068] S301: Segment the text information into multiple text units.
[0069] For example, when evaluating a trained large language model, the model input is text information, which can be a long text. When the server 22 receives the text information, it divides the text information into multiple text units, such as multiple words, multiple characters or multiple phrases.
[0070] S302: Determine the symbol identifier corresponding to each text unit in a preset vocabulary, and the symbol identifiers corresponding to the multiple text units respectively constitute the symbol sequence.
[0071] For example, the text information is divided into multiple words, and each word is further converted into a corresponding token id in a preset word list by a tagger. That is, each word in the text information corresponds to a token id, and the text information includes multiple words. Therefore, the token ids corresponding to the multiple words constitute a symbol sequence.
[0072] Optionally, the graphics processor is used to perform forward calculation on the symbol sequence through the model to be evaluated deployed on the graphics processor, including: the graphics processor is used to perform forward calculation on the symbol sequence through model weights corresponding to the model to be evaluated deployed on the graphics processor.
[0073] For example, the trained large language model is deployed in parallel on the first graphics processor 23 and the second graphics processor 24. For example, the trained large language model is deployed on the first graphics processor 23 and the second graphics processor 24 in tensor parallel mode. The model weight corresponding to the large language model on the first graphics processor 23 is recorded as W0, and the model weight corresponding to the large language model on the second graphics processor 24 is recorded as W1. The first graphics processor 23 performs a forward calculation on the symbol sequence through W0 to obtain logits0. The second graphics processor 24 performs a forward calculation on the symbol sequence through W1 to obtain logits1. Specifically, logits0 and logits1 represent logits slices on the first graphics processor 23 and the second graphics processor 24. In the embodiment of the present disclosure, the first graphics processor 23 can be recorded as GPU0, and the second graphics processor 24 can be recorded as GPU1.
[0074] Optionally, calculating the loss value according to the logical value and the symbol sequence includes: performing cross entropy calculation according to the logical value and the symbol sequence to obtain the loss value.
[0075] For example, a cross entropy function is deployed on GPU0. GPU0 substitutes logits0 and the symbol sequence into the cross entropy function, and uses the cross entropy function to calculate the cross entropy of logits0 and the symbol sequence to obtain the loss value loss0. Similarly, a cross entropy function is also deployed on GPU1. GPU1 substitutes logits1 and the symbol sequence into the cross entropy function, and uses the cross entropy function to calculate the cross entropy of logits1 and the symbol sequence to obtain the loss value loss1. Specifically, the cross entropy function deployed on GPU0 and the cross entropy function deployed on GPU1 are equivalent distributed cross entropy or parallel cross entropy. loss0 and loss1 are distributed model loss values.
[0076] Optionally, the loss values calculated respectively by the multiple graphics processors are aggregated into one graphics processor to obtain an aggregated loss value, including: aggregating the loss values calculated respectively by the multiple graphics processors into a target graphics processor among the multiple graphics processors to obtain an aggregated loss value.
[0077] For example, in the first graphics processor 23 and the second graphics processor 24, the first graphics processor 23 is the main GPU and the second graphics processor 24 is the auxiliary GPU. The loss values calculated by the first graphics processor 23 and the second graphics processor 24 respectively can be aggregated to the main GPU, that is, the main GPU is the target GPU, and the main GPU aggregates loss0 and loss1 into loss.
[0078] Calculating the perplexity of the model to be evaluated according to the aggregated loss value, including: the target graphics processor calculating the perplexity of the model to be evaluated according to the aggregated loss value.
[0079] For example, the main GPU aggregates loss0 and loss1 into loss, and then calculates the perplexity of the large language model to be evaluated based on the loss.
[0080] The disclosed embodiment calculates logical values and loss values in parallel by multiple GPUs, so that multiple GPUs can support PPL evaluation of large language models with large parameter scales and long sequence lengths, solving the problem that a single GPU card with small video memory cannot support PPL evaluation of the large language model. In addition, multiple GPUs not only calculate logical values in parallel, but also calculate loss values in parallel, thereby improving the resource utilization of video memory of multiple GPU cards. In addition, since the video memory resources of multiple GPU cards are more sufficient, when evaluating large language models, longer model inputs are supported, that is, longer long text evaluations are supported, solving the problem that the length of model input is limited due to limited video memory resources of a single GPU card.
[0081] Figure 4 The flowchart of the perplexity calculation method provided by the embodiment of the present disclosure is as follows. Figure 4As shown, input represents the model input, and tokenizer represents the tokenizer, which converts the model input into an input token id, that is, a symbol sequence as described above. Further, the input token id is forward calculated by the model weight W0 corresponding to the large language model deployed on GPU0 to obtain logits0. The input token id is forward calculated by the model weight W1 corresponding to the large language model deployed on GPU1 to obtain logits1. Further, the cross entropy (Cross Entropy) of logits0 and input token id is calculated on GPU0 to obtain loss0. The cross entropy (Cross Entropy) of logits1 and input token id is calculated on GPU1 to obtain loss1. Then, loss0 and loss1 are aggregated into loss on GPU0, and PPL is calculated based on the loss.
[0082] The disclosed embodiment deploys the model to be evaluated on multiple graphics processors in parallel, so that the multiple graphics processors respectively perform forward calculations on the symbol sequence converted from the model input through the model to be evaluated to obtain a logical value, and calculate a loss value based on the logical value and the symbol sequence, that is, the multiple graphics processors calculate the loss value in parallel, and finally aggregate the loss values calculated by the multiple graphics processors to one graphics processor to obtain the aggregated loss value, and calculate the perplexity of the model to be evaluated based on the aggregated loss value, that is, the video memory consumed for calculating the perplexity can be distributed among the multiple graphics processors, so as to achieve the purpose of using multiple graphics processors to calculate the perplexity in parallel, thereby solving the problem that the PPL cannot be calculated due to the small video memory of a single graphics processor.
[0083] Figure 5 Schematic diagram of the structure of the perplexity calculation device provided in the embodiment of the present disclosure. The perplexity calculation device may be the server 22 as described in the above embodiment, or the perplexity calculation device may be a component or assembly in the server 22. The perplexity calculation device provided in the embodiment of the present disclosure may execute the processing flow provided in the perplexity calculation method embodiment, such as Figure 5 As shown, the perplexity calculation device 50 includes:
[0084] A conversion module 51, for converting a model input into a symbol sequence;
[0085] A sending module 52 is used to send the symbol sequence to multiple graphics processors, on which the models to be evaluated are deployed in parallel, and the graphics processors are used to perform forward calculations on the symbol sequence through the models to be evaluated deployed on the graphics processors to obtain a logic value, and calculate a loss value based on the logic value and the symbol sequence;
[0086] An aggregation module 53, configured to aggregate the loss values respectively calculated by the multiple graphics processors into one graphics processor to obtain an aggregated loss value;
[0087] The calculation module 54 is used to calculate the perplexity of the model to be evaluated according to the aggregated loss value.
[0088] Optionally, the model input is text information; when the conversion module 51 converts the model input into a symbol sequence, it is specifically used to:
[0089] dividing the text information into a plurality of text units;
[0090] The symbol identifier corresponding to each text unit in the preset vocabulary is determined, and the symbol identifiers corresponding to the multiple text units respectively constitute the symbol sequence.
[0091] Optionally, the graphics processor is used to perform forward calculation on the symbol sequence through the model to be evaluated deployed on the graphics processor, including:
[0092] The graphics processor is used to perform forward calculation on the symbol sequence through the model weights corresponding to the model to be evaluated deployed on the graphics processor.
[0093] Optionally, when the graphics processor calculates the loss value according to the logic value and the symbol sequence, it is specifically used to:
[0094] A loss value is obtained by performing cross entropy calculation based on the logic value and the symbol sequence.
[0095] Optionally, the aggregation module 53 aggregates the loss values respectively calculated by the multiple graphics processors into one graphics processor to obtain the aggregated loss value, specifically for: aggregating the loss values respectively calculated by the multiple graphics processors into a target graphics processor among the multiple graphics processors to obtain the aggregated loss value;
[0096] The target graphics processor is also used to calculate the perplexity of the model to be evaluated based on the aggregated loss value.
[0097] Figure 5 The perplexity calculation device of the illustrated embodiment can be used to execute the technical solution of the above-mentioned method embodiment, and its implementation principle and technical effect are similar and will not be repeated here.
[0098] Figure 6 Schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure. The electronic device may be the server 22 as described in the above embodiment. The electronic device provided in an embodiment of the present disclosure may execute the processing flow provided in the embodiment of the perplexity calculation method, such as Figure 6 As shown, the electronic device 60 includes: a memory 61, a processor 62, a computer program and a communication interface 63; wherein the computer program is stored in the memory 61 and is configured so that the processor 62 executes the data acquisition method as described above.
[0099] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored. The computer program is executed by a processor to implement the data acquisition method described in the above embodiment.
[0100] In addition, an embodiment of the present disclosure further provides a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the data acquisition method as described above is implemented.
[0101] It should be noted that the computer-readable medium disclosed above may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In an embodiment of the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, device or device. In an embodiment of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a computer-readable program code is carried. This propagated data signal may take a variety of forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium may also be any computer readable medium other than a computer readable storage medium, which may send, propagate or transmit a program for use by or in conjunction with an instruction execution system, apparatus or device. The program code contained on the computer readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0102] In some embodiments, the client and the server may communicate using any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol), and may be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any currently known or future developed network.
[0103] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0104] It should be noted that it should be understood that when one or more programs stored in the computer-readable medium are executed by the electronic device, the electronic device may also be enabled to execute other data acquisition methods provided by the examples of the present disclosure.
[0105] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" language or similar programming languages. The program code may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0106] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0107] The modules or units involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of a module or unit does not, in some cases, constitute a limitation on the module or unit itself.
[0108] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0109] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0110] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for calculating perplexity, characterized in that: The method comprises: Convert model input to a sequence of symbols; The symbol sequence is sent to multiple graphics processors, and the models to be evaluated are deployed in parallel on the multiple graphics processors. The graphics processors are used to perform forward calculations on the symbol sequence through the models to be evaluated deployed on the graphics processors to obtain a logic value, and calculate a loss value based on the logic value and the symbol sequence; Aggregating the loss values respectively calculated by the multiple graphics processors into one graphics processor to obtain an aggregated loss value; According to the aggregated loss value, the perplexity of the model to be evaluated is calculated.
2. The method according to claim 1, characterized in that: The model input is text information; Convert the model input to a sequence of symbols, including: dividing the text information into a plurality of text units; The symbol identifier corresponding to each text unit in the preset vocabulary is determined, and the symbol identifiers corresponding to the multiple text units respectively constitute the symbol sequence.
3. The method according to claim 1, characterized in that The graphics processor is used to perform forward calculation on the symbol sequence through the model to be evaluated deployed on the graphics processor, including: The graphics processor is used to perform forward calculation on the symbol sequence through the model weights corresponding to the model to be evaluated deployed on the graphics processor.
4. The method according to claim 1, characterized in that The loss value is calculated according to the logic value and the symbol sequence, including: A loss value is obtained by performing cross entropy calculation based on the logic value and the symbol sequence.
5. The method according to claim 1, characterized in that: Aggregating the loss values respectively calculated by the multiple graphics processors into one graphics processor to obtain an aggregated loss value, including: Aggregating the loss values respectively calculated by the multiple graphics processors to a target graphics processor among the multiple graphics processors to obtain an aggregated loss value; Calculating the perplexity of the model to be evaluated according to the aggregated loss value, including: The target graphics processor calculates the perplexity of the model to be evaluated according to the aggregated loss value.
6. A perplexity calculation device, characterized in that: include: A conversion module, which converts the model input into a sequence of symbols; A sending module, used for sending the symbol sequence to multiple graphics processors, on which the models to be evaluated are deployed in parallel, and the graphics processors are used for performing forward calculation on the symbol sequence through the models to be evaluated deployed on the graphics processors to obtain a logic value, and calculating a loss value based on the logic value and the symbol sequence; An aggregation module, used to aggregate the loss values respectively calculated by the multiple graphics processors into one graphics processor to obtain an aggregated loss value; A calculation module is used to calculate the perplexity of the model to be evaluated according to the aggregated loss value.
7. The perplexity calculation device according to claim 6, characterized in that: The model input is text information; When the conversion module converts the model input into a symbol sequence, it is specifically used to: dividing the text information into a plurality of text units; The symbol identifier corresponding to each text unit in the preset vocabulary is determined, and the symbol identifiers corresponding to the multiple text units respectively constitute the symbol sequence.
8. The perplexity calculation device according to claim 6, characterized in that: The graphics processor is used to perform forward calculation on the symbol sequence through the model to be evaluated deployed on the graphics processor, including: The graphics processor is used to perform forward calculation on the symbol sequence through the model weights corresponding to the model to be evaluated deployed on the graphics processor.
9. The perplexity calculation device according to claim 6, characterized in that: When the graphics processor calculates the loss value according to the logic value and the symbol sequence, it is specifically used to: A loss value is obtained by performing cross entropy calculation based on the logic value and the symbol sequence.
10. The perplexity calculation device according to claim 6, characterized in that: The aggregation module aggregates the loss values respectively calculated by the multiple graphics processors into one graphics processor to obtain the aggregated loss value, specifically used to: aggregate the loss values respectively calculated by the multiple graphics processors into a target graphics processor among the multiple graphics processors to obtain the aggregated loss value; The target graphics processor is also used to calculate the perplexity of the model to be evaluated based on the aggregated loss value.
11. An electronic device, characterized in that: include: Memory; processor; as well as Computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to any one of claims 1 to 5.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.