Data screening method and device, large language model output acceleration method and device, medium and product

By employing a hardware-software co-processing method that utilizes multi-threaded beam parallel comparison of probability values ​​and baseline values ​​in an AI acceleration computing chip, the problem of high computational complexity in the large language model data screening process is solved, achieving efficient output acceleration.

CN121071124AActive Publication Date: 2025-12-05SHANGHAI SUIYUAN TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511613057.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2025-12-05
Estimated Expiration
2045-11-06

AI Technical Summary

Technical Problem

Existing large language models require full sorting when using TopK or TopP algorithms during data filtering, resulting in high computational complexity, long processing time, and reduced output efficiency.

Method used

By using multi-threaded parallel comparison of probability values ​​with baseline values ​​in an AI-accelerated computing chip, avoiding full sorting, and employing a hardware-software co-processing data filtering method, data that meets the criteria can be quickly filtered out.

Benefits of technology

It effectively reduces the complexity of data filtering algorithms, improves the output efficiency of large language models, and reduces waiting time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121071124A_ABST
    Figure CN121071124A_ABST
Patent Text Reader

Abstract

The invention discloses a data screening method and device, a large language model output acceleration method and device, a medium and a product. The data screening method comprises the following steps: selecting a thread bundle and determining a local comparison value of the thread bundle; triggering each thread bundle to execute the operation of comparing the value between each current probability value and the own baseline value, and updating the own key value by using each current probability value greater than the own baseline value; comparing the screening threshold of the target screening algorithm with the key value of each thread bundle; if the screening threshold is not matched with the key value of each thread bundle, updating a target set and a current probability value set according to the identified first thread bundle and second thread bundle; and if the screening threshold is matched with the key value of the third thread bundle, updating the target set as a screening result according to the third thread bundle. According to the technical scheme of the embodiment of the invention, the calculation complexity of a data screening algorithm is effectively reduced in a software-hardware cooperation mode, and parallel calculation hardware in an artificial intelligence acceleration calculation chip is utilized to the maximum extent.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence (AI), in particular to a data screening method, a large language model output acceleration method, equipment, a medium and a product. BACKGROUND

[0002] In natural language tasks, we usually use a pre-trained large language model to generate output text (such as an answer) according to the given input text (such as a question). In order to generate the output text, we need to let the model predict each word element (also known as token) of the required output step by step until the output termination condition is reached.

[0003] In the prior art, the large language model gives the probability distribution of all candidate words at each prediction step, indicating its prediction of the next word element. Then, through a preset data screening algorithm, such as TopK or TopP algorithm, a plurality of candidate words that meet the screening condition are selected from all candidate words, and then one of the selected candidate words is selected as a word element for output through a probability random selection method.

[0004] The inventors have found that, whether TopK or TopP algorithm is used in data screening, full sorting of each candidate word is required, the calculation complexity of the screening algorithm is high, the implementation of the entire data screening process is time-consuming, and thus the waiting time of the text output process of the large language model is long. SUMMARY

[0005] The embodiments of the present application provide a data screening method, a large language model output acceleration method, equipment, a medium and a product, which efficiently reduce the implementation complexity of the data screening algorithm in a soft and hard cooperative manner, and thus effectively accelerate the output of the large language model.

[0006] According to an aspect of the embodiments of the present application, a data screening method is provided, which comprises:

[0007] In response to a request for screening each data according to the probability value of the data using a target screening algorithm, a plurality of thread bundles are selected in an artificial intelligence acceleration computing chip according to a single comparison bit width, and a local comparison value corresponding to each thread bundle is determined;

[0008] A current probability value set is initialized according to the probability value of each data, and a baseline value of each thread bundle at a current iteration round is set according to each local comparison value, wherein the current probability value set contains at least one current probability value, and each baseline value has the same data bit width as each current probability value;

[0009] The parallel threads execute according to the data parallelism of the threads, compare the numerical values between the current probability values in the current probability value set and the baseline value of the threads, and update the key value of the threads using the current probability values greater than the baseline value of the threads;

[0010] After the end operation, the filtering threshold of the target filtering algorithm is compared with the key value of each thread bundle;

[0011] If the filtering threshold does not match the key value of each thread bundle, a first thread bundle corresponding to the minimum upper bound key value of the filtering threshold and a second thread bundle corresponding to the maximum lower bound key value of the filtering threshold are identified;

[0012] After the target data whose probability value satisfies the filtering condition is identified according to the first thread bundle and the second thread bundle and added to the target set and a new current probability value set is determined, the next iteration round is started as a new current iteration round, and the operation of setting the baseline value of each thread bundle in the current iteration round according to the local comparison value is performed.

[0013] If the filtering threshold matches the key value of the third thread bundle, the target data whose probability value satisfies the filtering condition is identified according to the third thread bundle and added to the target set, and the target set is taken as the filtering result.

[0014] According to another aspect of the embodiments of the present application, a large language model output acceleration method is also provided, which comprises:

[0015] Obtaining the normalized probability values corresponding to all candidate words obtained by the current inference of the large language model;

[0016] According to the target filtering algorithm preset by the large language model, the data filtering method described in any one of the embodiments of the present application is used to filter the target set that meets the target filtering algorithm from the candidate words;

[0017] After resetting the normalized probability values of the candidate words in the target set, the target candidate word is obtained by sampling in the target set according to the preset sampling rule, and the target candidate word is taken as the current round output of the model.

[0018] According to another aspect of the embodiments of the present application, a data filtering device is also provided, which comprises:

[0019] A thread bundle parameter determination module is configured to select a plurality of thread bundles in an artificial intelligence acceleration computing chip according to a single comparison bit width in response to a request for filtering data according to the probability values of the data using a target filtering algorithm, and determine a local comparison value corresponding to each thread bundle, respectively.

[0020] The baseline value determination module is configured to initialize a current probability value set according to probability values of the data, and set baseline values of the thread bundles in the current iteration round according to the local comparison values, wherein the current probability value set comprises at least one current probability value, and the baseline values have the same data bit width as the current probability values;

[0021] The thread bundle parallel execution module is configured to trigger the thread bundles to perform operations of comparing the current probability values in the current probability value set with the baseline values of the thread bundles, and updating key values of the thread bundles by using the current probability values greater than the baseline values of the thread bundles, in parallel according to the data parallelism of the thread bundles.

[0022] The key value comparison module is configured to compare the filtering threshold of the target filtering algorithm with the key values of the thread bundles after the operations are completed.

[0023] The adjacent thread bundle identification module is configured to identify a first thread bundle corresponding to a minimum upper bound key value of the filtering threshold and a second thread bundle corresponding to a maximum lower bound key value of the filtering threshold if the filtering threshold does not match the key values of the thread bundles.

[0024] The repeated iteration module is configured to start a next iteration round as a new current iteration round after the target data whose probability values satisfy the filtering condition are identified from the first thread bundle and the second thread bundle and added to the target set and a new current probability value set is determined, and return to perform the operation of setting the baseline values of the thread bundles in the current iteration round according to the local comparison values.

[0025] The end iteration module is configured to add the target data whose probability values satisfy the filtering condition to the target set according to the third thread bundle if the filtering threshold matches the key value of the third thread bundle, and take the target set as a filtering result.

[0026] According to another aspect of the embodiments of the present application, there is also provided a large language model output acceleration device, which comprises:

[0027] The filtering information acquisition module is configured to acquire normalized probability values corresponding to all candidate words obtained by current inference of the large language model.

[0028] The target set filtering module is configured to filter a target set meeting a target filtering algorithm from the candidate words by using the data filtering method according to the target filtering algorithm pre-set by the large language model.

[0029] The target candidate word output module is configured to reset the normalized probability values of the candidate words in the target set, sample target candidate words from the target set according to a preset sampling rule, and take the target candidate words as current round output of the model.

[0030] According to another aspect of the embodiments of the present application, an electronic device is also provided, which comprises:

[0031] at least one processor; and

[0032] a memory connected with the at least one processor in communication; wherein,

[0033] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the data screening method according to any one of the embodiments of the present application, or to perform the large language model output acceleration method according to any one of the embodiments of the present application.

[0034] According to another aspect of the embodiments of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the data screening method according to any one of the embodiments of the present application, or to implement the large language model output acceleration method according to any one of the embodiments of the present application when the computer instructions are executed by the processor.

[0035] According to another aspect of the embodiments of the present application, a computer program product is also provided, which comprises a computer program for enabling a processor to implement the steps of the data screening method or the large language model output acceleration method according to any one of the embodiments of the present application when the computer program is executed by the processor.

[0036] The technical solution of the embodiments of the present application can make full use of the rich thread bundle resources in the artificial intelligence acceleration computing chip, implement the data screening process based on the target screening algorithm, and through the algorithm improvement of comparing the probability value of each data with the baseline value of the thread bundle itself by multiple thread bundles in parallel, the target data meeting the data screening condition can be quickly screened out after one or several iterations. This new data screening method with software and hardware collaboration can effectively avoid full sorting of a large number of data to be screened, thereby effectively reducing the complexity of the existing data screening algorithm, and can maximize the use of the computing resources in the artificial intelligence acceleration computing chip. In addition, by applying the above data screening improvement scheme in the screening process of the token output of the large language model, the output of the large language model can be efficiently accelerated, and the token generation efficiency of the large language model can be improved.

[0037] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to make the technical solutions in the embodiments of the present application clearer, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and all other drawings obtained by those skilled in the art without any creative effort should be within the protection scope of the present application.

[0039] Figure 1 is a flow chart of a data screening method according to an embodiment of the present application;

[0040] Figure 2 is a schematic diagram of storing register data scattered according to a special data scattering instruction in continuous shared memory according to an embodiment of the present application;

[0041] Figure 3 is a schematic diagram of a token output process performed by a large language model in each output step in the prior art;

[0042] Figure 4 is a flow chart of a large language model output acceleration method according to an embodiment of the present application;

[0043] Figure 5 is a structural diagram of a data screening device according to an embodiment of the present application;

[0044] Figure 6 is a structural diagram of a large language model output acceleration device according to an embodiment of the present application;

[0045] Figure 7 is a structural diagram of an electronic device implementing the data screening or large language model output acceleration method according to an embodiment of the present application. DETAILED DESCRIPTION

[0046] In order to make the technical solutions in the embodiments of the present application clearer, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and all other drawings obtained by those skilled in the art without any creative effort should be within the protection scope of the present application.

[0047] It is to be understood that the terms "first", "second", and the like, used in the description and the claims of the application and the above drawings, are used to distinguish between similar objects, and are not necessarily used to describe a particular sequential or chronological order. It is to be understood that the use of the terms so used herein is interchangeable under appropriate circumstances such that the embodiments of the application described herein are, for example, capable of efficient implementation in other than the order illustrated or other than the order described herein. Moreover, the terms "comprise", "have" and any variations thereof are intended to cover a non-exclusive inclusion, for example, a process, method, system, product or apparatus that comprises a list of steps or units is not necessarily limited to those steps or units that are clearly recited, but can include other steps or units that are not expressly listed or inherent to such process, method, product or apparatus.

[0048] First, the scheme of each embodiment of the application is mainly used for the cooperative optimization of software and hardware of the calculation logic of a specific data screening algorithm (hereinafter referred to as a target screening algorithm). The target screening algorithm is specifically TopK or TopP algorithm.

[0049] Among them, the TopK algorithm refers to sorting all the data to be sorted according to a preset sorting value (for example, the probability value of the data), and selecting the top K data as the data screening result in the order of the sorting value from large to small. The TopP algorithm also sorts all the data to be sorted according to a preset sorting value, and starts to accumulate the sum of the sorting values of each data in the order of the sorting value from large to small. When the accumulated sum is greater than or equal to P value, each data participating in the accumulation is taken as the data screening result.

[0050] In order to facilitate understanding, first, the improvement scheme of each embodiment of the application for the above-mentioned target data screening algorithm is described.

[0051] Taking the sorting value as the probability value and the target screening algorithm as the TopK algorithm as an example, the inventors consider using the following algorithm to realize the data screening logic in order to avoid full sorting: assuming that the probability values of each data have X kinds of selectable numerical values, and then X kinds of baseline values can be set. Then, the probability value of each data can be compared with the above-mentioned X baseline values respectively, and the number of probability values greater than each baseline value is counted respectively. Then, the target baseline value matching the K value can be obtained from the above-mentioned number values. Finally, all the data whose probability value is greater than the target baseline value can be taken as the screening result meeting the screening condition of the TopK algorithm.

[0052] Figure 1A flowchart of a data screening method provided for an embodiment of the present application. The embodiment can be applicable to the case of optimizing the calculation complexity of a specific target screening algorithm in a soft and hard cooperative manner. The method can be executed by a data screening device, which can be realized in the form of hardware and / or software, and can generally be configured in an electronic device adapted to an artificial intelligence acceleration computing chip, such as a cloud server, an edge computing node, or various intelligent terminal devices. Correspondingly, as shown in Figure 1 the method can include:

[0053] S110, in response to a request for screening each data according to the probability value of the data using a target screening algorithm, selecting a plurality of thread bundles in an artificial intelligence acceleration computing chip according to a single comparison bit width, and determining a local comparison value corresponding to each thread bundle.

[0054] In the embodiment, the request for screening each data according to the probability value of the data using a target screening algorithm can be understood as a data screening request for screening one or more target data satisfying a screening condition from the plurality of data based on the probability value of each data to be sorted using a TopK algorithm or a TopP algorithm.

[0055] Each probability value of each data has a preset data bit width, such as 32 bits or 64 bits, etc. The data bit width is generally adapted to the hardware parameters of the artificial intelligence acceleration computing chip selected when screening data. The artificial intelligence acceleration computing chip can be a computing hardware specially designed for AI calculation, such as a GPU (Graphics Processing Unit, graphics processor), TPU (Tensor Processing Unit, tensor processor), or NPU (Neural Processing Unit, neural network processor) dedicated computing chip, etc.

[0056] In the embodiment, the inventor proposes an implementation manner of comparing the probability value of each data with various possible baseline values in parallel, and improves the implementation logic of the target screening algorithm. Based on the above improvement, the rich parallel computing resources of the artificial intelligence acceleration computing chip, i.e., the thread bundle, can be fully utilized, and the complex full sorting operation involved in the execution of the existing target screening algorithm can be avoided.

[0057] Among them, the thread bundle (referred to as Warp) is the basic execution unit in the artificial intelligence acceleration computing chip, which refers to a logical grouping composed of a fixed number of threads (typically 32). Among them, the threads in the same thread bundle can be executed synchronously, and the thread bundles in the artificial intelligence acceleration computing chip can also be executed synchronously. By utilizing the rich thread bundle resources in the artificial intelligence acceleration computing chip, the numerical comparison process between the probability value of each data and the baseline value can be efficiently realized.

[0058] In addition, the inventors further found that if all the baseline values are directly set for the complete data bit width of the probability value, the number of baseline values that need to be set at one time is too large. For example, for a 32-bit probability value, a total of baseline values are needed to cover all the selectable values of the probability value, and thread bundles are needed for parallel comparison, which is obviously difficult to implement. Based on this, the inventors consider that the complete probability value is compared multiple times locally based on a preset single comparison bit width. Correspondingly, the single comparison bit width can be understood as the bit width selected for comparison in a probability value at each time of local comparison.

[0059] In a specific example, when the single comparison bit width is 4 bits, the 4-bit data bits in each probability value that match the single comparison bit width can be compared with the baseline value in the order of comparison from high bit to low bit. In order to further simplify the comparison operation, the value of each probability value can be ensured not to change when the single comparison bit width is set, and the baseline value can be updated each time the comparison is performed.

[0060] Further, the local comparison value corresponding to each thread bundle can be set for the single comparison bit width. Among them, the selectable values of all local comparison values correspond to all selectable values of the single comparison bit width in the probability value.

[0061] Correspondingly, in an optional implementation of the embodiment, selecting a plurality of thread bundles in the artificial intelligence acceleration computing chip according to the single comparison bit width and determining the local comparison value corresponding to each thread bundle can include:

[0062] According to the single comparison bit width B and the formula , M thread bundles are selected in the artificial intelligence acceleration computing chip;

[0063] According to the selectable values of the binary number string matching the single comparison bit width B, the local comparison value corresponding to each thread bundle is determined.

[0064] In the optional embodiment, the number of thread bundles can be determined according to the single comparison bit width, for example, if the single comparison bit width is 4 bits, there are 16 optional values (i.e., local comparison values) corresponding to the single comparison bit width, and then 16 thread bundles can be set to simultaneously complete the value size comparison of each probability value and the baseline value corresponding to each local comparison value.

[0065] Then, as long as the baseline value of each thread bundle is updated based on the value of the single comparison bit width and the local comparison value of the different thread bundles in different iteration rounds, the baseline value of each thread bundle can be updated. For example, assuming that the probability value of each data is 8 bits, the single comparison bit width is B=2, and the number of thread bundles M is 4, then 4 thread bundles can be selected in the artificial intelligence acceleration computing chip, and the local comparison values of the above 4 thread bundles are set to 00, 01, 10 and 11 respectively.

[0066] At this time, for a thread bundle A with a local comparison value of 11, in the first value comparison (i.e., in the first iteration round), the baseline value of the thread bundle A can be set to 11000000, at this time, it is equivalent to comparing the value size of the baseline value with the 7th-8th bit of each probability value, then in the second value comparison, the baseline value of the thread bundle A can be updated and set to 11110000, at this time, it is equivalent to continuing to compare the value size of the baseline value with the 5th-6th bit of each probability value, and so on.

[0067] S120, initializing a current probability value set according to the probability values of each data, wherein the current probability value set contains at least one current probability value.

[0068] In the embodiment, since the single comparison bit width is set, multiple rounds of iterations may be required to finally obtain each item of data that meets the data screening condition. Then, the current probability value set can be initialized based on the probability values of all data, and after updating each current probability value in the current probability value set in one iteration, a new iteration round is started until the entire data screening process is completed.

[0069] S130, setting the baseline value of each thread bundle in the current iteration round according to each local comparison value, wherein each baseline value has the same data bit width as each current probability value.

[0070] In an optional embodiment of the embodiment, setting the baseline value of each thread bundle in the current iteration round according to each local comparison value can include:

[0071] The local comparison value of each thread bundle is left shifted (S-i*B) bits, and is accumulated and summed with the baseline value updated in the last iteration round of each thread bundle to obtain the baseline value of each thread bundle in the current i th iteration round.

[0072] The baseline value of each thread bundle is initialized as a full 0 binary string, and S is the data bit width of each current probability value.

[0073] As described above, after setting the local comparison bit width, each item of data meeting the target screening algorithm can be filtered out in sequence in a manner of comparing from high bits to low bits. Furthermore, the baseline value of each thread bundle can be updated by shifting based on the local comparison value, so that each baseline value matches the local data required to be compared between each probability value and each baseline value in each iteration round.

[0074] In a specific example, the data bit width of each current probability value is 8, the single comparison bit width is 4, the local comparison value of a certain thread bundle is 1101, in the first iteration round, the local comparison value of the thread bundle B needs to be left shifted by 4 bits (8-1*4) to obtain 11010000, and the 11010000 is accumulated and summed with the initialized 8-bit full 0 binary string 00000000 to obtain the baseline value of the thread bundle B in the first iteration round as 11010000.

[0075] Through the above setting, the numerical size relationship between the first 4 bits of each probability value and the baseline value of the thread bundle B can be compared in the first iteration round. Then, in the second iteration round, the local comparison value of the thread bundle B needs to be left shifted by 0 bits (8-2*4) to obtain 00001101, and the 00001101 is accumulated and summed with the baseline value 11010000 updated in the last iteration round to obtain the baseline value in the second iteration round as 11011101. Since the first four bits of the baseline value have been compared with each probability value, in the second iteration round, the numerical size relationship between the last 4 bits of each probability value and the new baseline value of the thread bundle B is compared.

[0076] S140, trigger each thread bundle to perform an operation of comparing the numerical size between each current probability value in the current probability value set and the baseline value of the thread bundle, and using each current probability value greater than the baseline value of the thread bundle to update the key value of the thread bundle.

[0077] In the embodiment, since the number of thread bundles matches the number of baseline values, the comparison of the numerical values between each current probability value and each baseline value can be started in parallel by triggering each thread bundle to execute simultaneously. Further, since each thread bundle includes multiple threads, the number of the multiple threads is the data parallelism of each thread bundle, for example, 16 or 32. The data parallelism is associated with a hardware parameter of the artificial intelligence acceleration computing chip, and the embodiment does not limit this.

[0078] Since multiple data tasks can be executed in parallel within a thread bundle, for N current probability values, the comparison of the numerical values between each current probability value and each baseline value can be completed in multiple internal iterations based on the data parallelism M of the thread bundle, where the number of internal iterations is the upward rounding value of (N / M).

[0079] Based on the above embodiments, the operation of triggering each thread bundle to execute the comparison of the numerical values between each current probability value and each baseline value according to the data parallelism of the thread bundle and updating the key value of the thread bundle using each current probability value greater than the baseline value can include:

[0080] The operation of triggering each thread bundle to execute the comparison of the numerical values between each current probability value and each baseline value according to the data parallelism of the thread bundle and updating the key value of the thread bundle using each current probability value greater than the baseline value can include:

[0081] S1401, sequentially obtaining multiple current probability values matching the data parallelism of the thread bundle from the current probability value set.

[0082] S1402, comparing the numerical values between each current probability value and each baseline value in parallel, and updating the key value of the thread bundle using each current probability value greater than the baseline value.

[0083] S1403, determining whether the processing of all current probability values in the current probability value set is completed, if yes, determining that the operation is ended; otherwise, returning to execute S1401.

[0084] In the alternative embodiment, each thread bundle can be regarded as a current thread bundle. From the perspective of each current thread bundle, the implementation process can include: the current thread bundle sequentially obtains multiple current probability values matching the data parallelism of the thread bundle from the current probability value set, the current thread bundle compares the numerical values between each current probability value and each baseline value in parallel, and updates the key value of the current thread bundle using each current probability value greater than the baseline value, the current thread bundle returns to execute the operation of sequentially obtaining multiple current probability values matching the data parallelism of the thread bundle from the current probability value set until the processing of all current probability values in the current probability value set is completed.

[0085] In one specific example, if the total number of current probability values contained in the current probability value set is 128, and the number of data parallelism within each thread bundle is 32, then within the thread bundle, the value comparison between 32 current probability values and the thread bundle's own baseline value can be performed at one time in each internal iteration, and then in one current iteration round, the value comparison between the thread bundle's own baseline value and all current probability values in the current probability value set needs to be completed through 4 internal iterations within each thread bundle.

[0086] In the embodiment, through the implementation of multi-thread bundle parallelism and multi-data parallelism within a single thread bundle, the rich parallel computing units in the artificial intelligence acceleration computing chip can be fully utilized to quickly compare the value between each current probability value in the current probability value set and each baseline value under each single comparison bit width.

[0087] After each thread bundle completes the value comparison between all current probability values and its own baseline value, the current probability values greater than its own baseline value can be counted accordingly, and the key value (Value) corresponding to each thread bundle can be updated according to each current probability value.

[0088] As described above, the technical solutions of the embodiments of the present application can be used to optimize the algorithm logic of the TopK algorithm or the TopP algorithm. The way of updating the key value corresponding to each thread bundle according to each current probability value is different for different data filtering algorithms.

[0089] In one optional implementation of the embodiment, the target filtering algorithm is the TopK algorithm that filters the top K data in the order of probability value from large to small; wherein K is the filtering threshold in the TopK algorithm.

[0090] Correspondingly, the current thread bundle updates its own key value using each current probability value greater than its own baseline value, which can include:

[0091] The current thread bundle accumulatively updates its own key value using the number of each current probability value greater than its own baseline value.

[0092] Specifically, since the data filtering target of the TopK algorithm is to filter the top K data in the order of probability value from large to small, then after each thread bundle compares its own baseline value updated at the current time with each current probability value respectively, the number of each current probability value greater than its own baseline value can be counted, and the key value of the thread bundle can be accumulatively updated according to the number.

[0093] In one specific example, when the single comparison bit width is 2 and the data bit width of each current probability value is 8, in the first iteration round, the value between each current probability value and the self baseline value of each thread bundle needs to be compared for the 7th-8th bit of each current probability value in the current probability value set. Assuming that the thread bundle C counts 10 current probability values greater than the self baseline value in the first iteration round, the self key value of the thread bundle C can be updated to 10 after the first iteration round.

[0094] Further, in the second iteration round, the value between each current probability value and the self baseline value of each thread bundle needs to be compared for the 5th-6th bit of each current probability value in the current probability value set. Assuming that the thread bundle C further counts 8 current probability values greater than the self baseline value in the second iteration round, the self key value of the thread bundle C can be updated to 10+8=18 after the second iteration round.

[0095] In another optional implementation of the embodiment, the target screening algorithm is a TopP algorithm that accumulatively sums the probability values of each data in descending order of the probability values until the accumulated probability value reaches P, where P is the screening threshold in the TopP algorithm.

[0096] Correspondingly, the current thread bundle uses each current probability value greater than the self baseline value to update the self key value, which can include:

[0097] The current thread bundle accumulatively sums each current probability value greater than the self baseline value and uses the accumulative sum result to accumulatively update the self key value.

[0098] Specifically, since the data screening target of the TopP algorithm is to screen one or more data whose accumulated probability value reaches P in descending order of the probability values, each thread bundle can count the accumulative sum result of each current probability value greater than the self baseline value in each iteration based on the single comparison bit width, and accumulatively update the self key value according to the accumulative sum result.

[0099] In another specific example, when the single comparison bit width is 2 and the data bit width of each current probability value is 8, in the first iteration round, the value of each current probability value is compared with the value of the self baseline value of each thread bundle for the 7th-8th bit of each current probability value in the current probability value set. Assuming that the thread bundle C counts the accumulated summation result of 10 current probability values greater than the self baseline value as 0.75 in the first iteration round, the self key value of the thread bundle C can be updated as 0.75 after the first iteration round. Further, in the second iteration round, the value of each current probability value is compared with the value of the self baseline value of each thread bundle for the 5th-6th bit of each current probability value in the current probability value set. Assuming that the thread bundle C counts the accumulated summation result of 8 current probability values greater than the self baseline value as 0.12 in the second iteration round, the self key value of the thread bundle C can be updated as 0.75+0.12=0.87 after the second iteration round.

[0100] In the embodiments of the present application, an implementation mode is given that in different iteration rounds, each thread bundle updates the new self key value based on the key value generated in the previous iteration round, and compares the new key value with the same filtering threshold (the aforementioned K value or P value) after each update. In fact, there is another optional implementation mode, that is, after each iteration round, the self key value of each thread bundle is cleared, and the filtering threshold used by each thread bundle in each iteration round is updated as the difference between the filtering threshold used in the previous iteration round and the statistical value (number or accumulated summation value) of each probability value in the target set screened out after the previous iteration round, which also achieves the comparison effect and obtains each item of data meeting the TopK or TopP filtering requirement.

[0101] S150, after the end operation, comparing the filtering threshold of the target filtering algorithm with the key value of each thread bundle: when the filtering threshold and the key value of each thread bundle do not match, performing S160; when the filtering threshold and the key value of the third thread bundle match, performing S180.

[0102] The judgment criteria of the end operation are that each thread bundle completes the complete process of comparing each current probability value in the current probability value set with the self baseline value and updating the self key value using all current probability values greater than the self probability value in the current probability value set.

[0103] After the above operation is completed, the filtering threshold of the target filtering algorithm can be compared with the key value of each thread bundle. Specifically, for the TopK algorithm, the filtering threshold is the K value, and for the TopP algorithm, the filtering threshold is the P value.

[0104] In the optional embodiment, if the screening threshold does not match the key value of each thread bundle, it means that in the numerical comparison of the current probability value and the baseline value for a single comparison bit width, the current comparison data accuracy is limited to the single comparison bit width, and the complete data screening process cannot be truly completed in the current iteration round. At this time, by executing S160-S170, the data that definitely satisfies the data screening condition is screened out and added to the target set, and after the current probability value set that needs to be iterated next time is re-determined, a new iteration round is started to compare the numerical size of each current probability value and the updated baseline value at a higher data accuracy.

[0105] In addition, if the screening threshold matches the key value of a certain specific thread bundle (i.e., the third thread bundle), for example, when the K value is 10, the key value of the thread bundle D is exactly 10, it means that in the current iteration round, the complete data screening process is realized. For another example, when the P value is 0.9, the numerical difference between the key value of the thread bundle D and 0.9 is less than or equal to the preset difference threshold, for example: 0.05%, at this time, it can also be explained that in the current iteration round, the complete data screening process is realized.

[0106] S160, identify the first thread bundle corresponding to the minimum upper bound key value of the screening threshold and the second thread bundle corresponding to the maximum lower bound key value of the screening threshold, and execute S170.

[0107] As described above, when the screening threshold does not match the key value of each thread bundle, it means that the screening threshold falls between the key values corresponding to the two thread bundles. It can be understood that the smaller the local comparison value of a thread bundle, the larger the key value corresponding to the thread bundle after one iteration round.

[0108] For example, assuming that the screening threshold is 0.75, by traversing the key values of each thread bundle, the local comparison values and the matching key values of the four thread bundles are respectively: thread bundle A: local comparison value 00, key value 0.9, thread bundle B: local comparison value 01, key value 0.8, thread bundle C: local comparison value 10, key value 0.7. Thread bundle D, local comparison value 11, key value 0.5.

[0109] By comparing the screening threshold with the key values of the above-mentioned thread bundles, it is found that the screening threshold falls between the key value 0.8 corresponding to the thread bundle B and the key value 0.7 corresponding to the thread bundle C, and then the key value 0.8 can be understood as the minimum upper bound key value of the screening threshold, and 0.7 can be understood as the maximum lower bound key value of the screening threshold. That is, the thread bundle C is the first thread bundle, and the thread bundle B is the second thread bundle.

[0110] Specifically, the maximum lower bound key value can be understood as the maximum value of all the key values less than the filtering threshold among all the key values corresponding to each thread bundle. Similarly, the minimum upper bound key value can be understood as the minimum value of all the key values greater than the filtering threshold among all the key values corresponding to each thread bundle.

[0111] S170, after the target data whose probability value satisfies the filtering condition is identified according to the first thread bundle and the second thread bundle and added to the target set and the new current probability value set is determined, the next iteration round is started as a new current iteration round, and the execution of S130 is returned.

[0112] In the optional embodiment, after the first thread bundle and the second thread bundle are identified, it can be known that the key value of the first thread bundle is the value closest to the filtering threshold among all the key values less than the filtering threshold. Therefore, all the current probability values greater than the first thread bundle baseline value satisfy the filtering condition matched with the target filtering algorithm, and at this time, all the current probability values greater than the first thread bundle baseline value can be added to the target set.

[0113] Meanwhile, since the key value of the second thread bundle is the value closest to the filtering threshold among all the key values greater than the filtering threshold. Therefore, in the subsequent filtering in the new iteration round, only the refined filtering needs to be performed on all the current probability values greater than the second thread bundle baseline value and less than the first thread bundle baseline value, and at this time, after all the current probability values greater than the second thread bundle baseline value and less than the first thread bundle baseline value are updated as the new current probability value set, the new iteration round is started, and the new baseline values of the thread bundles are determined again in the new iteration round, so that the data filtering process is finally ended when the filtering threshold is accurately matched with the key value of a specific thread bundle.

[0114] That is, in an optional embodiment of the present embodiment, identifying the target data whose probability value satisfies the filtering condition according to the first thread bundle and the second thread bundle and adding the target data to the target set and determining the new current probability value set can include:

[0115] According to the value size between each current probability value obtained by comparison and the first baseline value, each first type probability value greater than the first baseline value is obtained and added to the target set.

[0116] According to the value size between each current probability value obtained by comparison and the second baseline value, each second type probability value greater than the second baseline value is obtained, and the result obtained by filtering the first type probability value from each second type probability value is taken as the new current probability value set.

[0117] S180, the target data whose probability value meets the screening condition is identified according to the third thread bundle and is added to the target set, and the target set is taken as the screening result.

[0118] As described above, the complete data screening process can be completed based on the third thread bundle, at this time, each current probability value greater than the third thread bundle baseline value in the current probability value set can be acquired first, and each item of data matched with the above-mentioned current probability value is added to the target set, and is taken as the final data screening result together with the original data in the target set.

[0119] Correspondingly, in an optional implementation of the embodiment, the target data whose probability value meets the screening condition is identified according to the third thread bundle and is added to the target set, which can include:

[0120] According to the numerical value between each current probability value compared by the third thread bundle and the third baseline value, each third type probability value greater than the third baseline value is acquired and is added to the target set.

[0121] The technical scheme of the embodiment of the application can make full use of the rich thread bundle resources in the artificial intelligence acceleration computing chip, realize the data screening process based on the target screening algorithm, and through the algorithm improvement of the parallel comparison of each data probability value and the baseline value of the thread bundle by multiple thread bundles, the target data meeting the data screening condition can be quickly screened after one or several iterations, the new data screening method of software and hardware cooperation effectively avoids the full sorting of a large number of to-be-screened data, and thus the complexity of data screening is effectively reduced, and the computing resources in the artificial intelligence acceleration computing chip can be maximized.

[0122] After the software and hardware improvement based on the target screening algorithm, the inventors further find that the entire data screening process needs to frequently use registers for auxiliary execution. Further, the inventors creatively construct a special instruction for improving the register execution efficiency to maximize the above-mentioned data screening efficiency. Correspondingly, the register operation involved in the entire data screening process is further described below in combination with the use of the register.

[0123] Correspondingly, after the baseline value of each thread bundle in the current iteration round is set according to each local comparison value, the method can further include:

[0124] The baseline value of each thread bundle in the current iteration round is stored in the matched baseline register respectively, wherein the baseline register is a scalar register or a vector register.

[0125] In addition, the current thread bundle sequentially acquires multiple current probability values matched with the data parallel number of the current thread bundle from the current probability value set, which can include:

[0126] The current thread bundle sequentially obtains a plurality of current probability values in the current probability value set according to the data parallelism of the current thread bundle, and stores each current probability value obtained in the probability value register, wherein the probability value register is a vector register.

[0127] Through the above setting, in each iteration, the baseline register and the probability value register can be directly operated by using the specially constructed special instruction to achieve the purpose of quickly comparing data.

[0128] On the basis of the above embodiments, the current thread bundle compares the numerical values between each current probability value and the baseline value of the current thread bundle in parallel, and uses each current probability value greater than the baseline value of the current thread bundle to update the key value of the current thread bundle, which can include:

[0129] The current thread bundle obtains a target baseline register for storing the baseline value of the current thread bundle;

[0130] The current thread bundle calls a special comparison setting instruction matched with the artificial intelligence acceleration computing chip, compares the numerical values between each current probability value stored in the probability value register and the baseline value of the current thread bundle stored in the target baseline register in parallel, and stores the comparison result in a target mask register;

[0131] The target mask register is a scalar register, and when the target current probability value is greater than the baseline value of the current thread bundle, the storage location matched with the target current probability value in the target mask register is set to 1, otherwise, the storage location is set to 0.

[0132] The current thread bundle obtains each target current probability value matched with each storage location set to 1 in the target mask register, and updates the key value of the current thread bundle according to each target current probability value.

[0133] In this embodiment, according to the hardware characteristics of the artificial intelligence acceleration computing chip, a special comparison setting instruction is constructed, which can realize the fast comparison between a plurality of current probability values and the baseline value of the same thread bundle at a time, so as to effectively improve the algorithm execution efficiency.

[0134] Specifically, the instruction format of the special comparison setting instruction can be as shown in Table 1 or Table 2.

[0135] Table 1

[0136]

[0137] Table 2

[0138]

[0139] The first row in Table 1 and Table 2 refers to the instruction bits of [31:0], which represents that the special comparison set instruction is a 32-bit instruction. The second row refers to the physical meaning of the filled data under different instruction bits, for example, the filled "0x40" at the position of [6:0] represents that the first 7 bits of the special comparison set instruction are the instruction identifier of the special comparison set instruction. SRD[6:0] represents a mask register described by 7-bit data, which is a scalar register, VRS[7:0] represents a probability value register described by 8-bit data, which is a vector register, VRT[7:0] represents a baseline register described by 8-bit data, which is a vector register, SRT[6:0] represents a baseline register described by 7-bit data, which is a scalar register. NA represents meaningless data bits.

[0140] Further, according to the numerical size between each current probability value obtained by comparing the first thread bundle and the first baseline value, each first type probability value greater than the first baseline value is obtained and added to the target set, which can further include:

[0141] The first target mask register is updated after the first thread bundle compares the numerical size between each current probability value in the current probability value set and the baseline value of itself;

[0142] Each first type probability value matched with each storage position in which the first target mask register is set to 1 is added to the target set;

[0143] Correspondingly, according to the numerical size between each current probability value obtained by comparing the second thread bundle and the second baseline value, each second type probability value greater than the second baseline value can be obtained, which can include:

[0144] The second target mask register is updated after the second thread bundle compares the numerical size between each current probability value in the current probability value set and the baseline value of itself;

[0145] Each second type probability value matched with each storage position in which the second target mask register is set to 1 is obtained.

[0146] Further, after obtaining each second type probability value greater than the second baseline value, the result obtained by filtering each first type probability value in each second type probability value can be taken as a new current probability value set.

[0147] In the optional embodiment, the operation of filtering the first-type probability values from the second-type probability values can be implemented by directly performing a simple XOR operation on the data stored in the first target mask register and the second target mask register. After the XOR operation is completed, the current probability values that are set to 1 can be used to form a new set of current probability values. Through the above arrangement, the data processing logic can be further simplified, and the data screening efficiency can be improved.

[0148] On the basis of the above embodiments, after the results obtained by filtering the first-type probability values from the second-type probability values are used as a new set of current probability values, the following operations can also be included:

[0149] A special data scattering instruction matched with the artificial intelligence acceleration computing chip is called to store the results obtained by filtering the first-type probability values from the second-type probability values in the shared memory in a continuous manner, so that in the subsequent iteration rounds, the new set of current probability values can be loaded in an address-continuous manner by using linear loading, thereby reducing the data access delay and accelerating the data screening process.

[0150] In the optional embodiment, it is considered that the probability values corresponding to all the data obtained at the beginning are stored in the shared memory in a continuous manner. Further, the set of current probability values used in the first iteration round can be loaded into the probability value register in a continuous address space in a linear loading manner. However, from the second iteration round, the current probability values used as the new set of current probability values have completed a data screening process, and the next round of iteration data is no longer stored in the shared memory in a continuous manner. Therefore, when the operation of obtaining the current probability values in the updated set of current probability values is performed, linear loading can no longer be performed, which will cause a large data access delay.

[0151] Based on the above problem, the inventors further construct a special data scattering instruction matched with the hardware features of the artificial intelligence acceleration computing chip. After each iteration round determines the new current probability values used as the new set of current probability values, the special data scattering instruction is called to store the address-discrete new current probability values in the register in a scattered storage manner in the continuous shared memory. Then, the new set of current probability values composed of the new current probability values can continue to be loaded in an address-continuous manner by using linear loading, so as to quickly store the new set of current probability values into the probability value register, thereby effectively reducing the data access delay and accelerating the data screening process. The instruction format of the special data scattering instruction can be as shown in Table 3.

[0152] Table 3

[0153]

[0154] Specifically, the above-mentioned special data scatter instruction is also a 32-bit instruction, the first 7 bits of the instruction are identified by 0x30 as the instruction type of the special data scatter instruction. Specifically, in Figure 2 The schematic diagram of storing the register data scatter based on the special data scatter instruction in the continuous shared memory is shown in

[0155] As shown in Figure 2 , the special data scatter instruction is used to locate the matching current probability value (i.e., T0, T3 and T5) in the VRT[7:0] vector register according to the storage position identified as 1 in the SRS[6:0] scalar register, and to continuously store the located current probability values T0, T3 and T5 at the matching shared memory address according to the shared memory address (which can also be represented as an address offset) specified in the SRD[6:0] scalar register.

[0156] Further, the shared memory address specified in the SRD[6:0] can be calculated by incrementing the address by the total number of current probability values of each data scatter in byte units, and the calculated address increment is added to the shared memory address used at the last data scatter to serve as the starting address for the next time. Wherein, INC can be 1 or 0, when INC is 1, it means that the accumulation of the new shared memory address needs to be updated.

[0157] Further, after improving the above-mentioned data screening method, the improved data screening algorithm can be directly applied to the screening stage of each candidate word generated by the large language model. Wherein, the token output flowchart of the existing large language model at each output step is shown in Figure 3

[0158] As shown in Figure 3 , taking the TopP algorithm as an example, the token output process of the large language model is as follows:

[0159] 1. The large language model obtains the probability value (also referred to as logits) of each candidate word at each output step;

[0160] 2. Normalize the probability value of each candidate word between 0 and 1;

[0161] 3. Fully sort the normalized probability values in descending order;

[0162] 4. Based on the algorithm logic of the TopP algorithm, the cumulative value of each probability value is counted in descending order, and the probability value that meets the condition of reaching P is marked with a mask value "1" according to the screening threshold P.​

[0163] 5. Take each probability value marked as "1" in the mask value as a data screening result based on the TopP algorithm;

[0164] 6. Re-normalize each probability value in the data screening result to between 0 and 1;

[0165] 7. Sample and obtain a target probability value from the re-normalized probability values in a random selection manner;

[0166] 8. The large language model outputs a target candidate word corresponding to the target probability value as an output token.

[0167] It can be understood that steps 3-5 (the dashed box part in Figure 3 ) in the above operations are the TopP algorithm logic. Further, steps 3-5 can be optimized and improved using the data screening method described in embodiments of the present application. Since the full sorting operation is completely avoided in the above optimization and improvement process, the algorithm complexity of the TopP algorithm can be reduced from O(n*log(n)) to O(n) to achieve efficient large language model output acceleration.

[0168] Correspondingly, Figure 4 a flowchart of a large language model output acceleration method provided by an embodiment of the present application. The embodiment can be applicable to efficiently accelerating the process of outputting a token of a large language model in a soft and hard cooperative manner. The method can be executed by a large language model output acceleration device. The device can be realized in the form of hardware and / or software and can generally be configured in an electronic device adapted to an artificial intelligence acceleration computing chip, such as a cloud server, an edge computing node, or various intelligent terminal devices. Correspondingly, as Figure 4 shown, the method can include:

[0169] S410. Obtain normalized probability values corresponding to all candidate words obtained by current reasoning of a large language model.

[0170] Among them, all candidate words obtained by current reasoning of the large language model are all candidate words that can be used as alternative output tokens obtained by the large language model at the current output step. At the current output step, the large language model finally selects one candidate word from the above all candidate words as an output token for model output.

[0171] Specifically, the large language model reasons a matching probability value for each candidate word. By normalizing the above probability values, the normalized probability value of each candidate word can be obtained.

[0172] S420, according to the target filtering algorithm preset by the large language model, the data filtering method is adopted, and the target set meeting the target filtering algorithm is filtered from the candidate words.

[0173] The normalized probability value of each candidate word is equivalent to the probability value of each data used in the data filtering method of each embodiment, and the normalized probability value of each candidate word is used as the original input, so that each candidate word meeting the data filtering condition of the target filtering algorithm can be filtered from all candidate words and added to the target set.

[0174] S430, after resetting the normalized probability value of each candidate word in the target set, the target candidate word is sampled according to the preset sampling rule, and the target candidate word is used as the current round output of the model.

[0175] The target candidate word is the token output by the large language model at the current output step, that is, token. In addition, the sampling rule can be a random sampling rule.

[0176] The technical scheme of the embodiment of the application can efficiently realize the output acceleration of the large language model by applying the data filtering method improved by the cooperation of software and hardware in the filtering process of the output token of the large language model, and effectively improves the token generation efficiency of the large language model.

[0177] It needs to be emphasized again that the focus of each embodiment of the application is to propose a hardware and software architecture for the large language model generation strategy TopK and TopP, which can be integrated into the existing deep learning framework to realize the innovation. The point is that without full sorting operation, each candidate word meeting the TopK or TopP data filtering condition can be obtained, so that the decoding sampling time is greatly reduced, and the efficiency is greatly improved compared with the existing acceleration card. At the same time, through the cooperation of software and hardware, there are two innovations in algorithm optimization and instruction design: 1. Using hardware multithreading technology, combining efficient comparison setting and efficient data scattering, the hardware resource utilization rate of candidate sampling is improved, and the delay of candidate word filtering is reduced.

[0178] More specifically, the main innovations include:

[0179] 1. The embodiments of the application propose an acceleration hardware and software architecture for TopK or TopP, which uses multithreading technology, and each thread bundle is fixedly responsible for a specific baseline value. One iteration can determine the B-bit baseline value, where the size of B is determined by the number of thread bundles participating in the calculation. From high to low, iteration is performed in turn, and finally the TopP or TopK set meeting the requirements is filtered out.

[0180] 2. Each embodiment of the present application also designs a high-efficiency comparison and setting instruction, each thread bundle can compare a plurality of data parallel to the data of the thread bundle itself with the baseline value of the thread bundle itself at a time, and then write the mask result of the corresponding bit into the corresponding bit of the scalar register, thereby saving the storage occupation, and combining with the subsequent instructions to realize the efficient use of hardware efficiency.

[0181] 3. Each embodiment of the present application also constructs a high-efficiency data scattering method, which scatters the discrete address data of the marked specific mask to the continuous storage space in the shared memory, realizes the continuous storage of the screening data, guarantees the linear loading of the specific data in the next round of iteration, and reduces the delay.

[0182] Figure 5 A structural schematic diagram of a data screening device provided by an embodiment of the present application is shown in FIG. 1. Figure 5 As shown in the figure, the device comprises a thread bundle parameter determination module 510, a baseline value determination module 520, a thread bundle parallel execution module 530, a key value comparison module 540, an adjacent thread bundle identification module 550, a repeated iteration module 560, and an end iteration module 570, wherein:

[0183] The thread bundle parameter determination module 510 is configured to select a plurality of thread bundles in an artificial intelligence acceleration calculation chip according to a single comparison bit width in response to a request for screening each data according to a probability value of the data, and determine a local comparison value corresponding to each thread bundle.

[0184] The baseline value determination module 520 is configured to initialize a current probability value set according to the probability value of each data, and set a baseline value of each thread bundle in a current iteration round according to each local comparison value, wherein the current probability value set contains at least one current probability value, and each baseline value has the same data bit width as each current probability value.

[0185] The thread bundle parallel execution module 530 is configured to trigger each thread bundle to perform an operation of comparing a numerical value between each current probability value in the current probability value set and a baseline value of the thread bundle itself according to a data parallel number of the thread bundle itself in parallel, and updating a key value of the thread bundle itself using each current probability value greater than the baseline value of the thread bundle itself.

[0186] The key value comparison module 540 is configured to compare a screening threshold of the target screening algorithm with the key value of each thread bundle after the end operation.

[0187] The adjacent thread bundle identification module 550 is configured to identify a first thread bundle corresponding to a maximum lower bound key value of the screening threshold and a second thread bundle corresponding to a minimum upper bound key value of the screening threshold if the screening threshold does not match the key value of each thread bundle.

[0188] The repeated iteration module 560 is configured to, after identifying the target data whose probability value meets the screening condition according to the first thread bundle and the second thread bundle and adding the target data to the target set and determining the new current probability value set, start the next iteration round as a new current iteration round, and return to perform the operation of setting the baseline value of each thread bundle in the current iteration round according to each local comparison value.

[0189] The end iteration module 570 is configured to, if the screening threshold matches the key value of the third thread bundle, identify the target data whose probability value meets the screening condition according to the third thread bundle, add the target data to the target set, and take the target set as the screening result.

[0190] The technical scheme of the embodiment of the application can make full use of the rich thread bundle resources in the artificial intelligence acceleration computing chip, implement the data screening process based on the target screening algorithm, and improve the algorithm by comparing the probability value of each data with the baseline value of the thread bundle itself in parallel through multiple thread bundles. After one or several iterations, the target data meeting the data screening condition can be quickly screened out. The new data screening method of software and hardware cooperation effectively avoids full sorting of a large amount of data to be screened, thereby effectively reducing the complexity of data screening, and can maximize the use of the computing resources in the artificial intelligence acceleration computing chip.

[0191] On the basis of the above embodiments, the thread bundle parameter determination module 510 can be specifically configured to:

[0192] determine the local comparison value corresponding to each thread bundle according to the single comparison bit width B and the formula select M thread bundles in the artificial intelligence acceleration computing chip;

[0193] determine the local comparison value corresponding to each thread bundle according to the single comparison bit width B and the formula

[0194] On the basis of the above embodiments, the baseline value determination module 520 can be specifically configured to:

[0195] left shift the local comparison value of each thread bundle by (S-i*B) bits, and then add and sum the baseline value of each thread bundle updated in the last iteration round to obtain the baseline value of each thread bundle in the current i th iteration round;

[0196] wherein the baseline value of each thread bundle is initialized as a full-0 binary number string, and S is the data bit width of each current probability value.

[0197] On the basis of the above embodiments, the repeated iteration module 560 can be specifically configured to:

[0198] According to the numerical value between each current probability value obtained by comparison of the first thread bundle and the first baseline value, each first type probability value greater than the first baseline value is obtained and added to the target set;

[0199] According to the numerical value between each current probability value obtained by comparison of the second thread bundle and the second baseline value, each second type probability value greater than the second baseline value is obtained, and the result obtained by filtering each first type probability value in each second type probability value is taken as a new current probability value set;

[0200] Correspondingly, the end iteration module 570 can be specifically used for:

[0201] According to the numerical value between each current probability value obtained by comparison of the third thread bundle and the third baseline value, each third type probability value greater than the third baseline value is obtained and added to the target set.

[0202] On the basis of each of the above embodiments, the thread bundle parallel execution module 530 can be specifically used for:

[0203] Parallelly triggering each thread bundle to respectively perform the following operations:

[0204] Sequentially obtaining, from the current probability value set, a plurality of current probability values matching the data parallel number of itself;

[0205] Parallelly comparing the numerical value between each current probability value and the baseline value of itself, and using each current probability value greater than the baseline value of itself to update the key value of itself;

[0206] Returning to perform the operation of sequentially obtaining, from the current probability value set, a plurality of current probability values matching the data parallel number of itself until the processing of all current probability values in the current probability value set is completed.

[0207] On the basis of each of the above embodiments, the target screening algorithm can be a TopK algorithm for screening the top K data in the order of probability value from large to small; wherein K is a screening threshold in the TopK algorithm.

[0208] Correspondingly, the thread bundle parallel execution module 530 can be further used for:

[0209] The current thread bundle uses the number of each current probability value greater than the baseline value of itself to accumulate and update the key value of itself.

[0210] On the basis of each of the above embodiments, the target screening algorithm can be a TopP algorithm for accumulating and summing the probability values of each data in the order of probability value from large to small until the accumulated probability value reaches P; wherein P is a screening threshold in the TopP algorithm.

[0211] Correspondingly, the thread bundle parallel execution module 530 can be further used for:

[0212] The current thread bundle accumulates and sums each current probability value greater than the baseline value of the current thread bundle, and accumulatively updates the key value of the current thread bundle using the accumulated sum result.

[0213] On the basis of each of the above embodiments, the baseline register storage module can be further used for:

[0214] After setting the baseline value of each thread bundle at the current iteration round according to each local comparison value, the baseline value of each thread bundle at the current iteration round is stored in the matched baseline register respectively, wherein the baseline register is a scalar register or a vector register;

[0215] Correspondingly, the thread bundle parallel execution module 530 can be further used for:

[0216] The current thread bundle sequentially obtains a plurality of current probability values in the current probability value set according to the data parallelism degree of the current thread bundle, and stores each current probability value obtained in the probability value register, wherein the probability value register is a vector register.

[0217] On the basis of each of the above embodiments, the thread bundle parallel execution module 530 can be further used for:

[0218] The current thread bundle obtains a target baseline register for storing the baseline value of the current thread bundle;

[0219] The current thread bundle calls a special comparison and setting instruction matched with the artificial intelligence acceleration calculation chip, compares the numerical value between each current probability value stored in the probability value register and the baseline value of the current thread bundle stored in the target baseline register in parallel, and stores the comparison result in the target mask register;

[0220] The target mask register is a scalar register, and when the target current probability value is greater than the baseline value of the current thread bundle, the storage location matched with the target current probability value in the target mask register is set to 1, otherwise, the storage location is set to 0.

[0221] The current thread bundle obtains each target current probability value matched with each storage location set to 1 in the target mask register, and updates the key value of the current thread bundle according to each target current probability value.

[0222] On the basis of each of the above embodiments, the repeated iteration module 560 can be further used for:

[0223] Obtain the first target mask register updated by the first thread bundle after comparing the numerical value between each current probability value in the current probability value set and the baseline value of the first thread bundle;

[0224] Match each first type of probability value matched with each storage location set to 1 in the first target mask register to the target set;

[0225] After the second thread bundle obtains the numerical value after comparing each current probability value in the current probability value set with the baseline value of itself, the second target mask register is updated;

[0226] Obtain each second type of probability value matched with each storage location set to 1 in the second target mask register.

[0227] On the basis of each of the above embodiments, the data spreading module can also be included, which is used to, after the results obtained after filtering each first type of probability value in each second type of probability value are used as a new current probability value set, call a special data spreading instruction matched with the artificial intelligence acceleration computing chip to continuously store the results obtained after filtering each first type of probability value in each second type of probability value in the shared memory, so as to continuously load the new current probability value set in the subsequent round iteration in a linear manner, reduce the data access delay, and accelerate the data screening process.

[0228] The data screening device provided by the embodiments of the present application can execute the data screening method provided by any of the embodiments of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0229] Figure 6 A structural schematic diagram of a large language model output acceleration device provided by an embodiment of the present application is shown in FIG. 6. Figure 6 As shown in the figure, the device includes a screening information acquisition module 610, a target set screening module 620, and a target candidate word output module 630, wherein:

[0230] The screening information acquisition module 610 is used to acquire normalized probability values corresponding to all candidate words obtained by the current reasoning of the large language model;

[0231] The target set screening module 620 is used to screen a target set that meets a target screening algorithm from each candidate word according to the target screening algorithm pre-set by the large language model by using the data screening method described in any one of the embodiments of the present application.

[0232] The target candidate word output module 630 is used to sample a target candidate word from the target set according to a pre-set sampling rule after resetting the normalized probability values of each candidate word in the target set, and output the target candidate word as the current round output of the model.

[0233] The technical solution of the embodiment of the present application can efficiently realize the output acceleration of the large language model and effectively improve the word generation efficiency of the large language model by applying the data screening algorithm improved through the cooperation of software and hardware in the screening process of the output word element of the large language model.

[0234] The large language model output acceleration device provided by the embodiment of the present application can execute the large language model output acceleration method provided by any embodiment of the present application, and has the corresponding function modules and beneficial effects of the execution method.

[0235] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0236] Figure 7 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.

[0237] As shown in Figure 7 The electronic device 10 includes at least one artificial intelligence acceleration computing chip 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is in communication connection with the at least one artificial intelligence acceleration computing chip 11, wherein the memory stores a computer program that can be executed by the at least one artificial intelligence acceleration computing chip, and the artificial intelligence acceleration computing chip 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The artificial intelligence acceleration computing chip 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0238] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.

[0239] The artificial intelligence acceleration computing chip 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the artificial intelligence acceleration computing chip 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various artificial intelligence acceleration computing chips running machine learning model algorithms, a digital signal artificial intelligence acceleration computing chip (DSP), and any appropriate artificial intelligence acceleration computing chip, controller, microcontroller, etc. The artificial intelligence acceleration computing chip 11 performs various methods and processes described above, such as performing the data screening method or the large language model output acceleration method as described in embodiments of the present application.

[0240] In some embodiments, the data screening method or the large language model output acceleration method as described in embodiments of the present application can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the artificial intelligence acceleration computing chip 11, one or more steps of the data screening method or the large language model output acceleration method as described in embodiments of the present application described above can be performed. Alternatively, in other embodiments, the artificial intelligence acceleration computing chip 11 can be configured to perform the data screening method or the large language model output acceleration method as described in embodiments of the present application by any other appropriate means (e.g., by means of firmware), wherein:

[0241] The data screening method comprises: in response to a request for screening each data according to a probability value of the data using a target screening algorithm, selecting a plurality of thread bundles in the artificial intelligence acceleration computing chip according to a single comparison bit width, and determining a local comparison value corresponding to each thread bundle, respectively;

[0242] Initializing a current probability value set according to the probability values of the data, and setting baseline values of each thread bundle at a current iteration round according to the local comparison values, wherein the current probability value set contains at least one current probability value, and each baseline value has the same data bit width as each current probability value;

[0243] Triggering each thread bundle to perform, in parallel, an operation of comparing a numerical value between each current probability value in the current probability value set and the baseline value of itself according to a data parallelism number of itself, and using each current probability value greater than the baseline value of itself to update a key value of itself;

[0244] After the end operation, comparing a screening threshold of the target screening algorithm with the key values of each thread bundle;

[0245] If the screening threshold does not match the key value of each thread bundle, a first thread bundle corresponding to the maximum lower bound key value of the screening threshold and a second thread bundle corresponding to the minimum upper bound key value of the screening threshold are identified;

[0246] After each target data whose probability value meets the screening condition is identified according to the first thread bundle and the second thread bundle and added to the target set and a new current probability value set is determined, a next iteration round is started as a new current iteration round, and the operation of setting the baseline value of each thread bundle under the current iteration round according to each local comparison value is returned to be executed.

[0247] If the screening threshold matches the key value of the third thread bundle, the target data whose probability value meets the screening condition is identified according to the third thread bundle and added to the target set, and the target set is taken as the screening result.

[0248] In addition, the large language model output acceleration method comprises:

[0249] obtaining normalized probability values corresponding to all candidate words obtained by current inference of the large language model;

[0250] According to a target screening algorithm preset by the large language model, a data screening method as described in any one of the embodiments of the present application is used to screen a target set meeting the target screening algorithm from the candidate words;

[0251] After resetting the normalized probability values of the candidate words in the target set, a target candidate word is sampled from the target set according to a preset sampling rule, and the target candidate word is taken as the current round output of the model.

[0252] The various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), an application specific standard product (ASSP), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being embodied in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable artificial intelligence acceleration computing chip, which can be a special-purpose or general-purpose programmable artificial intelligence acceleration computing chip, can receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.

[0253] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a human operator of a general purpose computer, a special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed, enables the functions / operations specified in the flow charts and / or block diagrams to be implemented. The computer program can execute entirely on a machine, partly on the machine, as a stand-alone software package, partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0254] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0255] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0256] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0257] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0258] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0259] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method of data screening, characterized by, The method comprises the following steps: In response to a request for screening data according to the probability value of the data using a target screening algorithm, a plurality of thread bundles are selected in an artificial intelligence acceleration computing chip according to a single comparison bit width, and a local comparison value corresponding to each thread bundle is determined; A current probability value set is initialized according to the probability value of each data, and a baseline value of each thread bundle in a current iteration round is set according to each local comparison value, wherein the current probability value set contains at least one current probability value, and the data bit width of each baseline value is the same as that of each current probability value; Each thread bundle is triggered in parallel to perform an operation of comparing the numerical value between each current probability value in the current probability value set and the baseline value of the thread bundle itself according to the data parallelism of the thread bundle itself, and updating the key value of the thread bundle itself using each current probability value greater than the baseline value of the thread bundle itself; After the operation is completed, the screening threshold of the target screening algorithm is compared with the key value of each thread bundle; If the screening threshold does not match the key value of each thread bundle, a first thread bundle corresponding to the maximum lower bound key value of the screening threshold and a second thread bundle corresponding to the minimum upper bound key value of the screening threshold are identified; After the target data whose probability value meets the screening condition is identified according to the first thread bundle and the second thread bundle and added to a target set and a new current probability value set is determined, a next iteration round is started as a new current iteration round, and the operation of setting the baseline value of each thread bundle in the current iteration round according to each local comparison value is performed again; If the screening threshold matches the key value of the third thread bundle, the target data whose probability value meets the screening condition is identified according to the third thread bundle and added to the target set, and the target set is taken as a screening result.

2. The method of claim 1, wherein, Selecting a plurality of thread bundles in an artificial intelligence acceleration computing chip according to a single comparison bit width and determining a local comparison value corresponding to each thread bundle comprises: According to single comparison bit width B, and formula Select M thread bundles in the artificial intelligence acceleration computing chip; Determining the local comparison value corresponding to each thread bundle according to each selectable value of a binary number string matching the single comparison bit width B.

3. The method of claim 2, wherein, Setting the baseline value of each thread bundle in a current iteration round according to each local comparison value comprises: After the local comparison value of each thread bundle is left shifted by (S-i*B) bits, the left shifted value is added to the baseline value of each thread bundle updated in a previous iteration round to obtain the baseline value of each thread bundle in the current i-th iteration round; Wherein, the baseline value of each thread bundle is initialized as a binary number string of all 0s, and S is the data bit width of each current probability value.

4. The method of claim 1, wherein, Identifying the target data whose probability value meets the screening condition according to the first thread bundle and the second thread bundle and adding the target data to a target set and determining a new current probability value set comprises: Obtaining each first type probability value greater than the first baseline value according to the numerical value between each current probability value and the first baseline value obtained by the comparison of the first thread bundle, and adding the first type probability value to the target set; Obtaining each second type probability value greater than the second baseline value according to the numerical value between each current probability value and the second baseline value obtained by the comparison of the second thread bundle, and taking the result obtained by filtering the first type probability value from the second type probability value as the new current probability value set; Correspondingly, adding the target data whose probability value meets the screening condition to the target set according to the third thread bundle comprises: According to the numerical size between each current probability value in the third thread bundle comparison and the third baseline value, each third type probability value greater than the third baseline value is obtained and added to the target set.

5. The method of claim 4, wherein, The parallel triggering of each thread bundle performs the operation of comparing the numerical size between each current probability value in the current probability value set and the baseline value of the thread bundle itself, and updating the key value of the thread bundle itself using each current probability value greater than the baseline value of the thread bundle itself, including: The parallel triggering of each thread bundle respectively performs the following operations: Sequentially obtaining, from the current probability value set, multiple current probability values matching the data parallel number of the thread bundle itself; Parallel comparison of the numerical size between each current probability value and the baseline value of the thread bundle itself, and updating the key value of the thread bundle itself using each current probability value greater than the baseline value of the thread bundle itself; Returning to the operation of sequentially obtaining, from the current probability value set, multiple current probability values matching the data parallel number of the thread bundle itself until the processing of all current probability values in the current probability value set is completed.

6. The method of claim 5, wherein, The target screening algorithm is a TopK algorithm that screens the top K data in descending order of the probability values of the data; wherein K is the screening threshold in the TopK algorithm. Correspondingly, the current thread bundle updates the key value of the thread bundle itself using each current probability value greater than the baseline value of the thread bundle itself, including: The current thread bundle accumulatively updates the key value of the thread bundle itself using the number of each current probability value greater than the baseline value of the thread bundle itself.

7. The method of claim 5, wherein, The target screening algorithm is a TopP algorithm that accumulatively sums the probability values of each data in descending order of the probability values of the data until the accumulated probability value reaches P; wherein P is the screening threshold in the TopP algorithm. Correspondingly, the current thread bundle updates the key value of the thread bundle itself using each current probability value greater than the baseline value of the thread bundle itself, including: The current thread bundle accumulatively sums each current probability value greater than the baseline value of the thread bundle itself, and accumulatively updates the key value of the thread bundle itself using the accumulative sum result.

8. The method of claim 5, wherein, After setting the baseline value of each thread bundle in the current iteration round according to each local comparison value, it further includes: Storing the baseline value of each thread bundle in the current iteration round into the matching baseline register, wherein the baseline register is a scalar register or a vector register; Correspondingly, the current thread bundle sequentially obtains, from the current probability value set, multiple current probability values matching the data parallel number of the thread bundle itself, including: The current thread bundle sequentially obtains multiple current probability values in the current probability value set according to the data parallel degree of the thread bundle itself, and stores each obtained current probability value into the probability value register, wherein the probability value register is a vector register.

9. The method of claim 8, wherein, The current thread bundle parallel compares the numerical size between each current probability value and the baseline value of the thread bundle itself, and updates the key value of the thread bundle itself using each current probability value greater than the baseline value of the thread bundle itself, including: The current thread bundle obtains a target baseline register for storing the baseline value of the thread bundle itself; The current thread bundle calls a special comparison setting instruction matched with the artificial intelligence acceleration computing chip, parallel compares the numerical size between each current probability value stored in the probability value register and the baseline value of the thread bundle itself stored in the target baseline register, and stores the comparison result into a target mask register; The target mask register is a scalar register, and when the target current probability value is greater than the baseline value, the storage location in the target mask register that matches the target current probability value is set to 1, otherwise, the storage location is set to 0. The current thread bundle obtains each target current probability value that matches each storage location in the target mask register that is set to 1, and updates the key value according to each target current probability value.

10. The method of claim 9, wherein, According to the numerical size between each current probability value obtained by the first thread bundle comparison and the first baseline value, each first type probability value greater than the first baseline value is obtained and added to the target set, including: The first target mask register is obtained after the first thread bundle compares the numerical size between each current probability value in the current probability value set and the baseline value of itself and updates; Each first type probability value that matches each storage location in the first target mask register that is set to 1 is added to the target set. Correspondingly, according to the numerical size between each current probability value obtained by the second thread bundle comparison and the second baseline value, each second type probability value greater than the second baseline value is obtained, including: The second target mask register is obtained after the second thread bundle compares the numerical size between each current probability value in the current probability value set and the baseline value of itself and updates; Each second type probability value that matches each storage location in the second target mask register that is set to 1 is obtained.

11. The method of claim 10, wherein, After the result obtained by filtering each first type probability value in each second type probability value is taken as a new current probability value set, it further includes: A special data dissemination instruction matched with the artificial intelligence acceleration computing chip is called to continuously store the result obtained by filtering each first type probability value in each second type probability value in the shared memory, so that in subsequent round iterations, the new current probability value set is loaded in an address-continuous manner to reduce data access delay and accelerate the data screening process.

12. A large language model output acceleration method, characterized in that, It includes: Obtain the normalized probability value corresponding to each candidate word obtained by the large language model in the current reasoning; According to the target screening algorithm preset by the large language model, the data screening method of any one of claims 1-11 is used to screen a target set that meets the target screening algorithm from each candidate word; After resetting the normalized probability value of each candidate word in the target set, a target candidate word is sampled from the target set according to a preset sampling rule, and the target candidate word is taken as the current round output of the model.

13. An electronic device, comprising: The electronic device includes: at least one processor; and a memory connected in communication with the at least one processor; wherein The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the data screening method of any one of claims 1-11 or the large language model output acceleration method of claim 12.

14. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the data screening method of any one of claims 1-11 or the large language model output acceleration method of claim 12 when executed.

15. A computer program product, characterised in that, The computer program product comprises a computer program which, when executed by a processor, implements the data screening method according to any one of claims 1-11, or implements the large language model output acceleration method of claim 12.

Citation Information

Patent Citations

  • Multi-thread arithmetic coding circuit and method based on standard JPEG 2000

    CN102523455A

  • Parallel SLAM (Simultaneous Localization and Mapping) method in dynamic scene based on detection and segmentation

    CN115451939A

  • Decoding processing method and device, and storage medium

    CN116245088A

  • Register sharing method, device and equipment for universal graphics processor, and medium

    CN118672654A

  • Question processing method and device, electronic equipment and storage medium

    CN120780801A