Data processing method and device and electronic equipment
By converting the exponential operation with the natural constant as the base of the Softmax function into the exponential operation with the rational number as the base, and through splitting and target table query, the calculation complexity is reduced and the data processing efficiency and accuracy of the classification model are improved.
Patent Information
- Application Number
- CN202510405640.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-11
Smart Images

Figure CN120296516A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of data processing and artificial intelligence. Specifically, this application relates to a data processing method, apparatus, and electronic device. Background Art
[0002] Artificial intelligence technology is a technology that simulates human intelligence and is mainly implemented by constructing neural networks. By training various models through artificial intelligence technology, various trained models can be used to execute various services. Among them, the classification model is a relatively important one and is widely used in various fields. For example, an image classification model can identify the categories of various objects in an image, and an audio classification model can identify the categories of emotions expressed in audio, etc.
[0003] Currently, the classification model can use the Softmax function as the activation function, which not only provides probability output, supports binary classification tasks or multi-classification tasks, but also simplifies the gradient calculation and enhances the stability and interpretability of the classification model. However, the Softmax function is relatively complex, resulting in a high operation delay, which seriously affects the efficiency of the classification model in data processing. Summary of the Invention
[0004] Embodiments of this application provide a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. In order to achieve the purpose of improving the operation efficiency and further improving the efficiency of the classification model in data processing, the technical solutions provided by the embodiments of this application are as follows: According to one aspect of the embodiments of this application, a data processing method is provided. The method includes: Obtain data to be processed, where the data to be processed includes at least one of text, audio, and images; Input the data to be processed into a trained classification model, and through the classification model, perform the following operations to obtain the classification result of the data to be processed: Extract features from the data to be processed to obtain a feature vector of the data to be processed. The feature vector includes multiple feature values, and the multiple feature values correspond one-to-one to multiple candidate categories; For each feature value in the feature vector , based on this feature value , calculate a feature output value that the category of the data to be processed is the candidate category corresponding to this feature value , where ; Based on the feature output values of the respective candidate categories corresponding to the data to be processed, obtain the classification result of the data to be processed.
[0005] According to another aspect of the embodiments of the present application, a data processing device is provided, and the device includes: A to-be-processed data acquisition module, configured to acquire to-be-processed data, where the to-be-processed data includes at least one of text, audio, and images; A classification result determination module, configured to input the to-be-processed data into a trained classification model, and perform the following operations through the classification model to obtain a classification result of the to-be-processed data: Extract features from the to-be-processed data to obtain a feature vector of the to-be-processed data, where the feature vector includes a plurality of feature values, and the plurality of feature values correspond to a plurality of candidate categories one by one; For each feature value in the feature vector , based on this feature value , calculate a feature output value that the category of the to-be-processed data is the candidate category corresponding to this feature value , where ; Based on the feature output values of the respective candidate categories corresponding to the to-be-processed data, obtain the classification result of the to-be-processed data.
[0006] Optionally, the classification result determination module may be configured to calculate a first exponent corresponding to this feature value based on this feature value , where ; ; Determine the integer part of the first exponent and the decimal part ; ; Determine a first calculation result of a power operation with the integer part as the exponent and 2 as the base; Based on the decimal part , determine a query index, and based on the query index, obtain a second calculation result corresponding to the query index through a preset target table, where the target table includes a plurality of indexes and a calculation result corresponding to each index, each index in the target table identifies a decimal, and the calculation result corresponding to each index is a calculation result of a power operation with 2 as the base and the decimal identified by this index as the exponent; Calculate a first product of the first calculation result and the second calculation result, and use the first product as the feature output value of the candidate category corresponding to this feature value .
[0007] Optionally, the target table includes M first indexes, and each first index identifies a binary sequence with a first bit width , where M = +1; The classification result determination module can be used to determine the fractional part of the binary representation; Determine the high-order sequence of the first bit width in the binary representation, and determine the query index corresponding to the high-order sequence; Determine the calculation result corresponding to the index in the target table that is the same as the query index as the first exponent calculation result, and based on the first exponent calculation result, obtain the second calculation result corresponding to the query index.
[0008] Optionally, the bit width of the binary representation is greater than the bit width of the index in the target table; The classification result determination module can be used to determine the calculation result corresponding to the second index in the target table as the second exponent calculation result, where the second index is the index closest to the query index among the indexes greater than the query index; use the low-order bits other than the high-order bits of the binary representation as the offset, and determine the binary fraction corresponding to the offset; determine the difference between the first exponent calculation result and the second exponent calculation result; calculate the second product of the difference and the binary fraction; and determine the sum of the second product and the first exponent calculation result as the second calculation result corresponding to the query index.
[0009] Optionally, the target table includes multiple sub-tables, each sub-table includes multiple indexes and the calculation result corresponding to each index, and the bit widths of the binary sequences identified by the indexes in different sub-tables are different; The classification result determination module can be used to determine, from the multiple sub-tables included in the target table, the target sub-table whose index bit width is the first bit width; and determine the calculation result corresponding to the index in the target sub-table that is the same as the query index as the first exponent calculation result.
[0010] Optionally, the first bit width is determined by the following method: Obtain the data processing accuracy requirement for the data to be processed; According to the data processing accuracy requirement, determine the first bit width from multiple candidate bit widths.
[0011] According to another aspect of the embodiments of the present application, an electronic device is provided, including a processor configured to execute the steps of the method provided in any optional embodiment of the present application.
[0012] Optionally, the processor includes a first multiplier, a second multiplier, a converter, a multiplexer, and a first flip-flop. The first multiplier is connected to the converter, the converter is connected to the multiplexer, the converter is connected to the first flip-flop, and the first flip-flop is connected to the second multiplier; The first multiplier is configured to calculate, for each eigenvalue in the eigenvector of the data to be processed , based on the eigenvalue , a first exponent corresponding to the eigenvalue , where ; The converter is configured to determine the integer part and the fractional part of the first exponent , and based on the fractional part determine a query index, and send the integer part of the first exponent to the first flip-flop; The first flip-flop is configured to receive and store the integer part of the first exponent sent by the converter , and when the multiplexer determines a second calculation result, send the integer part to the second multiplier; The multiplexer is configured to obtain, based on the query index, through a preset target table, a second calculation result corresponding to the query index, where the target table includes a plurality of indexes and a calculation result corresponding to each index, each index in the target table identifies a decimal, and the calculation result corresponding to each index is the calculation result of a power operation with 2 as the base and the decimal identified by the index as the exponent; The second multiplier is configured to receive the integer part sent by the first flip-flop, determine a first calculation result of a power operation with 2 as the base and the integer part as the exponent; and calculate a first product of the first calculation result and the second calculation result, and use the first product as the characteristic output value of the candidate category corresponding to the eigenvalue .
[0013] Optionally, the target table includes M first indexes, each first index identifies a binary sequence with a first bit width , where M = +1; the processor further includes an arithmetic unit; the second multiplier is connected to the multiplexer through the arithmetic unit; The converter is specifically configured to determine the binary representation of the fractional part ; determine the high-order sequence with the first bit width in the binary representation, and determine the query index corresponding to the high-order sequence; The multiplexer is specifically configured to determine, as the first exponent calculation result, the calculation result corresponding to the index in the target table that is the same as the query index; The arithmetic unit is specifically configured to obtain a second calculation result corresponding to the query index based on the first exponent calculation result.
[0014] Optionally, the bit width of the binary representation is greater than the bit width of the index in the target table; the processor further includes a second flip-flop; the second flip-flop is connected to the arithmetic unit, and the second flip-flop is also connected to the converter; The converter is specifically configured to use the low bits of the binary representation except the high bits as a bias, determine a binary fraction corresponding to the bias; and send the binary fraction to the second flip-flop; The second flip-flop is configured to receive the binary fraction sent by the converter, and send the binary fraction to the arithmetic unit when the multiplexer determines a second exponent calculation result; The multiplexer is specifically configured to determine a calculation result corresponding to a second index in the target table as a second exponent calculation result, where the second index is the index closest to the query index among the indexes greater than the query index; The arithmetic unit is specifically configured to receive the binary fraction sent by the second flip-flop, determine a difference between the first exponent calculation result and the second exponent calculation result; calculate a second product of the difference and the binary fraction; and determine the sum of the second product and the first exponent calculation result as the second calculation result corresponding to the query index.
[0015] Optionally, the target table includes multiple sub-tables, each sub-table includes multiple indexes and calculation results corresponding to each index, and the bit widths of the binary sequences identified by the indexes in different sub-tables are different; The multiplexer is specifically configured to determine a target sub-table with a first bit width of the index bit width from the multiple sub-tables included in the target table; and determine a calculation result corresponding to an index identical to the query index in the target sub-table as the first exponent calculation result.
[0016] Optionally, the first bit width is determined by the converter in the following manner: Obtain a data processing precision requirement for the data to be processed; Determine the first bit width from multiple candidate bit widths according to the data processing precision requirement.
[0017] According to still another aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method provided in any optional embodiment of the present application are implemented.
[0018] According to one aspect of the embodiments of the present application, a computer program product is provided, including a computer program which, when executed by a processor, implements the steps of the method provided in any optional embodiment of the present application.
[0019] The beneficial effects brought by the technical solution provided by the embodiments of the present application are as follows: by transforming the exponential operation with the natural constant as the base in the softmax activation function into an exponential operation with a rational constant as the base, it is possible to effectively reduce the computational complexity while ensuring the computational accuracy required for the calculation, improve the computational efficiency, and further improve the efficiency of the classification model in performing classification services. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for the description in the embodiments of the present application.
[0021] Figure 1 It is a schematic flowchart of implementing a data processing method provided by an embodiment of the present application; Figure 2 It is a schematic flowchart of obtaining a classification result provided by an embodiment of the present application; Figure 3 It is a schematic flowchart of determining the feature output value of the candidate category corresponding to the eigenvalue provided by an embodiment of the present application; Figure 4 It is a schematic diagram of a target provided by an embodiment of the present application; Figure 5 It is a schematic flowchart of non-linear fitting provided by an embodiment of the present application; Figure 6 It is a schematic diagram of the principle of non-linear fitting provided by the present application; Figure 7 It is another schematic diagram of the principle of non-linear fitting provided by the present application; Figure 8 It is a schematic diagram of numerical splitting provided by an embodiment of the present application; Figure 9 It is a schematic flowchart of obtaining a classification result based on eigenvalues provided by an embodiment of the present application; Figure 10 It is a schematic diagram of a processor provided by an embodiment of the present application; Figure 11 It is a schematic diagram of the first transformed exponent provided by an embodiment of the present application; Figure 12 It is a schematic diagram of operation decomposition provided by an embodiment of the present application; Figure 13 It is another schematic flowchart of non-linear fitting provided by the specification of the present application; Figure 14Schematic structural diagram of a data processing device provided by an embodiment of the present application; Figure 15 Schematic structural diagram of an electronic device corresponding to data processing provided by an embodiment of the present application. Detailed implementation manners
[0022] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0023] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "comprising" and "including" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude being implemented as other features, information, data, steps, operations, elements, components and / or their combinations supported by the art of the present technology. It should be understood that when we say that an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" or "A, B" indicates being implemented as "A", or being implemented as "B", or being implemented as "A and B". When describing multiple (two or more) items, if the relationship between the multiple items is not clearly defined, the multiple items can refer to one, multiple or all of the multiple items. For example, for the description of "parameter A includes A1, A2, A3", it can be implemented that parameter A includes A1 or A2 or A3, and it can also be implemented that parameter A includes at least two of the three items of parameter A1, A2, A3.
[0024] When designing a classification model, a corresponding activation function can be selected according to subsequent service requirements. There are various activation functions. If the activation function includes an exponential operation with the natural constant as the base, since the exponential operation is relatively complex to calculate, it will affect the inference efficiency of the classification model using this activation function. Among them, the activation function of the classification model can be the Softmax function, and the expression is:
[0025] Among them, is the feature vector of the data to be processed, is the total number of categories of the data to be processed, and this total number of categories is The number of eigenvalues, with each candidate category corresponding to one eigenvalue. and respectively represent the ith eigenvalue (the eigenvalue corresponding to the category with category index i) and the jth eigenvalue (the eigenvalue corresponding to the category with category index i) in the eigenvector, represents the feature output value indicating that the category of the data to be processed belongs to the category with category index i.
[0026] The following formula is an optimized Softmax function expression:
[0027] where is the largest eigenvalue in the eigenvector z. Due to the fast growth rate of the exponential function, directly calculating may result in a value that is too large and exceeds the floating-point number range that can be represented by the computing device. By subtracting the largest eigenvalue in the eigenvector, the result of the exponential operation can be effectively kept within the floating-point number range that can be represented by a computing device, avoiding overflow problems during the operation.
[0028] It can be seen that the exponential operation included in the Softmax function is relatively complex to implement, and the computing device cannot obtain the operation result through simple operations. In the existing methods for implementing the exponential operation with the natural constant e as the base, intermediate data needs to be obtained first, and there are dependencies between multiple intermediate data, resulting in a large operation delay and reducing the efficiency of the classification model in outputting classification results.
[0029] The embodiments of the present application provide a data processing method. The embodiments of the present application can transform the exponential operation with the natural constant as the base in the activation function into an exponential operation with a rational constant as the base, reduce operation errors, and improve the accuracy of the classification result of the data to be processed obtained through the classification model. Through the base conversion, the dependency relationship of the intermediate data is also eliminated, the computational complexity is reduced, the computational efficiency is improved, and further the efficiency of the classification model in performing classification services is improved.
[0030] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the embodiments of the present application and the technical effects produced by the technical solutions of the present application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can be referred to, learned from, or combined with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0031] Figure 1The figure is a schematic flowchart of a method for implementing data processing provided by an embodiment of the present application. The method is executed by an electronic device with computing capabilities that can execute the embodiments of the present application, such as a server deployed with a classification model, a personal computer (PC), etc., and can also be a computing device for training the classification model. The embodiments of the present application do not limit this. For the sake of convenience of description, the embodiments of the present application are described with a server as the execution subject. As Figure 1 shown, the method includes S101 to S102: S101: Obtain the data to be processed, where the data to be processed includes at least one of text, audio, and images.
[0032] When the server executes a classification task using the classification model, it can first obtain the data to be processed. The types of the data to be processed are diverse and can include at least one of text, audio, and images. It can be understood that the embodiments of the present application do not limit the specific classification tasks executed by the classification model and the specific types of the data to be processed. Multiple types of data can also appear in combination. At this time, the classification model can execute a multi-classification task for multiple data types. For example, the image may include text. The classification model can classify the objects in the image, such as any one of a cat and a dog, and can also classify the text in the image, such as any one of a person's name and a place name for the named entity in the text.
[0033] S102: Input the data to be processed into the trained classification model, and execute S201 to S203 through the classification model to obtain the classification result of the data to be processed.
[0034] It should be noted that the data processing method provided by the embodiments of the present application can also be used when training the classification model to improve the training efficiency of the classification model.
[0035] Among them, Figure 2 is a schematic flowchart of obtaining a classification result provided by an embodiment of the present application. As Figure 2 shown, the operations executed through the classification model include S201 to S203: S201: Extract features from the data to be processed to obtain a feature vector of the data to be processed. The feature vector includes multiple feature values, and the multiple feature values correspond to multiple candidate categories one by one.
[0036] The classification model includes multiple neural network layers, including a feature extraction layer (which can be one or more hidden layers) for extracting features of the data to be processed. Input the data to be processed into the trained classification model, and the feature extraction layer can extract the features of the output to be processed to obtain a feature vector of the data to be processed. Based on the feature vector, the class prediction result of the data to be processed can be obtained through the classification layer.
[0037] For the specific network structure of the classification model, the embodiments of the present application do not make any limitations. In theory, it can be any classification model. For example, it can be a binary classification model or a multi-classification model.
[0038] The feature vector may include multiple feature values, and the multiple feature values correspond one-to-one to multiple candidate categories. The specific number of candidate categories can also be set as needed. For example, for a three-class classification model, since this classification model is a three-class classification model, the feature vector of this classification model has three feature values, which can be f1, f2, and f3 respectively. That is, the feature vector is [f1, f2, f3].
[0039] S202: For each feature value in the feature vector, based on this feature value, calculate the feature output value that the category of the data to be processed is the candidate category corresponding to this feature value.
[0040] Among them, the feature value is , and the feature output value that the category of the data to be processed is the candidate category corresponding to this feature value is . Each in the embodiments of the present application is in the expression of the aforementioned Softmax function.
[0041] Continuing with the above example, assuming that the feature vector includes three feature values, then for each feature value, the feature output value of the corresponding candidate category can be obtained through the Softmax function. Specifically, as follows:
[0042]
[0043]
[0044] The denominator of the Softmax function expression is the sum of the feature output values of each candidate category. Divide the feature output value of each candidate category by this denominator respectively, and the normalized probability of each candidate category can be obtained. Since the denominator part of the Softmax function expression is the same when determining the normalized probability of each candidate category corresponding to each feature value, therefore, the numerator part of the Softmax function expression calculated based on this feature value can be used as the probability of the candidate category corresponding to this feature value. However, in the embodiments of the present application, the numerator part of the Softmax function expression calculated based on this feature value is determined as the feature output value of the candidate category corresponding to this feature value. The candidate category is the category that the classification model can recognize, and can be all the categories preset during the training of the classification model.
[0045] In the embodiment of the present application, the in the Softmax function is transformed into . The natural constant is an irrational number. Transforming an exponential operation with an irrational number as the base into an exponential operation with a rational number as the base greatly reduces the calculation difficulty, lowers the computational complexity, improves the calculation efficiency, and further improves the efficiency of the classification model in performing classification operations.
[0046] S203: Based on the feature output values of each candidate category corresponding to the data to be processed, obtain the classification result of the data to be processed.
[0047] Continuing with the above example, suppose the corresponding categories are cat, dog, and pig respectively. is the largest. Therefore, the candidate category cat corresponding to is determined as the category of the data to be processed, and the classification result of the processed data is obtained. If the classification result also needs to include the probability corresponding to the category, the probability value obtained through the Softmax function can be output. This probability value is obtained through the following formula:
[0048] To further reduce the computational complexity and achieve operation pipelining, facilitating subsequent increases in the frequency and peak performance of the Neural network Processing Unit (NPU), in an alternative embodiment of the present application, for step S202, for each eigenvalue in the feature vector, the feature output value of the candidate category corresponding to the eigenvalue for the data to be processed can be calculated through the following method : Figure 3 is a schematic flowchart of a process for determining the feature output value of the candidate category corresponding to the eigenvalue provided by the embodiment of the present application. As Figure 3 shown, it includes steps S301 to S306.
[0049] S301: Based on the eigenvalue, calculate the first exponent corresponding to the eigenvalue.
[0050] Among them, .
[0051] It should be noted that the first exponent can be a floating-point number.
[0052] S302: Determine the integer part and the fractional part of the first exponent.
[0053] Among them, the integer part can be called , and the fractional part can be called It is understandable that the first exponent can be a non-integer. In this case, the non-integer can include an integer part and a fractional part, and the fractional part is less than 1.
[0054] For example, if is -2.5, then can be split into an integer part and a fractional part , then .
[0055] Among them, the fractional part is a positive number because the query index required for subsequent table look-up is a positive index.
[0056] S303: Determine a first calculation result of a power operation with the integer part as the exponent and 2 as the base.
[0057] For exponentiation operations, there are ordinary exp operation methods and fast exp operation methods. For the integer part, either the ordinary exp operation or the fast exp operation method can be used to determine the first calculation result. Among them, using the fast exp operation method is more efficient. Using the ordinary exp operation method and the fast exp operation method are well-known methods to those skilled in the art and will not be elaborated here.
[0058] Continuing with the above example, the first calculation result is .
[0059] S304: Determine a query index based on the fractional part.
[0060] The server can directly use the fractional part as the query index. Correspondingly, the indexes included in the preset target table are preset decimals, and then query is performed based on the query index. Of course, the fractional part can also be subjected to number system conversion to obtain a converted value, and the query index is determined based on the converted value, and then the second calculation result is determined. The embodiments of the present application do not limit this.
[0061] S305: Based on the query index, obtain a second calculation result corresponding to the query index through a preset target table.
[0062] Among them, the target table includes multiple indexes and the calculation results corresponding to each index. Each index in the target table identifies a decimal, and the calculation result corresponding to each index is the calculation result of a power operation with 2 as the base and the decimal identified by the index as the exponent.
[0063] Taking the fractional part of the binary decimal corresponding to the index corresponding value as an example, Table 1 is a schematic diagram of a target table provided by an embodiment of the present application, as shown in Table 1.
[0064] Table 1
[0065] In Table 1, the indexes are 0001, 0010, 0011, 0100, and 0101 respectively, and the corresponding calculation results are 1 to 5. For the convenience of understanding, the corresponding values of the indexes are also shown in Table 1, which are 0.0625, 0.125, 0.1875, 0.25, and 0.3125 respectively. The corresponding values of the indexes are decimal numbers, and the fractional part of the binary decimal corresponding to the decimal number is the index.
[0066] When determining the second calculation result, the query index can be determined according to the fractional part, and then the second calculation result can be obtained by querying the target table according to the query index. Suppose , then the second calculation result is obtained by looking up the table is result 4, that is .
[0067] Obtaining the second calculation result by querying the target table can improve the efficiency of determining the second calculation result, and further improve the efficiency of determining the classification result.
[0068] S306: Calculate the first product of the first calculation result and the second calculation result, and output the first product as the feature value corresponding to the candidate category.
[0069] Continuing with the above example, calculate , the feature value corresponding to the candidate category is .
[0070] In the embodiment of the present application, the exponent in the exponentiation operation is split into an integer exponent and a fractional exponent, the dependence relationship of some intermediate data in the exponentiation operation is removed, the exponentiation operation is pipelined, the calculation complexity is reduced, the operation delay is reduced, and further the efficiency of the classification model performing the classification task is improved. Pipelining the exponentiation operation also facilitates improving the frequency and peak performance of the NPU subsequently.
[0071] Specifically, the frequency of the NPU refers to the number of pulses emitted by the NPU clock signal per second. The clock cycle is a complete waveform cycle of the clock signal, and the time length is the reciprocal of the frequency, usually measured in Hertz (Hz). Within each clock cycle, the NPU can perform a preset number of operations, such as multiplication operations, number system conversions, etc. A higher clock frequency means that more operations can be performed per unit time. Therefore, increasing the frequency can improve the efficiency of determining the classification result. Although increasing the frequency allows more operations to be performed per unit time, it may be the case that an exponential operation cannot be completed within a unit time. In the embodiments of the present application, through pipeline design, the exponential operation is decomposed into multiple sub-operations that can be completed in a relatively short time, such as determining the first calculation result of the integer part and determining the second calculation result of the decimal part, so that each calculation operation can be completed within one clock cycle, reducing the operation delay, facilitating subsequent increase of the NPU frequency, and optimizing the peak performance of the NPU.
[0072] In an alternative embodiment of the present application, for step S304, in order to reduce the error of the second calculation result obtained by querying and improve the accuracy of the obtained classification result, when the server uses the decimal part as the query index, if there is an index in the target table that is the same as the query index, the calculation result corresponding to the index that is the same as the query index is determined as the second calculation result; then, if there is no index in the target table that is the same as the query index, the third index closest to the query index is determined from the target table, and according to the third index, the second calculation result corresponding to the query index is determined. Here, "closest" means having the smallest numerical difference from the query index.
[0073] For example, Table 2 shows another schematic diagram of the target provided by the embodiments of the present application, as shown in Table 2.
[0074] Table 2
[0075] In Table 2, there are indexes 00110, 01100, 10011, and 11001 respectively, and the calculation results 6 to 9 corresponding to each index. For ease of understanding, the index corresponding values are also shown in Table 2, which are 0.2, 0.4, 0.6, and 0.8 respectively. The index corresponding value is a decimal number, and the decimal part of the binary decimal corresponding to this decimal number is the index. When querying, if the decimal part is 0.25, then there is no index in the target table that is the same as the index 00100 (the index corresponding value is 0.25). The third index 00110 (the index corresponding value is 0.2) closest to the query index can be selected, and then the second calculation result corresponding to the query index is obtained. The second calculation result is result 6, that is .
[0076] That is, when there is no index in the target table that is the same as the query index, the third index closest to the query index can also be determined, and then the second calculation result corresponding to the query index can be obtained. The third index refers to the index with the smallest numerical difference from the query index. By determining the third index closest to the query index, the error of the determined second calculation result is reduced.
[0077] For ease of understanding, the index corresponding values are used for illustration. Actually, when determining the third index closest to the query index, it is determined based on the numerical difference between the query indexes, that is, the index with the smallest numerical difference from the query index in the target table is determined as the third index. Suppose the precision of the index corresponding value is 0.1, then the numerical difference between the two closest indexes is 0.1, specifically including 0.1, 0.2, 0.3, 0.4... 0.9, 1.0. If the precision is 0.05, then the numerical difference between the two closest indexes in the target table is 0.05. Among them, the precision is positively correlated with the number of indexes. Generally speaking, the higher the precision, the more indexes there are. The precision can be preset, and the precision can also be related to the bit width, which will be described later and will not be elaborated here.
[0078] Optionally, when determining the second calculation result corresponding to the query index according to the third index, the server can determine the calculation result corresponding to the third index as the second calculation result corresponding to the query index. The fourth index can also be determined from the target table. If the third index is greater than the query index, the fourth index is the index closest to the query index among the indexes smaller than the query index. If the third index is smaller than the query index, the fourth index is the index closest to the query index among the indexes greater than the query index. The calculation result corresponding to the fourth index and the calculation result corresponding to the third index are added and averaged to obtain the second calculation result corresponding to the query index.
[0079] For example, the index corresponding values included in the target table are 0.25, 0.3, 0.35, and 0.4, and the calculation results corresponding to each index corresponding value are calculation results 1 to 4 respectively. Suppose the query index is 0.325 and the third index is 0.35, then the fourth index can be 0.3. Suppose the third index is 0.3, then the fourth index can be 0.35. The calculation results 2 and 4 corresponding to 0.35 and 0.3 are added and averaged to obtain the average result, and this average result is used as the second calculation result corresponding to the query index.
[0080] By querying the target table, the second calculation result can be obtained without real-time calculation, thereby improving the efficiency of the second calculation result. Although there is an error in the second calculation result obtained by querying the target table, the error can be reduced by determining the third index closest to the query index. The error of the determined second calculation result can also be reduced by improving the accuracy of the target table when presetting the target table. It can be understood that the higher the accuracy of the target table, the more likely it is to query an index identical to the query index when querying through the target table, and the smaller the error of the obtained second calculation result.
[0081] In addition, when presetting the target table, each index can be set as a binary number. Then, the target table includes M first indexes, and each first index identifies a binary sequence with a first bit width where M = + 1.
[0082] Among them, the first bit width can be preset.
[0083] For example, the preset first bit width is 4, = 4, M = = 17. Therefore, the target table includes 17 first indexes.
[0084] When determining the query index based on the fractional part , the binary representation of the fractional part can be determined; the high-order sequence of the first bit width in the binary representation is determined, and the query index corresponding to the high-order sequence is determined.
[0085] The server can determine the high-order sequence as the query index. For example, if the fractional part = 0.28125 and the bit width is 8, then the binary representation of the fractional part is 0.01001000, the first bit width is 4, and the high-order sequence of the first bit width is 0100. Then, the query index can be 0100.
[0086] After determining the query index, the server can obtain the second calculation result corresponding to the query index through the preset target table based on the query index.
[0087] Specifically, the server can determine the calculation result corresponding to the index identical to the query index in the target table as the first exponential calculation result, and obtain the second calculation result corresponding to the query index based on the first exponential calculation result.
[0088] For example, if the query index is 0100, in the preset target table, the second calculation result corresponding to 0100 is determined to be 1.189, and the first exponential calculation result is obtained. This first exponential calculation result can be used as the second calculation result corresponding to the query index.
[0089] Since the server performs calculations using binary number systems, when presetting the target table, the index values are in binary. During actual queries, the fractional part is converted to binary to facilitate server queries, avoid additional number system conversion overhead, improve query efficiency, and thus improve the efficiency of obtaining the second calculation result.
[0090] In addition, when presetting the target table, different bit widths of the index can be set according to different precision requirements to flexibly select the corresponding sub-table with the appropriate bit width later. The target table can include multiple sub-tables, and each sub-table includes multiple indexes and the calculation results corresponding to each index. The bit widths of the binary sequences identified by the indexes in the same sub-table are different.
[0091] When determining the calculation result corresponding to the index in the target table that is the same as the query index as the first exponential calculation result, the embodiments of the present application can first determine the target sub-table and then query the calculation result in the target sub-table.
[0092] In other words, the embodiments of the present application can determine, from the multiple sub-tables included in the target table, the target sub-table whose index bit width is the first bit width; and determine the calculation result corresponding to the index in the target sub-table that is the same as the query index as the first exponential calculation result.
[0093] Figure 4 This is a schematic diagram of a target table provided by the embodiments of the present application, as Figure 4 shown.
[0094] The target table includes Sub-table 1, Sub-table 2 to Sub-table n. The index bit widths in each sub-table are different. The server can determine the sub-table with the index bit width of the first bit width as the target sub-table for subsequent querying of calculation results. It should be noted that the number of indexes in each sub-table of the target table can be set as needed. Specifically, it can be set as the operation result with 2 as the base and the first bit width as the exponent. It can also be set as the operation result with 2 as the base and the first bit width as the exponent plus 1. For example, if the first bit width is 4, then the number of indexes in each sub-table of the target table is , generally speaking, when the bit width is 4, the number of binary numbers that can be represented is 16. The calculation result corresponding to one more index can be used to compensate for the deviation. The numerical difference between every two closest indexes in each sub-table of the target table is the reciprocal of the operation result with 2 as the base and the first bit width as the exponent. Continuing with the above example, if Index 1 and Index 2 are the closest indexes, then Index 1 - Index 2 = 1 / 16.
[0095] Corresponding sub-tables are set according to different bit widths. After the server queries the corresponding target sub-table according to the bit width, it can determine the calculation result in the target sub-table without querying the indexes of non-first bit widths, greatly reducing the amount of data queried and improving the query efficiency.
[0096] Among them, when determining the query index of the fractional part it is possible to determine the high-order sequence of the first bit width in the binary representation of the fractional part and the first bit width is determined in the following way: Obtain the data processing precision requirement for the data to be processed; According to the data processing precision requirement, determine the first bit width from multiple candidate bit widths.
[0097] The candidate bit widths can be the bit widths of the indexes in the target table. It can be understood that the target table can include multiple sub-tables, and the bit width of the index in each sub-table is the candidate bit width. Therefore, there are multiple candidate bit widths.
[0098] It is also possible to pre-determine the correspondence between data processing precision and bit width, so as to determine the first bit width from multiple candidate bit widths according to the data processing precision requirement. For example, when the data processing precision is the first precision, the bit width is 4, and when the data processing precision is the second precision, the bit width is 8, etc.
[0099] The data processing precision is positively correlated with the bit width, that is, the higher the data processing precision, the larger the bit width. Selecting the bit width according to the requirements has high flexibility, and still can quickly determine the second calculation result, improving the efficiency of determining the classification result.
[0100] When determining the query index, there may be a situation where the bit width of the binary representation is greater than the first bit width. At this time, there may also be corresponding calculation results for the other bit widths in the binary representation except for the first bit width.
[0101] Continuing with the above example, the fractional part = 0.28125, the binary representation is 0.01001000, the bit width is 8, the first bit width is 4, then the query index is 0100, the bit width of the index in the target table is 4, and the calculation result corresponding to the index consistent with 0100 is 1.189. Then the second calculation result obtained by querying is 1.189, that is , the low-order bits of the actual binary representation are 1000, and there are also corresponding calculation results, that is .
[0102] In other words, the bit width of the index in the target table is the first bit width, and the bit width of the binary representation is greater than the first bit width. Then, the bit width of the binary representation is greater than the bit width of the index in the target table. To further improve the accuracy of the determined second calculation result, non-linear fitting can be used to obtain the calculation result of the low-order bits of the binary representation.
[0103] Figure 5 This is a schematic diagram of the non - linear fitting process provided by the embodiments of this application, as Figure 5 shown.
[0104] Specifically, the server can use the non - linear fitting method to determine the calculation result corresponding to the low - order data, and combine it with the calculation result corresponding to the high - order data to obtain the second calculation result of the fractional part.
[0105] The non - linear fitting includes the following operations: The server can determine the calculation result corresponding to the second index in the target table as the second exponential calculation result. The second index is the index closest to the query index among the indices greater than the query index. Selecting the calculation result corresponding to the index greater than the query index is because the first exponential calculation result corresponding to the query index obtained based on the high - order bits of the binary representation ignores the calculation result corresponding to the low - order bits of the binary representation. Then, the true calculation result of the fractional part is greater than the first exponential calculation result. Therefore, in order to improve the accuracy of the second calculation result, the server can select the calculation result corresponding to the index larger than the query index.
[0106] Figure 6 This is a schematic diagram of the principle of non - linear fitting provided by this application, as Figure 6 shown.
[0107] The exponential function is a function with base 2 and variable x. As x increases, y grows rapidly. Generally speaking, the true calculation result of the fractional part should be less than the second exponential calculation result, that is, the true calculation result of the fractional part is between the first exponential calculation result and the second exponential calculation result. Figure 6 In it, the query index is index, the second index is index + 1, and the first exponential calculation result corresponding to the query index is , and the second exponential calculation result corresponding to the second index , and the second calculation result is .
[0108] For example, if the fractional part = 0.28125 and the bit width is 8, then the binary representation of the fractional part is 0.01001000. The first bit width is 4, so the first exponential calculation result corresponding to the query index 0100 is 1.189. The bit width is 4, index + 1 is 0101, then the first index closest to and greater than the query index is 0101, and the calculation result corresponding to 0101 is 1.248, that is, the second exponential calculation result is 1.248.
[0109] Figure 7Another schematic diagram of the non - linear fitting principle provided by the embodiment of the present application is as follows Figure 7 as shown.
[0110] The proportionality coefficient bias (the binary number corresponding to the offset) can be obtained from the lower bits of the binary representation, and is used to multiply the difference between the first exponential calculation result and the second exponential calculation result . This is equivalent to taking a value between the first exponential calculation result and the second exponential calculation result, and the obtained value represents the calculation result corresponding to the lower bits of the binary representation. And it is added to the first exponential calculation result to obtain the second calculation result of the fractional part.
[0111] Specifically, Figure 8 a schematic diagram of numerical splitting provided by the embodiment of the present application is as follows Figure 8 as shown.
[0112] If the bit - width of the binary representation is L bits, the binary representation can be split into a query index of LH bits and a bias of LL bits. In other words, the server can use the lower bits of the binary representation except the higher bits as the bias to determine the binary fraction corresponding to the bias.
[0113] Continuing with the above example, the bias is 1000, and the binary fraction corresponding to this bias is 0.1000.
[0114] When obtaining the second calculation result corresponding to the query index based on the first exponential calculation result, the server can determine the difference between the first exponential calculation result and the second exponential calculation result, calculate the second product of this difference and the binary fraction, and determine the sum of the second product and the first exponential calculation result as the second calculation result corresponding to the query index.
[0115] Continuing with the above example, the difference is 1.248 - 1.189 = 0.059. The second product is = 0.0295, and the sum of the second product and the first exponential calculation result is 0.0295 + 1.189 = 1.2185. Then the second calculation result corresponding to the query index 0.0100 is 1.2185, that is ≈1.2185. If the first calculation result is = 0.125, then ≈ ≈0.1523.
[0116] By splitting the fractional part Convert it into a binary representation, split it into high and low bits, which are used as the query index and offset respectively, further splitting the calculation process, reducing the dependency relationship of intermediate data, enabling the calculation of the fractional part to be more pipelined, reducing the calculation complexity, facilitating the increase of the frequency and peak performance of the NPU. The peak performance refers to the maximum calculation ability that the NPU can achieve under ideal conditions, usually measured by the number of operations that can be executed per second. Use the target table to fit When fitting the numerical value in the interval [0,1], there will be a certain error. However, due to the optimized Softmax function the input of which is less than or equal to 0, so is a number less than or equal to 1, that is, the numerically value obtained by non-linear fitting will be multiplied by a value less than 1, reducing the error of the final calculation result, making the efficiency of obtaining the classification result higher and the accuracy of the obtained classification result higher.
[0117] Of course, a query table for the low bits of the binary representation can also be created in advance. The query table includes multiple fifth indexes, and each fifth index identifies a binary decimal with a second bit width. Then, after determining the binary decimal corresponding to the offset, based on this binary decimal, determine the lookup table index. For example, use the fractional part of this binary decimal as the lookup table index. Based on the lookup table index, through the preset query table, obtain the third exponential calculation result corresponding to the lookup table index. Add the first exponential calculation result and the third exponential calculation result to obtain the second calculation result corresponding to this query index.
[0118] It should be noted that the embodiments of the present application can be used to optimize ordinary power operations and fast power operations, improving the calculation efficiency of power operations.
[0119] To better understand and illustrate the practical value of the solution provided by the embodiments of the present application, the optional implementation manners of the present application will be described below in combination with specific scenario embodiments.
[0120] Figure 9 is a schematic flowchart of obtaining a classification result based on an eigenvalue provided by an embodiment of the present application, as Figure 9 shown.
[0121] When obtaining a classification result through an eigenvalue, taking the eigenvalue as a floating-point number as an example, the server can execute four steps, namely floating-point multiplication A, number system conversion, non-linear fitting, and floating-point multiplication B. A and B represent different data for floating-point multiplication operations.
[0122] Figure 10 is a schematic diagram of a processor provided by an embodiment of the present application, as Figure 10As shown. The processor includes a floating-point multiplier A (a first multiplier), a number system converter (a converter), a multiplexer, an arithmetic unit, a floating-point multiplier B (a second multiplier), a second flip-flop, and a first flip-flop. The second flip-flop is used to store the binary fraction obtained by the number system converter, and the binary fraction is the binary fraction corresponding to the bias of the fractional part. The first flip-flop is used to store the integer part obtained by the number system converter.
[0123] Specifically, after obtaining the eigenvalue, the eigenvalue is compared with Input to the floating-point multiplier A, and the floating-point multiplier A can determine the first exponent corresponding to the eigenvalue. The first exponent can also be converted in number system to become a two's complement type number with a built-in decimal point.
[0124] Figure 11 A schematic diagram of the first exponent after conversion provided by an embodiment of the present application is as Figure 11 shown.
[0125] The number system converter splits the first exponent into an integer part with an H-bit width and a fractional part with an L-bit width.
[0126] If H = 5, L = 10, assuming the first exponent = -10, it is represented as 10110.0000000000 in two's complement type data with a built-in decimal point. Assuming the first exponent = -2.5, it is represented as 11101.1000000000 in two's complement type data with a built-in decimal point.
[0127] Figure 12 A schematic diagram of arithmetic decomposition provided by an embodiment of the present application is as Figure 12 shown.
[0128] After that, the floating-point multiplier A sends the first exponent to the number system converter, and the number system converter receives the first exponent and performs arithmetic decomposition on the first exponent. Specifically, the first exponent can be split into an integer part and a fractional part, and the fractional part is less than 1. For example, if the converted first exponent is 11101.1000000000, then the obtained fractional part is 0.1000000000. The first pipelining operation for the exponent operation is performed.
[0129] Furthermore, the number system converter can also determine the query index corresponding to the fractional part, send the query index to the multiplexer, and the multiplexer receives the query index and determines the second operation result corresponding to the fractional part by querying a preset target table based on the query index.
[0130] The number system converter uses the lower bits of the binary representation except the highest bit as the bias, determines the binary fraction corresponding to the bias, and sends the binary fraction to the second flip-flop so that the second flip-flop stores the binary fraction. A flip-flop can automatically execute a predefined operation or function when a specific event occurs. In the embodiment of the present application, when the multiplexer determines the first exponential calculation result corresponding to the query index and the second exponential calculation result corresponding to the second index, the second flip-flop sends the binary fraction to the arithmetic unit so that the arithmetic unit can perform subsequent operations.
[0131] The reason for storing the binary fraction in the flip-flop is to ensure the data consistency of the data input to the arithmetic unit, that is, the query index corresponding to the fractional part of the same exponent and the binary fraction of the bias corresponding to the fractional part are input to the arithmetic unit.
[0132] For example, there is a first exponent m, the fractional part is m2, the integer part is m1, the query index corresponding to m2 is m3, the lower bits of the binary representation of m2 except the highest bit are the bias m4, and the binary fraction corresponding to m4 is m5. Another first exponent is r, the fractional part is r2, the integer part is r1, the query index corresponding to r2 is r3, the lower bits of the binary representation of r2 except the highest bit are the bias r4, and the binary fraction corresponding to r4 is r5. Suppose the number system converter first determines m1 and m2, and then determines r1 and r2. The multiplexer is determining the first exponential calculation result corresponding to m3. To ensure that the data input to the arithmetic unit is all related to m, rather than including data related to both m and r at the same time, therefore, m5 is stored in the second flip-flop. When the multiplexer determines the first exponential calculation result corresponding to m3 and the second exponential calculation result corresponding to the second index, m5 is sent from the flip-flop to the arithmetic unit.
[0133] After that, the arithmetic unit can determine the difference between the first exponential calculation result and the second exponential calculation result, calculate the second product of the difference and the binary fraction, and determine the sum of the second product and the first exponential calculation result as the second calculation result corresponding to the query index.
[0134] Figure 13 Another schematic diagram of the non-linear fitting process provided in the specification of the present application is as Figure 13 shown.
[0135] The width of the first digit is 8, and the target table includes 257 indexes, which can be indexed and numbered from 0 to 256 respectively. Specifically, each binary number that can be represented by the width of the first digit can be determined, which are 00000000, 00000001, 00000010... 11111110, 11111111, obtaining 257 indexes in the target table. The 257 indexes respectively represent the binary decimals corresponding to 0.00000000 to 1. For example, label 1 represents 0.00000001. Of course, the binary index 00000001 corresponding to label 1 also represents 0.00000001.
[0136] For ease of understanding, Figure 13 In the example, each binary query index is numbered to obtain the decimal number labels corresponding to each binary query index, which are 0 to 256 respectively.
[0137] For ease of understanding and explanation, Figure 13 In the example, the 257 indexes respectively representing the binary decimals corresponding to 0.00000000 to 1.00000000 are called the first values, such as 0.00000000, 0.00000001, and 0.00000010, etc., which are the first values. The decimal numbers corresponding to the first values are represented in fractional form to obtain the index corresponding values. The denominator of this fraction is the result of the power operation with 2 as the base and the width of the first digit as the exponent, and the numerator is the product of the decimal number of the first value and the result of the power operation with 2 as the base and the width of the first digit as the exponent.
[0138] For example, the first value is 0.01000000, the corresponding decimal number is 0.25, and the width of the first digit is 8. Then the numerator is , and the denominator is . Then the index corresponding values in the target table are (0 / 256) to (256 / 256), such as the index corresponding value can be (64 / 256). The foregoing result can also be used as the exponent, 2 as the base, and the obtained operation formula is used as the index corresponding value, that is .
[0139] After that, the multiplexer determines the first exponent calculation result and the second exponent calculation result of the second index in the target table through the input query index. The arithmetic unit performs an interpolation operation, that is, the arithmetic unit determines the difference between the first exponent calculation result and the second exponent calculation result, calculates the product of the difference and the second binary decimal, and determines the sum of the second product and the first exponent calculation result as the second calculation result with 2 as the base and the fractional part as the exponent.
[0140] Similar to the aforementioned second trigger, to ensure that the second calculation result input to the subsequent floating-point multiplier B and the integer part are related data with the same first exponent, the integer part can be first stored in the first trigger. When the multiplexer determines the second calculation result, in other words, since this second calculation result is actually output by the arithmetic unit, it can also be understood that when the arithmetic unit outputs the second calculation result, the first trigger sends the integer part to the floating-point multiplier B, so that the floating-point multiplier B can determine the first calculation result corresponding to the integer part, multiply the first calculation result by the second calculation result, and obtain the feature output value whose category of the data to be processed corresponds to the candidate category corresponding to this eigenvalue.
[0141] Of course, the number system converter can also first split the first exponent to obtain the integer part and the fractional part, and then determine the two's complement type data of the fractional part. The embodiments of the present application do not limit this.
[0142] Embodiments of the present application provide a data processing device, as Figure 14 shown. The device 140 includes a data to be processed acquisition module 1401 and a classification result determination module 1402, where: The data to be processed acquisition module 1401 is used to acquire the data to be processed, where the data to be processed includes at least one of text, audio, and images; The classification result determination module 1402 is used to input the data to be processed into the trained classification model, and perform the following operations through the classification model to obtain the classification result of the data to be processed: Extract features from the data to be processed to obtain a feature vector of the data to be processed. The feature vector includes multiple feature values, and the multiple feature values correspond to multiple candidate categories one by one; For each eigenvalue in the feature vector , based on this eigenvalue , calculate the feature output value whose category of the data to be processed corresponds to the candidate category corresponding to this eigenvalue , where ; Based on the feature output values corresponding to the respective candidate categories of the data to be processed, obtain the classification result of the data to be processed.
[0143] Optionally, the classification result determination module 1402 can be used to calculate the first exponent corresponding to this eigenvalue , where ; Determine the integer part of the first exponent and the fractional part ; ; Determine the integer part as the exponent, and the first calculation result of the power operation with 2 as the base; Based on the fractional part determine a query index, and based on the query index, obtain a second calculation result corresponding to the query index through a preset target table, where the target table includes multiple indexes and the calculation result corresponding to each index, and each index in the target table identifies a decimal number, and the calculation result corresponding to each index is the calculation result of the power operation with 2 as the base and the decimal number identified by the index as the exponent; Calculate the first product of the first calculation result and the second calculation result, and use the first product as the feature output value of the candidate category corresponding to the feature value .
[0144] Optionally, the target table includes M first indexes, and each first index identifies a binary sequence with a first bit width , where M = + 1; The classification result determination module 1402 can be used to determine the binary representation of the fractional part ; Determine the high-order sequence of the first bit width in the binary representation, and determine the query index corresponding to the high-order sequence; Determine the calculation result corresponding to the index in the target table that is the same as the query index as the first exponent calculation result, and based on the first exponent calculation result, obtain the second calculation result corresponding to the query index.
[0145] Optionally, the bit width of the binary representation is greater than the bit width of the index in the target table; The classification result determination module 1402 can be used to determine the calculation result corresponding to the second index in the target table as the second exponent calculation result, where the second index is the index closest to the query index among the indexes greater than the query index; use the low-order part of the binary representation except the high-order part as the offset, and determine the binary decimal corresponding to the offset; determine the difference between the first exponent calculation result and the second exponent calculation result; calculate the second product of the difference and the binary decimal; and determine the sum of the second product and the first exponent calculation result as the second calculation result corresponding to the query index.
[0146] Optionally, the target table includes multiple sub-tables, each sub-table includes multiple indexes and the calculation result corresponding to each index, and the bit widths of the binary sequences identified by the indexes in different sub-tables are different; The classification result determination module 1402 can be used to determine, from multiple sub-tables included in the target table, a target sub-table whose index bit width is a first bit width; and determine a calculation result corresponding to an index identical to the query index in the target sub-table as a first exponent calculation result.
[0147] Optionally, the first bit width is determined by the following method: Obtain the data processing accuracy requirement for the data to be processed; Determine the first bit width from multiple candidate bit widths according to the data processing accuracy requirement.
[0148] The device according to the embodiments of the present application can execute the method provided by the embodiments of the present application, and its implementation principle is similar and has corresponding technical effects. The actions performed by each module in the device according to the embodiments of the present application correspond to the steps in the method according to the embodiments of the present application. For the detailed function descriptions of each module of the device, reference can be specifically made to the descriptions in the corresponding methods shown above, and details are not described herein again.
[0149] An electronic device is provided in an embodiment of the present application, including a processor configured to execute the steps of the method provided in any optional embodiment of the present application above.
[0150] In an optional embodiment, an electronic device is provided, as Figure 15 shown Figure 15 The electronic device 4000 shown includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, and the transceiver 4004 can be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiments of the present application.
[0151] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in connection with the present disclosure. The processor 4001 may also be a combination that implements computing functions, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0152] Optionally, the processor 4001 includes a first multiplier, a second multiplier, a converter, a multiplexer, and a first flip-flop. The first multiplier is connected to the converter, the converter is connected to the multiplexer, the converter is connected to the first flip-flop, and the first flip-flop is connected to the second multiplier; The first multiplier is configured to, for each eigenvalue in the feature vector of the data to be processed , based on this eigenvalue , calculate a first exponent corresponding to this eigenvalue , where ; The converter is configured to determine the integer part and the fractional part of the first exponent , and based on the fractional part , determine a query index, and send the integer part of the first exponent to the first flip-flop; The first flip-flop is configured to receive and store the integer part of the first exponent sent by the converter , and when the multiplexer determines a second calculation result, send the integer part to the second multiplier; The multiplexer is configured to obtain a second calculation result corresponding to the query index through a preset target table based on the query index, where the target table includes a plurality of indexes and a calculation result corresponding to each index. Each index in the target table identifies a decimal, and the calculation result corresponding to each index is the result of a power operation with 2 as the base and the decimal identified by this index as the exponent; The second multiplier is configured to receive the integer part sent by the first flip-flop , and determine a first calculation result of a power operation with the integer part as the exponent and 2 as the base; and calculate a first product of the first calculation result and the second calculation result, and use the first product as the feature output value of the candidate category corresponding to the feature .
[0153] Optionally, the target table includes M first indexes, and each first index identifies a binary sequence with a first bit width , where M = +1; the processor further includes an arithmetic unit; the second multiplier is connected to the multiplexer through the arithmetic unit; The converter is specifically configured to determine the binary representation of the fractional part ; determine the high-order sequence of the first bit width in the binary representation, and determine the query index corresponding to the high-order sequence; The multiplexer is specifically configured to determine the calculation result corresponding to the index in the target table that is the same as the query index as the first exponent calculation result; The arithmetic unit is specifically configured to obtain a second calculation result corresponding to the query index based on the first exponent calculation result.
[0154] Optionally, the bit width of the binary representation is greater than the bit width of the indexes in the target table; the processor further includes a second flip-flop; the second flip-flop is connected to the arithmetic unit, and the second flip-flop is also connected to the converter; The converter is specifically configured to use the low-order bits other than the high-order bits of the binary representation as a bias, determine the binary fraction corresponding to the bias; and send the binary fraction to the second flip-flop; The second flip-flop is configured to receive the binary fraction sent by the converter, and when the multiplexer determines the second exponent calculation result, send the binary fraction to the arithmetic unit; The multiplexer is specifically configured to determine the calculation result corresponding to the second index in the target table as the second exponent calculation result, where the second index is the index closest to the query index among the indexes greater than the query index; The arithmetic unit is specifically configured to receive the binary fraction sent by the second flip-flop, determine the difference between the first exponent calculation result and the second exponent calculation result; calculate a second product of the difference and the binary fraction; and determine the sum of the second product and the first exponent calculation result as the second calculation result corresponding to the query index.
[0155] Optionally, the target table includes a plurality of sub-tables, each sub-table includes a plurality of indexes and the calculation results corresponding to each index, and the bit widths of the binary sequences identified by the indexes in different sub-tables are different; The multiplexer is specifically configured to determine, from the plurality of sub-tables included in the target table, a target sub-table whose index bit width is the first bit width; and determine the calculation result corresponding to the index identical to the query index in the target sub-table as the first exponential calculation result.
[0156] Optionally, the first bit width is determined by the converter in the following manner: Obtain the data processing precision requirement for the data to be processed; Determine the first bit width from a plurality of candidate bit widths according to the data processing precision requirement.
[0157] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 15 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0158] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.
[0159] The memory 4003 is used to store the computer program for implementing the embodiments of the present disclosure and is controlled by the processor 4001 for execution. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0160] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0161] An embodiment of the present application further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0162] It should be understood that although the flowchart of the embodiments of the present application indicates various operation steps by arrows, the execution order of these steps is not limited to the order indicated by the arrows. Unless otherwise clearly stated in this article, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage among these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.
[0163] The above are only optional implementation manners of some implementation scenarios of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the embodiments of the present application, adopting other similar implementation means based on the technical idea of the present disclosure also belongs to the protection scope of the embodiments of the present application.
Claims
1. A data processing method, characterized in that Including: Obtain data to be processed, where the data to be processed includes at least one of text, audio, and images; Input the data to be processed into a trained classification model, and perform the following operations through the classification model to obtain the classification result of the data to be processed: Extract features from the data to be processed to obtain a feature vector of the data to be processed, where the feature vector includes multiple feature values, and the multiple feature values correspond one-to-one to multiple candidate categories; For each eigenvalue in the feature vector , based on this eigenvalue , calculate the feature output value of the category of the data to be processed as the candidate category corresponding to this eigenvalue , where ; Obtain the classification result of the data to be processed based on the feature output values of the respective candidate categories corresponding to the data to be processed.
2. The method according to claim 1, wherein For each eigenvalue in the eigenvector , based on this eigenvalue , calculate the feature output value of the category of the data to be processed as the candidate category corresponding to this eigenvalue , including: Based on this eigenvalue , calculate the first exponent corresponding to this eigenvalue , where ; Determine the first exponent of the integer part and the fractional part ; Determine the integer part as the exponent, and the first calculation result of the power operation with base 2; Based on the fractional part Determine a query index. Based on the query index, obtain a second calculation result corresponding to the query index through a preset target table, where the target table includes multiple indexes and the calculation result corresponding to each index. Each index in the target table identifies a decimal number, and the calculation result corresponding to each index is the calculation result of a power operation with 2 as the base and the decimal number identified by the index as the exponent; Calculate a first product of the first calculation result and the second calculation result, and use the first product as the feature output value of the candidate category corresponding to the eigenvalue .
3. The method according to claim 2, wherein The target table includes M first indexes, and each first index identifies a binary sequence with a first bit width, where M = ; +1; Based on the fractional part Determine a query index, and based on the query index, obtain a second calculation result corresponding to the query index through a preset target table, including: Determine the fractional part in binary representation; Determine the high-order sequence of the first bit width in the binary representation, and determine the query index corresponding to the high-order sequence; Determine the calculation result corresponding to the index in the target table that is the same as the query index as the first exponent calculation result, and based on the first exponent calculation result, obtain the second calculation result corresponding to the query index.
4. The method according to claim 3, characterized in that The bit width of the binary representation is greater than the bit width of the index in the target table, and the method further includes: Determine the calculation result corresponding to the second index in the target table as the second exponent calculation result, where the second index is the index closest to the query index among the indices greater than the query index; Use the low-order bits other than the high-order bits of the binary representation as the offset, and determine the binary fraction corresponding to the offset; The obtaining the second calculation result corresponding to the query index based on the first exponent calculation result includes: Determine the difference between the first exponent calculation result and the second exponent calculation result; Calculate the second product of the difference and the binary fraction; Determine the sum of the second product and the first exponent calculation result as the second calculation result corresponding to the query index.
5. The method according to claim 3, characterized in that, The target table includes multiple sub-tables, each sub-table includes multiple indices and the calculation result corresponding to each index, and the bit widths of the binary sequences identified by the indices in different sub-tables are different; The determining the calculation result corresponding to the index in the target table that is the same as the query index as the first exponent calculation result includes: Determine a target sub-table with an index bit width of the first bit width from the multiple sub-tables included in the target table; Determine the calculation result corresponding to the index in the target sub-table that is the same as the query index as the first exponent calculation result.
6. The method according to claim 5, wherein The first bit width is determined by the following method: Obtain the data processing accuracy requirement for the data to be processed; Determine the first bit width from multiple candidate bit widths according to the data processing accuracy requirement.
7. An electronic device, comprising a processor, characterized in that, The processor is configured to execute the steps of the method according to any one of claims 1 to 6.
8. The processor according to claim 7, wherein The processor includes a first multiplier, a second multiplier, a converter, a multiplexer, and a first flip-flop. The first multiplier is connected to the converter, the converter is connected to the multiplexer, the converter is connected to the first flip-flop, and the first flip-flop is connected to the second multiplier; The first multiplier is used for each eigenvalue in the eigenvector of the data to be processed , and based on the eigenvalue , calculate the first exponent corresponding to the eigenvalue , where ; The converter is configured to determine the integer part and the fractional part of the first exponent, and determine a query index based on the fractional part , and send the integer part of the first exponent to a first flip-flop; The first trigger is used to receive the integer part of the first exponent sent by the converter and store it. When the multiplexer determines the second calculation result, the integer part is sent to the second multiplier; The multiplexer is configured to obtain a second calculation result corresponding to the query index based on the query index through a preset target table, where the target table includes a plurality of indexes and calculation results corresponding to each index. Each index in the target table identifies a decimal number, and the calculation result corresponding to each index is the result of a power operation with base 2 and exponent being the decimal number identified by the index; The second multiplier is configured to receive the integer part sent by the first flip-flop , and determine a first calculation result of a power operation with base 2 and the integer part as the exponent; and calculate a first product of the first calculation result and the second calculation result, and output the first product as the feature value corresponding to the candidate category of the feature .
9. The processor according to claim 8, wherein The target table includes M first indexes, and each first index identifies a first bit width of a binary sequence, where M = +1; The processor further includes an arithmetic unit; The second multiplier is connected to the multiplexer through the arithmetic unit; The converter is specifically configured to determine the fractional part of the binary representation; determine the high-order sequence of the first bit width in the binary representation, and determine the query index corresponding to the high-order sequence; The multiplexer is specifically configured to determine the calculation result corresponding to the index identical to the query index in the target table as the first exponent calculation result; The arithmetic unit is specifically configured to obtain a second calculation result corresponding to the query index based on the first exponent calculation result.
10. The processor according to claim 9, wherein The bit width of the binary representation is greater than the bit width of the indexes in the target table; the processor further includes a second flip-flop; the second flip-flop is connected to the arithmetic unit, and the second flip-flop is further connected to the converter; The converter is specifically configured to use the low bits other than the high bits of the binary representation as a bias, determine the binary decimal corresponding to the bias; and send the binary decimal to the second flip-flop; The second flip-flop is configured to receive the binary decimal sent by the converter and send the binary decimal to the arithmetic unit when the multiplexer determines the second exponent calculation result; The multiplexer is specifically configured to determine the calculation result corresponding to the second index in the target table as the second exponent calculation result, where the second index is the index closest to the query index among the indexes greater than the query index; The arithmetic unit is specifically configured to receive the binary decimal sent by the second flip-flop and determine the difference between the first exponent calculation result and the second exponent calculation result; Calculate a second product of the difference and the binary decimal; Determine the sum of the second product and the first exponent calculation result as the second calculation result corresponding to the query index.
11. The processor according to claim 9, wherein, The target table includes a plurality of sub-tables, each sub-table includes a plurality of indexes and calculation results corresponding to each index, and the bit widths of the binary sequences identified by the indexes in different sub-tables are different; The multiplexer is specifically configured to determine a target sub-table with the index bit width being the first bit width from the plurality of sub-tables included in the target table; and determine the calculation result corresponding to the index identical to the query index in the target sub-table as the first exponent calculation result.
12. The processor according to claim 11, wherein, The first bit width is determined by the converter in the following manner: Obtain the data processing precision requirement for the data to be processed; Determine the first bit width from a plurality of candidate bit widths according to the data processing precision requirement.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.
14. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.