Business processing device, method, system, electronic device and storage medium

By designing a service processing device that includes parallel internal product units and basic internal product units, the problems of low hardware utilization and insufficient computing efficiency in the prior art are solved, and higher hardware utilization and service processing efficiency are achieved.

CN114528982BActive Publication Date: 2025-05-27NANJING UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111628128.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-05-27
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

In the prior art, the hardware utilization rate is low, the hardware resources are wasted, the computing efficiency is low, the equipment cannot achieve the ideal throughput rate, and the service processing efficiency is low.

Method used

A service processing device is designed, including a service processing unit, which consists of a parallel internal product unit and a basic internal product unit. The parallel internal product unit improves calculation efficiency by performing matrix multiplication operations in parallel; the basic internal product unit flexibly responds to input sequences of different lengths through hierarchical layout and the use of accumulators.

Benefits of technology

By improving the flexibility and hardware utilization of hardware, reducing hardware resource waste, improving computing efficiency and business processing efficiency, and thus improving device throughput.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114528982B_ABST
    Figure CN114528982B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a service processing device, method, system, electronic device, and storage medium. The device includes a service processing unit configured to perform at least one service operation corresponding to a service processing network; the service processing unit includes a first number of parallel inner product units, and the parallel inner product unit includes a second number of basic inner product units; the basic inner product unit includes a third number of multipliers and a first adder; any one of the basic inner product units is configured to perform a multiplication operation between any column of a left multiplication service matrix and a right multiplication service matrix corresponding to any one service operation; the left multiplication service matrix is split into a fourth number of sub-service matrices, and the fourth number of sub-service matrices corresponding to any one service operation are sequentially input into the third number of multipliers in the order of execution time; the parallel inner product unit is configured to perform the multiplication operation between the second number of columns of the left multiplication service matrix and the right multiplication service matrix in parallel. By using the embodiments of the present disclosure, the hardware utilization rate and the computing efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a service processing device, method, system, electronic device, and storage medium. Background Art

[0002] Artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Since its birth, the theory and technology of artificial intelligence have become increasingly mature, and the application fields have also been continuously expanded.

[0003] With the development of artificial intelligence technology, the requirements for hardware devices are also getting higher and higher. In related technologies, during the design process of hardware devices, the operations involved in the related service processing of the machine learning network are fixedly mapped to the corresponding hardware units; however, in the process of related service processing of some machine learning networks, multiple operations are often involved, and the length of the input data is often variable in different service scenarios, resulting in problems such as low hardware utilization rate, waste of hardware resources, low computing efficiency, and the device being unable to achieve the ideal throughput rate and low service processing efficiency in related technologies. Summary of the Invention

[0004] The present disclosure provides a service processing device, method, system, electronic device, and storage medium to at least solve the problems such as low hardware utilization rate, waste of hardware resources, low computing efficiency, and the device being unable to achieve the ideal throughput rate and low service processing efficiency in related technologies. The technical solution of the present disclosure is as follows:

[0005] According to the first aspect of the embodiments of the present disclosure, a service processing device is provided, including: a service processing unit;

[0006] The service processing unit is configured to perform at least one service operation during the service processing of the service processing network corresponding to the target service; the service processing unit includes a first number of parallel inner product units, and any one of the parallel inner product units includes a second number of basic inner product units, and the second number of basic inner product units are connected in parallel; any one of the basic inner product units includes a third number of multipliers and a first adder; the third number of multipliers are respectively connected in series with the first adder;

[0007] Any one of the basic inner product units is configured to perform a multiplication operation between any column of the left multiplication service matrix and the right multiplication service matrix corresponding to any one of the service operations; the left multiplication service matrix is split into a fourth number of sub-service matrices, and the fourth number of sub-service matrices corresponding to any one of the service operations are sequentially input into the third number of multipliers according to the execution time sequence; the fourth number is the number of rows of the left multiplication service matrix;

[0008] Any one of the parallel inner product units is configured to perform the multiplication operation between the second number of columns in the left multiplication service matrix and the right multiplication service matrix in parallel.

[0009] In an alternative embodiment, when the at least one service operation is multiple service operations, the first network parameter in the service processing network is an integer multiple of the first number, and the first network parameter represents the amount of operations corresponding to the multiple service operations;

[0010] The second network parameter in the service processing network is an integer multiple of the second number; the second network parameter represents the number of columns of the right multiplication service matrix corresponding to the at least one operation;

[0011] The third network parameter in the service processing network is an integer multiple of the third number, and the third network parameter represents the number of columns of the left multiplication service matrix corresponding to the at least one operation.

[0012] In an alternative embodiment, when the at least one service operation is one service operation, the first network parameter in the service processing network is an integer multiple of the first number, and the first network parameter represents the amount of operations corresponding to the multiple service operations;

[0013] The second network parameter in the service processing network is an integer multiple of the second number; the second network parameter represents the number of columns of the right multiplication service matrix corresponding to the at least one operation;

[0014] The third network parameter in the service processing network is an integer multiple of the third number, and the third network parameter is an integer multiple of the product of the first number and the second number, and the third network parameter represents the number of columns of the left multiplication service matrix corresponding to the at least one operation.

[0015] In an alternative embodiment, when the fourth number is greater than the third number, any one of the basic inner product units further includes: an accumulator; the accumulator is connected in series with the first adder;

[0016] The accumulator is configured to perform an accumulation process on the outputs of the first adder for the fifth number of times;

[0017] The fifth quantity is equal to the value obtained by rounding up the quotient of dividing the fourth quantity by the third quantity.

[0018] In an alternative embodiment, the service processing unit further includes: a second adder, a control unit, and a quantization unit;

[0019] The second adder is connected in series with the first quantity of parallel inner product units respectively;

[0020] The first quantity of parallel inner product units and the second adder are respectively connected in series with the input end of the control unit, and the output end of the control unit is connected in series with the quantization unit;

[0021] The control unit is configured to control the data source transmitted to the quantization unit;

[0022] The quantization unit is configured to perform quantization processing on the data controlled and transmitted by the control unit.

[0023] In an alternative embodiment, the service processing device further includes: a dynamic random access memory, a static random access memory, and a local register;

[0024] The dynamic random access memory is configured to store input service data, network parameters, and output data during the service processing of the service processing network;

[0025] The static random access memory is configured to store the input service data, the network parameters, and intermediate processing result data during the service processing of the service processing network;

[0026] The local register is configured to store target data, where the target data is data whose usage times during the service processing of the service processing network are greater than or equal to a preset number of times.

[0027] According to a second aspect of the embodiments of the present disclosure, there is provided a service processing method, including:

[0028] Obtaining service data of a target service;

[0029] Controlling the service processing device according to any one of the first aspect to execute inputting the service data into the service processing network corresponding to the target service for service processing to obtain a service processing result of the target service.

[0030] According to a third aspect of the embodiments of the present disclosure, there is provided a service processing system, including:

[0031] A service data acquisition module configured to execute obtaining service data of a target service;

[0032] A service processing device is configured to perform service processing by inputting the service data into a service processing network corresponding to the target service, and obtain a service processing result of the target service.

[0033] According to a fourth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the method as described in the second aspect above.

[0034] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the method described in the second aspect of the embodiments of the present disclosure.

[0035] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer program product containing instructions, when it runs on a computer, enabling the computer to execute the method described in the second aspect of the embodiments of the present disclosure.

[0036] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0037] Combining a basic inner product unit, a parallel inner product unit, and a service processing unit, at least one service operation in the service processing process of the network can be mapped to hardware in a hierarchical manner, that is, at least one service operation is mapped to a service processing device including a hierarchically arranged basic inner product unit, a parallel inner product unit, and a service processing unit, greatly improving the flexibility of hardware mapping; and by splitting the network layer input service matrix (left-multiplying service matrix) according to the number of rows, and mapping the split sub-service matrices to the chronological order of execution, different lengths of input sequences can be effectively handled, realizing parallel processing of variable input sequences, greatly improving the hardware utilization rate, reducing waste of hardware resources, improving the computing efficiency, and thus also improving the throughput rate and service processing efficiency of the service processing device.

[0038] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. Description of the Drawings

[0039] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments in line with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.

[0040] Figure 1 It is a schematic diagram of a service processing device shown according to an exemplary embodiment;

[0041] Figure 2Schematic diagram of a basic inner product unit provided according to an exemplary embodiment;

[0042] Figure 3 Schematic diagram of a parallel inner product unit provided according to an exemplary embodiment;

[0043] Figure 4 Schematic diagram of another basic inner product unit provided according to an exemplary embodiment;

[0044] Figure 5 Schematic diagram of a service processing unit provided according to an exemplary embodiment;

[0045] Figure 6 Flowchart of a service processing method shown according to an exemplary embodiment;

[0046] Figure 7 Schematic diagram of a service processing device based on an exemplary embodiment, which performs inputting service data into a service processing network corresponding to a target service for service processing to obtain a service processing result of the target service;

[0047] Figure 8 Schematic diagram of a service processing system block diagram shown according to an exemplary embodiment;

[0048] Figure 9 Block diagram of an electronic device for service processing shown according to an exemplary embodiment. Detailed implementation

[0049] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0050] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar user accounts, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0051] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties.

[0052] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a service processing device shown according to an exemplary embodiment. As Figure 1 shown, it may include: a service processing unit 100;

[0053] In a specific embodiment, the above-mentioned service processing unit 100 is configured to perform at least one service operation during the service processing of the service processing network corresponding to the target service.

[0054] In the embodiments of this specification, the target service may vary according to different actual application scenario requirements. Optionally, the target service may be a multimedia resource recommendation service (for example, recommending multimedia resources in combination with the user's interest preferences for multimedia resources). Correspondingly, the service processing network may be a network for interest recognition. This service processing network may be obtained by pre-training a machine learning network based on sample service data. The sample service data may be user account attribute data of a sample user account, resource attribute data of multimedia resources corresponding to the sample user account (resource attribute data of the first multimedia resource that has performed a preset operation and resource attribute data of the second multimedia resource that has not performed a preset operation). Optionally, the multimedia resources may include static resources such as text and images, and may also include dynamic resources such as short videos. In a specific embodiment, the resource attribute data may be information for describing multimedia resources. Taking the multimedia resource as a video as an example, the resource attribute data may include data such as the publisher information, resource identifier, release date, video frame image, audio information, playing duration, title information, etc. of the multimedia resource. The user account attribute data may be data for describing the interest preferences of the sample user account. Specifically, the user account attribute data may include, but is not limited to, data such as the user gender, age, education level, and region corresponding to the user account. The above-mentioned preset operations may include, but are not limited to, browsing, clicking, conversion (for example, purchasing related products based on multimedia resources, or downloading related applications based on multimedia resources), etc. Optionally, the target service may be a classification service (for example, identifying the object category in an image). Correspondingly, the service processing network may be a network for object recognition. This service processing network may be obtained by pre-training a machine learning network based on sample object images with object category annotations (such as category labels of objects such as cats and dogs).

[0055] In a specific embodiment, at least one service operation may be at least one matrix multiplication operation. Optionally, when the sizes of the left multiplication matrix and / or the right multiplication matrix corresponding to two matrix multiplication operations are different, these two matrix multiplication operations correspond to different matrix multiplication operations.

[0056] In an alternative embodiment, the above-mentioned service processing unit 100 may include a first number of parallel inner product units 101; specifically, the first number may be an integer greater than or equal to one, and the first number may be set in combination with the amount of operations in the service operations in practical applications.

[0057] In an alternative embodiment, any one of the above-mentioned parallel inner product units 101 may include a second number of basic inner product units 102; specifically, the second number may be an integer greater than or equal to one, and the second number may be set in combination with the matrix size corresponding to the service operations in practical applications. Optionally, the above-mentioned second number of basic inner product units may be connected in parallel.

[0058] In an alternative embodiment, any one of the above-mentioned basic inner product units 102 may include a third number of multipliers and a first adder; specifically, the third number may be an integer greater than or equal to one, and the third number may be set in combination with the matrix size corresponding to the service operations in practical applications. Specifically, the above-mentioned third number of multipliers may be respectively connected in series with the above-mentioned first adder; the first adder may be an adder for adding the output data of the third number of multipliers.

[0059] In an alternative embodiment, any one of the above-mentioned basic inner product units 102 is configured to perform a multiplication operation between any column of the left multiplication service matrix and the right multiplication service matrix corresponding to any service operation; specifically, during the process of participating in the service operation, the left multiplication service matrix may be split into a fourth number of sub-service matrices, the fourth number may be the number of rows of the left multiplication service matrix, and the fourth number of sub-service matrices corresponding to any service operation may be sequentially input into the third number of multipliers in the order of execution time;

[0060] In a specific embodiment, the left multiplication service matrix corresponding to each service operation may be the input matrix corresponding to the network layer where the service operation is located. For example, the size of the input matrix corresponding to a certain network layer is: S * dmodel. Correspondingly, the input matrix may be split into S sub-service matrices of 1 * dmodel; further, the S sub-service matrices may be sequentially input into the third number of multipliers in the order of execution time, and the third number of multipliers processes the operations between the S sub-service matrices of 1 * dmodel and a certain column of the corresponding right multiplication service matrix in the order of execution time. That is, by splitting the input matrix into multiple sub-service matrices according to the number of rows, it is possible to avoid the fixed mapping between service operations and hardware devices and flexibly handle service operations corresponding to input service matrices of different lengths.

[0061] In an alternative embodiment, any one of the above-mentioned parallel inner product units 101 may be configured to perform in parallel the multiplication operation between the left multiplication service matrix and the second number of columns in the right multiplication service matrix corresponding to a certain service operation.

[0062] In practical applications, during the process of business processing in a business processing network, the business operations to be processed at a certain moment can be one or more business operations. Optionally, at least one parallel inner product unit can be selected according to the actual application to process one business operation.

[0063] In a specific embodiment, assume that the third quantity is v_l, as Figure 2 shown Figure 2 is a schematic diagram of a basic inner product unit provided according to an exemplary embodiment. Specifically, the inputs of the v_l multipliers can sequentially be the elements of a certain row of v_l columns in the left multiplication service matrix corresponding to any business operation, and a certain column in the right multiplication service matrix corresponding to this business operation.

[0064] In a specific embodiment, assume that the second quantity is v_n and the third quantity is v_l; as Figure 3 shown Figure 3 is a schematic diagram of a parallel inner product unit provided according to an exemplary embodiment. Specifically, v_n basic inner product units are connected in parallel. The v_n basic inner product units share the left multiplication service matrix, and each basic inner product unit corresponds to a column of the right multiplication service matrix. Figure 3 In it, left and right respectively represent the corresponding elements of the left multiplication service matrix and the corresponding elements of the right multiplication service matrix. x, y, and z respectively represent the coordinate position information of the elements in the corresponding service matrix. Taking left(x,y) as an example, left(x,y) is the element in the x-th row and y-th column of the left multiplication service matrix; taking right(y,z) as an example, right(y,z) is the element in the y-th row and z-th column of the right multiplication service matrix. Optionally, taking basic inner product unit 0 as an example, basic inner product unit 0 can be used to perform the operation between the elements of the x-th row and the y to y + v_l - 1 columns of the left multiplication service matrix and the elements of the y to y + v_l - 1 rows and z-th column of the right multiplication service matrix; correspondingly, the output output(x,z) of basic inner product unit 0 can be the element in the x-th row and z-th column of the result matrix after multiplying the left multiplication service matrix and the right multiplication service matrix; correspondingly, the results output by the v_n basic inner product units: output(x,z:z + v_n - 1) can be the elements in the x-th row and the z to z + v_n - 1 columns of the result matrix after multiplying the left multiplication service matrix and the right multiplication service matrix.

[0065] As can be seen from the technical solutions provided in the embodiments of this specification above, the service processing device in the embodiments of this specification, in combination with the basic inner product unit, the parallel inner product unit, and the service processing unit, can perform hardware mapping on at least one service operation in the service processing process of the network in a hierarchical manner, that is, map at least one service operation to a service processing device including a hierarchically arranged basic inner product unit, a parallel inner product unit, and a service processing unit, greatly improving the flexibility of hardware mapping; and by splitting the network layer input service matrix (left-multiplying service matrix) according to the number of rows, and mapping the split sub-service matrices to the order of execution time, it can effectively handle input sequences of different lengths, achieve parallel processing of variable input sequences, greatly improve the hardware utilization rate, reduce the waste of hardware resources, improve the computing efficiency, and thus can also improve the throughput rate and service processing efficiency of the service processing device.

[0066] In an alternative embodiment, when the above at least one service operation is multiple service operations, the first network parameter in the above service processing network can be an integer multiple of the first quantity, the second network parameter in the service processing network is an integer multiple of the second quantity; the third network parameter in the service processing network is an integer multiple of the third quantity.

[0067] In a specific embodiment, the first network parameter can represent the amount of computation corresponding to multiple service operations; the second network parameter represents the number of columns of the right-multiplying service matrix corresponding to at least one operation; the third network parameter represents the number of columns of the left-multiplying service matrix corresponding to at least one operation.

[0068] In an alternative embodiment, taking the service processing network as a service network trained based on a multi-head attention network as an example, the first network parameter can be the number of heads of the attention network, the second network parameter can be the network parameter used to limit the number of columns of the right-multiplying service matrix in the service processing network, and the second network parameter can be the network parameter used to limit the number of columns of the left-multiplying service matrix in the service processing network.

[0069] In a specific embodiment, the multiple relationship between the first network parameter and the first quantity, the multiple relationship between the second network parameter and the second quantity, and the multiple relationship between the third network parameter and the third quantity can be randomly selected, or can be set in combination with requirements such as the demand for computing efficiency, the number of hardware devices, and the size of the device in actual applications.

[0070] In the above embodiments, by setting the first network parameter to an integer multiple of the first quantity, the second network parameter to an integer multiple of the second quantity, and the third network parameter to an integer multiple of the third quantity, in each service operation process, the hardware device can be fully utilized, enabling all the hardware devices in the service processing device to operate, thereby better improving the hardware utilization rate, reducing the waste of hardware resources, and enhancing the computing efficiency, throughput rate, and service processing efficiency of the service processing device.

[0071] In an alternative embodiment, when at least one service operation is a single service operation, the first network parameter in the above service processing network is an integer multiple of the first quantity, and the second network parameter in the service processing network is an integer multiple of the second quantity; the third network parameter in the service processing network is an integer multiple of the third quantity, and the third network parameter is an integer multiple of the product of the first quantity and the second quantity.

[0072] In the above embodiments, when at least one service operation is a single service operation, by setting the first network parameter to an integer multiple of the first quantity, the second network parameter to an integer multiple of the second quantity, the third network parameter to an integer multiple of the third quantity, and the third network parameter to an integer multiple of the product of the first quantity and the second quantity, in the service operation process, the hardware device can be fully utilized, enabling all the hardware devices in the service processing device to operate, thereby better improving the hardware utilization rate, reducing the waste of hardware resources, and enhancing the computing efficiency, throughput rate, and service processing efficiency of the service processing device.

[0073] In an alternative embodiment, when the fourth quantity is greater than the third quantity, any one of the above basic inner product units may further include: an accumulator; the accumulator is connected in series with the first adder;

[0074] Specifically, the accumulator is configured to perform an accumulation process on the outputs of the first adder for the fifth quantity of times;

[0075] The fifth quantity is equal to the value obtained by rounding up the value of dividing the fourth quantity by the third quantity.

[0076] In practical applications, when the fourth quantity is greater than the third quantity, the operation cannot be completed in one round. Correspondingly, the accumulator can be combined to accumulate the outputs of each adder for multiple rounds (the fifth quantity of times).

[0077] In a specific embodiment, assuming that the third quantity is equal to v_l, as Figure 4 shown, Figure 4 is a schematic diagram of another basic inner product unit provided according to an exemplary embodiment.

[0078] In the above embodiments, in the case where the basic inner product unit cannot complete the operation in one round, the accumulator is combined to accumulate the outputs of the first adder in multiple rounds (the fifth number of times), so as to realize the inner product of vectors of any length based on the basic inner product unit and improve the flexibility of operation processing.

[0079] In an alternative embodiment, the above service processing unit may further include: a second adder, a control unit, and a quantization unit;

[0080] In a specific embodiment, the above second adder is connected in series with the first number of parallel inner product units respectively;

[0081] The first number of parallel inner product units and the second adder are respectively connected in series with the input end of the control unit, and the output end of the control unit is connected in series with the quantization unit;

[0082] The control unit is used to control the data source transmitted to the quantization unit; specifically, the data source transmitted to the quantization unit can be the second adder or the first number of parallel inner product units. Specifically, according to the actual application requirements, the processor can send a corresponding data source control instruction to the control unit; optionally, in the case of needing to process multiple service operations, the processor can send a first control instruction, and this first control instruction can indicate that the first number of parallel inner product units is the data source. Correspondingly, when the control unit receives the first control instruction, it can transmit the outputs of the first number of parallel inner product units to the quantization unit; optionally, in the case of needing to process one service operation, the processor can send a second control instruction, and this second control instruction can indicate that the second adder is the data source. Correspondingly, when the control unit receives the second control instruction, it can transmit the output of the second adder to the quantization unit.

[0083] The quantization unit is used to perform quantization processing on the data controlled and transmitted by the control unit. Specifically, the quantization degree of the quantization unit can be configured in combination with the actual application.

[0084] In a specific embodiment, assume that the first number is equal to v_g, as Figure 5 shown, Figure 5 is a schematic diagram of a service processing unit provided according to an exemplary embodiment. Optionally, v_g control units can be selected to control the data source transmitted to the quantization unit, or only one control unit can be set to control the data source transmitted to the quantization unit.

[0085] In a specific embodiment, when the above-mentioned service processing device is used to process multiple service operations, the inputs of each parallel inner product unit come from different left-multiplication service matrices and right-multiplication service matrices. Correspondingly, in combination with the control unit, the outputs of the first number of parallel inner product units can skip the second adder and be directly input into the quantization unit.

[0086] In a specific embodiment, when the above-mentioned service processing device is used to process one service operation, the inputs of the first number of parallel inner product units come from the same left-multiplication service matrix and right-multiplication service matrix. Correspondingly, in combination with the control unit, the outputs of the first number of parallel inner product units can be added by the second adder and then input into the quantization unit.

[0087] In the above embodiments, by combining the second adder and the control unit, it is convenient to flexibly meet the processing requirements of one or more service operations, and on the basis of ensuring the hardware utilization rate and calculation efficiency, the application scenarios of the service processing device are made more extensive.

[0088] In an alternative embodiment, the above-mentioned service processing device may further include: a dynamic random access memory, a static random access memory, and a local register;

[0089] Specifically, the above-mentioned dynamic random access memory can be used to store the input service data, network parameters, and output data during the service processing of the service processing network;

[0090] Specifically, the above-mentioned static random access memory can be used to store the input service data, network parameters, and intermediate processing result data during the service processing of the service processing network;

[0091] Specifically, the above-mentioned local register can be used to store target data, where the target data is data whose usage times during the service processing of the service processing network are greater than or equal to a preset number of times. Specifically, the preset number of times can be set in advance in combination with the actual application.

[0092] In practical applications, during the operation processing, data will be read from the local register for operation processing; optionally, in the case where there is no corresponding data in the local register, the data required for operation processing can be read from the static random access memory to the local memory and then operation processing can be performed; optionally, in the case where there is no corresponding data in the static random access memory, the data required for operation processing can be read from the dynamic random access memory to the static random access memory, and then read from the static random access memory to the local memory for operation processing.

[0093] In a specific embodiment, the dynamic random access memory has a large storage space although its read efficiency is low; and although the read efficiency of the static random access memory is higher than that of the dynamic random access memory, when the power supply stops, the data stored in the static random access memory will disappear; correspondingly, the input service data, network parameters, and output data in the service processing process can be stored in the dynamic random access memory; and the input service data, network parameters, and intermediate processing result data in the service processing process of the service processing network can be stored in the static random access memory; in addition, although the local register has a high read efficiency (higher than both the dynamic random access memory and the static random access memory), its storage space is small, and correspondingly, the data with the usage times greater than or equal to the preset times (data with a high usage frequency) in the service processing process can be stored in the local register.

[0094] In a specific embodiment, the input service data, network parameters, and intermediate processing result data in the service processing process of the service processing network can be divided into a left multiplication service matrix and a right multiplication service matrix. Among them, the input service data and the intermediate processing result are the left multiplication service matrix during the operation process; the network parameter is the right multiplication service matrix during the operation process.

[0095] In an alternative embodiment, two static random access memories can be set, which are respectively used to store the left multiplication service matrix and the right multiplication service matrix. Optionally, the static random access memory for storing the left multiplication service matrix can be v_g (the first quantity) double-ended storage units, the bit width of each storage unit can be v_l×q bit (that is, the data volume transmitted each time), and the depth of each storage unit (how many times the static random access memory can allow the data volume of v_l×q bit to be transmitted) needs to consider the maximum supported input sequence length (the number of rows of the input data); specifically, the depth can be the product of a storage depth unit and the number of storage units.

[0096] In a specific embodiment, taking the service processing network as an example of a network trained based on the BERT-base network, a storage depth unit can be [dmodel / (v_g×v_l)]×max_s, and the minimum storage unit can be max{max_s×head, dff} / dmodel storage units. Among them, both dmodel and dff are the third network parameters; dmodel is the network parameter corresponding to the attention network in the BERT-base network for limiting the number of columns of the left-multiplied service matrix; dff is the network parameter corresponding to the feed-forward network in the BERT-base network for limiting the number of columns of the left-multiplied service matrix. head is the first network parameter (i.e., the number of heads corresponding to the attention network); v_g is the first quantity, v_l is the third quantity; max_s is the maximum input sequence length.

[0097] In a specific embodiment, the number of storage units in the static random access memory for storing the left-multiplied service matrix is equal to the number of parallel inner product units. Correspondingly, the storage units and the parallel inner product units can be connected in a one-to-one correspondence. Optionally, the data storage method can be sequential stacking.

[0098] In a specific embodiment, the static random access memory for storing the right-multiplied service matrix can be v_g dual-port SRAM units, and the bit width of each unit can be: v_l×v_n×q bit, and the depth only needs to meet the requirement of being able to store all the network parameters.

[0099] In a specific embodiment, the local register can be designed in the basic inner product unit to facilitate the multiplier in the basic inner product unit to read the operation data. Specifically, the local register temporarily stores the frequently used network parameters from the static random access memory, realizes the local reuse of the frequently used network parameters, reduces the memory access to the static random access memory, and reduces the overall power consumption.

[0100] In a specific embodiment, taking the service processing network as an example of a network trained based on the BERT-base network, the storage capacity of the local register can be max(dmodel / v_l, dff / (v_n×v_l)).

[0101] In the above embodiments, on the basis of storing the input service data, network parameters, and output data in a dynamic random access memory, the input service data, network parameters, and intermediate processing result data during the service processing of the service processing network are stored in a static random access memory, which can effectively ensure the security of data and the speed of data reading. Moreover, the data used more than or equal to a preset number of times during the service processing of the service processing network is stored in a local register, which can combine the characteristic of frequent reuse of some network data during the service operation process, effectively reduce the reading power consumption, improve the data speed and efficiency, and thus greatly improve the service processing efficiency.

[0102] Figure 6 is a flowchart of a service processing method shown according to an exemplary embodiment, as Figure 6 shown, this service processing method is used in a terminal electronic device and includes the following steps.

[0103] In step S101, obtain the service data of the target service;

[0104] In step S103, based on the service processing device, execute inputting the service data into the service processing network corresponding to the target service for service processing to obtain the service processing result of the target service.

[0105] In a specific embodiment, the service data may be data objectively existing in the scenario corresponding to the target service. Optionally, taking the multimedia resource recommendation service as an example, the service data may be the user account attribute data of the target user account (the user account that needs to recommend multimedia resources) and the resource attribute data of the multimedia resources to be recommended; correspondingly, the service processing result may be the interest index of the target user account for the multimedia resources to be recommended, and this interest index may represent the degree of interest of the target user account in the multimedia resources to be recommended. Taking the classification service as an example, the service data may be an object image including a certain object, and correspondingly, the service processing result may be the category information of the object included in the object image.

[0106] In an optional embodiment, as Figure 7 shown, Figure 7 is a schematic diagram of executing inputting the service data into the service processing network corresponding to the target service for service processing based on the service processing device to obtain the service processing result of the target service shown according to an exemplary embodiment. Specifically, the sub-service matrices in the service data (left multiplying the service matrix) and the corresponding network parameters in the service processing network (right multiplying the service matrix) can be sequentially read in the order of execution time and input into each basic inner product unit in the corresponding parallel inner product unit for matrix multiplication processing.

[0107] As can be seen from the technical solutions provided in the embodiments of this specification, in this specification, during the business processing based on the business processing network, in combination with the business processing device with a three-layer structure, at least one business operation during the business processing of the network can be mapped to hardware hierarchically, greatly improving the flexibility of the hardware mapping; and by splitting the input data (business data) of the business processing network by rows and mapping the split sub-business matrices to the business processing device in the order of execution time, it can effectively handle input sequences of different lengths, achieve parallel processing of variable input sequences, greatly improve the hardware utilization rate, reduce the waste of hardware resources, improve the computing efficiency, and thus greatly improve the business processing efficiency on the basis of improving the throughput rate of the business processing device.

[0108] Figure 8 is a block diagram of a business processing system shown according to an exemplary embodiment. Referring to Figure 8 , the system includes:

[0109] A service data acquisition module 810, configured to execute the acquisition of service data of a target service;

[0110] A service processing device 820, configured to execute inputting the service data into a service processing network corresponding to the target service for service processing to obtain a service processing result of the target service.

[0111] Regarding the system in the above embodiments, the specific manners in which each module and device perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0112] Figure 9 is a block diagram of an electronic device for business processing shown according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as Figure 9 shown. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements a business processing method. The display screen of the electronic device may be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device may be a touch layer covering the display screen, or may be a button, a trackball, or a touchpad provided on the housing of the electronic device, or may also be an external keyboard, a touchpad, or a mouse, etc.

[0113] Those skilled in the art can understand,Figure 9 The structure shown is only a block diagram of some of the structures related to the present disclosure, and does not constitute a limitation on the electronic device to which the present disclosure is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0114] In an exemplary embodiment, an electronic device is further provided, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the service processing method in the embodiments of the present disclosure.

[0115] In an exemplary embodiment, a computer-readable storage medium is further provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the service processing method in the embodiments of the present disclosure.

[0116] In an exemplary embodiment, a computer program product containing instructions is further provided. When it runs on a computer, the computer is enabled to execute the service processing method in the embodiments of the present disclosure.

[0117] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application may include non-volatile and / or volatile memories. Non-volatile memories may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0118] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0119] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A service processing device, characterized in that, it includes: a service processing unit; The service processing unit is configured to perform at least one service operation during the service processing of a service processing network corresponding to a target service; the at least one service operation is at least one matrix multiplication operation; the service processing unit includes a first number of parallel inner product units, and any one of the parallel inner product units includes a second number of basic inner product units, and the second number of basic inner product units are connected in parallel; any one of the basic inner product units includes a third number of multipliers and a first adder; the third number of multipliers are respectively connected in series with the first adder; Any one of the basic inner product units is configured to perform a multiplication operation between any column of a left multiplication service matrix and a right multiplication service matrix corresponding to any one of the service operations; the left multiplication service matrix is split into a fourth number of sub-service matrices, and the fourth number of sub-service matrices corresponding to any one of the service operations are sequentially input into the third number of multipliers in the order of execution time; the fourth number is the number of rows of the left multiplication service matrix; the left multiplication service matrix corresponding to each service operation is the input matrix corresponding to the network layer where each service operation is located; Any one of the parallel inner product units is configured to perform a multiplication operation between the left multiplication service matrix and a second number of columns of the right multiplication service matrix in parallel.

2. The service processing device according to claim 1, characterized in that, in the case where the at least one service operation is multiple service operations, a first network parameter in the service processing network is an integer multiple of the first number, and the first network parameter characterizes the amount of operations corresponding to the multiple service operations; a second network parameter in the service processing network is an integer multiple of the second number; the second network parameter characterizes the number of columns of the right multiplication service matrix corresponding to the at least one operation; a third network parameter in the service processing network is an integer multiple of the third number, and the third network parameter characterizes the number of columns of the left multiplication service matrix corresponding to the at least one operation.

3. The service processing device according to claim 1, characterized in that, in the case where the at least one service operation is one service operation, a first network parameter in the service processing network is an integer multiple of the first number, and the first network parameter characterizes the amount of operations corresponding to the multiple service operations; a second network parameter in the service processing network is an integer multiple of the second number; the second network parameter characterizes the number of columns of the right multiplication service matrix corresponding to the at least one operation; a third network parameter in the service processing network is an integer multiple of the third number, and the third network parameter is an integer multiple of the product of the first number and the second number, and the third network parameter characterizes the number of columns of the left multiplication service matrix corresponding to the at least one operation.

4. The service processing device according to any one of claims 1 to 3, characterized in that, in the case where the fourth number is greater than the third number, any one of the basic inner product units further includes: an accumulator; the accumulator is connected in series with the first adder; The accumulator is configured to perform an accumulation process on the outputs of the fifth number of the first adders; The fifth number is equal to the value obtained by rounding up the value of the fourth number divided by the third number.

5. The service processing device according to any one of claims 1 to 3, characterized in that, The service processing unit further includes: a second adder, a control unit, and a quantization unit; The second adder is serially connected to the first number of parallel inner product units respectively; The first number of parallel inner product units and the second adder are serially connected to the input end of the control unit respectively, The output end of the control unit is serially connected to the quantization unit; The control unit is used to control the data source transmitted to the quantization unit; The quantization unit is used to perform quantization processing on the data transmitted under the control of the control unit.

6. The service processing device according to any one of claims 1 to 3, characterized in that, The service processing device further includes: a dynamic random access memory, a static random access memory, and a local register; The dynamic random access memory is used to store the input service data, network parameters, and output data during the service processing of the service processing network; The static random access memory is used to store the input service data, the network parameters, and the intermediate processing result data during the service processing of the service processing network; The local register is used to store target data, where the target data is data whose usage times during the service processing of the service processing network are greater than or equal to a preset number of times.

7. A service processing method, characterized in that, including: Obtaining service data of a target service; Based on the service processing device according to any one of claims 1 to 6, performing inputting the service data into a service processing network corresponding to the target service for service processing to obtain a service processing result of the target service.

8. A service processing system, characterized in that, including: A service data acquisition module, configured to perform obtaining service data of a target service; The service processing device according to any one of claims 1 to 6, configured to perform inputting the service data into a service processing network corresponding to the target service for service processing to obtain a service processing result of the target service.

9. An electronic device, characterized in that, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the service processing method according to claim 7.

10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the service processing method according to claim 7.

11. A computer program product, including computer instructions, characterized in that, The computer instructions, when executed by a processor, implement the service processing method according to claim 7.

Citation Information

Patent Citations

  • Digital inner product calculator based on first moment

    CN101957738A