Operator fusion method and device, equipment and medium
By determining and merging operators in the neural network processing unit NPU, the problem of low operator fusion efficiency is solved, efficient operator fusion and audio data processing are realized, and sound field expansion and listening experience are improved.
Patent Information
- Application Number
- CN202410167615.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-05
- Publication Date
- 2025-07-25
AI Technical Summary
In the prior art, operator fusion has problems such as low fusion efficiency, poor applicability and excessive calculation amount when fusion operators are fusion, which is difficult to meet the general universality, ease of use and performance optimization requirements of the computing framework.
By obtaining multiple operators to be fused, based on the processing function of the neural network processing unit NPU, the fused target operator is determined and merged into a combined operator, and the processing is performed using the NPU to avoid repeated calculations of the mutually exclusive functional units.
It realizes improving operator fusion efficiency under low computing volume, improving sound field expansion and listening experience, and meeting user listening needs.
Smart Images

Figure CN120371493A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of neural network technology, and in particular to an operator fusion method, device, equipment and medium. Background Art
[0002] With the advancement of technology, the demand for high-performance parallel computing in the consumer electronics market is exploding, especially in machine vision, artificial intelligence, cloud computing, augmented reality / virtual reality, software-defined radio and other emerging fields, which have great demands for heterogeneous computing systems. Heterogeneous computing enables computing units on the same system-on-chip to complete their respective computing tasks through programming, thus achieving higher efficiency and lower power consumption than a single computing unit.
[0003] However, in the related technologies, when performing operator fusion, there are problems such as low fusion efficiency, poor applicability, or excessive computational consumption of the fusion itself. Therefore, there is an urgent need for an efficient operator fusion method that not only meets the universal applicability of the computing framework, but also meets the flexibility requirements of ease of use, and also meets the demand for continuous performance optimization. Summary of the invention
[0004] The present invention provides an operator fusion method, device, equipment and medium to solve the technical problems in the related art that when performing operator fusion, there are low fusion efficiency, poor applicability, or excessive amount of calculation consumed by the fusion itself.
[0005] In a first aspect, an embodiment of the present invention provides an operator fusion method, the method comprising:
[0006] Obtain multiple operators to be fused;
[0007] For any operator, based on the processing function of the preset neural network processing unit NPU, at least one target operator that can be fused with the operator is determined from multiple operators;
[0008] The operator is fused with at least one target operator into a combined operator and processed using the NPU.
[0009] In a possible implementation, in the method provided by the embodiment of the present invention, the NPU includes multiple functional units arranged in sequence, and obtaining multiple operators to be fused includes:
[0010] Get multiple operators;
[0011] Based on the operation type to be performed by each operator, the functional unit corresponding to each operator and the mutually exclusive functional units are determined.
[0012] In a possible implementation manner, in the method provided by the embodiments of the present invention, for any operator, based on the processing functions of a pre-set neural network processing unit (NPU), at least one target operator that can be fused with the operator is determined from multiple operators, including:
[0013] For any operator and the functional unit corresponding to the operator, at least one target operator that can be fused with the operator is determined from multiple operators, and the functional units corresponding to any two of the target operator and the operator are not mutually exclusive.
[0014] In a possible implementation manner, in the method provided by the embodiments of the present invention, for any operator and the functional unit corresponding to the operator, at least one target operator that can be fused with the operator is determined from multiple operators, including:
[0015] For any operator and the functional unit corresponding to the operator, the first target functional unit that is not mutually exclusive with the operator is determined;
[0016] The first target operator corresponding to the target functional unit is sequentially determined from multiple operators.
[0017] In a possible implementation manner, in the method provided by the embodiments of the present invention, fusing the operator with at least one target operator into a combined operator and processing it using the NPU includes:
[0018] Fusing the operator with at least one target operator into a combined operator;
[0019] Removing the operator and the target operator from multiple operators, and adding the combined operator to multiple operators;
[0020] Processing the operators in multiple operators that do not have mutually exclusive functional units using the NPU.
[0021] In a possible implementation manner, in the method provided by the embodiments of the present invention, after fusing the operator with at least one target operator into a combined operator, the method further includes:
[0022] Determining the functional units that are mutually exclusive to the fused operator.
[0023] In a possible implementation manner, in the method provided by the embodiments of the present invention, after obtaining multiple operators to be fused, the method further includes:
[0024] Removing matrix multiplication operators from multiple operators.
[0025] In a second aspect, an operator fusion device provided by the embodiments of the present invention includes:
[0026] An acquisition unit, configured to acquire multiple operators to be fused;
[0027] A determining unit, configured to determine, for any operator, at least one target operator that can be fused with the operator from multiple operators based on the processing functions of a pre-set neural network processing unit (NPU).
[0028] A processing unit, configured to fuse the operator with at least one target operator into a combined operator and process the combined operator using the NPU.
[0029] In a possible implementation manner, in the device provided by the embodiment of the present invention, if the NPU includes multiple function units arranged in sequence, the obtaining unit is specifically configured to:
[0030] Obtain multiple operators;
[0031] Based on the operation types to be run by each operator, determine the function units corresponding to each operator and the mutually exclusive function units.
[0032] In a possible implementation manner, in the device provided by the embodiment of the present invention, the determining unit is specifically configured to:
[0033] For any operator and the function unit corresponding to the operator, determine at least one target operator that can be fused with the operator from multiple operators, and the function units corresponding to any two of the target operator and the operator are not mutually exclusive.
[0034] In a possible implementation manner, in the device provided by the embodiment of the present invention, the determining unit is specifically configured to:
[0035] For any operator and the function unit corresponding to the operator, determine the first target function unit that is not mutually exclusive with the operator;
[0036] Sequentially determine the first target operator corresponding to the target function unit from multiple operators.
[0037] In a possible implementation manner, in the device provided by the embodiment of the present invention, the processing unit is specifically configured to:
[0038] Fuse the operator with at least one target operator into a combined operator;
[0039] Remove the operator and the target operator from multiple operators, and add the combined operator to multiple operators;
[0040] Use the NPU to process the operators in multiple operators that do not have mutually exclusive function units.
[0041] In a possible implementation manner, in the device provided by the embodiment of the present invention, the processing unit is further configured to:
[0042] Determine the function units that are mutually exclusive to the fused operator.
[0043] In a possible implementation manner, in the device provided by the embodiment of the present invention, the obtaining unit is further configured to:
[0044] Remove the matrix multiplication operator among multiple operators.
[0045] In a third aspect, an embodiment of the present invention provides an electronic device, including: at least one processor, at least one memory, and computer program instructions stored in the memory. When the computer program instructions are executed by the processor, the method provided in the first aspect of the embodiment of the present invention is implemented.
[0046] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by the processor, the method provided in the first aspect of the embodiment of the present invention is implemented.
[0047] In the embodiment of the present invention, first, multiple operators to be fused are obtained, then for any operator, at least one target operator that can be fused with the operator is determined among the multiple operators based on the processing function of a pre-set neural network processing unit (NPU), and finally, the operator and the at least one target operator are fused into a combined operator and processed using the NPU. Compared with the related art, by processing audio data to obtain early reflections and late reverberations, and using the early reflections and late reverberations to generate the final output audio output data, improvements in sound field expansion, naturalness of reverberation, and listening perception are achieved with a relatively low computational complexity, meeting the listening needs of users. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 It is a schematic flowchart of an operator fusion method provided by an embodiment of the present invention;
[0049] Figure 2 It is a specific schematic flowchart of an operator fusion method provided by an embodiment of the present invention;
[0050] Figure 3 It is a schematic structural diagram of a neural network processing unit provided by an embodiment of the present invention;
[0051] Figure 4 It is another schematic structural diagram of a neural network processing unit provided by an embodiment of the present invention;
[0052] Figure 5 It is a schematic execution flowchart of an operator fusion method provided by an embodiment of the present invention;
[0053] Figure 6 It is a schematic execution flowchart of an operator fusion method provided by an embodiment of the present invention;
[0054] Figure 7Schematic diagram of a structure of an operator fusion device provided by an embodiment of the present invention;
[0055] Figure 8 Schematic diagram of a structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0056] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0057] Some terms appearing in the text are explained below:
[0058] 1. In the embodiments of the present invention, the term "and / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0059] 2. In the embodiments of the present invention, the term "heterogeneous computing", that is, heterogeneous computing (Heterogeneous Computing) emerged in the mid-1980s, mainly referring to a computing method in which computing units of different types of instruction sets and architectures are used to form a system. Common categories of computing units include CPUs, GPUs, DSPs, ASICs, FPGAs, etc. A heterogeneous computing platform often includes processors using different instruction set architectures (ISAs), and these processors work together to complete computing tasks.
[0060] 3. In the embodiments of the present invention, the term "artificial neural network", that is, artificial neural networks (Artificial Neural Networks, ANNs) is also simply referred to as neural networks (NNs) or called a connection model (Connection Model). It is an algorithmic mathematical model that mimics the behavioral characteristics of animal neural networks and performs distributed parallel information processing. This network relies on the complexity of the system and adjusts the relationships between a large number of internal nodes to achieve the purpose of processing information.
[0061] With the progress of technology, the demand for high-performance parallel computing in the consumer electronics market is growing explosively. Especially in machine vision, artificial intelligence, cloud computing, AR / VR, software-defined radio, and other emerging fields, there is a very large demand for heterogeneous computing systems. Heterogeneous computing enables computing units on the same system-on-chip to complete the computing tasks they are good at through programming, thus achieving more efficient and lower-power results than a single computing unit.
[0062] To enable neural network models to perform efficient computations on heterogeneous platforms, in addition to a high-performance hardware computing platform, efficient compilation by an AI compiler is also required. The AI compiler needs to connect with the changes in the neural network model algorithm above to meet the research demands of algorithm developers who are constantly exploring, and also needs to meet the demands of diverse hardware for the final binary output below, satisfying the resource requirements of different deployment environments. It must satisfy the general universality of the framework, the flexibility requirements of usability, and the continuous optimization demands of performance. The AI compiler ensures the convenient expression and efficient execution of machine learning algorithms and has increasingly become an important part of the design of machine learning frameworks.
[0063] The operator fusion method of related technologies completes the configuration of post-processing through an exhaustive approach. For example, for the operator fusion of conv+relu, a sample matching this combination will be written, and relevant parameters and enable switches will be configured in this sample. However, when there are many post-processing functional modules, it will cause the problem of combinatorial explosion, that is, there will be thousands of samples. This leads to the need to increase manpower and material resources to implement these samples, and the anti-interference ability of the program is very poor. When a functional module changes, these samples also need to be modified. It is very difficult to ensure that each sample can correctly configure parameters and enable switches, and thousands of tests must be completed for verification.
[0064] In summary, when performing operator fusion, there are problems such as low fusion efficiency, poor applicability, or excessive computational consumption of the fusion itself. Therefore, there is an urgent need for an efficient operator fusion method that must satisfy the general universality of the computing framework, the flexibility requirements of usability, and the continuous optimization demands of performance.
[0065] The operator fusion method, device, equipment, and medium provided by the present invention will be described in more detail below with reference to the accompanying drawings and embodiments.
[0066] An embodiment of the present invention provides an operator fusion method, as Figure 1 shown, including:
[0067] Step S101, obtain multiple operators to be fused.
[0068] In an example of the present disclosure, a plurality of operators for calculation are obtained for subsequent processing of the operators in a Neural Network Processing Unit (NPU). The NPU includes a plurality of functional units that have been preset according to the operator type, quantity, etc. Each functional unit is used to calculate different types of operators. After the operators are obtained, based on the operation type to be run by each operator, the functional unit corresponding to each operator and the mutually exclusive functional units are determined.
[0069] Step S102: For any operator, based on the processing function of the preset Neural Network Processing Unit (NPU), determine at least one target operator that can be fused with the operator among the plurality of operators.
[0070] In an example of the present disclosure, for any operator and the functional unit corresponding to the operator, determine at least one target operator that can be fused with the operator among the plurality of operators. The functional units corresponding to any two of the target operator and the operator are not mutually exclusive. Exemplarily, for any operator and the functional unit corresponding to the operator, determine the first target functional unit that is not mutually exclusive with the operator, and then sequentially determine the target operator corresponding to the target functional unit among the plurality of operators.
[0071] Step S103: Fuse the operator and at least one target operator into a combined operator and process it using the NPU.
[0072] In an example of the present disclosure, fuse the operator and at least one target operator into a combined operator, then remove the operator and the target operator from the plurality of operators, add the combined operator to the plurality of operators, and finally process the operators in the plurality of operators that do not have mutually exclusive functional units using the NPU.
[0073] In this step, after fusing the operator and at least one target operator into a combined operator, the mutually exclusive functional units of the fused operator are also determined, so that when further fusing subsequently, it is clear which type of operator the fused operator can be further fused with.
[0074] As Figure 2 shown, the specific process of the operator fusion method provided by the embodiments of the present invention may include the following steps:
[0075] Step S201: Obtain a plurality of operators to be fused.
[0076] In an example of the present disclosure, a plurality of operators for calculation are obtained for subsequent processing of the operators in a Neural Network Processing Unit (NPU). To accelerate neural network inference, more and more Application Processors (APs) in application processors embed a specific hardware computing chip platform dedicated to neural network model inference, namely, a neural network processing unit. Since most of the operators in a neural network model consist of conv, matmul, and fully connected layers, and these operators all contain matrix multiplication operations, and matrix multiplication operations are very time-consuming in processors such as CPUs or GPUs, most NPUs include dedicated hardware transposes for calculating matrix multiplication.
[0077] In an example, as Figure 3 shown, it is an example of an NPU, which includes 4 functional units, namely, matrix multiplication and functional units 1, 2, and 3. Since matrix multiplication operations are very time-consuming in CPUs, most NPUs include dedicated hardware transposes for calculating matrix multiplication. However, a neural network model is not just matrix multiplication. For example, after conv (vector convolution operation) and fully connected layers perform matrix multiplication, a bias (addition operator) needs to be added, and activation operators such as relu (rectified linear unit) will follow these operators. To further improve the inference speed of the neural network model, in addition to the hardware acceleration unit for matrix multiplication, the NPU usually also includes functional modules related to post-processing. For example, in Figure 3 the functional units, the operation of adding bias can be performed in functional unit 1, the operations of relu-related operators can be performed in functional unit 2, and the operation of the add (adder) operator can be performed in functional unit 3. Then, the three operators (conv + relu + add) can be fused into a new operator, and finally, a single instruction can be generated for this new operator to complete the calculation of these three operators. Of course, the composition and order of the NPU can be adjusted according to actual requirements and the distribution of operator types to be calculated, and the embodiments of the present disclosure do not limit this.
[0078] Step S202: Based on the operation type to be performed by each operator, determine the functional unit corresponding to each operator and the mutually exclusive functional units.
[0079] In an example of the present disclosure, after obtaining the operator, determine the functional unit that can calculate the operator and the functional unit that cannot run other operators after calculating the operator based on the operation type of the operator, that is, the mutually exclusive functional units.
[0080] Step S203: For any operator and the corresponding functional unit, determine at least one target operator that can be fused with the operator among multiple operators.
[0081] In an example of the present disclosure, for any operator and the corresponding functional unit, determine the first target functional unit that is not mutually exclusive with the operator, and then sequentially determine the first target operator corresponding to the target functional unit among multiple operators. In this way, multiple operators can be fused into a new operator, and directly calculating the new operator is equivalent to processing the original two operators without conflict in the NPU simultaneously.
[0082] It should be noted that in the embodiments of the present disclosure, all operators can be traversed and all non-conflicting operators can be fused, or only two operators can be fused. Subsequently, the fused operator is further judged, and an operator that is not mutually exclusive with it is selected for fusion.
[0083] Step S204: Fuse the operator with at least one target operator into a combined operator and process it using the NPU.
[0084] In an example, the structure of the NPU is as Figure 4 shown. There are matrix multiplication and eight other functional units. After completing the matrix multiplication operation, the subsequent functional units can be used to continue completing other operations, so as to achieve the purpose of completing the calculations of multiple operators at one time through one calculation instruction. In the above functional modules, unused functional units can be configured with byPass to skip. When using the above functional units, relevant parameters and enable switches of the corresponding functional units need to be configured. For example, when using functional unit 1, the enable switch of functional unit 1 needs to be turned on, and at the same time, the value of another multiplier needs to be configured.
[0085] In the operation, each functional unit supports the fusion of some operators. For example, the LeakyRelu operator needs to be jointly combined by functional unit 3 and functional unit 4 to complete the calculation. Then, the LeakyRelu operator is attributed to functional unit 3 here. Since the LeakyRelu operator occupies the multiplier of functional unit 4, functional unit 4 cannot fuse other operators anymore. Therefore, it is recorded that the LeakyRelu operator is mutually exclusive with functional unit 4. For the quantization operator quant op, it needs a multiplier and an adder to jointly combine to complete the calculation. Then, the quant operator can be attributed to functional unit 1 and functional unit 7 here. Although the quant operator will occupy adder functional unit 2 or functional unit 8, functional unit 2 or functional unit 8 after fusing the quant operator can still fuse other operators. The specific reason is that multiple first-degree polynomials can be expanded into one first-degree polynomial:
[0086] x2 = a2×(a1×x1 + b1) + b2
[0087] After expanding the combination of the above two first-degree polynomials, it is:
[0088] x2 = (a2 × a1) × x1 + (a2 × b1 + b2)
[0089] As can be seen from the above formula, the combination of two first-degree polynomials can be reduced to a first-degree polynomial. Also, as known from step S202, for each functional unit, it is known which operators it supports for fusion. Then, in actual use, for each fusible operator of each functional unit, a corresponding action function is written. In the action function, the fusion of relevant parameters and the configuration of corresponding enable switches and other operations are completed.
[0090] Next, in combination with Figure 5 and Figure 6 The execution process of the operator fusion method in the embodiments of the present disclosure will be described in detail.
[0091] Step S501, obtain an operator.
[0092] Step S502, determine whether the operator is a matrix multiplication operator. If so, perform step S501; otherwise, perform step S503, and record the operator to be fused as op(m); and record the fusion process functional unit (1) as functional unit (n).
[0093] In an example of the present disclosure, it is determined whether the first operator belongs to an operator of matrix multiplication, such as conv, matmul, transposeconv, etc. If so, skip it and obtain the next operator; if not, perform fusion.
[0094] Step S503, fuse the operator.
[0095] Step S504, operate on the fused operator.
[0096] As Figure 6 shown, in step S503, it includes the following steps:
[0097] Step S5031, determine whether all functional units have been traversed. If so, perform step S504; otherwise, perform step S5032.
[0098] In an example of the present disclosure, it is determined whether n in functional unit (n) is greater than the number of functional units. Still using Figure 4 as an example, it is determined whether n is greater than 8. If n is greater than 8, the fusion process functional unit (n) has exceeded the maximum abstract functional unit, so the operator fusion ends.
[0099] Step S5032: Determine whether the current operator is mutually exclusive with the current functional unit. If so, increment n by 1 and proceed to Step S5031; otherwise, proceed to Step S5033.
[0100] Step S5033: Traverse all fusible operators, fuse with the traversed operator, increment n by 1, and proceed to Step S5031.
[0101] In an example of the present disclosure, the current fusion process functional unit (n) supports fusing op (m), triggers the action function for fusing op (m) in the functional unit (n), and configures the corresponding switches and parameters. Then the fusion process functional unit (n) jumps to the next functional unit (n + i) in the maximum mutually exclusive state. For example, if the swish operator is fused in functional unit 4, then it is necessary to jump to functional unit 8, record the new fusion process functional unit (n + i) as functional unit (n); obtain the next op (m + 1) of op (m), record op (m + 1) as op (m), and jump to Step S5031.
[0102] Using the operator fusion method disclosed in the embodiments of the present disclosure, the problem of combinatorial explosion during operator fusion can be efficiently solved. Only the action function for fusing ops under each functional unit needs to be written, and then all possible operator fusion combinations can be automatically enumerated through the abstract functional unit jump mechanism. It is easy to maintain. Even if there are changes in the functional units of the hardware, only the corresponding functional units and action functions need to be modified to complete the adjustment.
[0103] As Figure 7 shown, based on the same inventive concept of the operator fusion method, the present invention also provides an operator fusion device, including:
[0104] An acquisition unit 701, configured to acquire a plurality of operators to be fused;
[0105] A determination unit 702, configured to, for any operator, determine at least one target operator that can be fused with the operator among the plurality of operators based on the processing functions of the pre-set neural network processing unit NPU;
[0106] A processing unit 703, configured to fuse the operator with at least one target operator into a combined operator and process it using the NPU.
[0107] In a possible implementation manner, in the device provided in the embodiments of the present invention, if the NPU includes a plurality of sequentially arranged functional units, the acquisition unit 701 is specifically configured to:
[0108] Acquire a plurality of operators;
[0109] Based on the operation types to be performed by each operator, determine the functional unit corresponding to each operator and the mutually exclusive functional units.
[0110] In a possible implementation manner, in the apparatus provided by the embodiments of the present invention, the determining unit 702 is specifically configured to:
[0111] For any operator and the functional unit corresponding to the operator, determine at least one target operator that can be fused with the operator among multiple operators, and the functional units corresponding to any two operators among the target operator and the operator are not mutually exclusive.
[0112] In a possible implementation manner, in the apparatus provided by the embodiments of the present invention, the determining unit 702 is specifically configured to:
[0113] For any operator and the functional unit corresponding to the operator, determine the first target functional unit that is not mutually exclusive with the operator;
[0114] Sequentially determine the first target operator corresponding to the target functional unit among multiple operators.
[0115] In a possible implementation manner, in the apparatus provided by the embodiments of the present invention, the processing unit 703 is specifically configured to:
[0116] Fuse the operator with at least one target operator into a combined operator;
[0117] Remove the operator and the target operator from multiple operators, and add the combined operator to the multiple operators;
[0118] Use the NPU to process the operators in multiple operators that do not have mutually exclusive functional units.
[0119] In a possible implementation manner, in the apparatus provided by the embodiments of the present invention, the processing unit 703 is further configured to:
[0120] Determine the functional units that are mutually exclusive to the fused operator.
[0121] In a possible implementation manner, in the apparatus provided by the embodiments of the present invention, the obtaining unit 701 is further configured to:
[0122] Remove the matrix multiplication operator from multiple operators.
[0123] In addition, the operator fusion method and apparatus of the embodiments of the present invention described in combination with Figures 1-7 can be implemented by an electronic device. Figure 8 FIG. shows a schematic hardware structure diagram of an electronic device provided by an embodiment of the present invention.
[0124] Specifically refer to the following Figure 8 , which shows a schematic structural diagram of an electronic device 800 suitable for implementing the embodiments of the present disclosure. Figure 8 The electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0125] As shown Figure 8 in FIG. 4, the electronic device 800 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 801, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803 to implement the voice control method of the embodiments described in the present disclosure. In the RAM 803, various programs and data required for the operation of the electronic device 800 are also stored. The processing device 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0126] Generally, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device 800 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 8 FIG. 4 shows the electronic device 800 having various devices, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be alternatively implemented or had.
[0127] Specifically, according to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program codes for executing the method shown in the flowchart, so as to implement the voice control method as described above. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device 809, or installed from the storage device 808, or installed from the ROM 802. When the computer program is executed by the processing device 801, the above functions defined in the method of the embodiment of the present disclosure are executed.
[0128] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0129] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0130] The above computer-readable medium can be included in the above electronic device; or it can exist separately and not be assembled into the electronic device.
[0131] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to:
[0132] Obtain a plurality of operators to be fused;
[0133] For any operator, based on the processing functions of a pre-set neural network processing unit (NPU), determine at least one target operator that can be fused with the operator among multiple operators;
[0134] Fuse the operator with at least one target operator into a combined operator and process it using the NPU.
[0135] Optionally, when one or more of the above programs are executed by the electronic device, the electronic device may also perform the other steps described in the foregoing embodiments.
[0136] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through an Internet service provider via the Internet).
[0137] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0138] The units involved in the embodiments described in the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation on the unit itself in some cases.
[0139] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application Specific Integrated Circuits (ASICs), Application Specific Standard Products (ASSPs), Systems on Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0140] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disk read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0141] In an embodiment of the present invention, first, a plurality of operators to be fused are obtained. Then, for any one operator, based on the processing functions of a pre-set neural network processing unit (NPU), at least one target operator that can be fused with the operator is determined among the plurality of operators. Finally, the operator and the at least one target operator are fused into a combined operator and processed using the NPU. Compared with the related art, by processing audio data to obtain early reflections and late reverberations, and using the early reflections and late reverberations to generate the final output audio output data, improvements in sound field expansion, naturalness of reverberation, and listening perception are achieved with a relatively low computational load, meeting the listening needs of users.
[0142] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0143] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0144] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0145] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0146] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0147] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. An operator fusion method, characterized in that, including: obtaining a plurality of operators to be fused; for any operator, based on the processing function of a pre-set neural network processing unit (NPU), determining at least one target operator in the plurality of operators that can be fused with the operator; fusing the operator with the at least one target operator into a combined operator and processing the combined operator using the NPU.
2. The operator fusion method according to claim 1, wherein If the NPU includes a plurality of function units arranged in sequence, then the obtaining a plurality of operators to be fused includes: obtaining a plurality of the operators; based on the operation type to be run by each operator, determining the function unit corresponding to each operator and the mutually exclusive function units.
3. The operator fusion method according to claim 2, wherein The step of, for any operator, based on the processing function of a pre-set neural network processing unit (NPU), determining at least one target operator in the plurality of operators that can be fused with the operator includes: for any operator and the function unit corresponding to the operator, determining at least one target operator in the plurality of operators that can be fused with the operator, where the function units corresponding to any two of the target operator and the operator are not mutually exclusive.
4. The operator fusion method according to claim 3, wherein The step of, for any operator and the function unit corresponding to the operator, determining at least one target operator in the plurality of operators that can be fused with the operator includes: for any operator and the function unit corresponding to the operator, determining the first target function unit that is not mutually exclusive with the operator; sequentially determining in the plurality of operators the first target operator corresponding to the target function unit.
5. The operator fusion method according to claim 4, wherein The step of fusing the operator with the at least one target operator into a combined operator and processing the combined operator using the NPU includes: fusing the operator with the at least one target operator into a combined operator; removing the operator and the target operator from the plurality of operators and adding the combined operator to the plurality of operators; processing, using the NPU, the operators in the plurality of operators that do not have mutually exclusive function units.
6. The operator fusion method according to claim 5, wherein After fusing the operator with the at least one target operator into a combined operator, the method further includes: determining the mutually exclusive function units of the fused operator.
7. The operator fusion method according to claim 5, characterized in that After obtaining the plurality of operators to be fused, the method further includes: removing matrix multiplication operators from the plurality of operators.
8. An operator fusion device, characterized in that, including: an obtaining unit configured to obtain a plurality of operators to be fused; a determining unit configured to, for any operator, based on the processing function of a pre-set neural network processing unit (NPU), determine at least one target operator in the plurality of operators that can be fused with the operator; a processing unit configured to fuse the operator with the at least one target operator into a combined operator and process the combined operator using the NPU.
9. An electronic device, characterized in that, including: at least one processor, at least one memory, and computer program instructions stored in the memory, where when the computer program instructions are executed by the processor, the method according to any one of claims 1-7 is implemented.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Operator merging method and device, electronic device and storage medium
CN112270413A
Model data processing method and device, electronic equipment and storage medium
CN114168154A
Data processing method and device for neural network model, apparatus, and storage medium
WO2021164506A1