Method and device for reasoning by using open neural network switching ONNX model based on fusion operator, electronic equipment and medium
By fusion operators in the ONNX model, multiple ONNX operation operators are fused into fusion operators based on coefficient matrix, the complex problem of dynamic input shape transformation process in the ONNX model is solved, and a more efficient inference process is achieved.
Patent Information
- Application Number
- CN202510107510.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-27
AI Technical Summary
In the ONNX model, the shape transformation process of dynamic input is complex, requiring multiple operator combinations to occupy inference resources and reduce inference efficiency.
Through the fusion operator, multiple ONNX operation operators used in combination with the Shape operator are fused into fusion operators based on coefficient matrix to simplify the shape operation process.
The ONNX model is realized to respond to dynamic inputs more streamlinedly and efficiently, save inference resources, and improve inference efficiency.
Smart Images

Figure CN120046732A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of machine learning technology, and particularly to a method, apparatus, electronic device, and medium for performing inference using an Open Neural Network Exchange (ONNX) model based on a fusion operator. Background Art
[0002] Open Neural Network Exchange (ONNX) is an open file format designed for machine learning, providing a common language that different machine learning frameworks (such as Pytorch, TensorFlow, MXNet, etc.) can use to describe their models, enabling different artificial intelligence frameworks to store machine learning models in the same format and interact. ONNX uses a unified operator semantics, has good flexibility, and is convenient for users to transplant and perform secondary optimization on trained models.
[0003] In a machine learning model in ONNX format (also referred to as an ONNX model in this article), the Shape operator outputs the shape information of the input tensor. The output of the Shape operator does not directly participate in the core calculation of the inference, but is crucial for the entire inference process of the machine learning model. Summary of the Invention
[0004] In order to handle the dynamic input of a machine learning model (i.e., an input tensor with an unfixed shape), when exporting the model to the ONNX format, the operation process related to the shape of the input tensor will be converted into operators together. For example, it can be converted into a combination of the Shape operator and operators such as Gather, Concat, Add, Div, etc., indicating that first, the shape of the input tensor is obtained through the Shape operator, and then operations such as selection, splicing, addition, or division are performed on the shape through operators such as Gather, Concat, Add, Div, etc., thereby realizing the shape transformation of the input tensor. Although this method solves the problem of the unfixed shape of the input tensor, relatively speaking, when the shape operation process is relatively complex, multiple (for example, ten or even twenty) operators are often required to complete a specific shape transformation. Executing these multiple operators related to the shape transformation occupies inference resources and reduces the inference efficiency.
[0005] In view of this, the embodiments of the present disclosure provide a method, apparatus, electronic device, and medium for performing inference using an Open Neural Network Exchange (ONNX) model based on a fusion operator, so that the ONNX model can more concisely and efficiently handle dynamic input, save inference resources, and improve inference efficiency.
[0006] In a first aspect, an embodiment of the present disclosure provides a method for performing inference using an Open Neural Network Exchange (ONNX) model based on a fusion operator, including: obtaining shape information of an input tensor of the ONNX model, where the shape information is a vector of dimension N, and each element of the vector indicates each dimension of the input tensor, where N≥1; generating an input shape vector based on the shape information, where the input shape vector is a vector of dimension N+1; performing an operation on the input shape vector using the fusion operator to obtain a changed shape vector, where the fusion operator fuses at least one ONNX operation operator, and the fusion operator is based on a coefficient matrix, and one dimension of the coefficient matrix is N+1; and performing inference using the ONNX model based on the changed shape vector.
[0007] According to a specific implementation manner of an embodiment of the present disclosure, the ONNX operation operator is selected from at least one of an extraction operator, a splicing operator, and a numerical calculation operator. The extraction operator includes a Gather operator and a Slice operator. The splicing operator includes a Concat operator. The numerical calculation operator includes an Add operator, a Sub operator, a Mul operator, and a Div operator.
[0008] According to a specific implementation manner of an embodiment of the present disclosure, an initial value of the coefficient matrix is an initial coefficient matrix of N×(N+1). The first N columns of the initial coefficient matrix are an identity matrix, and elements of the (N+1)-th column are all 0. The coefficient matrix is obtained by performing an operation corresponding to the ONNX operation operator on the initial coefficient matrix.
[0009] According to a specific implementation manner of an embodiment of the present disclosure, when the ONNX operation operator is an extraction operator, corresponding rows in the initial coefficient matrix are extracted according to input coefficients of the extraction operator as the coefficient matrix.
[0010] According to a specific implementation manner of an embodiment of the present disclosure, when the ONNX operation operator is a splicing operator, a plurality of coefficient matrices obtained by extracting the initial coefficient matrix for a plurality of dimensions to be spliced are spliced to obtain the coefficient matrix.
[0011] According to a specific implementation manner of an embodiment of the present disclosure, when the ONNX operation operator is a numerical calculation operator, the coefficient matrix obtained by extracting the initial coefficient matrix participates in a calculation corresponding to the numerical calculation operator to obtain the coefficient matrix.
[0012] According to a specific implementation manner of an embodiment of the present disclosure, the fusion operator is the transpose of the coefficient matrix, and the operation of using the fusion operator on the input shape vector to obtain a changed shape vector includes: multiplying the input shape vector by the fusion operator to obtain a changed shape vector.
[0013] In a second aspect, an embodiment of the present disclosure provides an apparatus for performing inference using an Open Neural Network Exchange (ONNX) model based on a fusion operator, including: a shape information acquisition unit configured to acquire shape information of an input tensor of the model, where the shape information is a vector of dimension N, and each element of the vector indicates each dimension of the input tensor, where N≥1; a shape vector generation unit configured to generate an input shape vector based on the shape information, where the input shape vector is a vector of dimension N + 1; a fusion operator operation unit configured to perform an operation on the input shape vector using the fusion operator to obtain a changed shape vector, where the fusion operator incorporates at least one ONNX operation operator, and the fusion operator is based on a coefficient matrix, and one dimension of the coefficient matrix is N + 1; and an ONNX model inference unit configured to perform inference using the ONNX model based on the changed shape vector.
[0014] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method for performing inference using an Open Neural Network Exchange (ONNX) model based on a fusion operator as described in the embodiments of the present disclosure.
[0015] In a fourth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium, characterized in that the non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the method for performing inference using an Open Neural Network Exchange (ONNX) model based on a fusion operator as described in the embodiments of the present disclosure.
[0016] In summary, through the method of operator fusion, the present disclosure fuses the operators related to the shape operation of the input tensor into a fusion operator based on a coefficient matrix, simplifies the shape operation in the ONNX model, enables the ONNX model to more concisely and efficiently handle dynamic inputs, saves inference resources, and improves inference efficiency. Description of the Drawings
[0017] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the accompanying drawings required for the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0018] Figure 1 It is a schematic flowchart of the method for reasoning using an ONNX model based on a fusion operator provided by an embodiment of the present disclosure;
[0019] Figure 2 It is a schematic flowchart of the operation of the operator related to shape operation in the ONNX model;
[0020] Figure 3 It is a schematic structural diagram of the device for reasoning using an ONNX model based on a fusion operator provided by an embodiment of the present disclosure;
[0021] Figure 4 It shows an exemplary structural diagram of a device capable of implementing the method according to an embodiment of the present disclosure. Detailed implementation manners
[0022] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0023] The following uses specific specific examples to illustrate the implementation manners of the present disclosure. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only some embodiments of the present disclosure, rather than all embodiments. The present disclosure can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.
[0024] It should be noted that the following description pertains to various aspects of embodiments within the scope of the appended claims. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on this disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of the aspects set forth herein can be used to implement an apparatus and / or practice a method. Additionally, this apparatus and / or method can be implemented using other structures and / or functionality in addition to one or more of the aspects set forth herein.
[0025] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of this disclosure schematically. The diagrams only show the components related to this disclosure and are not drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0026] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the aspects can be practiced without these specific details.
[0027] In an ONNX model, the operation result output by the Shape operator is the shape information of the input tensor, such as the N-dimensional vector described later. To handle the dynamic input of a machine learning model (i.e., an input tensor with an unfixed shape), when exporting the model to the ONNX format, the operation process related to the shape of the input tensor will be converted into operators together. For example, it can be converted into a combination of the Shape operator and operators such as Gather, Concat, Add, Div, etc., which means first obtaining the shape of the input tensor through the Shape operator, and then performing operations such as selection, splicing, addition, or division on the shape through operators such as Gather, Concat, Add, Div, etc., thereby achieving the shape transformation of the input tensor. Although this method solves the problem of the unfixed shape of the input tensor, relatively speaking, when the shape operation process is relatively complex, multiple (e.g., ten or even twenty) operators are often required to complete a specific shape transformation. Executing these multiple operators related to shape transformation occupies inference resources and reduces the inference efficiency.
[0028] In view of this, the embodiments of this disclosure provide a method, apparatus, electronic device, and medium for performing inference using an Open Neural Network Exchange (ONNX) model based on a fused operator, so that the ONNX model can handle dynamic input more concisely and efficiently, save inference resources, and improve inference efficiency.
[0029] The method for inference using an ONNX model based on a fusion operator according to an embodiment of the present disclosure can be executed, for example, by an electronic device, and in particular, by one or more processors within the electronic device. In this embodiment, the electronic device can be a device with computing capabilities such as a laptop computer, a desktop computer, a tablet computer, a workstation, or a server, and the present disclosure does not make specific limitations. Please refer to Figure 1 , Figure 1 FIG. is a schematic flowchart of the method for inference using an ONNX model based on a fusion operator provided by an embodiment of the present disclosure. As Figure 1 shown, the method includes the following steps.
[0030] S101, obtaining shape information of an input tensor of an ONNX model, where the shape information is a vector of dimension N, and each element of the vector indicates each dimension of the input tensor, where N≥1.
[0031] In an embodiment of the present disclosure, the ONNX model can be, for example, a dynamic ONNX model in ONNX format exported from a machine learning framework such as Pytorch or other machine learning frameworks. "Dynamic" means that some dimensions of the input tensor of the model are dynamic and can be set independently by the user when using the model. The ONNX model in an embodiment of the present disclosure can be used, for example, in fields such as image recognition, speech recognition, computer vision, and natural language processing, but is not limited thereto. The input tensor can be obtained, for example, by extracting features from input data (such as images, languages, natural language texts, etc.) that are the objects of inference of the model.
[0032] In an embodiment of the present disclosure, the shape information of the input tensor of the ONNX model can be obtained, for example, through a Shape operator (or other feasible means). The shape information can be, for example, a vector of dimension N (N≥1), that is, the shape information can be a vector with N elements (such as a row vector), and each element of the vector indicates each dimension of the input tensor. For example, in the case of an input tensor in NCHW format, the shape information can be a four-dimensional row vector of (N, C, H, W), where N represents the number of batches, C represents the number of channels, H represents the height, and W represents the width. For example, when RGB color image data is used as the input tensor, N can represent the number of image sheets, C = 3 represents the three RGB channels, H represents the number of pixels in the height direction of the image, and W represents the number of pixels in the width direction of the image. In the case of dynamic input, the values of one or more elements in the shape information (N, C, H, W) may change, but the dimension N = 4 of the shape information itself remains unchanged.
[0033] S102. Generate an input shape vector based on the shape information, where the input shape vector is a vector with a dimension of N + 1.
[0034] In the embodiments of the present disclosure, for example, an element [1] can be added to the end of the vector representing the shape information to generate an input shape vector with a dimension of N + 1 based on the N-dimensional shape information. For example, when the shape information is (N, C, H, W), the input shape vector can be, for example, v = (N, C, H, W, 1). The input shape vector can be regarded as the input data or operand of the fusion operator described in detail later.
[0035] S103. Use the fusion operator to perform an operation on the input shape vector to obtain a changed shape vector. The fusion operator incorporates at least one ONNX operation operator, and the fusion operator is based on a coefficient matrix, where one dimension of the coefficient matrix is N + 1.
[0036] In the embodiments of the present disclosure, a fusion operator (e.g., named FusedShape) is set to perform an operation on the input shape vector instead of at least one ONNX operator related to shape transformation used in combination with the Shape operator, so as to simplify the shape operation process and improve the inference efficiency. In some embodiments, the fused ONNX operation operators are selected from at least one of an extraction operator, a splicing operator, and a numerical calculation operator. The extraction operator can include, for example, the Gather operator and the Slice operator. The splicing operator can include, for example, the Concat operator. The numerical calculation operators can include, for example, the Add operator, the Sub operator, the Mul operator, and the Div operator, etc. The Gather operator and the Slice operator are used to extract specific elements of the operand data (tensor) through the input coefficients indices (in the case of the Gather operator), starts, and ends (in the case of the Slice operator), etc. The Concat operator is used to splice two or more pieces of operand data (tensors) into a single tensor. The Add operator, the Sub operator, the Mul operator, and the Div operator are used to perform addition, subtraction, multiplication, and division operations on the operand data (tensor) and a constant specified by the input coefficient B.
[0037] The fusion operator in the embodiments of the present disclosure is based on a coefficient matrix, where one dimension of the coefficient matrix is N + 1, enabling matrix multiplication with the input shape vector with a dimension of N + 1. For example, the dimension of the coefficient matrix can be K × (N + 1), that is, K rows and N + 1 columns. In this case, the transpose of the K × (N + 1)-dimensional coefficient matrix can be used as the fusion operator, and multiplying the 1 × (N + 1)-dimensional input shape vector by the fusion operator can obtain a 1 × K-dimensional changed shape vector. In the coefficient matrix, the element A ij(i ≤ K, j ≤ N) indicates that the i-th dimension of the changed shape vector, which is the operation result of the fusion operator, uses the element of the j-th dimension of the input shape vector, and the element A i(N+1) (i ≤ K) indicates that the i-th dimension of the changed shape vector uses a constant term.
[0038] The initial value of the coefficient matrix can be, for example, an initial coefficient matrix of N×(N + 1). The first N columns of this initial coefficient matrix are the identity matrix, and the elements of the (N + 1)-th column are all 0. In the case of N = 4 in the previous example, the initial coefficient matrix can be, for example:
[0039]
[0040] Multiplying the input shape vector by the transpose of this initial coefficient matrix can obtain the same result as the operation result of the Shape operator. That is, when no other operators are fused, the initial coefficient matrix is equivalent to a single Shape operator.
[0041] v × M T =(N C H W
[0042] In the embodiments of the present disclosure, the coefficient matrix for the fusion operator can be obtained, for example, by performing an operation corresponding to the fused ONNX operation operator on the initial coefficient matrix. The fusion of operators, that is, the generation of the coefficient matrix, can be performed, for example, in the preprocessing stage before reasoning using the ONNX model. By generating the coefficient matrix in the preprocessing stage for operator fusion, in the reasoning stage, a single fused operator can be used to replace the one or more fused operators to perform operations related to shape transformation, thereby saving reasoning resources and improving reasoning efficiency.
[0043] In the embodiments of the present disclosure, the main idea of operator fusion is to convert "performing a series of operations on the operation result of the Shape operator through one or more operators to achieve shape transformation" into "performing corresponding operations on the initial coefficient matrix through the one or more operators to obtain a fused operator, and achieving shape transformation by simply multiplying the input shape vector by the single fused operator". The fused ONNX operation operator is one or more operators that are combined with the Shape operator and perform a series of operations on the operation result of the Shape operator to achieve shape transformation.
[0044] When the fused ONNX operation operator is an extraction operator, the corresponding rows in the initial coefficient matrix can be extracted as the coefficient matrix according to the input coefficients of the extraction operators Gather or Slice. It can also be understood as performing operations on the initial coefficient matrix using the extraction operators Gather or Slice without changing the input coefficients, and taking the operation result as the coefficient matrix. For example, when the operator combined with the Shape operator is the Gather operator, the operator to be fused is the Gather operator. Assume that when the input coefficient indices of the Gather operator is 1, it means to extract the second dimension of the shape information that is the operation result of the Shape operator. Assume the input shape vector v = (N, C, H, W, 1). In this case, when performing operator fusion, the initial coefficient matrix M is operated on using the Gather operator with the input coefficient indices = 1, and the second row of the aforementioned initial coefficient matrix M is extracted to obtain the coefficient matrix M 1 :
[0045] M 1 = Gather(data = M, indices = 1) = (0 1 0 0 0).
[0046] Taking the transpose of the coefficient matrix M 1 and multiplying it with the aforementioned input shape vector v = (N, C, H, W, 1), a changed shape vector that is the same as the combined operation result of the Shape operator and the Gather operator can be obtained.
[0047] When the fused ONNX operation operator is a concatenation operator, multiple coefficient matrices obtained by extracting the initial coefficient matrix for multiple dimensions to be concatenated are concatenated to obtain the coefficient matrix that is the basis of the fused operator. Assume that the second and third dimensions of the input tensor are to be concatenated. Then, it is necessary to first operate on the input tensor using the Shape operator to obtain the shape information, and then use two Gather operators with the input coefficients indices = 1 and indices = 2 respectively to extract the second and third dimensions of the operation result (shape information) of the Shape operator. Next, the Concat operator is used to concatenate these two dimensions to obtain the changed shape information. In this case, the operators combined with the Shape operator are the Gather operator and the Concat operator, that is, the operators to be fused are the Gather operator and the Concat operator. Still using the previous example, assume the input shape vector v = (N, C, H, W, 1) and the initial coefficient matrix is the aforementioned M. In this case, when performing operator fusion, the initial coefficient matrix M is first operated on using two Gather operators with the input coefficients indices = 1 and indices = 2 to obtain the coefficient matrix M 1 and M 2, and then use the Concat operator to convert the coefficient matrix M 1 and M 2 Splicing, get the coefficient matrix M 3 .
[0048] M 1 =Gather(data=M,indices=1)=(0 1 0 0 0)
[0049] M 2 =Gather(data=M,indices=2)=(0 0 1 0 0)
[0050]
[0051] The coefficient matrix M 3 The transpose of is used as the fusion operator and multiplied with the aforementioned input shape vector v = (N, C, H, W, 1) to obtain a changed shape vector that is the same as the combined operation result of the Shape operator and two Gather operators and the Concat operator.
[0052] When the fused ONNX operator is a numerical calculation operator, the coefficient matrix obtained by extracting the initial coefficient matrix is used in the calculation corresponding to the numerical calculation operator to obtain the coefficient matrix that serves as the basis of the fused operator. Figure 2 In the example, the operators used in combination with the Shape operator are the Gather operator, the Add operator, and the Div operator, so the operators to be fused are the Gather operator, the Add operator, and the Div operator. When the shape information is (N, C, H, W), the result of the operation process can be written as (C+2) / 2. Still using the previous example, assume that the input shape vector v = (N, C, H, W, 1) and the initial coefficient matrix is the aforementioned M. In this case, when performing operator fusion, first use the Gather operator with input coefficient indices = 1 to operate on the aforementioned initial coefficient matrix M to obtain the coefficient matrix M 1 , and then use the Add operator with input coefficient B = 2 to add the coefficient matrix M 1 Perform addition operation to obtain the coefficient matrix M 2 , and then use the Div operator with input coefficient B = 2 to apply to the coefficient matrix M 2 Perform division operation to obtain the coefficient matrix M 3 .
[0053] M 1 =Gather(data=M,indices=1)=(0 1 0 0 0)
[0054] M 2=Add(data = M, B = 2)=(0 1 0 0 2)
[0055] M 3 =Div(data = M, B = 2)=(0 1 / 2 0 0 2 / 2)
[0056] Take the transpose of the coefficient matrix M 3 as the fusion operator and multiply it with the aforementioned input shape vector v=(N, C, H, W, 1), and a changed shape vector that is the same as the combined operation result (C + 2) / 2 of the Shape operator, Gather operator, Add operator, and Div operator can be obtained.
[0057] S104, Based on the changed shape vector obtained by operating on the input shape vector using the fusion operator, perform inference using the ONNX model.
[0058] As described above, before performing inference using the ONNX model, preprocess the ONNX model with dynamic input, that is, fuse the operators related to shape transformation that are used in combination with the Shape operator into a fusion operator based on a coefficient matrix (such as the transpose of the coefficient matrix). When performing inference, operate on the input shape vector generated based on the shape information of the input tensor of the model using the fusion operator, and the same operation result as the combined operation of the Shape operator and one or more fused operators can be obtained. Thus, during the inference process, it is possible to replace the combined operation of multiple arithmetic operators and only perform a single simple matrix multiplication to multiply the input shape vector by the fusion operator to achieve shape transformation, enabling more concise and efficient handling of dynamic input, saving inference resources, and improving inference efficiency.
[0059] The embodiments of the present disclosure also provide a device for performing inference using an ONNX model based on a fusion operator, and an exemplary structural schematic diagram thereof can be seen in Figure 3 . As Figure 3 shown, the device for performing inference using an ONNX model based on a fusion operator may, for example, include:
[0060] A shape information acquisition unit 210, configured to acquire the shape information of the input tensor of the model, where the shape information is a vector of dimension N, and each element of the vector indicates each dimension of the input tensor, where N≥1;
[0061] A shape vector generation unit 220, configured to generate an input shape vector based on the shape information, where the input shape vector is a vector of dimension N + 1;
[0062] The fusion operator operation unit 230 is configured to operate on the input shape vector by using the fusion operator to obtain a changed shape vector. The fusion operator incorporates at least one ONNX operation operator, and the fusion operator is based on a coefficient matrix, where one dimension of the coefficient matrix is N + 1;
[0063] The ONNX model inference unit 240 is configured to perform inference by using the ONNX model based on the changed shape vector.
[0064] For the specific functions of each unit, for example, refer to the steps of the method for performing inference by using the ONNX model based on the fusion operator as described above, which will not be elaborated here.
[0065] The embodiment of the present disclosure further provides a non-transitory computer-readable storage medium, which stores computer instructions for causing a computer to execute the method for performing inference by using the ONNX model based on the fusion operator according to the embodiment of the present disclosure.
[0066] The embodiment of the present disclosure further provides a computer program product, which includes a computing program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer is caused to execute the method for performing inference by using the ONNX model based on the fusion operator according to the embodiment of the present disclosure.
[0067] Figure 4 A schematic diagram showing a method that can implement the embodiment of the present disclosure or a device 1000 that can implement the embodiment of the present disclosure is shown. In some embodiments, it may include more or fewer devices than shown. In some embodiments, it can be implemented by using a single or multiple devices. In some embodiments, it can be implemented by using cloud or distributed devices.
[0068] As Figure 4As shown, device 1000 includes a processor 1001, which can perform various appropriate operations and processes according to programs and / or data stored in a read-only memory (ROM) 1002 or programs and / or data loaded into a random access memory (RAM) 1003 from a storage section 1008. The processor 1001 can be a multi-core processor or can include multiple processors. In some embodiments, the processor 1001 can include a general-purpose main processor and one or more special coprocessors, such as, for example, a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), a digital signal processing unit (DSP), and so on. In the RAM 1003, various programs and data required for the operation of the device 1000 are also stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0069] The above-mentioned processor and memory are jointly used to execute the programs stored in the memory, and when the programs are executed by a computer, they can implement the methods, steps, or functions described in the above embodiments.
[0070] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, a touch screen, etc.; an output section 1007 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read from it can be installed into the storage section 1008 as needed. Figure 4 Only some components are schematically shown, and it does not mean that the device 1000 only includes Figure 4 the components shown.
[0071] The systems, devices, modules, or units illustrated in the above embodiments can be implemented by a computer or its associated components. The computer can be, for example, a mobile terminal, a smart phone, a personal computer, a laptop computer, an in-vehicle human-machine interaction device, a personal digital assistant, a media player, a navigation device, a game console, a tablet computer, a wearable device, a smart TV, an Internet of Things system, a smart home, an industrial computer, a server, or a combination thereof.
[0072] Although not shown, in an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, a method for performing inference using an ONNX model based on a fusion operator according to an embodiment of the present disclosure is implemented.
[0073] The storage medium in the embodiments of the present disclosure includes permanent and non-permanent, removable and non-removable articles that can implement information storage by any method or technology. Examples of storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0074] Although not shown, an embodiment of the present disclosure further provides a computer program product, including: a computer program / instructions, and when the computer program / instructions are executed by a processor, a method for performing inference using an ONNX model based on a fusion operator according to an embodiment of the present disclosure is implemented.
[0075] The methods, programs, systems, devices, etc. in the embodiments of the present disclosure can be executed or implemented in a single or multiple networked computers, and can also be practiced in a distributed computing environment. In the embodiments of this specification, in these distributed computing environments, tasks can be executed by remote processing devices connected through a communication network.
[0076] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, system, or computer program product. Therefore, those skilled in the art can think that the implementation of the functional modules / units or controllers and related method steps clarified in the above embodiments can be achieved in a manner combining software, hardware, and soft / hardware.
[0077] Unless explicitly stated, the actions or steps of the methods and programs recorded according to the embodiments of the present disclosure do not necessarily have to be executed in a specific order and can still achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0078] In this document, multiple embodiments of the present disclosure are described. For the sake of brevity, the descriptions of the embodiments are not exhaustive, and the same or similar features or parts between the embodiments may be omitted. In this document, "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean applicable to at least one embodiment or example according to the present disclosure, rather than all embodiments. The above terms do not necessarily refer to the same embodiment or example. Without contradiction, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples.
[0079] The exemplary systems and methods of the present disclosure have been specifically shown and described with reference to the above embodiments, which are only examples of the best mode for implementing the systems and methods. Those skilled in the art can understand that various changes can be made to the embodiments of the systems and methods described herein when implementing the systems and / or methods without departing from the spirit and scope of the present disclosure defined in the appended claims.
Claims
1. A method for reasoning using an open neural network exchange ONNX model based on a fusion operator, characterized in that: include: Obtain shape information of an input tensor of the ONNX model, where the shape information is a vector of dimension N, each element of which indicates each dimension of the input tensor, where N≥1; Based on the shape information, generate an input shape vector, wherein the input shape vector is a vector with a dimension of N+1; Using the fusion operator to operate the input shape vector to obtain a changed shape vector, the fusion operator is fused with at least one ONNX operation operator, and the fusion operator is based on a coefficient matrix, and a dimension of the coefficient matrix is N+1; Based on the changed shape vector, reasoning is performed using the ONNX model.
2. The method according to claim 1, characterized in that The ONNX operation operator is selected from at least one of an extraction operator, a splicing operator, and a numerical calculation operator. The extraction operators include a Gather operator and a Slice operator, the concatenation operator includes a Concat operator, and the numerical calculation operators include an Add operator, a Sub operator, a Mul operator, and a Div operator.
3. The method according to claim 2, characterized in that The initial value of the coefficient matrix is an initial coefficient matrix of N×(N+1), the first N columns of the initial coefficient matrix are unit matrices, and the elements of the N+1th column are all 0. The coefficient matrix is obtained by performing an operation corresponding to the ONNX operation operator on the initial coefficient matrix.
4. The method according to claim 3, characterized in that When the ONNX operator is an extraction operator, corresponding rows in the initial coefficient matrix are extracted as the coefficient matrix according to input coefficients of the extraction operator.
5. The method according to claim 3, characterized in that: When the ONNX operator is a concatenation operator, multiple coefficient matrices obtained by extracting the initial coefficient matrix for multiple dimensions to be concatenated are concatenated to obtain the coefficient matrix.
6. The method according to claim 3, characterized in that When the ONNX operator is a numerical calculation operator, the coefficient matrix obtained by extracting the initial coefficient matrix is involved in the calculation corresponding to the numerical calculation operator to obtain the coefficient matrix.
7. The method according to any one of claims 1 to 6, characterized in that The fusion operator is the transpose of the coefficient matrix, and the operation of the input shape vector using the fusion operator to obtain the changed shape vector includes: The input shape vector is multiplied by the fusion operator to obtain a changed shape vector.
8. A device for reasoning using an open neural network exchange ONNX model based on a fusion operator, characterized in that: include: A shape information acquisition unit is configured to acquire shape information of an input tensor of the model, wherein the shape information is a vector of dimension N, each element of the vector indicates each dimension of the input tensor, wherein N≥1; A shape vector generating unit, configured to generate an input shape vector based on the shape information, wherein the input shape vector is a vector with a dimension of N+1; A fusion operator operation unit is configured to operate the input shape vector using the fusion operator to obtain a changed shape vector, wherein the fusion operator is fused with at least one ONNX operation operator, and the fusion operator is based on a coefficient matrix, and a dimension of the coefficient matrix is N+1; The ONNX model inference unit is configured to perform inference using the ONNX model based on the changed shape vector.
9. An electronic device, characterized in that: The electronic device includes: at least one processor; and a memory communicatively coupled to the at least one processor, wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions for causing a computer to execute the method of any one of claims 1 to 7.
Citation Information
Cited By
Data processing method and device, storage medium and electronic equipment
CN121807315A
Data processing method and apparatus, storage medium, and electronic device
CN121807315B