Processing method and device of reasoning model and electronic equipment
By dynamically adjusting the data format in the inference model, the problem of inflexible memory arrangement in the existing technology is solved, and more efficient computing performance and inference model efficiency are achieved.
Patent Information
- Application Number
- CN202510100023.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
The existing inference model framework lacks flexibility in memory arrangement and cannot choose the optimal memory arrangement according to the characteristics of different node nodes, resulting in poor performance and reducing the efficiency of the inference model usage.
By obtaining the channel information of the operation node in the target inference model, we judge whether the input data format matches the channel information. If it does not match, a pre-configured conversion node is added to format the input data to match the performance of the operation node.
By dynamically adjusting the data format, the calculation time of the operation node is reduced, the performance of the operation node is improved, and the efficiency of the entire inference model is improved.
Smart Images

Figure CN120012934A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer applications, and in particular to a method, device and electronic equipment for processing an inference model. Background Art
[0002] In the current common inference model frameworks, the data is generally arranged in memory in one way. For example, the MNN (Mobile Neural Network) deep learning framework uses the NC4HW4 arrangement by default, or the NCNN (NcnnConvolutional Neural Network) deep learning framework uses the NCHW arrangement by default, etc. Different arrangements have their own advantages and disadvantages for different scenarios.
[0003] Taking convolution as an example, the NC4HW4 arrangement can achieve optimal performance when the number of channels is a multiple of the SIMD (Single Instruction Multiple Data) parallel number (such as 4 / 8), and its computing performance is generally higher than the NCHW arrangement with the same amount of computing. However, for sizes that are not multiples of the SIMD parallel number, such as when the SIMD parallel number is 4, the calculation time for data of 5, 6, 7, and 8 channel sizes is the same. At this time, the performance of algorithms such as im2col (image to column) of the NCHW arrangement is better than that of the NC4HW4 arrangement.
[0004] In common reasoning frameworks, generally only one memory arrangement method is used globally for calculation, and different memory arrangements are not used for calculations on different nodes. This makes it difficult to optimize the performance of each node, thus reducing the efficiency of the use of the reasoning model. Summary of the invention
[0005] In view of this, an object of the present invention is to provide a method, device and electronic device for processing an inference model to alleviate the above technical problems.
[0006] In a first aspect, an embodiment of the present invention provides a method for processing an inference model, the method comprising: obtaining a target inference model, the target inference model being a pre-built inference model library, the inference model being composed of a plurality of operation nodes connected together, and each of the operation nodes being configured with a data storage node, the data output by the operation node being stored in the data storage node according to a preset data format; extracting channel information of each of the operation nodes in the target inference model; determining whether a data format of current input data of the operation node matches the channel information; if not, obtaining a pre-configured conversion node; and adding the conversion node to the input end of the operation node to perform format conversion on the input data of the operation node.
[0007] In combination with the first aspect, an embodiment of the present invention provides a first possible implementation method of the first aspect, wherein the above-mentioned step of determining whether the data format of the current input data of the operation node matches the channel information includes: determining whether the channel information of the operation node supports the calculation method of the data format of the current input data of the operation node; if so, determining that the data format of the current input data of the operation node matches the channel information; if not, determining that the data format of the current input data of the operation node does not match the channel information.
[0008] In combination with the first aspect, an embodiment of the present invention provides a second possible implementation of the first aspect, wherein the above-mentioned reasoning model is an audio streaming reasoning model; the audio streaming reasoning model is also provided with a data cache node, and the data cache node is used to cache the audio frame data of the current operation cycle and serve as input data for the next operation cycle.
[0009] In combination with the second possible implementation of the first aspect, an embodiment of the present invention provides a third possible implementation of the first aspect, wherein the above method also includes: extracting channel information of a target operation node connected to the data cache node; wherein the target operation node is at least one operation node constituting the target reasoning model; determining whether the data format in the data cache node matches the channel information of the target operation node; if not, adding a pre-configured first data splitting node at the output end of the data cache node, splitting the data cached in the data cache node through the first data splitting node, and converting the data format of the split data into a data format that matches the channel information of the target operation node.
[0010] In combination with the third possible implementation of the first aspect, an embodiment of the present invention provides a fourth possible implementation of the first aspect, wherein the above method also includes: if the data format in the data cache node matches the channel information of the target operation node, then adding a pre-configured second data splitting node at the output end of the data cache node, and splitting the data cached in the data cache node through the second data splitting node.
[0011] In combination with the first aspect, an embodiment of the present invention provides a fifth possible implementation of the first aspect, wherein the above method also includes: adding the conversion node at the output end of the data storage node, and converting the data format of the data stored in the data storage node into a preset data format through the conversion node.
[0012] In combination with the third possible implementation of the first aspect, an embodiment of the present invention provides a sixth possible implementation of the first aspect, wherein the above method also includes: determining the target reasoning model after adding the node as an updated reasoning model; deploying the updated reasoning model to an embedded device to perform reasoning operations using the updated reasoning model through the embedded device.
[0013] In a second aspect, an embodiment of the present invention further provides a processing device for an inference model, the device comprising: an acquisition module, used to acquire a target inference model, the target inference model being a pre-constructed inference model library, the inference model being composed of a plurality of operation nodes connected together, and each of the operation nodes being configured with a data storage node, the data output by the operation node being stored in the data storage node in accordance with a preset data format; an extraction module, used to extract channel information of each of the operation nodes in the target inference model; a judgment module, used to judge whether the data format of the current input data of the operation node matches the channel information; an adding module, used to acquire a pre-configured conversion node when the judgment result of the judgment module is no; and adding the conversion node to the input end of the operation node to perform format conversion on the input data of the operation node.
[0014] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method described in the first aspect when executing the computer program.
[0015] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to perform the steps of the method described in the first aspect.
[0016] The embodiments of the present invention bring the following beneficial effects:
[0017] The processing method, device and electronic device of the reasoning model provided by the embodiments of the present invention can obtain the target reasoning model and extract the channel information of each operation node in the target reasoning model; determine whether the data format of the current input data of the operation node matches the channel information, and if not, obtain a pre-configured conversion node and add the conversion node to the input end of the operation node to convert the format of the input data of the operation node, so as to make the data input to the operation node more matched with the performance of the operation node, which can not only reduce the calculation time of the operation node, but also improve the performance of the operation node, thereby improving the efficiency of the entire target reasoning model.
[0018] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0019] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 A flowchart of a method for processing an inference model provided by an embodiment of the present invention;
[0022] Figure 2 A schematic diagram of processing results of an inference model provided by an embodiment of the present invention;
[0023] Figure 3 A schematic diagram of the structure of a processing device for an inference model provided by an embodiment of the present invention;
[0024] Figure 4 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0026] At present, in common reasoning frameworks, generally only one memory arrangement method is used globally for calculation, and different memory arrangements are not used for calculations for different nodes, which makes it impossible to optimize the performance of each node.
[0027] Based on this, an inference model processing method, device and electronic device provided in an embodiment of the present invention can effectively alleviate the above technical problems.
[0028] To facilitate understanding of this embodiment, a method for processing an inference model disclosed in an embodiment of the present invention is first introduced in detail.
[0029] In a possible implementation, the embodiment of the present invention provides a method for processing an inference model, such as Figure 1 A flowchart of a method for processing an inference model is shown, the method comprising:
[0030] Step S102, obtaining a target reasoning model;
[0031] The target reasoning model in the embodiment of the present invention is a reasoning model pre-built in the reasoning model library, the reasoning model is composed of a plurality of operation nodes connected, and each operation node is configured with a data storage node, and the data output by the operation node is stored in the data storage node according to a preset data format;
[0032] Generally, the above-mentioned reasoning model is also called an AI (Artificial Intelligence) reasoning model, and the framework of the reasoning model is usually used to deploy the software framework of the deep learning model. By importing the deep learning model and its input, the AI calculation results can be obtained.
[0033] Furthermore, the operation nodes of the inference model are also called Node nodes, such as Node nodes for convolution operations, deconvolution operations, etc. Furthermore, the above data storage nodes are also called Tensors, which are used for data in the inference model framework. Common ones include the input Tensor, output Tensor, and weight Tensor of the Node node, etc.
[0034] Step S104, extracting channel information of each operation node in the target reasoning model;
[0035] Step S106, determining whether the data format of the current input data of the operation node matches the channel information;
[0036] Step S108, if not, obtain a pre-configured conversion node;
[0037] Step S110: adding a conversion node to the input end of the operation node to perform format conversion on the input data of the operation node.
[0038] Specifically, the data format in the above step S106 usually refers to the number of channels of the data arrangement method. The arrangement method of the input data of the operation node is converted by the conversion node so that the number of channels of the input data matches the operation node, thereby improving the computing efficiency of the operation node and the computing performance of the entire reasoning model.
[0039] The processing method of the reasoning model provided by the embodiment of the present invention can obtain the target reasoning model and extract the channel information of each operation node in the target reasoning model; determine whether the data format of the current input data of the operation node matches the channel information, if not, obtain a pre-configured conversion node, and add the conversion node to the input end of the operation node to convert the format of the input data of the operation node, so as to make the data input to the operation node more matched with the performance of the operation node, which can not only reduce the calculation time of the operation node, but also improve the performance of the operation node, thereby improving the efficiency of the entire target reasoning model.
[0040] In actual use, in the above step S106, when judging whether the data format of the current input data of the operation node matches the channel information, it is performed based on the calculation method of the data format of the current input data, that is, judging whether the channel information of the operation node supports the calculation method of the data format of the current input data of the operation node; if so, determining that the data format of the current input data of the operation node matches the channel information; if not, determining that the data format of the current input data of the operation node does not match the channel information.
[0041] Furthermore, the above-mentioned reasoning model in the embodiment of the present invention is an audio streaming reasoning model; the audio streaming reasoning model is also provided with a data cache node, and the data cache node is used to cache the audio frame data of the current operation cycle and serve as input data for the next operation cycle.
[0042] In specific implementation, the above audio streaming inference model is usually embedded in a device that processes audio data in real time to periodically perform AI processing on the audio data. Usually, audio data has a correlation between the past and the future, that is, the AI inference of audio data in the current cycle needs to rely on the audio data in the previous cycle. Therefore, during the AI processing process, it is necessary to cache the data information of the previous few frames of audio data. This method of periodically performing AI calculations and storing the results of the previous cycle for the current cycle is called audio streaming inference, and the model that processes audio data can be called an audio streaming inference model.
[0043] Furthermore, the above-mentioned data cache node is generally also called a Storebuf node. In the audio AI inference scenario, part of the data cache of the current frame needs to be output to the Storebuf node and used as the input of the next frame.
[0044] Therefore, in an embodiment of the present invention, channel information of a target operation node connected to a data cache node can be extracted; wherein the target operation node is at least one operation node constituting a target reasoning model; then, a determination is made as to whether the data format in the data cache node matches the channel information of the target operation node; if not, a preconfigured first data splitting node is added to the output end of the data cache node, and the data cached in the data cache node is split through the first data splitting node, and the data format of the split data is converted into a data format that matches the channel information of the target operation node.
[0045] Further, if the data format in the data cache node matches the channel information of the target operation node, a preconfigured second data splitting node is added to the output end of the data cache node, and the data cached in the data cache node is split through the second data splitting node.
[0046] Furthermore, in the embodiment of the present invention, the above-mentioned conversion node may be added to the output end of the data storage node, and the data format of the data stored in the data storage node may be converted into a preset data format through the conversion node.
[0047] In addition, in an embodiment of the present invention, the target reasoning model after adding the node can also be determined as an updated reasoning model, and the updated reasoning model can be deployed to the embedded device so that the updated reasoning model can be used to perform reasoning operations through the embedded device.
[0048] In actual use, the conversion node in the embodiment of the present invention generally refers to a node that converts the format of the input data of the operation node. Taking the inference model as an audio streaming inference model as an example, the arrangement of the input data of the operation node in the data storage node is commonly NCHW, NWHC, NC4HW4, etc. In the framework of common inference models, the input Tensor of the inference model is generally fixed to the NCHW arrangement. For the streaming inference of the audio streaming inference model, it is necessary to synchronously output the intermediate data of this inference to the next frame. Therefore, for the scenario where the arrangement of the data cache node is NC4HW4, the framework of the inference model will be converted to NCHW output, and the next time it is used, it needs to be converted to NC4HW4, which adds a lot of extra operations in this process.
[0049] In an embodiment of the present invention, by analyzing the structure of the model to be inferred, the structure of the inference model is modified according to the input and output tensor information of each operation node, so that the data storage node of each operation node can use the most efficient memory arrangement for inference, thereby improving AI computing performance.
[0050] For ease of understanding, the processing of the inference model is described by taking the audio streaming inference model as an example, including the following steps:
[0051] (1) Use open source frameworks such as pytorch and tensorflow to train and generate inference models, and export the inference models in onnx format.
[0052] (2) Analyze the optimal calculation method for each operation node (Node node) in the reasoning model to add conversion nodes.
[0053] Specifically, in this process, the number of channels is analyzed, and conversion nodes are added where appropriate.
[0054] Furthermore, the conversion node in the embodiment of the present invention is generally a Layout-convert conversion node.
[0055] (3) Analyze the data cache node of the inference model, that is, the splitting scheme of the Storebuf node, to determine whether to add the first data splitting node or the second data splitting node at the output end of the Storebuf node.
[0056] In specific implementation, the first data splitting node and the second data splitting node in the process can both be referred to as data splitting nodes, which can be implemented by Storebuf-split nodes. Specifically, they can be set according to the operation node connected to the output end of the data cache node to configure the parameters of the corresponding data splitting node.
[0057] (4) Export the inference model to which the conversion node, the first data splitting node or the second data splitting node is added in (2) or (3).
[0058] (5) The inference model is deployed to the embedded device. When the inference model is inferring, the framework of the inference model will perform corresponding processing according to the data format of the current Node node.
[0059] further, Figure 2 A schematic diagram of the processing result of an inference model is shown, wherein it is assumed that the operation nodes of the inference model include convolution operation nodes Conv1, Conv2, Conv3, and connection nodes Concat1, Concat2, a data storage node Tensor of each operation node, and a data cache node Storebuf, wherein the dotted box shows the framework of the inference model, the data storage node Tensor1 and the data cache node Storebuf1 above the dotted box are the inputs of the inference model, and, it is assumed that the arrangement modes matched by the data storage node Tensor1 and the data cache node Storebuf1 are both NCHW, and, Figure 2 The multiple Tensors in represent the data storage nodes of the corresponding operation nodes.
[0060] Furthermore, assuming that the data output by the convolution operation node Conv1 is arranged in NCHW, the arrangement matched by the connection node Concat1 connected thereto is also NCHW, therefore, there is no need to add a conversion node between the convolution operation node Conv1 and the connection node Concat1.
[0061] Furthermore, assuming that the arrangement of the data output by the connection node Concat1 is NCHW, the arrangement matched by the subsequently connected convolution operation node Conv2 is NC4HW4. Therefore, it is necessary to add a conversion node Layout-convert1 between the connection node Concat1 and the convolution operation node Conv2.
[0062] Moreover, since the arrangement mode matched by the convolution operation node Conv2 is NC4HW4, the arrangement mode of its output data is also NC4HW4. Therefore, a first data splitting node Storebufconcat1 can be added at the output end of the convolution operation node Conv2 to arrange the output data of the convolution operation node Conv2 in NCHW and store it in the data storage node, which can be used as cache data for subsequent frames.
[0063] further, Figure 2In the figure, it is assumed that the data of the data cache node Storebuf1 needs to be input to the connection node Concat1 and the connection node Concat2 for processing respectively, and it is assumed that the connection node Concat2 and the subsequent convolution operation node Conv3 match the arrangement of NC4HW4. Therefore, it is necessary to add the first data splitting node Storebufconcat1 at the output end of the data cache node Storebuf1, so that part of the output data of the data cache node Storebuf1 is input to the connection node Concat1 in the arrangement of NCHW, and the other part is input to the connection node Concat2 in the arrangement of NC4HW4.
[0064] At the same time, since the arrangement matching the convolution operation node Conv3 is NC4HW4, there is no need to add a conversion node between the connection node Concat2 and the convolution operation node Conv3.
[0065] It is further assumed that the data between the convolution operation nodes Conv3 needs to be stored in the NCHW arrangement. Therefore, a conversion node Layout-convert2 needs to be added to the output node of the convolution operation node Conv3.
[0066] based on Figure 2 The processing result of the reasoning model shown can be deployed to the embedded device as an updated reasoning model to select the optimal calculation method according to the channel information of the operation node, such as the layout format, thereby reducing the calculation time of the reasoning model.
[0067] Further, based on the above embodiment, the embodiment of the present invention also provides a processing device for an inference model, such as Figure 3 A schematic diagram of the structure of a processing device for an inference model is shown, the device comprising the following structure:
[0068] An acquisition module 30 is used to acquire a target reasoning model, wherein the target reasoning model is a pre-built reasoning model in a reasoning model library, wherein the reasoning model is composed of a plurality of operation nodes connected together, and each of the operation nodes is configured with a data storage node, and the data output by the operation node is stored in the data storage node according to a preset data format;
[0069] An extraction module 32, used to extract channel information of each of the operation nodes in the target reasoning model;
[0070] A judgment module 34, used to judge whether the data format of the current input data of the operation node matches the channel information;
[0071] The adding module 36 is used to obtain a pre-configured conversion node when the judgment result of the judging module is no; add the conversion node to the input end of the operation node to perform format conversion on the input data of the operation node.
[0072] The processing device for the inference model provided in the embodiment of the present invention has the same technical features as the processing method for the inference model provided in the above embodiment, so it can also solve the same technical problems and achieve the same technical effects.
[0073] Furthermore, an embodiment of the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above method when executing the computer program.
[0074] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are executed.
[0075] Furthermore, an embodiment of the present invention also provides a structural diagram of an electronic device, such as Figure 4 As shown, it is a schematic diagram of the structure of the electronic device, wherein the electronic device includes a processor 41 and a memory 40, the memory 40 stores computer executable instructions that can be executed by the processor 41, and the processor 41 executes the computer executable instructions to implement the above method.
[0076] exist Figure 4 In the illustrated embodiment, the electronic device further includes a bus 42 and a communication interface 43 , wherein the processor 41 , the communication interface 43 and the memory 40 are connected via the bus 42 .
[0077] Among them, the memory 40 may include a high-speed random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 43 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 42 can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 42 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0078] The processor 41 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the hardware integrated logic circuit in the processor 41 or the instruction in the form of software. The above processor 41 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in the embodiment of the present invention can be directly embodied as a hardware decoding processor to execute, or it can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the field such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor 41 reads the information in the memory and completes the above method in combination with its hardware.
[0079] The computer program product of the inference model processing method, device and electronic device provided in the embodiments of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods described in the previous method embodiments. The specific implementation can be found in the method embodiments, which will not be repeated here.
[0080] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the device described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0081] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, and it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0082] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program codes.
[0083] In the description of the present invention, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc., indicating the orientation or positional relationship, are based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first", "second", and "third" are used for descriptive purposes only, and cannot be understood as indicating or implying relative importance.
[0084] Finally, it should be noted that the above embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can still modify the technical solutions recorded in the above embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A method for processing an inference model, characterized in that: The method comprises: Obtain a target reasoning model, wherein the target reasoning model is a pre-built reasoning model in a reasoning model library, wherein the reasoning model is composed of a plurality of operation nodes connected together, and each of the operation nodes is configured with a data storage node, and the data output by the operation node is stored in the data storage node according to a preset data format; Extracting channel information of each of the operation nodes in the target reasoning model; Determine whether the data format of the current input data of the operation node matches the channel information; If no, get the pre-configured conversion node; The conversion node is added to the input end of the operation node to perform format conversion on the input data of the operation node.
2. The method according to claim 1, characterized in that The step of determining whether the data format of the current input data of the operation node matches the channel information comprises: Determine whether the channel information of the operation node supports the calculation method of the data format of the current input data of the operation node; If yes, determining that the data format of the current input data of the operation node matches the channel information; If not, it is determined that the data format of the current input data of the operation node does not match the channel information.
3. The method according to claim 1, characterized in that The inference model is an audio streaming inference model; The audio streaming inference model is also provided with a data cache node, and the data cache node is used to cache the audio frame data of the current operation cycle and serve as input data for the next operation cycle.
4. The method according to claim 3, characterized in that The method further comprises: Extracting channel information of a target operation node connected to the data cache node; wherein the target operation node is at least one operation node constituting the target reasoning model; Determine whether the data format in the data cache node matches the channel information of the target operation node; If not, add a preconfigured first data splitting node at the output end of the data caching node, split the data cached in the data caching node through the first data splitting node, and convert the data format of the split data into a data format that matches the channel information of the target operation node.
5. The method according to claim 4, characterized in that The method further comprises: If the data format in the data cache node matches the channel information of the target operation node, a preconfigured second data splitting node is added to the output end of the data cache node, and the data cached in the data cache node is split through the second data splitting node.
6. The method according to claim 1, characterized in that The method further comprises: The conversion node is added at the output end of the data storage node, and the data format of the data stored in the data storage node is converted into a preset data format through the conversion node.
7. The method according to claim 4, characterized in that The method further comprises: determining the target reasoning model after adding the node as the updated reasoning model; The updated inference model is deployed to the embedded device so as to perform inference operations using the updated inference model through the embedded device.
8. A processing device for an inference model, characterized in that: The device comprises: An acquisition module is used to acquire a target reasoning model, wherein the target reasoning model is a pre-built reasoning model in a reasoning model library, wherein the reasoning model is composed of a plurality of operation nodes connected together, and each of the operation nodes is configured with a data storage node, and the data output by the operation node is stored in the data storage node according to a preset data format; An extraction module, used to extract channel information of each of the operation nodes in the target reasoning model; A judging module, used to judge whether the data format of the current input data of the operation node matches the channel information; The adding module is used to obtain a pre-configured conversion node when the judgment result of the judging module is no; and add the conversion node to the input end of the operating node to perform format conversion on the input data of the operating node.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 7 are executed.