Neural network quantitative perception training method, image acquisition device and storage medium

By inserting pseudo-quantization modules into the neural network model and performing grouping regularization training, the problems of additional operations and network depth reduction in quantization perception training are solved, and efficient quantization operations and high-precision neural network model are realized.

CN119940445APending Publication Date: 2025-05-06BEIJING WEIMAI MEDICAL EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411988522.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art In neural network quantization perception training, the insertion of pseudo-quantization modules leads to additional operations, loss of time performance, and pre-merging the network layer will reduce network depth and lower evaluation indicators.

Method used

By inserting the pseudo-quantization module at each target position of the converged neural network model, multiple pseudo-quantization modules with consistent quantization parameters of the pseudo-quantization module are divided into the same group, and regularization training is performed using the cost function of the regular term, the representative values ​​of the quantization parameters are determined and fixedly set, avoiding updating the quantization parameters in quantization perception training.

Benefits of technology

It effectively avoids additional operations caused by insertion of pseudo-quantization modules, maintains time performance, and ensures that the network depth does not decrease by fixed quantization parameters, thereby improving the quantization operation efficiency and prediction accuracy of neural network models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940445A_ABST
    Figure CN119940445A_ABST
Patent Text Reader

Abstract

The invention provides a neural network quantitative perception training method, an image acquisition device and a storage medium. The method comprises the following steps: training an initial neural network model to convergence based on a preset training strategy, and inserting a pseudo-quantization module at each target position of the converged neural network model to obtain a pre-training model; dividing a plurality of pseudo-quantization modules which can perform neural network layer merging only when quantization parameters of the pseudo-quantization modules are consistent into the same group; performing regularization training on the pre-training model on the basis of the cost function containing the regularization term; determining a quantization parameter representative value of each pseudo quantization module group, and fixedly setting the current quantization parameter of each pseudo quantization module as the quantization parameter representative value; and performing quantitative perception training based on the quantitative perception training model. Therefore, additional operation caused by insertion of a pseudo quantization module can be avoided, the efficiency of model operation is improved, and the prediction time of deep learning in the image acquisition device is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of deep learning technology, and in particular to a method for quantized perceptual training of a neural network, an image acquisition device, and a storage medium. Background Art

[0002] A neural network is a computational model that simulates the working principle of neurons in the human brain. It simulates complex nonlinear relationships through a large number of neurons and connection weights, and is widely used in image recognition, speech recognition, natural language processing, and other fields. Traditional neural network models usually use high-precision floating-point numbers to represent the connection weights between neurons, which leads to high storage and computing overhead for neural networks. As the scale of neural networks continues to expand and the number of application scenarios increases, the demand for storage and computing resources is also increasing. Quantization is one of the ways to solve this problem.

[0003] Quantization significantly reduces the storage and computational overhead of neural networks by compressing floating-point numbers into a representation with lower bits. However, quantization introduces quantization errors, resulting in loss of precision. A common method to mitigate quantization errors is to insert pseudo-quantization modules into neural networks for quantization-aware training, treating quantization errors as noise for neural network learning. However, the initial model used for quantization awareness often causes some network structures that can perform merge operations to generate additional operations during the inference process after quantization due to the insertion of pseudo-quantization modules, thereby losing time performance. Some solutions are to preferentially merge the mergeable layers of the network before quantization awareness, but the merged network is equivalent to reducing the depth of the network, and the resulting evaluation index will be lower than the non-pre-merged quantization-aware results. Summary of the invention

[0004] This application is proposed to solve the above technical problems in the prior art. This application aims to provide a method for neural network quantization perception training, an image acquisition device and a storage medium, which can avoid additional operations caused by the insertion of a pseudo quantization module and improve the prediction accuracy of the neural network model.

[0005] According to the first scheme of the present application, a method for quantization-aware training of a neural network is provided, the method comprising: training an initial neural network model to convergence based on a preset training strategy, inserting a pseudo quantization module at each target position of the converged neural network model to obtain a pre-trained model; grouping multiple pseudo quantization modules in the pre-trained model, wherein multiple pseudo quantization modules that require consistent quantization parameters of the pseudo quantization modules for neural network layer merging are grouped into the same group; continuing to perform regularization training on the grouped pre-trained models based on a cost function containing a regularization term to obtain a regularized pre-trained model; determining a representative value of the quantization parameter of each pseudo quantization module group, and fixing the current quantization parameter of each pseudo quantization module in each pseudo quantization module group to the representative value of the quantization parameter, to obtain a quantization-aware training model in which the quantization parameters of each pseudo quantization module in each group are consistent; performing quantization-aware training based on the quantization-aware training model, and during the quantization-aware training, not updating the quantization parameters of each pseudo quantization module in each pseudo quantization module group.

[0006] According to a second solution of the present application, a neural network-based image acquisition device is provided, wherein the image acquisition device includes a processor configured to execute the method of neural network quantized perceptual training described in each embodiment of the present application.

[0007] According to a third embodiment of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor executes the method for quantized perceptual training of a neural network as described in the various embodiments of the present application.

[0008] According to a fourth embodiment of the present application, a computer program product is provided, including a computer program, which, when executed by a processor, implements the method for quantized perceptual training of a neural network as described in the various embodiments of the present application.

[0009] Compared with the prior art, the beneficial effects of the embodiments of the present application are:

[0010] The method provided in the embodiment of the present application inserts a pseudo-quantization module at each target position of the converged neural network model, and groups multiple pseudo-quantization modules that require consistent quantization parameters of the pseudo-quantization modules for merging neural network layers into the same group, and then uses a cost function containing a regularization term to train the grouped pre-trained model to obtain a regularized pre-trained model, so that the quantization parameters of each pseudo-quantization module in each group of pseudo-quantization modules in the regularized pre-trained model are close, and the quantization parameters of each pseudo-quantization module in each group will not have a large deviation.

[0011] Determine the representative value of the quantization parameter of each pseudo-quantization module group, and fix the quantization parameter of each pseudo-quantization module in each pseudo-quantization module group to the representative value of the quantization parameter, so as to obtain a quantization-aware training model, and then perform quantization-aware training based on the quantization-aware training model. During the quantization-aware training, the quantization parameters of each pseudo-quantization module in the pseudo-quantization module group remain unchanged and will not be updated. In this way, no additional calculation process will be generated, thereby not losing time performance, and the depth of the network will not be reduced, thereby improving the efficiency of the quantization calculation while ensuring that the final neural network model has a high prediction accuracy.

[0012] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above description and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] In the drawings, which are not necessarily drawn to scale, the same reference numerals may describe similar parts in different views. Similar reference numerals with letter suffixes or different letter suffixes may represent different examples of similar components. The accompanying drawings generally illustrate various embodiments by way of example and not by way of limitation, and together with the description and claims, are used to illustrate the disclosed embodiments. Such embodiments are illustrative and exemplary and are not intended to be exhaustive or exclusive embodiments of the present method, apparatus, system, or non-transitory computer-readable medium having instructions for implementing the method.

[0014] Figure 1 A flow chart of a method for quantized perceptual training of a neural network according to an embodiment of the present application is shown.

[0015] Figure 2 A schematic diagram is shown of a situation in which the quantization parameters of a pseudo quantization module affect the operation efficiency according to an embodiment of the present application.

[0016] Figure 3 A schematic diagram of a neural network-based image acquisition device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0017] In order to enable those skilled in the art to better understand the technical solution of the present application, the present application is described in detail below in conjunction with the accompanying drawings and specific implementation methods. The embodiments of the present application are further described in detail below in conjunction with the accompanying drawings and specific implementation methods, but are not intended to limit the present application.

[0018] The words "first", "second" and similar words used in this application do not indicate any order, quantity or importance, but are only used to distinguish. The words "include" or "comprises" and similar words used in this application mean that the elements before the word include the elements listed after the word, and do not exclude the possibility of covering other elements. In this application, the arrows shown in the figures of each step are only examples of the execution order, not limitations. The technical solution of this application is not limited to the execution order described in the embodiments. The various steps in the execution order can be combined, decomposed, or the order can be changed, as long as the logical relationship of the execution content is not affected.

[0019] All terms (including technical terms or scientific terms) used in this application have the same meaning as those understood by ordinary technicians in the field to which this application belongs, unless otherwise specifically defined. It should also be understood that terms defined in general dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an idealized or extremely formal sense, unless explicitly defined here. Techniques and equipment known to ordinary technicians in the relevant field may not be discussed in detail, but where appropriate, the techniques and equipment should be considered as part of the specification.

[0020] Figure 1 The flowchart of the method for quantized perceptual training of a neural network according to an embodiment of the present application is shown. The method includes steps S101 to S105, and the arrows shown in the figure for each step are only examples of the execution order, not limitations. The technical solution of the present application is not limited to the execution order described in the embodiment, and each step in the execution order can be combined, decomposed, or changed in order, as long as it does not affect the logical relationship of the execution content.

[0021] In step S101, the initial neural network model is trained to convergence based on a preset training strategy, and a pseudo-quantization module is inserted at each target position of the converged neural network model to obtain a pre-trained model. The initial neural network model has a basic network architecture, but the internal parameters have not been optimized to adapt to specific tasks, and these parameters are usually randomly initialized. The initial neural network model, such as a multi-layer perceptron (MLP), a convolutional neural network (CNN), a recurrent neural network (RNN) and its variants (LSTM, GRU), etc., is not limited to this.

[0022] The preset training strategy is a series of rules and methods that are pre-set before training the initial neural network model to guide the training process of the initial neural network model. The setting of the preset training strategy includes the selection and parameter setting of the optimizer, the determination of the loss function, the learning rate adjustment strategy, the number of training rounds and the determination of the batch size. This is only used as an example and does not constitute a limitation on the specific solution. Users can set the preset training strategy according to actual needs.

[0023] Specifically, the initial neural network model is trained based on a preset training strategy. As the training progresses, the loss function value of the neural network model will gradually decrease. When the value of the loss function no longer decreases significantly or fluctuates within a very small range, the neural network model is trained to convergence.

[0024] Quantization is a method of converting floating-point operations in neural networks into low-precision integer operations. The pseudo-quantization module is a module that simulates the quantization process, mainly to reduce the computational workload and storage requirements of the neural network model. Quantization can make the neural network model run faster and take up less memory.

[0025] The target position can be determined according to the structure and task of the neural network model. For example, in a convolutional neural network, a pseudo-quantization module can be inserted after the convolution layer. This is only used as an example and does not constitute a limitation on the specific solution. The user can set the target position by himself.

[0026] In step S102, the plurality of pseudo-quantization modules in the pre-trained model are grouped, wherein the plurality of pseudo-quantization modules that require the quantization parameters of the pseudo-quantization modules to be consistent in order to merge the neural network layers are grouped into the same group. Specifically, after inserting the pseudo-quantization modules at each target position of the converged neural network model and obtaining the pre-trained model, the plurality of pseudo-quantization modules are grouped to obtain a plurality of pseudo-quantization module groups. The basis for grouping the plurality of pseudo-quantization modules is that the plurality of pseudo-quantization modules in the group need to have consistent quantization parameters, and the pseudo-quantization modules that require the same quantization parameters in order to merge are selected to form a group, while the pseudo-quantization modules that do not need to have consistent quantization parameters are not grouped.

[0027] In the neural network, the quantization parameter determines how the pseudo-quantization module converts high-bit neural network parameters into equivalent low-bit neural network parameters. When merging neural network layers, ensuring that the pseudo-quantization modules in the same group participating in the merger have consistent quantization parameters can avoid additional calculations caused by the inability to merge pseudo-quantization modules, which affects the accuracy and precision of the neural network model.

[0028] Continue to execute step S103, and continue to perform regularization training on the grouped pre-trained model based on the cost function containing the regularization term to obtain a regularized pre-trained model. Specifically, add the cost function containing the regularization term to the loss of the grouped pre-trained model and continue training, so as to ensure that the quantization parameters of each pseudo-quantization module in each pseudo-quantization module group are as close as possible.

[0029] Among them, a cost function containing a regularization term is added to the original loss function of the pre-trained model, and the cost function is shown in formula (1):

[0030] L=L o + R formula (1);

[0031] In formula (1), L is the cost function, Lo is the original loss function, R is the regularization term, and W is the weight of the regularization term;

[0032] The expression of R is shown in formula (2):

[0033]

[0034] In formula (2), G is the number of pseudo-quantization module groups, N i is the number of pseudo quantization modules in the i-th pseudo quantization module group, is the output value of the jth pseudo-quantization module in the i-th group, max i is the maximum value of the output values ​​of all pseudo-quantization modules in group i, min i It is the minimum value of the output values ​​of all pseudo quantization modules in the i-th group.

[0035] Based on the cost function containing the regularization term of the embodiment of the present application, the grouped pre-trained model continues to be regularized trained until convergence to obtain a regularized pre-trained model.

[0036] The quantization parameters are gain parameters and offset parameters. The gain parameter plays a scaling role in the quantization process and determines the scaling ratio of the quantized data range relative to the original data range. For example, assuming that the original data range is [0,10], if the gain parameter is 2, then the quantized data range may be scaled to [0,20]. In neural networks, the gain parameter can help adjust the distribution of data to make it more suitable for quantization operations.

[0037] The offset parameter is used to adjust the center position of the data after quantization, which is equivalent to adding an offset to the quantized value. For example, in a quantization process, the range of the original data after quantization is [0,10]. If the offset parameter is 3, the final quantized data range will become [3,13]. In neural networks, the offset parameter can be used to align data in different layers or make the data conform to specific quantization format requirements.

[0038] In step S104, the representative value of the quantization parameter of each pseudo-quantization module group is determined, and the current quantization parameter of each pseudo-quantization module in each pseudo-quantization module group is fixedly set to the representative value of the quantization parameter, so as to obtain a quantization perception training model in which the quantization parameters of each pseudo-quantization module in each group are consistent.

[0039] Based on the regularized training of the embodiment of the present application, the quantization parameters of each pseudo quantization module in each pseudo quantization module group are made close. As an exemplary illustration, the average value of the quantization parameters in each pseudo quantization module group can be used as the representative value of the quantization parameter, for example, the average value of the gain parameter in each pseudo quantization module group is calculated as the representative value of the gain parameter, and the average value of the offset parameter in each pseudo quantization module group is calculated as the representative value of the offset parameter.

[0040] Among them, the median value of the quantization parameter of each pseudo quantization module in each pseudo quantization module group can also be used as the representative value of the quantization parameter, which is only used as an exemplary description and does not constitute a limitation on the specific solution.

[0041] Specifically, each pseudo-quantization module group has a set quantization parameter representative value, and then the quantization parameter of each pseudo-quantization module in each pseudo-quantization module group is fixedly set to the quantization parameter representative value of the corresponding group. For example, for the first pseudo-quantization module group including 5 pseudo-quantization modules, its gain parameter representative value is 2, and its offset parameter representative value is 4, then the gain parameters of the 5 pseudo-quantization modules in the first pseudo-quantization module group are all fixedly set to 2, and the offset parameters are all fixedly set to 4.

[0042] Moreover, for each pseudo-quantization module group with fixed quantization parameters set, in the subsequent training process, the quantization parameters of each pseudo-quantization module in each pseudo-quantization module group will not be adjusted and updated, and only the quantization parameters of the pseudo-quantization modules that are not divided into pseudo-quantization module groups will be adjusted and updated.

[0043] In some other embodiments, the method for determining the representative value of the quantization parameter may further include obtaining the quantization range of each pseudo quantization module group, and determining the representative value of the quantization parameter of each pseudo quantization module group based on the quantization range of each pseudo quantization module group. Specifically, when calculating the quantization parameters of each pseudo quantization module group, the quantization range of each pseudo quantization module group is first calculated, and each pseudo quantization module group has the same quantization range.

[0044] The quantization range can be expressed as {x|x∈[α i ,β i ]}, where α i is the minimum value of the quantization range of the i-th pseudo-quantization module group, β i is the maximum value of the quantization range of the i-th pseudo quantization module group, x is the quantization parameter, where α and β are calculated as shown in formulas (3) and (4) respectively:

[0045]

[0046] Among them, N is a hyperparameter, max(,i) is the i-th largest value in the x matrix, and min(,i) is the i-th smallest value in the x matrix.

[0047] In general, the value of N is positively correlated with the number of pseudo-quantization modules in the pseudo-quantization module group, and can also be calculated according to the empirical formula. The empirical formula for the value of N is: Where n is the number of pseudo quantization modules in the current pseudo quantization module group, is the floor operator, and N is the value of the hyperparameter empirical formula.

[0048] Afterwards, the representative values ​​of the quantization parameters of each pseudo quantization module group are calculated, including the representative value S of the gain parameter and the representative value Z of the offset parameter. The calculation of the representative value S of the gain parameter and the representative value Z of the offset parameter of each pseudo quantization module group is shown in formula (5) and formula (6), respectively:

[0049]

[0050] Z i =-round(β i ·S i )-2 b-1 Formula (6);

[0051] In formula (5), S i and Z in formula (6) i is the gain parameter representative value S and offset parameter representative value Z of the i-th pseudo quantization module group, b is the number of quantization bits, α i and β i are respectively the minimum and maximum values ​​of the quantization range of the i-th group of pseudo quantization modules.

[0052] Then, the gain parameter of each pseudo quantization module in each pseudo quantization module group is fixedly set to the gain parameter representative value S corresponding to each group, and the offset parameter is fixedly set to the offset parameter representative value X corresponding to each group.

[0053] The fixed setting means that the gain parameter representative value S and the offset parameter representative value Z will not be modified or learned during the subsequent quantization perception training process.

[0054] Back to the embodiment of the present application, in step S105, quantization-aware training is performed based on the quantization-aware training model, and during the quantization-aware training, the quantization parameters of each pseudo-quantization module in each pseudo-quantization module group are not updated. That is, when the current quantization parameters of each pseudo-quantization module in each pseudo-quantization module group are fixedly set to the representative value of the quantization parameter, during the quantization-aware training based on the quantization-aware training model, the quantization parameters of the grouped pseudo-quantization modules are not updated and learned, while other parameters continue to be updated and optimized.

[0055] In addition, during the quantization-aware training based on the quantization-aware training model, the quantization parameters of each pseudo-quantization module that is not grouped are kept updated or the quantization parameters are fixed after being updated several times and are no longer updated to ensure the stability of the neural network model.

[0056] like Figure 2 As shown in the figure, the dotted box is the pseudo quantization module, based on Figure 2 It can be seen that the quantization modules Q and DQ form a pseudo-quantization module. The outputs after convolution calculation are input to different neural network structures respectively. The outputs need to be pseudo-quantized twice to obtain different outputs y1 and y2. These two pseudo-quantizations are the pseudo-quantization composed of Q1 and DQ1 and the pseudo-quantization composed of Q2 and DQ2. When the quantization parameters of the quantization modules Q1 and Q2 are the same, y1 and y2 are equivalent, then Q1 and Q2 can be fused and quantized with the preceding DQ. Otherwise, Q1 and Q2 can only be fused and quantized with the preceding DQ separately because the insertion of the pseudo-quantization module causes additional calculations.

[0057] That is to say, traditionally, when a neural network model with a complex structure is subjected to quantization perception, the insertion of a pseudo-quantization module will cause the structure that was originally capable of merging and accelerating to generate additional calculation processes, resulting in a loss of time performance.

[0058] The method provided in the embodiment of the present application groups the pseudo quantization modules, and sets the quantization parameters of each pseudo quantization module in each pseudo quantization module group to a fixed quantization parameter representative value, that is, Figure 2Q1 and Q2 in the method have the same quantization parameters, so that the obtained neural network model can be used for quantization-aware training. Since the structure is not merged before the pseudo-quantization module is inserted, the original network depth and nonlinear transformation capability can be maintained, and better accuracy can be obtained. In addition, the gain parameters and offset parameters of each pseudo-quantization module in each group that affects the merge are unified, so that the insertion of the pseudo-quantization module will not cause additional calculation process in the post-quantization reasoning process, and the optimal time performance can be maintained.

[0059] In this way, applying the neural network model obtained based on the method of the embodiment of the present application to medical image fields such as image acquisition can improve the real-time performance and prediction accuracy of the acquired images.

[0060] In some other embodiments of the present application, a neural network-based image acquisition device is provided, such as Figure 3 As shown, the image acquisition device 300 includes a processor 301, which is configured to execute the method of neural network quantization perception training described in various embodiments of the present application.

[0061] Specifically, for example, during radiotherapy, it is necessary to use an image acquisition device to collect medical images of the patient in real time so that the doctor can accurately locate the tumor, determine the radiation dose and irradiation range, etc. Since the patient's body may move slightly during the treatment, the location of the tumor may also change. In order to ensure the accuracy and safety of the treatment, it is required that the collected images can be quickly acquired and fed back to the doctor and treatment equipment in a timely manner so that the treatment plan can be adjusted in time.

[0062] The neural network model obtained by the method of neural network quantization perception training described in each embodiment of the present application is applied to an image acquisition device, so that during the real-time acquisition of medical images of patients by the image acquisition device, the insertion of a pseudo-quantization module that causes additional calculations can be avoided, thereby effectively reducing the inference time of the neural network model in the image acquisition device, thereby improving the real-time performance of image acquisition and ensuring the accuracy of image acquisition.

[0063] In addition, the image acquisition device 300 further includes a display 302, so that the medical images acquired by the image acquisition device 300 can be presented on the display 302. The display 302 can also be used to display medical reports, etc., which is not limited.

[0064] The processor 301 may be a processing device including one or more general-purpose processing devices, such as a microprocessor, a central processing unit (CPU), a graphics processing unit (GPU), etc. More specifically, the processor 301 may be a complex instruction set computing (CISC) microprocessor, a reduced instruction set computing (RISC) microprocessor, a very long instruction word (VLIW) microprocessor, a processor running other instruction sets, or a processor running a combination of instruction sets. The processor 301 may also be one or more special-purpose processing devices, such as an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a system on a chip (SoC), etc.

[0065] According to an embodiment of the present application, a computer-readable storage medium is also provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the steps of the method for quantized perceptual training of a neural network as described in the various embodiments of the present application.

[0066] The above-mentioned computer-readable storage medium may be, for example, read-only memory (ROM), random access memory (RAM), phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), electrically erasable programmable read-only memory (EEPROM), other types of random access memory (RAM), flash disks or other forms of flash memory, cache, registers, static memory, compact disk read-only memory (CD-ROM), digital versatile disks (DVD) or other optical storage, cassettes or other magnetic storage devices, or any other possible non-temporary medium used to store information or instructions that can be accessed by a computer device.

[0067] According to an embodiment of the present application, a computer program product is also provided, which includes a computer program. When the computer program is executed by a processor, the steps of the method for quantization-aware training of a neural network as described in each embodiment of the present application are implemented.

[0068] This application describes various operations or functions that may be implemented as software codes or instructions or defined as software codes or instructions. Such content may be source code or differential code ("incremental" or "patch" code) that can be directly executed ("object" or "executable" form). Software codes or instructions may be stored in a computer-readable storage medium, and when executed, may cause a machine to perform the described functions or operations, and include any mechanism for storing information in a form accessible to a machine (e.g., a computing device, an electronic system, etc.), such as recordable or non-recordable media (e.g., read-only memory (ROM), random access memory (RAM), magnetic disk storage media, optical storage media, flash memory devices, etc.).

[0069] In addition, although exemplary embodiments have been described herein, the scope includes any and all embodiments based on the present application with equivalent elements, modifications, omissions, combinations (e.g., various embodiments intersecting schemes), adaptations or changes. The elements in the claims will be interpreted broadly based on the language adopted in the claims, and are not limited to the examples described in this specification or during the implementation of the application, and the examples will be interpreted as non-exclusive. Therefore, this specification and examples are intended to be considered as examples only, and the true scope and spirit are indicated by the following claims and the full scope of their equivalents.

[0070] The above description is intended to be illustrative rather than restrictive. For example, the above examples (or one or more of them) can be used in combination with each other. For example, those of ordinary skill in the art can use other embodiments when reading the above description. In addition, in the above-mentioned specific embodiments, various features can be grouped together to simplify the application. This should not be interpreted as an intention that a disclosed feature that does not require protection is necessary for any claim. On the contrary, the subject matter of the present application may be less than all the features of a specific disclosed embodiment. Thus, the claims are incorporated into the specific embodiments as examples or embodiments, wherein each claim is independently used as a separate embodiment, and it is considered that these embodiments can be combined with each other in various combinations or arrangements. The scope of the present application should be determined with reference to the full scope of the equivalent forms of the attached claims and these claims.

[0071] The above embodiments are only exemplary embodiments of the present application and are not intended to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present application within the essence and protection scope of the present application, and such modifications or equivalent substitutions shall also be deemed to fall within the protection scope of the present application.

Claims

1. A method for quantized perceptual training of a neural network, characterized in that: The method comprises: The initial neural network model is trained to convergence based on a preset training strategy, and a pseudo-quantization module is inserted at each target position of the converged neural network model to obtain a pre-trained model; Grouping multiple pseudo quantization modules in the pre-trained model, wherein multiple pseudo quantization modules that require consistent quantization parameters of the pseudo quantization modules for neural network layer merging are grouped into the same group; Continue to perform regularization training on the grouped pre-trained model based on the cost function containing the regularization term to obtain a regularized pre-trained model; Determine a quantization parameter representative value of each pseudo quantization module group, and set the current quantization parameter of each pseudo quantization module in each pseudo quantization module group to the quantization parameter representative value, so as to obtain a quantization perception training model in which the quantization parameters of each pseudo quantization module in each group are consistent; Quantization-aware training is performed based on the quantization-aware training model. During the quantization-aware training, the quantization parameters of each pseudo-quantization module in each pseudo-quantization module group are not updated.

2. The method according to claim 1, characterized in that: The method comprises: adding a cost function containing a regularization term to the original loss function of the pre-trained model, wherein the cost function is as shown in formula (1): L=L o +WR formula (1); In formula (1), L is the cost function, Lo is the original loss function, R is the regularization term, and W is the weight of the regularization term; where the expression of R is shown in formula (2): In formula (2), G is the number of pseudo-quantization module groups, N i is the number of pseudo quantization modules in the i-th pseudo quantization module group, is the output value of the jth pseudo-quantization module in the i-th group, max i is the maximum value of the output values ​​of all pseudo-quantization modules in group i, min i It is the minimum value of the output values ​​of all pseudo quantization modules in the i-th group.

3. The method according to claim 1, characterized in that The method comprises: obtaining a quantization range of each pseudo quantization module group, and determining a quantization parameter representative value of each pseudo quantization module group based on the quantization range of each pseudo quantization module group.

4. The method according to any one of claims 1 to 3, characterized in that: The quantization parameters are a gain parameter and an offset parameter.

5. The method according to claim 3, characterized in that: The quantization range is expressed as {x|x∈[α i ,β i ]}, where α i is the minimum value of the quantization range of the i-th pseudo-quantization module group, β i is the maximum value of the quantization range of the i-th pseudo-quantization module group, where α and β i The calculations are shown in formulas (3) and (4) respectively: Among them, N is a hyperparameter, max(x,i) is the i-th largest value in the x matrix, and min(x,i) is the i-th smallest value in the x matrix.

6. The method according to claim 4, characterized in that The calculation of the gain parameter S and offset parameter Z of each pseudo-quantization module group is shown in formula (5) and formula (6) respectively: Z i =-round(β i ·S i )-2 b-1 Formula (6) In the above two formulas, S i and Z i is the gain parameter S and offset parameter Z of the i-th pseudo-quantization module group, b is the number of quantization bits, α i and β i are respectively the minimum and maximum values ​​of the quantization range of the i-th group of pseudo quantization modules.

7. The method according to claim 1, characterized in that The method comprises: In the process of performing quantization-aware training based on the quantization-aware training model, the quantization parameters of each pseudo-quantization module that is not grouped are kept updated or the quantization parameters are fixed after being updated several times and are no longer updated.

8. An image acquisition device based on a neural network, characterized in that: The image acquisition device includes a processor configured to execute the method for quantized perceptual training of a neural network as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to perform the method for quantization-aware training of a neural network as described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method for quantization-aware training of a neural network as described in any one of claims 1 to 7 is implemented.