An image processing method, apparatus, storage medium, and electronic device

By introducing a square standardized BN layer into the image processing neural network, the problems of slow convergence speed during initialization and poor training results are solved, and fast convergence and high-accuracy image detection are achieved, which is suitable for information processing of vehicle autonomous driving environments.

CN114494145BActive Publication Date: 2025-07-04CHINA AUTOMOTIVE INNOVATION CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111631866.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-29
Publication Date
2025-07-04
Estimated Expiration
2041-12-29

AI Technical Summary

Technical Problem

The image processing neural network converges slowly during initialization and poor training results, which affects the image processing effect. Especially in autonomous vehicle driving, the existing BN layer methods fail to minimize generalization errors.

Method used

BN layer is introduced into the image processing neural network for square standardization conversion. By obtaining initialization parameters and using training data for training, we ensure the uniformity and difference of the input data, and avoid gradient vanishing and data distribution offset during the training process.

Benefits of technology

It improves the training convergence speed and detection accuracy of image processing neural networks, enhances the generalization ability of the model, and is suitable for environmental information processing in vehicle autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494145B_ABST
    Figure CN114494145B_ABST
Patent Text Reader

Abstract

The present invention discloses an image processing method, apparatus, storage medium and electronic device, including: obtaining initialization parameters of an image processing neural network to be trained; training the image processing neural network to be trained with training data to obtain a trained image processing neural network; the training data includes image data; the image processing neural network includes at least one intermediate layer, and the intermediate layer includes a BN layer, and the BN layer is used to perform square normalization conversion on the input data of the intermediate layer; performing image processing through the trained image processing neural network to obtain an image processing result. Thereby, the convergence speed can be increased when training an image detection neural network, the accuracy of image detection, recognition, etc. of the trained image detection neural network is high, and it is beneficial to the generalization of the detection model of the image detection neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular, to an image processing method, apparatus, storage medium, and electronic device. Background Art

[0002] In vehicle autonomous driving, it is necessary to process the acquired image data to obtain environmental information of the traffic environment where the vehicle is located. Usually, an image processing model is used to process the acquired image data. An image processing neural network model is a commonly used image processing model. Generally, based on sample image data, an image processing neural network model with a preset structure is trained so that the trained image processing neural network model can automatically process image data.

[0003] Before the image processing neural network model is trained, it needs to be initialized to obtain an initialized neural network model, and the initialized neural network model is trained during training to obtain a trained image processing neural network model. Among them, the initialization of the image processing neural network model affects the training result of the neural network model, directly affects the convergence speed during training, and whether it can converge to an acceptable result within an acceptable time. In some cases, the initialization parameters may result in poor training results, causing the trained image processing neural network model to be unable to obtain satisfactory image processing results.

[0004] Adding a BN layer to the neural network model can accelerate the network convergence speed. However, this BN layer often does not bring the best effect, and there may be different reasons: (1) The nature of the initialization parameters imposed during initialization cannot be maintained after training to a certain stage; (2) This method does not minimize the generalization error to the greatest extent. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides an image processing method, apparatus, storage medium, and electronic device, which can achieve a fast convergence speed during the training of the image processing neural network and a high image detection accuracy of the trained image processing neural network.

[0006] To achieve the above object of the invention, the present invention provides an image processing method, which includes:

[0007] Obtain the initialization parameters of the image processing neural network to be trained;

[0008] Train the image processing neural network to be trained using training data to obtain a trained image processing neural network; the training data includes image data; the image processing neural network includes at least one intermediate layer, and the intermediate layer includes a BN layer, and the BN layer is used to perform square normalization conversion on the input data of the intermediate layer;

[0009] Perform image processing through the trained image processing neural network to obtain an image processing result.

[0010] Optionally, the intermediate layer includes a convolutional layer, the BN layer, and an activation function layer arranged in sequence.

[0011] Optionally, training the image processing neural network to be trained with training data to obtain a trained image processing neural network includes:

[0012] Obtain the input data of the intermediate layer;

[0013] The BN layer performs square normalization conversion on the input data to obtain normalized conversion data, and uses the normalized conversion data as the activation input data of the activation function layer.

[0014] Optionally, the BN layer performs square normalization conversion on the input data to obtain normalized conversion data, and uses the normalized conversion data as the activation input data of the activation function layer, including:

[0015] According to the input data of the intermediate layer, obtain the sign feature, square value, mean square value, and variance of the input data;

[0016] Translate and scale the input data according to the sign feature, square value, mean square value, and variance of the input data to obtain normalized conversion data, and use the normalized conversion data as the activation input data of the activation function layer.

[0017] Optionally, the intermediate layer includes a convolutional layer, the BN layer, an activation function layer, a pooling layer, and a fully connected layer arranged in sequence.

[0018] Optionally, the activation function layer includes a non-linear function.

[0019] Optionally, the initialization parameters include the weights and biases of the image detection neural network.

[0020] On the other hand, the present invention also provides an image processing device, and the device includes:

[0021] An initialization module for obtaining the initialization parameters of the image processing neural network to be trained;

[0022] A training module for training the image processing neural network to be trained with training data to obtain a trained image processing neural network; the training data includes image data; the image processing neural network includes at least one intermediate layer, and the intermediate layer includes a BN layer, and the BN layer is used to perform square normalization conversion on the input data of the intermediate layer;

[0023] A processing module, configured to perform image processing through the trained image processing neural network to obtain an image processing result.

[0024] On the other hand, the present invention also provides a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the image processing method as described in the above technical solution.

[0025] On the other hand, the present invention also provides an image processing electronic device, which includes a processor and a memory, and a program is stored in the memory, and the program is loaded and executed by the processor to implement the image processing method as described in the above technical solution.

[0026] Implementing the present invention has the following beneficial effects:

[0027] The present invention is directed to image processing performed in vehicle autonomous driving. The used image detection neural network includes at least one intermediate layer, and the intermediate layer includes a BN layer, which is used to perform square normalization conversion on the input data of the intermediate layer. Thereby, the convergence speed can be increased during the training of the image detection neural network, the accuracy of image detection, recognition, etc. of the trained image detection neural network is high, and it is beneficial to the generalization of the image detection neural network detection model. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the present invention, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1 is a schematic flowchart of the image detection method provided by an embodiment of the present invention;

[0030] Figure 2 is a schematic structural diagram of the image detection device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0032] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server comprising a series of steps or units does not necessarily have to be limited to those steps, units or modules clearly listed, but may include other steps, units or modules not clearly listed or inherent to these processes, methods, products or devices.

[0033] In order to implement the technical solution of the present invention and make it easier for more engineering and technical workers to understand and apply the present invention, the working principle of the present invention will be further elaborated in conjunction with specific embodiments.

[0034] The present invention provides an image processing system, which can at least include a vehicle and a server, and can be applied in the field of vehicle networking. For example, the cloud server is used to generate image processing data in real time to perform dynamic judgment on the vehicle environment information during autonomous driving.

[0035] In the embodiments of this specification, the server can be a cloud server, and the cloud server can include an independently operating server, or a distributed server, or a server cluster composed of multiple servers. The cloud server can provide background services for the vehicle. Specifically, the cloud server can be provided with a database, and one or more image processing models are stored in the database for image processing, for example, for image detection or image recognition.

[0036] The following introduces an embodiment of an image processing method of the present invention. The image processing method is used to detect and identify the traffic environment where the vehicle is located during autonomous driving of the vehicle. The image processing neural network used as the image processing model is used to process the obtained traffic environment image data, and perform detection, segmentation, recognition, etc. on the traffic environment image data to identify the objects around the vehicle, obtain the environmental information around the vehicle and use it for autonomous driving of the vehicle. Specifically, this method can be used for the training of visual perception models such as forward, surround view, and panoramic view during autonomous driving, especially for the perception model for obstacle recognition in environmental images in traffic road scenes.

[0037] In the embodiments of this specification, during the automatic driving of a vehicle, an image processing neural network is used to detect and identify the traffic environment where the vehicle is located. Correspondingly, training data is used to train the image processing neural network so that the trained image processing neural network can be used for image processing in automatic driving to obtain the environmental information where the vehicle is located based on the environmental image data of the vehicle.

[0038] Optionally, the training data includes a plurality of image data, which can be environmental image data obtained during the driving of the vehicle. The image data includes one or more detection objects or recognition objects. For example, people, vehicles, road marking lines, traffic lights, street lights, green belts, etc. Optionally, the detection objects or recognition objects are further subdivided. For example, people are divided into adults, children, etc., and vehicles are divided into bicycles, battery cars, motorcycles, sedans, small trucks, medium trucks, large trucks, etc. The training data also includes object category information corresponding to the detection objects or recognition objects in each image data, and the object category information is used to indicate the actual category of the detection objects or recognition objects. Exemplarily, the training data is obtained from the data collected by a data collection vehicle, which is equipped with a front view camera, a surround view camera, and a panoramic camera, and collects environmental image data through at least one of the front view camera, the surround view camera, and the panoramic camera.

[0039] Figure 1 It is a schematic flowchart of the image processing method provided by the embodiments of the present invention. This specification provides the method operation steps as described in the embodiments or flowcharts, but based on routine or non-creative labor, there may be more or fewer operation steps. The step order listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. Specifically, as Figure 1 shown, the image processing method includes:

[0040] S201: Obtain the initialization parameters of the image processing neural network to be trained.

[0041] The image processing neural network includes multiple neural network layers, and the neural network layers include an input layer, an output layer, and at least one intermediate layer. Before training the image processing neural network for image processing using the training data, it is necessary to initialize the image detection neural network to obtain the image processing neural network to be trained with initialization parameters.

[0042] When obtaining the initialization parameters of the image processing neural network to be trained, it also includes initializing the image processing neural network. Exemplarily, the initialization parameters include the weights and biases of the image detection neural network. In the embodiments of this specification, the weights and biases of the image detection neural network are initialized to obtain the initial weights and initial biases of the image detection neural network.

[0043] In order to enable the initialized image processing neural network to learn and converge quickly during training and converge to an acceptable result, usually the initialization parameters need to meet the preset initialization conditions. Exemplarily, the initialization parameters make the variances of the output values of each intermediate layer of the image processing neural network consistent, and during the backpropagation process, for each intermediate layer, the variances of the derivatives of the loss function Cost with respect to the state z are consistent, where the state z is the input of the activation function layer in the intermediate layer. However, during the training process of the image processing neural network, as the training progresses, the characteristics of the initialization parameters cannot be maintained, and the parameters of the image processing neural network no longer meet the preset initialization conditions, which affects the subsequent training of the image processing neural network. For example, it causes the learning convergence speed of the subsequent training of the image processing neural network to slow down.

[0044] S203: Use the training data to train the image processing neural network to be trained, and obtain the trained image processing neural network; the training data includes image data; the image processing neural network includes at least one intermediate layer, and the intermediate layer includes a BN layer, and the BN layer is used to perform squared normalization conversion on the input data of the intermediate layer.

[0045] Optionally, each intermediate layer of the image processing neural network includes a convolutional layer, the BN layer, and an activation function layer arranged in sequence. Preferably, each intermediate layer of the image processing neural network includes a convolutional layer, the BN layer, an activation function layer, a pooling layer, and a fully connected layer arranged in sequence.

[0046] The activation function layer is used to perform a non-linear transformation, and the activation function layer includes a non-linear function. In the embodiments of the present invention, the BN layer is used to perform squared normalization conversion on the input data of the intermediate layer. Thus, it is possible to avoid the deviation of the input data distribution, avoid the input data values being too large, avoid the disappearance of gradients during backpropagation in the training process, strengthen the differences between the input data, and avoid the differences between the input data being too large, thereby improving the processing accuracy of the trained image processing neural network.

[0047] In a possible implementation manner, using the training data to train the image processing neural network to be trained and obtain the trained image processing neural network includes:

[0048] Obtain the input data of the intermediate layer;

[0049] The BN layer performs squared normalization conversion on the input data to obtain normalized conversion data, and uses the normalized conversion data as the activation input data of the activation function layer.

[0050] In a possible implementation, the BN layer performs square normalization transformation on the input data to obtain normalized transformation data, and uses the normalized transformation data as the activation input data of the activation function layer, including:

[0051] According to the input data of the intermediate layer, obtain the sign feature, square value, mean square value, and variance of the input data; optionally, obtain the sign feature, square value, mean square value, and variance of the input data x of the intermediate layer, where the sign feature of the input data x is used to represent whether the input data x is positive or negative;

[0052] Translate and scale the input data according to the sign feature, square value, mean square value, and variance of the input data to obtain normalized transformation data, and use the normalized transformation data as the activation input data of the activation function layer; optionally, use the mean square value of the input data to translate the input data, and use the variance to scale the input data; Exemplarily, for an image processing neural network including R intermediate layers, the transformation formula for the square normalization transformation of the i-th intermediate layer is:

[0053]

[0054] where the value range of i is [1, R], is the normalized transformation data corresponding to the input data of the i-th intermediate layer, x i is the input data of the i-th intermediate layer, is the sign feature of the input data of the i-th intermediate layer, is the square value of the input data, is the mean square value of the input data of the i-th intermediate layer, Var(x i ) is the variance of the input data of the i-th intermediate layer.

[0055] Thus, by using this BN layer to perform square normalization transformation on the input data of the intermediate layer, it is possible to avoid the deviation of the input data distribution, the large value of the input data, while strengthening the differences between the input data, and avoiding the excessive differences between the input data, and it is also beneficial to maintain the characteristics of the parameters at the initialization.

[0056] S205: Perform image processing through the trained image processing neural network to obtain an image processing result.

[0057] In this embodiment, for the image processing performed in vehicle autonomous driving, for the used image processing neural network, an improved BN layer is adopted, which can avoid the deviation of the input data distribution, the large value of the input data, while strengthening the differences between the input data, and avoiding the excessive differences between the input data, and it is also beneficial to maintain the characteristics of the parameters at the initialization, improve the convergence speed of the image processing neural network, and is beneficial to the generalization of the image processing neural network.

[0058] An embodiment of this specification also provides an image processing apparatus. Figure 2 It is a schematic structural diagram of the image processing apparatus provided by an embodiment of the present invention. The apparatus includes:

[0059] An initialization module, configured to obtain initialization parameters of the image processing neural network to be trained;

[0060] A training module, configured to train the image processing neural network to be trained using training data to obtain a trained image processing neural network; the training data includes image data; the image processing neural network includes at least one intermediate layer, and the intermediate layer includes a BN layer, and the BN layer is configured to perform squared normalization conversion on the input data of the intermediate layer;

[0061] A processing module, configured to perform image processing through the trained image processing neural network to obtain an image processing result.

[0062] An embodiment of this specification also provides a computer-readable storage medium. A program is stored in the storage medium, and the program is loaded and executed by a processor to implement the image processing method as described in the above technical solution.

[0063] An embodiment of this specification also provides an image processing electronic device. The electronic device includes a processor and a memory. A program is stored in the memory, and the program is loaded and executed by the processor to implement the image processing method as described in the above technical solution.

[0064] In the specification provided here, a large number of specific details are set forth. However, it can be understood that the embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures, and technologies have not been shown in detail so as not to obscure the understanding of this specification.

[0065] Similarly, it should be understood that, in order to streamline this disclosure and assist in understanding one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting the intention that: the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected by the claims of the present invention, the inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present invention.

[0066] Those skilled in the art can understand that the modules in the devices in the embodiments can be adaptively changed and set in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise explicitly stated, each feature disclosed in this specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature that provides the same, equivalent or similar purpose.

[0067] In addition, those skilled in the art can understand that although the embodiments described herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the claims of the present invention, any one of the claimed embodiments can be used in any combination.

[0068] The present invention can also be implemented as a device or system program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, can also be provided on a carrier signal, or can be provided in any other form.

[0069] It should be noted that the above embodiments are illustrative of the present invention rather than limiting the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the claims listing several units of a system, several of these systems can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order and these words can be interpreted as names.

Claims

1. An image processing method, characterized in that, The method includes: Obtaining initialization parameters of an image processing neural network to be trained; the image processing neural network includes at least one intermediate layer, the intermediate layer includes a BN layer and an activation function layer, and the BN layer is used to perform squared normalization transformation on the input data of the intermediate layer; Obtaining the input data of the intermediate layer; According to the input data of the intermediate layer, obtaining the sign feature, squared value, mean squared value, and variance of the input data; Performing translation and scaling on the input data according to the sign feature, squared value, mean squared value, and variance of the input data to obtain normalized transformation data, and using the normalized transformation data as the activation input data of the activation function layer to obtain the trained image processing neural network; Wherein, the mean squared value is used to perform translation on the input data, the variance is used to perform scaling on the input data, and the conversion formula for the squared normalization transformation of the i-th intermediate layer is: The value range of i is [1, R], where R is the number of intermediate layers. is the standardized conversion data corresponding to the input data of the i-th intermediate layer, x i is the input data of the i-th intermediate layer, is the symbol feature of the input data of the i-th intermediate layer; is the square value of the input data of the i-th intermediate layer; is the mean square value of the input data of the i-th intermediate layer; Var(x i ) is the variance of the input data of the i-th intermediate layer; Performing image processing through the trained image processing neural network to obtain an image processing result.

2. The image processing method according to claim 1, wherein The intermediate layer includes a convolutional layer, the BN layer, and the activation function layer arranged in sequence.

3. The image processing method according to claim 2, characterized in that, The intermediate layer includes a convolutional layer, the BN layer, the activation function layer, a pooling layer, and a fully connected layer arranged in sequence.

4. The image processing method according to claim 2, wherein The activation function layer includes a non-linear function.

5. The image processing method according to claim 1, wherein The initialization parameters include the weights and biases of the image processing neural network.

6. An image processing apparatus, characterized in that, The device includes: An initialization module, configured to obtain initialization parameters of an image processing neural network to be trained; the image processing neural network includes at least one intermediate layer, the intermediate layer includes a BN layer and an activation function layer, and the BN layer is used to perform squared normalization transformation on the input data of the intermediate layer; A training module, configured to obtain the input data of the intermediate layer; according to the input data of the intermediate layer, obtaining the sign feature, squared value, mean squared value, and variance of the input data; performing translation and scaling on the input data according to the sign feature, squared value, mean squared value, and variance of the input data to obtain normalized transformation data, and using the normalized transformation data as the activation input data of the activation function layer to obtain the trained image processing neural network; wherein, the mean squared value is used to perform translation on the input data, the variance is used to perform scaling on the input data, and the conversion formula for the squared normalization transformation of the i-th intermediate layer is: The value range of i is [1, R], where R is the number of intermediate layers. is the normalized conversion data corresponding to the input data of the i-th intermediate layer, x i is the input data of the i-th intermediate layer, is the symbol feature of the input data of the i-th intermediate layer; is the square value of the input data of the i-th intermediate layer; is the mean square value of the input data of the i-th intermediate layer; Var(x i ) is the variance of the input data of the i-th intermediate layer; A processing module, configured to perform image processing through the trained image processing neural network to obtain an image processing result.

7. A computer-readable storage medium, characterized in that, A program is stored in the storage medium, The program is loaded and executed by a processor to implement the image processing method according to any one of claims 1 to 5.

8. An image processing electronic device, characterized in that, The electronic device includes a processor and a memory, a program is stored in the memory, and the program is loaded and executed by the processor to implement the image processing method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Breast ultrasound image tumor segmentation method based on full convolution network

    CN108776969A