Training Method, Image Processing Method, Apparatus, and Device for Image Processing Model
Through the cross-domain training joint image processing model, the problem that the inverse ISP algorithm is difficult to obtain real original images is solved, and an image processing model with better performance under a small amount of data is realized, which is suitable for scenarios such as autonomous driving and assisted driving.
Patent Information
- Application Number
- CN202310176197.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-17
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-02-17
AI Technical Summary
In the prior art, the simulated image training model obtained based on the inverse ISP algorithm has poor performance in real application scenarios, because the inverse ISP algorithm is difficult to obtain real original images, resulting in poor training results.
By constructing a joint image processing model, cross-domain training is performed using feature extraction network, feature format conversion network and discriminator network, feature conversion and update from the first image format to the second image format, and features that are more in line with the real original image.
The performance of the image processing model in real application scenarios is improved, the problems of small data volume, high acquisition cost and low labeling efficiency are solved, and better model generalization ability and robustness are obtained.
Smart Images

Figure CN116311119B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to computer vision technology, and in particular to a training method, an image processing method, an apparatus and a device for an image processing model. Background Art
[0002] In scenarios such as autonomous driving and assisted driving, it is usually necessary to perform image processing based on image processing models such as object detection models, semantic segmentation models, and image classification models to obtain image processing results for behavior decision-making and control of vehicle driving. In related technologies, the original image (RAW image) collected by an image sensor is usually converted into a target image in RGB, YUV, etc. format based on the ISP (Image Signal Processing) algorithm, and the deep neural network is trained based on the target image to obtain an image processing model. However, converting the original image into a target image will lose the detailed information of the original image, resulting in poor performance of the obtained image processing model. To solve this problem, a method of replacing the ISP algorithm with a neural network has been proposed. This method requires training a neural network model based on the original image. In related technologies, the collected and labeled target images are usually processed by the inverse ISP algorithm to obtain simulated images simulating the original images, and these simulated images are used for training the image processing model to obtain an image processing model suitable for the original image. However, due to the complexity of the ISP algorithm, it is often difficult to obtain a real original image through inverse ISP algorithm processing, resulting in poor performance of the image processing model obtained by training with the simulated images of non-real original images in real application scenarios. Summary of the Invention
[0003] To solve the technical problems such as the poor performance of the image processing model obtained by training with simulated images in real application scenarios, the present disclosure is proposed. Embodiments of the present disclosure provide a training method, an image processing method, an apparatus and a device for an image processing model.
[0004] According to one aspect of the embodiments of the present disclosure, a training method for an image processing model is provided, including: determining an initial joint image processing model; using a first feature extraction network of the initial joint image processing model to process first training image data in a first image format to obtain first training features in a first feature format corresponding to the first training image data, where the network parameters of the first feature extraction network are pre-based
[0005] Network parameters obtained by training with image data in a first image format; using the feature format conversion network of the initial joint image processing model to perform format conversion on the first training features to obtain second training features in a second feature format; using the second feature extraction network of the initial joint image processing model to process second training image data in a second image format to obtain third training features in the second feature format corresponding to the second training image data; based on the second training features, the first label data corresponding to the second training features, the third training features, and the second label data corresponding to the third training features, updating the network parameters of the initial joint image processing model to obtain an updated first joint image processing model; in response to the first joint image processing model meeting a preset condition, determining a target image processing model suitable for processing images in the second image format based on the first joint image processing model.
[0006] According to another aspect of the embodiments of the present disclosure, there is provided an image processing method, including: obtaining an image to be processed, where the image to be processed is an image in a second image format; using the target image processing model to process the image to be processed to obtain an image processing result corresponding to the image to be processed; the target image processing model is trained based on the image processing model training method provided in any of the above embodiments.
[0007] According to another aspect of the embodiments of the present disclosure, there is provided a training device for an image processing model, including: a first acquisition module configured to determine an initial joint image processing model; a first processing module configured to process first training image data in a first image format by using a first feature extraction network of the initial joint image processing model to obtain first training features in a first feature format, where network parameters of the first feature extraction network are network parameters pre-trained based on image data in the first image format; a second processing module configured to perform format conversion on the first training features by using a feature format conversion network of the initial joint image processing model to obtain second training features in a second feature format; a third processing module configured to process second training image data in a second image format by using a second feature extraction network of the initial joint image processing model to obtain third training features in the second feature format corresponding to the second training image data; a fourth processing module configured to update network parameters of the initial joint image processing model based on the second training features, first label data corresponding to the second training features, the third training features, and second label data corresponding to the third training features to obtain an updated first joint image processing model; a fifth processing module configured to, in response to the first joint image processing model meeting a preset condition, determine a target image processing model suitable for processing images in the second image format based on the first joint image processing model.
[0008] According to another aspect of the embodiments of the present disclosure, there is provided an image processing device, including: a second acquisition module configured to acquire an image to be processed, where the image to be processed is an image in a second image format; a sixth processing module configured to process the image to be processed by using a target image processing model to obtain an image processing result corresponding to the image to be processed; the target image processing model is trained based on the image processing model training method described in any one of the above embodiments.
[0009] According to yet another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program, where the computer program is used to execute the image processing model training method described in any one of the above embodiments of the present disclosure; or, the computer program is used to execute the image processing method described in any one of the above embodiments of the present disclosure.
[0010] According to yet another aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing executable instructions of the processor; the processor is configured to read the executable instructions from the memory and execute the instructions to implement the image processing model training method described in any one of the above embodiments of the present disclosure; or, execute the instructions to implement the image processing method described in any one of the above embodiments of the present disclosure.
[0011] Based on the training method, image processing method, device and equipment of the image processing model provided in the above embodiments of the present disclosure, by determining an initial joint image processing model, the processing of training image data in a first image format and a second image format is realized. The first training features corresponding to the first image format are converted from a first feature format to a second feature format through a feature format conversion network for updating network parameters, thereby realizing cross-domain joint training of training image data in the first image format and the second image format to solve problems such as less data volume, high acquisition cost, and low annotation efficiency of the second image format (such as RAW images). Compared with the method of obtaining the original image by using the inverse ISP algorithm in the related art, the present disclosure can make the features in the second feature format obtained by the feature format conversion network more conform to the features extracted from the real original image through joint training, so that the finally obtained target image processing model has better performance.
[0012] The technical solution of the present disclosure will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] By describing the embodiments of the present disclosure in more detail with reference to the drawings, the above and other objects, features and advantages of the present disclosure will become more obvious. The drawings are used to provide a further understanding of the embodiments of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the present disclosure and do not constitute a limitation to the present disclosure. In the drawings, the same reference numerals generally represent the same components or steps.
[0014] Figure 1 is an exemplary application scenario of the training method of the image processing model provided by the present disclosure;
[0015] Figure 2 is a schematic flowchart of the training method of the image processing model provided by an exemplary embodiment of the present disclosure;
[0016] Figure 3 is a schematic flowchart of the training method of the image processing model provided by another exemplary embodiment of the present disclosure;
[0017] Figure 4 is a schematic diagram of the network structure of the first image processing model provided by an exemplary embodiment of the present disclosure;
[0018] Figure 5 is a schematic diagram of the network structure of the initial joint image processing model provided by an exemplary embodiment of the present disclosure;
[0019] Figure 6 is a schematic flowchart of step 2055 provided by an exemplary embodiment of the present disclosure;
[0020] Figure 7 is a schematic flowchart of an image processing method provided by an exemplary embodiment of the present disclosure;
[0021] Figure 8 is a schematic structural diagram of a training device for an image processing model provided by an exemplary embodiment of the present disclosure;
[0022] Figure 9 is a schematic structural diagram of a training device for an image processing model provided by another exemplary embodiment of the present disclosure;
[0023] Figure 10 is a schematic structural diagram of an image processing device provided by an exemplary embodiment of the present disclosure;
[0024] Figure 11 is a schematic structural diagram of an application embodiment of an electronic device according to the present disclosure. Detailed Embodiments
[0025] Next, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.
[0026] It should be noted that: Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.
[0027] Those skilled in the art can understand that terms such as "first", "second", etc. in the embodiments of the present disclosure are only used to distinguish different steps, devices or modules, etc., and neither represent any specific technical meaning nor indicate an inevitable logical order between them.
[0028] It should also be understood that in the embodiments of the present disclosure, "a plurality" may refer to two or more, and "at least one" may refer to one, two or more.
[0029] It should also be understood that for any component, data or structure mentioned in the embodiments of the present disclosure, unless otherwise clearly defined or given a contrary indication in the context, it is generally understood to be one or more.
[0030] In addition, the term " / and" in the present disclosure is only a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in the present disclosure generally represents an "or" relationship between the associated objects before and after.
[0031] It should also be understood that the descriptions of the various embodiments in the present disclosure emphasize the differences between the various embodiments, and their similarities or resemblances can be referred to each other. For the sake of brevity, they will not be elaborated one by one.
[0032] Meanwhile, it should be understood that, for the sake of description convenience, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships.
[0033] The following description of at least one exemplary embodiment is actually merely illustrative and in no way restricts the present disclosure and its application or use.
[0034] Techniques, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the said techniques, methods, and devices should be regarded as part of the specification.
[0035] It should be noted that: like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, it does not need to be further discussed in subsequent drawings.
[0036] Embodiments of the present disclosure can be applied to electronic devices such as terminal devices, computer systems, servers, etc., which can operate together with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, servers, etc. include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments including any of the above systems, and so on.
[0037] Terminal devices, computer systems, servers and other electronic devices can be described in the general context of computer system-executable instructions (such as program modules) executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logics, data structures, etc., which perform specific tasks or implement specific abstract data types. The computer system / server can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.
[0038] Overview of the present disclosure
[0039] In the process of implementing the present disclosure, the inventors found that in scenarios such as autonomous driving and assisted driving, it is usually necessary to perform image processing based on image processing models such as object detection models, semantic segmentation models, and image classification models to obtain image processing results for behavior decision-making and control of vehicle driving. In the related art, the original image (RAW image) collected by an image sensor is usually converted into a target image in formats such as RGB and YUV based on the ISP (Image Signal Processing) algorithm, and the deep neural network is trained based on the target image to obtain an image processing model. However, when converting the original image into a target image, the detailed information of the original image will be lost, resulting in poor performance of the obtained image processing model. To solve this problem, a method of replacing the ISP algorithm with a neural network has been proposed. This method requires training a neural network model based on the original image. In the related art, the collected and labeled target images are usually processed through the inverse ISP algorithm to obtain simulated images that simulate the original image, and these simulated images are used for training the image processing model to obtain an image processing model suitable for the original image. However, due to the complexity of the ISP algorithm, it is often difficult to obtain a real original image through inverse ISP algorithm processing, resulting in poor performance of the image processing model obtained by training with simulated images of non-real original images in real application scenarios.
[0040] Exemplary overview
[0041] Figure 1 It is an exemplary application scenario of the training method of the image processing model provided by the present disclosure.
[0042] In the scenarios of autonomous driving or assisted driving, in order to obtain a target image processing model suitable for processing images in a second image format (such as the original image format, that is, the RAW image format, and the RAW image is the unprocessed original data collected by an image sensor), such as an image processing model like a target detection model, a semantic segmentation model, an image classification model, etc., by using the training method of the image processing model of the present disclosure, it is possible to perform cross-image domain joint training based on the training image data in the first image format (such as RGB (a color space composed of red, green, and blue), YUV (representing pixel colors with luminance and chrominance, Y represents luminance, that is, the grayscale value, and U, V represent chrominance, describing hue and saturation), etc.) and the training image data in the second image format. When the training image data in the second image format is difficult to collect, difficult to annotate, or has a small sample size, the features of the second image format can be augmented by the training image data in the first image format for model training. Specifically, a joint image processing network can be constructed, and an initial joint image processing model is determined through initialization. The initial joint image processing model includes a first feature extraction network for extracting the first training features in the first feature format of the first training image data in the first image format, a feature format conversion network for performing feature format conversion, a second feature extraction network for extracting the second training features in the second feature format of the second training image in the second image format, and may also include a head network related to the image processing task. The first feature extraction network can be used to extract the first training features in the first feature format corresponding to the first training image data. The feature format conversion network can perform format conversion on the first training features to convert the first training features into the second training features in the second feature format. The second feature extraction network can be used to extract the third training features in the second feature format of the second training image data. Then, the head network is used to predict the second training features and the third training features respectively, and the prediction results corresponding to the second training features and the third training features are obtained. The prediction results are compared with the first label data corresponding to the second training features and the second label data corresponding to the third training features to determine the network loss for updating the network parameters. Through continuous iterative updates, when the network loss meets the preset conditions or reaches the preset iteration number threshold, the training is ended, and a trained joint image processing model is obtained. The second feature extraction network and the head network in the joint image processing model are used as the target image processing model suitable for processing images in the second image format.The present disclosure enables the features in the second feature format obtained by the feature format conversion network through joint training to be more consistent with the features extracted from real original images, so that the finally obtained target image processing model has better performance. Based on the cross-image domain joint training, an image processing model suitable for the second image format with better performance is obtained with fewer training samples in the second image format, thereby solving problems such as high RAW image data acquisition cost and low annotation efficiency, and avoiding the situation where the model performance is poor due to using non-real original images obtained by the inverse ISP algorithm for model training.
[0043] In practical applications, in order to enable the feature format conversion network to convert the first training feature into a second training feature that is more consistent with the third training feature, the discriminator network can also be used to discriminate the sources of the second training feature and the third training feature, and the discriminative loss is used to guide the learning of the feature format conversion network, thereby further improving the performance of the obtained target image processing model.
[0044] The training method for the image processing model provided by the embodiments of the present disclosure can be applied to any scenario that requires image processing, not limited to the above-mentioned autonomous driving or assisted driving scenarios.
[0045] Exemplary method
[0046] Figure 2 It is a schematic flowchart of the training method for the image processing model provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices, specifically such as servers, terminals and other electronic devices, as Figure 2 shown, and includes the following steps:
[0047] Step 201, determine an initial joint image processing model.
[0048] Among them, the initial joint image processing model can be initialized based on the network parameters of the image processing model suitable for the first image format obtained by pre-training.
[0049] For example, the first image format is RGB format or YUV format, and an image processing model suitable for the first image format (which can be called the first image processing model) can be obtained by training based on the image data in the first image format. The first image processing model can include a feature extraction network (which can be called the third feature extraction network) and a head network (which can be called the second head network), and the specific structure can be set according to actual needs. For example, for the first image processing model for object detection tasks, the second head network is a head network for predicting the detection box of the region where the object is located in the image and the type to which the object belongs, and the specific details are not limited. The feature extraction network can be implemented based on a convolutional neural network, for example. A joint image processing network is constructed. The joint image processing network can include a feature extraction network for the first image format (the first feature extraction network), a feature extraction network for the second image format (such as the RAW image format) (the second feature extraction network), a feature format conversion network, and a head network. The joint image processing network is initialized based on the network parameters of the first image processing model. For example, the network parameters of the third feature extraction network of the first image processing model are used as the initial network parameters of the first feature extraction network of the joint image processing network, and the network parameters of the second head network of the first image processing model are used as the network parameters of the head network of the joint image processing network. For other parts of the network in the joint image processing network, they can be initialized based on a preset initialization rule. After the initialization is completed, an initial joint image processing model is obtained.
[0050] Step 202: Process the first training image data in the first image format using the first feature extraction network of the initial joint image processing model to obtain the first training feature in the first feature format corresponding to the first training image data.
[0051] Among them, the network parameters of the first feature extraction network are network parameters obtained by training in advance based on the image data in the first image format. The first image format can be image formats such as RGB format and YUV format, and can be specifically set according to actual needs. The network structure of the first feature extraction network can adopt any implementable feature extraction network structure, such as a convolutional neural network, a feature extraction network based on a transformer, etc., and the specific details are not limited.
[0052] Step 203: Use the feature format conversion network of the initial joint image processing model to perform format conversion on the first training feature to obtain the second training feature in the second feature format.
[0053] Among them, the network structure of the format conversion network can be set according to actual needs. For example, a convolutional neural network, a network based on a transformer, etc. can be adopted, and the specific details are not limited.
[0054] Step 204: Process the second training image data in the second image format using the second feature extraction network of the initial joint image processing model to obtain third training features in the second feature format corresponding to the second training image data.
[0055] Among them, the second image format can be the RAW image format or other formats different from the first image format, and no specific limitation is made. The network structure of the second feature extraction network can also adopt a convolutional neural network, a feature extraction network based on transformer, etc., and no specific limitation is made.
[0056] Step 205: Update the network parameters of the initial joint image processing model based on the second training features, the first label data corresponding to the second training features, the third training features, and the second label data corresponding to the third training features to obtain an updated first joint image processing model.
[0057] Among them, the first label data can be set according to the requirements of the image processing task. For example, for the object detection task, the first label data can include the labels related to the object detection task corresponding to each image in the first training image data, such as the detection box labels and type labels of each target object in the image. The second label data is similar to the first label data and is the label corresponding to the second training image data.
[0058] In an optional embodiment, for the update of the network parameters of the initial joint image processing model, the network loss can be determined based on the network prediction result and the label data according to a preset loss function, and then the network parameters can be updated using the gradient descent algorithm based on the network loss. The specific loss function can be set according to actual requirements.
[0059] Step 206: In response to the first joint image processing model meeting the preset conditions, determine a target image processing model suitable for processing images in the second image format based on the first joint image processing model.
[0060] Among them, the preset conditions can be set according to actual requirements. For example, it can include that the network loss converges or the number of iterations reaches a preset number threshold, and no specific limitation is made. The target image processing model suitable for the second image format includes the second feature extraction network and the head network in the first joint image processing model. The other parts of the network are obtained by cross-image domain training to assist in obtaining the target image processing model suitable for the second image format.
[0061] The training method of the image processing model provided in this embodiment processes the training image data of the first image format and the second image format by determining an initial joint image processing model, and converts the first training feature corresponding to the first image format from the first feature format to the second feature format through a feature format conversion network for updating network parameters, thereby realizing cross-domain joint training of the training image data based on the first image format and the second image format, so as to solve problems such as less data volume, high acquisition cost, and low annotation efficiency of the second image format (such as RAW images). Compared with the method of obtaining the original image by using the inverse ISP algorithm in the related art, the present disclosure can make the features of the second feature format obtained by the feature format conversion network more conform to the features extracted from the real original image through joint training, so that the finally obtained target image processing model has better performance.
[0062] Figure 3 It is a schematic flowchart of the training method of the image processing model provided in another exemplary embodiment of the present disclosure.
[0063] In an optional embodiment, determining the initial joint image processing model in step 201 includes:
[0064] Step 2011, obtain a first image processing model trained by using the image data based on the first image format. The first image processing model includes a third feature extraction network and a second head network.
[0065] Among them, the network structure of the third feature extraction network is the same as that of the first feature extraction network, and the structure of the second head network is the same as that of the head network of the initial joint image processing model, which will not be elaborated here.
[0066] Step 2012, initialize the network parameters of the pre-established initial joint image processing network based on the network parameters of the first image processing model and the preset initialization rules to obtain an initial joint image processing model. Among them, the network parameters of the first feature extraction network in the initial joint image processing model are initialized to the network parameters of the third feature extraction network, and the network parameters of the first head network of the initial joint image processing model are initialized to the network parameters of the second head network.
[0067] Among them, the network parameters of the feature format conversion network, the second feature extraction network, etc. in the initial joint image processing model can be initialized through preset initialization rules, such as random initialization, which is not specifically limited.
[0068] Exemplarily, Figure 4It is a schematic diagram of the network structure of the first image processing model provided by an exemplary embodiment of the present disclosure. The third feature extraction network can extract features from an image in the first image format to obtain the extracted features, and the second head network makes predictions on the extracted features for related tasks to obtain prediction results, such as object detection results, semantic segmentation results, and so on.
[0069] In an alternative embodiment, the first image processing model can also be a multi-task model, such as a multi-task model that can simultaneously implement at least two of tasks such as object detection, semantic segmentation, and image classification. Then, the second head network can include head networks corresponding to multiple tasks respectively, and the specific form is not limited. Correspondingly, the initial joint image processing model is also a multi-task model with the same tasks as the first image processing model.
[0070] In this embodiment, the initial joint image processing model is obtained by initializing the first image processing model in the first image format obtained through pre-training, which helps to improve the convergence speed of the joint image processing model and improve the training efficiency during the training process.
[0071] In an alternative embodiment, step 205 may specifically include the following steps:
[0072] Step 2051: Process the second training feature using the first head network of the initial joint image processing model to obtain the first training image processing result corresponding to the second training feature.
[0073] Among them, the first training image processing result may include at least one of object detection results (the object detection results may include the regressed object positions and types), image classification results, semantic segmentation results (which may include the probabilities that each pixel of the image belongs to each type), etc., and can be specifically set according to actual requirements.
[0074] Step 2052: Process the third training feature using the first head network to obtain the second training image processing result corresponding to the third training feature.
[0075] Among them, the second training image processing result is similar to the first training image processing result and will not be elaborated here.
[0076] Step 2053: Process the second training feature using the discriminator network of the initial joint image processing model to obtain the first discrimination result corresponding to the second training feature. The first discrimination result includes the first discrimination probability that the second training feature comes from the training image data in the first image format or the training image data in the second image format.
[0077] Among them, the discriminator network can be set according to actual needs. For example, it can be implemented using structures such as convolutional neural networks, transformers, etc., without specific limitations. The discriminator network is used to guide the learning of the feature format conversion network, making the features output by the feature format conversion network more consistent with the features directly extracted from the image in the second image format, so as to achieve the purpose of expanding the feature samples of the image in the second image format based on the image in the first image format. The first discrimination probability represents the probability that the second training feature comes from the first training image data in the first image format, or represents the probability that the second training feature comes from the second training image data in the second image format. For example, label 1 represents the first training image data in the first image format, label 0 represents the second training image data in the second image format, and the first discrimination probability is a probability value within the range of (0, 1), such as 0.2, 0.5, 0.8, etc.
[0078] Step 2054: Use the discriminator network to process the third training feature to obtain a second discrimination result corresponding to the third training feature. The second discrimination result includes the second discrimination probability that the third training feature comes from the training image data in the first image format or the training image data in the second image format.
[0079] Among them, the second discrimination probability is similar to the above-mentioned first discrimination probability and will not be elaborated here.
[0080] Step 2055: Based on the first training image processing result, the first discrimination probability, the first label data, the second training image processing result, the second discrimination probability, the second label data, and a preset update rule, update the network parameters of the initial joint image processing model to obtain a first joint image processing model.
[0081] Among them, the preset update rule can be set according to actual needs. For example, the network loss can be determined based on the first training image processing result, the first discrimination probability, the first label data, the second training image processing result, the second discrimination probability, the second label data, and a preset loss function, and the network parameters are updated using the gradient descent method based on the network loss. When the model includes a discriminator network, the first label data, in addition to including the first processing result label corresponding to the first training image processing result, also needs to include the first discrimination label corresponding to the first discrimination probability. The first discrimination label indicates the true source of the training feature corresponding to the first discrimination probability. For example, the first discrimination label being 1 indicates that the training feature corresponding to the first discrimination probability is the second training feature from the feature format conversion network. Similarly, the second label data includes the second processing result label and the second discrimination label, which will not be elaborated here.
[0082] It should be noted that although the discriminator network processes the second training feature and the third training feature separately, the discriminator network itself cannot directly know whether the input feature specifically comes from the second training feature of the feature format conversion network or the third training feature of the second feature extraction network. Instead, it infers and predicts the probability of the feature source through the discriminator network. Based on this, the discriminant loss can be determined through the output of the discriminator network and the discriminant label, and participate in the update of the network parameters, thereby guiding the continuous learning of the feature format conversion network until the discriminator network cannot distinguish the source of the input feature, which means that the feature converted by the feature format conversion network is consistent with the feature of the second image format image directly extracted. Thus, the feature samples of the second image format image can be extended based on the first image format image for training the first head network, making the prediction performance of the first head network better, and thus obtaining a target image processing model with stronger generalization ability and better performance suitable for the second image format image. Since both the second training feature and the third training feature need to be input into the discriminator network, their dimensions need to be the same.
[0083] In practical applications, when updating the network parameters, for the first feature extraction network, since it has been pre-trained and does not need to update the network parameters, only the feature format conversion network, the second feature extraction network, the first head network, and the discriminator network need to be updated. Since the feature format conversion network and the discriminator network are adversarial to each other during the training process, when updating the network parameters, the two need to be updated alternately, and different loss functions can be used to update the parameters of the feature format conversion network and the discriminator network respectively.
[0084] In an optional example, Figure 5 is a schematic diagram of the network structure of the initial joint image processing model provided by an exemplary embodiment of the present disclosure. Among them, the outputs of the feature format conversion network and the second feature extraction network need to be separately sent to the discriminator network for discrimination, but the discrimination results need to be used together to determine the network loss.
[0085] Through the discriminator network in this embodiment, the parameter learning of the feature format conversion network can be better guided, so that the feature converted by the feature format conversion network is more consistent with the feature of the second image format image directly extracted. Thus, the feature samples of the second image format image can be extended based on the first image format image for training the first head network, making the prediction performance of the first head network better, and thus obtaining a target image processing model with stronger generalization ability and better performance suitable for the second image format image.
[0086] Figure 6 is a schematic flowchart of step 2055 provided by an exemplary embodiment of the present disclosure.
[0087] In an alternative embodiment, the first label data includes a first processing result label and a first discrimination label; the second label data includes a second processing result label and a second discrimination label; the updating of the network parameters of the initial joint image processing model based on the first training image processing result, the first discrimination probability, the first label data, the second training image processing result, the second discrimination probability, the second label data and a preset updating rule to obtain a first joint image processing model includes:
[0088] Step 20551: Determine a first loss based on the first training image processing result, the first processing result label, the second training image processing result, the second processing result label and a first loss function.
[0089] Among them, the first loss function can be set according to actual requirements. For example, a cross-entropy loss function, a focal loss function, a GIoU loss function, etc. can be used, and specifically can be set according to task requirements. The focal loss function is a dynamically scaled cross-entropy loss function, and the GIoU loss function is a loss function proposed to solve the problems existing in the IoU loss function, and will not be elaborated here.
[0090] Exemplarily, for an object detection model, the first training image processing result includes the detection box and classification of the regressed target object, then the first loss function can include a classification loss function and a regression loss function. For example, the classification loss function can use the focal loss function, and the regression loss function can use the GIoU loss function. The determined first loss can be expressed as L GT = l1 + l2, where l1 represents the classification loss and l2 represents the regression loss.
[0091] Step 20552: Determine a second loss based on the first discrimination probability, the first discrimination label and a second loss function.
[0092] Among them, the second loss function can be expressed as follows:
[0093] l adv = log D(X y )
[0094] Among them, l adu represents the second loss, which can also be called the adversarial loss. X y represents the output feature of the feature format conversion network, that is, the second training feature. D represents the discriminator network, and D(X y ) represents the processing result of the discriminator network on X y , that is, the first discrimination result.
[0095] Step 20553: Determine the third loss based on the first discrimination probability, the first discrimination label, the second discrimination probability, the second discrimination label, and the third loss function.
[0096] Among them, the third loss function can be expressed as follows:
[0097] l D = logD(X r ) + log(1 - D(X y ))
[0098] Among them, l D represents the third loss, which can also be called the discrimination loss. X r represents the output features of the second feature extraction network, that is, the third training feature. D(X r ) represents the processing result of the discriminator network on X r , that is, the second discrimination result.
[0099] Step 20554: Update the network parameters of the feature format conversion network, the second feature extraction network, and the first head network of the initial joint image processing model based on the first loss and the second loss, and obtain the updated second joint image processing model.
[0100] Among them, the comprehensive loss can be determined based on the first loss and the second loss. For example, the comprehensive loss can be expressed as follows:
[0101] l total = L GT + l adu
[0102] Use the gradient descent method to update the network parameters of the feature format conversion network, the second feature extraction network, and the first head network, and obtain the updated second joint image processing model.
[0103] Step 20555: Update the network parameters of the discriminator network of the second joint image processing model based on the third loss, and obtain the first joint image processing model.
[0104] Specifically, based on the third loss, the gradient descent method can be used to update the network parameters of the discriminator network of the second joint image processing model, and the updated model is used as the first joint image processing model.
[0105] During the training process, a part of the network (including the feature format conversion network, the second feature extraction network, and the first head network) and the discriminator network can be alternately updated according to the above process until convergence or the number of iterations reaches the preset number threshold, and the trained joint image processing model is obtained.
[0106] In this embodiment, the network parameters are alternately updated by a partial network including a feature format conversion network, a second feature extraction network, and a first head network and a discriminator network to achieve adversarial training, which improves the overall performance of the model while ensuring that the feature format conversion network has good conversion performance.
[0107] In an alternative embodiment, step 206 of determining, in response to the first joint image processing model satisfying a preset condition, a target image processing model suitable for processing an image in a second image format based on the first joint image processing model includes:
[0108] Step 2061, in response to the network loss of the first joint image processing model satisfying a preset end condition, determining the target image processing model based on the second feature extraction network and the first head network of the first joint image processing model.
[0109] Among them, only the second feature extraction network and the first head network related to the image processing of the second image format are included in the first joint image processing model. The functions of the other parts of the network are to assist in the training of the second feature extraction network and the first head network. After the training is completed, the second feature extraction network and the first head network and their network parameters can be directly determined as the target image processing model.
[0110] The target image processing model obtained in this embodiment, because the feature samples of the image in the first image format are expanded, makes the generalization ability of the first head network stronger and the robustness better, so that it can adapt to the features extracted from various images in the second image format.
[0111] In an alternative embodiment, the method of the present disclosure further includes:
[0112] Step 207, in response to the first joint image processing model not satisfying a preset condition, using the first joint image processing model as the initial joint image processing model, and repeating the step of processing the first training image data in the first image format by the first feature extraction network of the initial joint image processing model to obtain the first training feature in the first feature format corresponding to the first training image data.
[0113] Among them, if the first joint image processing model does not satisfy the preset condition, the first joint image processing model is used as the initial joint image processing model, and the training process of the foregoing embodiment is repeated until the model converges or the number of iterations reaches a preset number threshold, and the training ends. The specific process is referred to the foregoing content and will not be elaborated here.
[0114] In this embodiment, the performance of the model is continuously improved through the iterative optimization of the joint image processing model. When the training ends, a model with stronger generalization ability and better robustness can be obtained.
[0115] Each of the above embodiments or optional examples of the present disclosure can be implemented alone or in any combination without conflict, and can be specifically set according to actual needs. The present disclosure does not make any limitations.
[0116] Any training method of an image processing model provided by an embodiment of the present disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices, servers, etc. Alternatively, any training method of an image processing model provided by an embodiment of the present disclosure can be executed by a processor. For example, the processor executes any training method of an image processing model mentioned in an embodiment of the present disclosure by calling corresponding instructions stored in a memory. This will not be elaborated further below.
[0117] Figure 7 It is a schematic flowchart of an image processing method provided by an exemplary embodiment of the present disclosure. This embodiment can be applied to electronic devices, specifically, such as electronic devices like servers, terminals, in-vehicle computing platforms, etc. As Figure 7 shown, it includes the following steps:
[0118] Step 301, obtain an image to be processed, where the image to be processed is an image in a second image format.
[0119] Step 302, use the target image processing model to process the image to be processed to obtain an image processing result corresponding to the image to be processed.
[0120] Among them, the target image processing model is trained based on the training method of the image processing model provided by any of the above embodiments.
[0121] For the specific structure and training process of the target image processing model in this embodiment, refer to the foregoing embodiments and will not be elaborated further here.
[0122] For the image processing method provided in this embodiment, since it processes an image in a second image format based on the target image processing model obtained by the above joint training, the robustness of the image processing result can be effectively improved.
[0123] Any image processing method provided by an embodiment of the present disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices, servers, etc. Alternatively, any image processing method provided by an embodiment of the present disclosure can be executed by a processor. For example, the processor executes any image processing method mentioned in an embodiment of the present disclosure by calling corresponding instructions stored in a memory. This will not be elaborated further below.
[0124] Exemplary device
[0125] Figure 8It is a schematic structural diagram of a training device for an image processing model provided by an exemplary embodiment of the present disclosure. The device of this embodiment can be used to implement the corresponding method embodiment for training the image processing model of the present disclosure, such as Figure 8 The shown device includes: a first acquisition module 501, a first processing module 502, a second processing module 503, a third processing module 504, a fourth processing module 505, and a fifth processing module 506.
[0126] The first acquisition module 501 is used to determine an initial joint image processing model.
[0127] The first processing module 502 is used to process the first training image data in the first image format by using the first feature extraction network of the initial joint image processing model to obtain first training features in the first feature format corresponding to the first training image data, and the network parameters of the first feature extraction network are network parameters pre-trained based on the image data in the first image format.
[0128] The second processing module 503 is used to convert the format of the first training features by using the feature format conversion network of the initial joint image processing model to obtain second training features in the second feature format.
[0129] The third processing module 504 is used to process the second training image data in the second image format by using the second feature extraction network of the initial joint image processing model to obtain third training features in the second feature format corresponding to the second training image data.
[0130] The fourth processing module 505 is used to update the network parameters of the initial joint image processing model based on the second training features, the first label data corresponding to the second training features, the third training features, and the second label data corresponding to the third training features to obtain an updated first joint image processing model.
[0131] The fifth processing module 506 is used to, in response to the first joint image processing model meeting a preset condition, determine a target image processing model suitable for processing images in the second image format based on the first joint image processing model.
[0132] Figure 9 It is a schematic structural diagram of a training device for an image processing model provided by another exemplary embodiment of the present disclosure.
[0133] In an optional embodiment, the first acquisition module 501 includes:
[0134] The first acquisition unit 5011 is used to acquire a first image processing model trained based on the image data in the first image format, and the first image processing model includes a third feature extraction network and a second head network.
[0135] An initialization unit 5012, configured to initialize network parameters of a pre-established initial joint image processing network based on network parameters of a first image processing model and a preset initialization rule, to obtain an initial joint image processing model, where network parameters of a first feature extraction network are initialized to network parameters of a third feature extraction network, and network parameters of a first head network of the initial joint image processing model are initialized to network parameters of a second head network.
[0136] In an optional embodiment, the fourth processing module 505 includes: a first processing unit 5051, a second processing unit 5052, a third processing unit 5053, a fourth processing unit 5054, and a fifth processing unit 5055.
[0137] The first processing unit 5051 is configured to process second training features by using a first head network of the initial joint image processing model, to obtain a first training image processing result corresponding to the second training features.
[0138] The second processing unit 5052 is configured to process third training features by using the first head network, to obtain a second training image processing result corresponding to the third training features.
[0139] The third processing unit 5053 is configured to process second training features by using a discriminator network of the initial joint image processing model, to obtain a first discrimination result corresponding to the second training features, where the first discrimination result includes a first discrimination probability that the second training features come from training image data of a first image format or training image data of a second image format.
[0140] The fourth processing unit 5054 is configured to process third training features by using the discriminator network, to obtain a second discrimination result corresponding to the third training features, where the second discrimination result includes a second discrimination probability that the third training features come from training image data of a first image format or training image data of a second image format.
[0141] The fifth processing unit 5055 is configured to update network parameters of the initial joint image processing model based on the first training image processing result, the first discrimination probability, first label data, the second training image processing result, the second discrimination probability, second label data, and a preset update rule, to obtain a first joint image processing model.
[0142] In an optional embodiment, the first label data includes a first processing result label and a first discrimination label; the second label data includes a second processing result label and a second discrimination label; specifically, the fifth processing unit 5055 is configured to:
[0143] Determine the first loss based on the first training image processing result, the first processing result label, the second training image processing result, the second processing result label, and the first loss function; determine the second loss based on the first discrimination probability, the first discrimination label, and the second loss function; determine the third loss based on the first discrimination probability, the first discrimination label, the second discrimination probability, the second discrimination label, and the third loss function; update the network parameters of the feature format conversion network, the second feature extraction network, and the first head network of the initial joint image processing model based on the first loss and the second loss to obtain an updated second joint image processing model; update the network parameters of the discriminator network of the second joint image processing model based on the third loss to obtain a first joint image processing model.
[0144] In an alternative embodiment, the fifth processing module 506 includes:
[0145] The first determination unit 5061 is configured to, in response to the network loss of the first joint image processing model satisfying a preset end condition, determine a target image processing model based on the second feature extraction network and the first head network of the first joint image processing model.
[0146] In an alternative embodiment, the apparatus of the present disclosure further includes:
[0147] The sixth processing module 507 is configured to, in response to the first joint image processing model not satisfying a preset condition, use the first joint image processing model as the initial joint image processing model, and repeat the step of processing the first training image data in the first image format by using the first feature extraction network of the initial joint image processing model to obtain the first training feature in the first feature format corresponding to the first training image data.
[0148] Figure 10 It is a schematic structural diagram of an image processing apparatus provided by an exemplary embodiment of the present disclosure. The apparatus of this embodiment can be used to implement the corresponding image processing method embodiment of the present disclosure, such as Figure 10 The shown apparatus includes: a second acquisition module 601 and a sixth processing module 602.
[0149] The second acquisition module 601 is configured to acquire an image to be processed, and the image to be processed is an image in the second image format.
[0150] The sixth processing module 602 is configured to process the image to be processed by using the target image processing model to obtain an image processing result corresponding to the image to be processed; the target image processing model is trained based on the image processing model training method provided in any of the above embodiments.
[0151] For the specific operations of each module and each unit in the apparatus embodiments of the present disclosure, please refer to the corresponding method embodiments described above, and details are not described herein again.
[0152] Exemplary electronic device
[0153] An embodiment of the present disclosure also provides an electronic device, including: a memory for storing a computer program;
[0154] a processor for executing the computer program stored in the memory, and when the computer program is executed, implementing the training method of the image processing model described in any one of the above embodiments of the present disclosure; or implementing the image processing method described in any one of the above embodiments of the present disclosure.
[0155] Figure 11 FIG. 11 is a schematic structural diagram of an application embodiment of the electronic device of the present disclosure. In this embodiment, the electronic device 10 includes one or more processors 11 and a memory 12.
[0156] The processor 11 may be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 10 to perform desired functions.
[0157] The memory 12 may include one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may run the program instructions to implement the methods of the various embodiments of the present disclosure described above and / or other desired functions. Various contents such as input signals, signal components, noise components, etc. may also be stored in the computer-readable storage medium.
[0158] In one example, the electronic device 10 may further include: an input device 13 and an output device 14, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).
[0159] For example, the input device 13 may be the above-mentioned microphone or microphone array for capturing the input signal of the sound source.
[0160] In addition, the input device 13 may further include, for example, a keyboard, a mouse, and so on.
[0161] The output device 14 may output various information to the outside, including the determined distance information, direction information, etc. The output device 14 may include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.
[0162] Of course, for simplicity, Figure 11 only some of the components related to the present disclosure in the electronic device 10 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 10 may further include any other appropriate components.
[0163] Exemplary computer program product and computer-readable storage medium
[0164] In addition to the above methods and devices, an embodiment of the present disclosure may also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the methods according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0165] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0166] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0167] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0168] The basic principles of the present disclosure have been described in connection with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are merely examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. Additionally, the specific details disclosed above are only for illustrative and facilitating understanding purposes and not limitations. The above details do not limit the present disclosure to necessarily adopting the above specific details for implementation.
[0169] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For system embodiments, since they basically correspond to method embodiments, they are described relatively simply. For relevant parts, reference can be made to the corresponding descriptions in the method embodiments.
[0170] The block diagrams of the devices, apparatuses, equipment, and systems involved in the present disclosure are only illustrative examples and do not intend to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.
[0171] The methods and apparatuses of the present disclosure can be implemented in many ways. For example, the methods and apparatuses of the present disclosure can be implemented through software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the method is only for illustration purposes. The steps of the method of the present disclosure are not limited to the specific order described above, unless otherwise specifically stated. Additionally, in some embodiments, the present disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers the recording medium storing the programs for executing the methods according to the present disclosure.
[0172] It should also be noted that in the apparatuses, equipment, and methods of the present disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of the present disclosure.
[0173] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the present disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the aspects shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0174] The above description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present disclosure to the forms disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A training method for an image processing model, comprising: Determining an initial joint image processing model; Processing first training image data in a first image format using a first feature extraction network of the initial joint image processing model to obtain first training features in a first feature format corresponding to the first training image data, where the network parameters of the first feature extraction network are network parameters pre-trained based on image data in the first image format; Converting the format of the first training features using a feature format conversion network of the initial joint image processing model to obtain second training features in a second feature format; Processing second training image data in a second image format using a second feature extraction network of the initial joint image processing model to obtain third training features in the second feature format corresponding to the second training image data; Updating the network parameters of the initial joint image processing model based on the second training features, first label data corresponding to the second training features, the third training features, and second label data corresponding to the third training features to obtain an updated first joint image processing model; In response to the first joint image processing model satisfying a preset condition, determining a target image processing model suitable for processing images in the second image format based on the first joint image processing model.
2. The method according to claim 1, wherein, The updating the network parameters of the initial joint image processing model based on the second training features, first label data corresponding to the second training features, the third training features, and second label data corresponding to the third training features to obtain an updated first joint image processing model includes: Processing the second training features using a first head network of the initial joint image processing model to obtain a first training image processing result corresponding to the second training features; Processing the third training features using the first head network to obtain a second training image processing result corresponding to the third training features; Processing the second training features using a discriminator network of the initial joint image processing model to obtain a first discrimination result corresponding to the second training features, where the first discrimination result includes a first discrimination probability that the second training features are from training image data in the first image format or from training image data in the second image format; Processing the third training features using the discriminator network to obtain a second discrimination result corresponding to the third training features, where the second discrimination result includes a second discrimination probability that the third training features are from training image data in the first image format or from training image data in the second image format; Updating the network parameters of the initial joint image processing model based on the first training image processing result, the first discrimination probability, the first label data, the second training image processing result, the second discrimination probability, the second label data, and a preset update rule to obtain the first joint image processing model.
3. The method according to claim 2, wherein The first label data includes a first processing result label and a first discrimination label; the second label data includes a second processing result label and a second discrimination label; Updating the network parameters of the initial joint image processing model based on the first training image processing result, the first discrimination probability, the first label data, the second training image processing result, the second discrimination probability, the second label data, and a preset update rule to obtain the first joint image processing model, including: Determining a first loss based on the first training image processing result, the first processing result label, the second training image processing result, the second processing result label, and a first loss function; Determining a second loss based on the first discrimination probability, the first discrimination label, and a second loss function; Determining a third loss based on the first discrimination probability, the first discrimination label, the second discrimination probability, the second discrimination label, and a third loss function; Updating the network parameters of the feature format conversion network, the second feature extraction network, and the first head network of the initial joint image processing model based on the first loss and the second loss to obtain an updated second joint image processing model; Updating the network parameters of the discriminator network of the second joint image processing model based on the third loss to obtain the first joint image processing model.
4. The method according to claim 2, wherein Responding to the first joint image processing model satisfying a preset condition, and determining a target image processing model suitable for processing an image in the second image format based on the first joint image processing model, including: Responding to the network loss of the first joint image processing model satisfying a preset end condition, and determining the target image processing model based on the second feature extraction network and the first head network of the first joint image processing model.
5. The method according to claim 1, wherein Determining the initial joint image processing model, including: Obtaining a first image processing model trained based on image data in the first image format, where the first image processing model includes a third feature extraction network and a second head network; Initializing the network parameters of a pre-established initial joint image processing network based on the network parameters of the first image processing model and a preset initialization rule to obtain the initial joint image processing model, where the network parameters of the first feature extraction network are initialized to the network parameters of the third feature extraction network, and the network parameters of the first head network of the initial joint image processing model are initialized to the network parameters of the second head network.
6. The method according to any one of claims 1-5, further comprising: Responding to the first joint image processing model not satisfying the preset condition, using the first joint image processing model as the initial joint image processing model, and repeating the step of processing the first training image data in the first image format by the first feature extraction network of the initial joint image processing model to obtain the first training feature in the first feature format corresponding to the first training image data.
7. An image processing method, including: Obtain an image to be processed, where the image to be processed is an image in a second image format; Process the image to be processed by using a target image processing model to obtain an image processing result corresponding to the image to be processed; the target image processing model is trained based on the training method of the image processing model according to any one of claims 1-6 above.
8. An apparatus for training an image processing model, comprising: A first obtaining module, configured to determine an initial joint image processing model; A first processing module, configured to process first training image data in a first image format by using a first feature extraction network of the initial joint image processing model to obtain first training features in a first feature format corresponding to the first training image data, where network parameters of the first feature extraction network are network parameters pre-trained based on image data in the first image format; A second processing module, configured to perform format conversion on the first training features by using a feature format conversion network of the initial joint image processing model to obtain second training features in a second feature format; A third processing module, configured to process second training image data in a second image format by using a second feature extraction network of the initial joint image processing model to obtain third training features in the second feature format corresponding to the second training image data; A fourth processing module, configured to update network parameters of the initial joint image processing model based on the second training features, first label data corresponding to the second training features, the third training features, and second label data corresponding to the third training features to obtain an updated first joint image processing model; A fifth processing module, configured to, in response to the first joint image processing model meeting a preset condition, determine a target image processing model suitable for processing an image in the second image format based on the first joint image processing model.
9. An image processing apparatus, comprising: A second obtaining module, configured to obtain an image to be processed, where the image to be processed is an image in a second image format; A sixth processing module, configured to process the image to be processed by using a target image processing model to obtain an image processing result corresponding to the image to be processed; the target image processing model is trained based on the training method of the image processing model according to any one of claims 1-6 above.
10. A computer-readable storage medium storing a computer program, where the computer program is used to execute the training method of the image processing model according to any one of claims 1-6 above; or, the computer program is used to execute the image processing method according to claim 7 above.
11. An electronic device, where the electronic device includes: A processor; A memory for storing executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the training method of the image processing model according to any one of claims 1-6 above; or, implement the image processing method according to claim 7 above.
Citation Information
Patent Citations
Image description generation method based on depth LSTM network
CN106650789A
Active sampling and Gaussian mixture model-based image super-resolution reconstruction method
CN107845064A