Image processing method and device and electronic equipment

By decomposing, extracting and fusion images of multiple video source images, the problem of inconsistent video quality in live broadcast scenes is solved, and the stability of picture quality and user experience is improved.

CN120259151APending Publication Date: 2025-07-04LENOVO (BEIJING) LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510369353.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In live broadcast scenarios of multiple video data sources, inconsistent video quality leads to poor user experience.

Method used

Each video source image is broken down into light maps and reflective maps through an intelligent model, feature data is extracted, adjusted and fused, and the target image is generated to maintain consistent image quality.

Benefits of technology

When switching images from multiple data sources, the stability of image quality is maintained and the user experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259151A_ABST
    Figure CN120259151A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and device and electronic equipment, and the method comprises the steps: obtaining source images transmitted by at least two data sources respectively, and enabling the at least two source images to be different in image quality; performing decomposition operation on each source image through an intelligent model to form multiple groups of illumination images and reflection images; feature extraction is carried out on each group of illumination images and reflection images, corresponding first feature data and second feature data are generated, the first feature data and / or the second feature data are / is adjusted based on adjustment parameters, and the adjustment parameters are preset parameters or parameters determined based on at least one source image; fusing the first feature data and the second feature data of each group through an intelligent model to generate a fused feature map; and constructing a corresponding target image based on the fused feature map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and particularly to an image processing method, apparatus, and electronic device. Background Art

[0002] During the processing in multiple scenarios, since the video data captured by multiple data streams may have different image qualities, parameters, etc. When presenting multiple video data to users, the consistency of the image quality of multiple video data cannot be accurately maintained. For example, in a live broadcast scenario, if multi-site linked live broadcast is involved, the image qualities and video parameters of multiple video sources may be different. If not properly processed, it will cause unstable output image quality of the video during the process of presenting multiple video sources to users, resulting in a poor viewing experience for users. Summary of the Invention

[0003] An embodiment of this application provides an image processing method, including:

[0004] Obtain source images respectively sent by at least two data sources, where the image qualities of at least two of the source images are different;

[0005] Through an intelligent model, perform decomposition operations on each of the source images respectively to form multiple groups of illumination maps and reflection maps;

[0006] Extract features from each group of the illumination maps and the reflection maps respectively to generate corresponding first feature data and second feature data, and adjust the first feature data and / or the second feature data based on adjustment parameters, where the adjustment parameters are preset parameters or parameters determined based on at least one of the source images;

[0007] Through the intelligent model, fuse each group of the first feature data and the second feature data respectively to generate a fused feature map;

[0008] Construct a corresponding target image based on the fused feature map.

[0009] Optionally, the intelligent model has a convolutional neural network. The step of, through the intelligent model, performing decomposition operations on each of the source images respectively to form multiple groups of illumination maps and reflection maps includes:

[0010] Input the source image into the convolutional neural network, so that the convolutional neural network outputs corresponding illumination images and reflection images according to the learned features and their mapping relationships with related images, where the loss function of the convolutional neural network is determined based on the reflectance consistency of the reflection map and the smoothness of the illumination map.

[0011] Optionally, the step of respectively fusing the first feature data and the second feature data of each group through the intelligent model to generate a fused feature map includes:

[0012] Performing a preset operation on a first element in the first feature data and a second element in the corresponding second feature data to obtain a corresponding fused element, where the preset operation includes at least one of the following: addition operation, weighted operation, splicing operation, and multiplication operation;

[0013] Generating the fused feature map based on the fused element.

[0014] Optionally, the step of adjusting the first feature data and / or the second feature data based on adjustment parameters includes:

[0015] Adjusting the first feature data and / or the second feature data corresponding to all the source images based on the adjustment parameters; or,

[0016] Adjusting the first feature data and / or the second feature data corresponding to the source images that do not meet the first image quality requirement based on the adjustment parameters, where the adjustment parameters include at least one of the following: brightness parameter, chrominance parameter, and coding parameter.

[0017] Optionally, after constructing a corresponding target image based on the adjusted fused feature map, the method further includes:

[0018] Evaluating the image quality of the target image based on an evaluation metric;

[0019] Adjusting the parameters of the intelligent model based on the evaluation result to optimize the intelligent model.

[0020] Optionally, the method further includes:

[0021] Determining a training data set based on illumination map training data and reflection map training data;

[0022] Training the feature extraction strategy and the fusion strategy of the intelligent model based on the training data set.

[0023] Optionally, the method further includes:

[0024] Selecting one of the multiple source images as a reference image based on a second image quality requirement;

[0025] Obtaining the image parameters of the reference image;

[0026] Determining the image parameters of the reference image as the adjustment parameters.

[0027] Optionally, the method further includes:

[0028] In response to a switching instruction for switching the source image;

[0029] Obtain the corresponding target image from the signal output path indicated by the switching instruction and output it.

[0030] An embodiment of the present application also provides an image processing apparatus, including:

[0031] An acquisition module configured to acquire source images respectively sent by at least two data sources, wherein the image qualities of at least two of the source images are different;

[0032] A decomposition module configured to perform a decomposition operation on each of the source images respectively through an intelligent model to form multiple groups of illumination maps and reflection maps;

[0033] An extraction module configured to perform feature extraction on each group of the illumination maps and the reflection maps respectively to generate corresponding first feature data and second feature data, and adjust the first feature data and / or the second feature data based on adjustment parameters, wherein the adjustment parameters are preset parameters or parameters determined based on at least one of the source images;

[0034] A fusion module configured to perform fusion on each group of the first feature data and the second feature data respectively through the intelligent model to generate a fusion feature map;

[0035] A construction module configured to construct a corresponding target image based on the fusion feature map.

[0036] An embodiment of the present application also provides an electronic device, including a processor and a memory, where an executable program is stored in the memory, and the memory executes the executable program to perform the steps of the method as described above.

[0037] The image processing method of the embodiment of the present application can extract features of each source image obtained from multiple data sources, and adjust the features based on adjustment parameters, so that the image quality of each obtained target image is kept consistent. During the process of switching the target images corresponding to multiple data sources, the stability of the image quality is maintained, and the user experience is improved. Description of the Drawings

[0038] Figure 1 It is a flowchart of the image processing method of the embodiment of the present application;

[0039] Figure 2 For the embodiment of the present application Figure 1 It is a flowchart of an embodiment of step S400 therein;

[0040] Figure 3 Flow chart of the first embodiment of the image processing method according to the embodiments of the present application;

[0041] Figure 4 Flow chart of the second embodiment of the image processing method according to the embodiments of the present application;

[0042] Figure 5 Flow chart of the third embodiment of the image processing method according to the embodiments of the present application;

[0043] Figure 6 Schematic diagram of the internal connection relationship of the electronic device according to the embodiments of the present application;

[0044] Figure 7 Block diagram of the structure of the electronic device according to the embodiments of the present application. Detailed implementation manners

[0045] Various solutions and features of the present application are described herein with reference to the accompanying drawings.

[0046] It should be understood that various modifications can be made to the embodiments applied herein. Therefore, the above description should not be regarded as a limitation, but only as an example of the embodiments. Those skilled in the art will think of other modifications within the scope and spirit of the present application.

[0047] The accompanying drawings included in the specification and forming a part of the specification illustrate the embodiments of the present application, and together with the general description of the present application given above and the detailed description of the embodiments given below are used to explain the principles of the present application.

[0048] These and other features of the present application will become apparent from the following description of the preferred forms of the embodiments given by way of non - limiting examples with reference to the accompanying drawings.

[0049] It should also be understood that although the present application has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present application.

[0050] When combined with the accompanying drawings, the above and other aspects, features and advantages of the present application will become more apparent in view of the following detailed description.

[0051] Specific embodiments of the present application are described hereinafter with reference to the accompanying drawings; however, it should be understood that the embodiments applied are only examples of the present application, and it can be implemented in various ways. Well - known and / or repetitive functions and structures are not described in detail to avoid obscuring the present application with unnecessary or redundant details. Therefore, the specific structural and functional details applied herein are not intended to be limiting, but only as a basis for the claims and a representative basis for teaching those skilled in the art to use the present application in substantially any suitable detailed structure in various ways.

[0052] This specification may use phrases such as "in one embodiment", "in another embodiment", "in yet another embodiment", or "in other embodiments", which may each refer to one or more of the same or different embodiments according to the present application.

[0053] An image processing method according to an embodiment of the present application, which processes source images sent by at least two data sources by using an intelligent model, can make the quality of the processed images the same or similar, so that during the process of switching between using at least two source images, the image quality of the processed output remains stable.

[0054] The following will describe in detail the image processing method of the present application in conjunction with specific embodiments. Figure 1 is a flowchart of the image processing method according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:

[0055] S100, obtain source images respectively sent by at least two data sources, where the image quality of at least two of the source images is different.

[0056] Exemplarily, the data source may be an image acquisition device for acquiring an image of a target object, and the target object is the object to be photographed, such as an object in various types of live broadcast scenes, a monitored object, a production product object, etc.

[0057] In one embodiment, at least two data sources may be data sources at different shooting angles, or data sources shot at different time periods, or data sources for different target objects. For example, in a live racing scene or a live sports event scene, multiple data sources may be set at different positions in the racing scene to capture different source images.

[0058] In one embodiment, multiple data sources respectively send their own source images. Since the device performance, shooting conditions, and shooting time of each data source may be different, the image quality of the source images will be different. For example, the source image captured by the first data source and sent to the electronic device is of the first image quality, and the source image captured by the second data source and sent to the electronic device is of the second image quality, and the first image quality is higher than the second image quality.

[0059] The electronic device obtains multiple source images from multiple data sources respectively. Then, each source image is processed separately.

[0060] S200, through the intelligent model, perform a decomposition operation on each of the source images to form multiple sets of illumination maps and reflection maps.

[0061] Exemplarily, the intelligent model can be a neural network model capable of processing images, such as a deep learning model. The intelligent model includes a convolutional neural network and an enhancement network, or is the convolutional neural network itself. A convolutional neural network (CNN for short) is a deep learning model designed to process data with a grid structure (such as images and audio), including convolutional layers, pooling layers, and fully connected layers. Among them, the convolutional layer: performs a convolution operation by sliding a convolution kernel over the input data to extract local features of the data. The parameters in the convolution kernel are shared, which greatly reduces the number of model parameters, reduces the computational amount, and the risk of overfitting. The pooling layer: is usually used for downsampling, that is, compressing the feature map extracted by the convolutional layer. Common pooling methods include max pooling and average pooling, which can reduce the data dimension while retaining the main features, improving the robustness and computational efficiency of the model. The fully connected layer: After passing through multiple convolutional layers and pooling layers, the extracted features are integrated, and the features are mapped to the output space in a fully connected manner for the final classification or regression task. The enhancement network can enrich the training dataset of the CNN by generating high-quality images and further improve its performance.

[0062] In this embodiment, the intelligent model is used to perform a decomposition operation on each source image. Each source image can be decomposed into a corresponding illumination map and reflection map, thereby forming multiple groups of illumination maps and reflection maps. Among them, one source image corresponds to a group of illumination maps (illumination images) and reflection maps (reflection images). The illumination map is an image that can reflect the illumination situation of an object or a scene. There are obvious bright and dark parts in the formation of the illumination map. The direction, intensity, and color of the light will cause different light and shadow effects on the object surface, forming a rich sense of hierarchy and three-dimensionality. The reflection image is formed based on the principle of light reflection. When light irradiates the object surface, part of the light is absorbed and part of the light is reflected. The reflected light enters the human eye or the imaging device, and then the reflection image is formed.

[0063] In this embodiment, the intelligent model performs a decomposition operation on the source image respectively based on the formation principles of the illumination map and the reflection map. The illumination map of the formed source image depicts the bright and dark parts, the direction, intensity, and color of the light of the source image. The reflection map of the formed source image depicts the light irradiation and reflection characteristics of the source image. The intelligent model decomposes each group of source images, forming multiple groups of illumination maps and reflection maps. Each group of illumination maps and reflection maps corresponds to a source image.

[0064] S300, perform feature extraction on the illumination map and the reflection map of each group respectively to generate corresponding first feature data and second feature data, and adjust the first feature data and / or the second feature data based on the adjustment parameter, where the adjustment parameter is a preset parameter or a parameter determined based on at least one of the source images.

[0065] Exemplarily, the intelligent model performs feature extraction on each group of illumination maps and reflection maps. Specifically: within the same group, feature extraction is performed on the illumination map to generate corresponding first feature data, so that the first feature data can characterize the features of the current illumination map; feature extraction is performed on the reflection map to generate corresponding second feature data, so that the second feature data can characterize the features of the current reflection map. In another group, feature extraction is performed on the illumination map again to generate corresponding first feature data, and feature extraction is performed on the reflection map to generate corresponding second feature data.

[0066] In one embodiment, performing feature extraction on the illumination map includes: color feature extraction, texture feature extraction, edge feature extraction, and deep learning-based feature extraction. For example, on the one hand, in the RGB color space of the illumination map, histogram statistics are respectively performed on the red, green, and blue channels to obtain three histograms, which are combined to describe the color features of the illumination map. On the other hand, by calculating the occurrence frequency of different gray-level pixel pairs in the illumination map, a gray-level co-occurrence matrix is generated, and multiple texture features, such as contrast, correlation, energy, and entropy, can be extracted from this matrix.

[0067] Performing feature extraction on the reflection map includes: geometric feature extraction, texture feature extraction, color feature extraction, and deep learning-based feature extraction. For example, on the one hand, the edges and contours of objects in the reflection image are important geometric features. Edge detection algorithms, such as the Canny operator and the Sobel operator, can be used to extract the edges of objects in the reflection image. By connecting these edge points, the contours of the objects can be obtained. On the other hand, by calculating the occurrence frequency of different gray-level pixel pairs in a specific direction and distance, a gray-level co-occurrence matrix is generated. Features such as contrast, correlation, energy, and entropy can be extracted from the gray-level co-occurrence matrix, and these features can reflect information such as the roughness and texture direction of the reflection surface.

[0068] Adjustment parameters are used to adjust the first feature data and / or the second feature data, thereby ensuring that the fused image obtained by fusing the first feature data and the second feature data meets the image quality requirements. On the one hand, the adjustment parameters are preset parameters. Specifically, the adjustment parameters can be preset based on historical data or empirical data, so that when adjusting each group of the first feature data and the second feature data, the adjustment parameters are used for the adjustment. On the other hand, the adjustment parameters are determined based on at least one source image. For example, they can be determined according to one or more of the multiple source images, such as determined according to a source image with the best image quality. Thus, after using the adjustment parameters to adjust the first feature data and the second feature data, a target image with a higher image quality than the original image quality of other source images can be obtained through processing. In this way, each image obtained by the electronic device from multiple data sources can reach the required image quality, and the image quality of each image can also be kept consistent. For example, during a live broadcast, when switching the screen for multiple data sources at the live broadcast site, the image quality of each obtained target image can be kept consistent.

[0069] S400, through the intelligent model, fuse the first feature data and the second feature data of each group respectively to generate a fused feature map.

[0070] Exemplarily, fusing different types of feature data can make full use of the advantages of each feature and improve the data representation ability and the performance of the model. In this embodiment, the intelligent model is used to fuse each group of the first feature data and the second feature data after parameter adjustment respectively to generate corresponding fused feature maps, and these fused feature maps can more fully represent the features of the corresponding source images. The specific fusion method can be realized by various fusion means. For example, each data unit in the first feature data can be subjected to addition operation, weighted operation, splicing operation, and / or multiplication operation, etc. with the corresponding data unit in the second feature data to obtain a fused data unit, and then the fused feature map is obtained based on the fused data unit.

[0071] In one embodiment, since the image quality of each source image is different, the image quality corresponding to the fused feature maps obtained by fusing the first feature data and the second feature data of each group is also different.

[0072] S500, construct a corresponding target image based on the fused feature map.

[0073] Exemplarily, when reconstructing an image based on the fused feature map, the generator in the intelligent model can use the fused feature map as input, and by continuously adjusting the parameters, make the generated image as close as possible to the real image. At the same time, the discriminator evaluates the generated image and feeds back to the generator to optimize the generation process. The fused feature map can be normalized so that its numerical range is within a specific interval. The fused feature map can also be feature-adjusted, such as adding additional layers to further extract or transform features, or performing operations such as cropping and splicing on the fused feature map. The features can be decoded using a decoder to achieve image reconstruction, and then the target image can be obtained.

[0074] The image processing method of the embodiments of the present application can extract features from each source image obtained from multiple data sources, and adjust the features based on the adjusted parameters, so that the image quality of each obtained target image is consistent. The stability of the image quality is maintained during the process of switching the target images corresponding to multiple data sources, improving the user experience.

[0075] In an embodiment of the present application, the intelligent model has a convolutional neural network. Through the intelligent model, each of the source images is respectively decomposed to form multiple groups of illumination maps and reflection maps, including:

[0076] The source image is input into the convolutional neural network, so that the convolutional neural network outputs the corresponding illumination image and reflection image according to the learned features and their mapping relationship with the related images, where the loss function of the convolutional neural network is determined based on the reflectance consistency of the reflection map and the smoothness of the illumination map.

[0077] Exemplarily, the intelligent model can first perform some preprocessing on the input source image, such as normalization, color space conversion, etc., to reduce the influence of illumination and reflection changes on the image. Then, the preprocessed source image is input into the convolutional neural network. The convolutional neural network learns the features of the image and establishes a mapping relationship from the input source image to the illumination map and the reflection map, so as to determine the illumination image and the reflection image according to this mapping relationship. In terms of the network structure, the convolutional neural network can adopt an encoder-decoder architecture. The encoder is used to extract the high-level features of the image, and the decoder generates the illumination map and the reflection map according to these features. In order to better conform to the physical model, the design of the loss function is usually determined based on the physical properties of the image, such as energy conservation, the value range of the reflectance, etc.

[0078] Preferably, the loss function of the convolutional neural network in the embodiments of the present application can be determined based on the reflectance consistency of the reflection map and the smoothness of the illumination map. As an evaluation index in the convolutional neural network, the loss function is used to measure the difference or error between the output of the convolutional neural network and the true label. Therefore, the parameters of the intelligent model can be optimized by minimizing the loss function. The loss function can be a non-negative real-valued function, denoted as L(Y, f(X)), where Y is the actual value (also called the label or true value), f(X) is the predicted value of the model (also called the output value or estimated value), and X is the input data. The smaller the value of the loss function, the closer the predicted result of the model is to the actual value, and the better the performance of the intelligent model. In this embodiment, the loss function is determined based on the reflectance consistency of the reflection map and the smoothness of the illumination map, so that the convolutional neural network can keep the smoothness of the illumination map within the corresponding preferred range and the reflectance consistency within the corresponding preferred range during the process of processing images.

[0079] In one embodiment of the present application, through the intelligent model, the first feature data and the second feature data in each group are respectively fused to generate a fused feature map, as Figure 2 shown, including the following steps:

[0080] S410, perform a preset operation on the first element in the first feature data and the second element in the corresponding second feature data to obtain a corresponding fused element, where the preset operation includes at least one of the following: addition operation, weighted operation, splicing operation, and multiplication operation;

[0081] S420, generate the fused feature map based on the fused element.

[0082] Exemplarily, during the fusion of images, corresponding elements of two feature maps with the same shape in the first feature data and the second feature data can be subjected to a preset operation to obtain corresponding fused elements.

[0083] When the preset operation is an addition operation, the corresponding feature values from different feature sources (the first feature data and the second feature data) are directly added. For example, the feature vector A=(a1, a2,..., a n ) in the first feature data and B=(b1, b2,..., b n ) in the second feature data, and the fused feature vector (fused element) F=(a1 + b1, a2 + b2,..., a n + b n ) is obtained through the addition operation. This addition operation can, to a certain extent, retain the information of each feature. When different features are complementary and their contribution degrees to the result are similar, the addition operation can effectively combine these features.

[0084] When the preset operation is a weighted operation, weights are assigned to the features from different feature sources, and then the weighted feature values are added together. Given feature vectors A = (a1, a2, …, a n ) and B = (b1, b2, …, b n ), and the corresponding weight vector W = (w1, w2, …, w n ), then the fused feature vector F = (w1a1 + w1b1, w2a2 + w2b2, …, w n a n + w n b n ). The magnitude of the weights reflects the importance of the corresponding features, which can be learned from training data or set according to prior knowledge. The weighted operation can weight features according to their importance, thus more flexibly fusing different features. It can effectively highlight the features that have a greater impact on the result, suppress unimportant features, and improve the quality and representativeness of the fused features.

[0085] When the preset operation is a concatenation operation, the feature vectors from different feature sources are concatenated in a certain order to form a longer feature vector. For example, given feature vectors A = (a1, a2, …, a m ) and B = (b1, b2, …, b n ), the fused feature vector F after concatenation = (a1, a2, …, a m , b1, b2, …, b n ). The concatenation operation can retain all the information of the original features without losing any details.

[0086] When the preset operation is a multiplication operation, the corresponding feature values from different feature sources are multiplied. For feature vectors A = (a1, a2, …, a n ) and B = (b1, b2, …, b n ), the fused feature vector F = (a1 × b1, a2 × b2, …, a n × b n ). The multiplication operation is usually used to emphasize a certain interaction or correlation between features. When two features both have relatively large values in certain dimensions, multiplication will make the feature values in these dimensions more prominent, thus highlighting the synergistic effect between features and discovering potential relationships between different features.

[0087] After obtaining each fusion element, the fused feature vector F as described above, the fusion elements are combined to form a fused feature map.

[0088] In one embodiment of the present application, the adjustment of the first feature data and / or the second feature data based on the adjustment parameter includes:

[0089] Adjusting the first feature data and / or the second feature data corresponding to all the source images based on the adjustment parameter; or,

[0090] Adjusting the first feature data and / or the second feature data corresponding to the source images that do not meet the first image quality requirement based on the adjustment parameter, where the adjustment parameter includes at least one of the following: brightness parameter, chroma parameter, and coding parameter.

[0091] Exemplarily, the adjustment parameter can be a preset parameter. Specifically, it can be preset based on historical data or empirical data to complete the adjustment parameter, and it can also be a parameter determined based on at least one source image. On the one hand, the adjustment parameter can be used to adjust the first feature data and / or the second feature data corresponding to each source image. For example, by adjusting the brightness parameter, chroma parameter, and / or coding parameter related to the first feature data and / or the second feature data, the target image can be adjusted, thereby improving the image quality of the target image corresponding to each source image and ensuring that the image quality of each target image can meet the preset first image quality requirement.

[0092] On the other hand, some of the source images may meet the first image quality requirement, and some may not. To improve the adjustment efficiency, only the first feature data and / or the second feature data corresponding to the source images that do not meet the first image quality requirement can be adjusted, so as to only adjust the target images corresponding to some source images. It is ensured that the target images corresponding to each data source can meet the first image quality requirement. Thus, the image quality of the current target image can be maintained stable during the process of switching and using multiple target images.

[0093] In one embodiment of the present application, after constructing the corresponding target image based on the adjusted fusion feature map, as Figure 3 shown, the method further includes the following steps:

[0094] S600, evaluating the image quality of the target image based on an evaluation index;

[0095] S700, adjusting the parameters of the intelligent model based on the evaluation result to optimize the intelligent model.

[0096] Exemplarily, after constructing the target image based on the intelligent model, the quality of the target image may meet the user's requirements or there may be errors. In this embodiment, the image quality of the target image will be evaluated based on evaluation metrics to determine whether the generated target image meets the preset requirements. If it meets the preset requirements, it indicates that the parameter configuration of the current intelligent model is appropriate; if it does not meet the preset requirements, it indicates that the parameters of the current intelligent model need to be adjusted, so that the parameters of the intelligent model can be adjusted accordingly based on the evaluation results to optimize the intelligent model.

[0097] In one embodiment of the present application, the method further includes the following steps:

[0098] Determine a training data set based on the illumination map training data and the reflection map training data;

[0099] Based on the training data set, train the feature extraction strategy and the fusion strategy of the intelligent model.

[0100] Exemplarily, the intelligent model needs to be trained to optimize itself, including data preparation, network construction, selection of appropriate loss functions and optimizers, model training and evaluation. When training, a training data set is prepared in advance to train the intelligent model. Among them, a training data set is determined based on the illumination map training data and the reflection map training data. The illumination map training data is used for the intelligent model to learn to decompose the source image into an illumination map, and the reflection map training data is used for the intelligent model to learn to decompose the source image into a reflection map. Both the illumination map training data and the reflection map training data can be historical data used by the intelligent model or determined based on relevant empirical data.

[0101] After training the intelligent model based on the training data set, the intelligent model can improve the intelligence and accuracy of the feature extraction strategy and the fusion strategy. The feature extraction strategy is a strategy for extracting features from the illumination map and the reflection map, and the fusion strategy is a strategy for fusing the first feature data and the second feature data.

[0102] In one embodiment of the present application, as Figure 4 shown, the method further includes the following steps:

[0103] S10, select one of the multiple source images as a reference image based on the second image quality requirement.

[0104] Exemplarily, the image qualities of multiple source images are different. Among them, relatively speaking, the image qualities of some source images are higher and meet the second image quality requirement. Therefore, one of the source images that meet the second image quality requirement can be selected as the reference image. The reference image is a reference for adjusting other source images. The second image quality requirement can be a pre-determined image quality requirement, such as determining the second image quality requirement based on the source image with the highest image quality among all source images, or selecting one from multiple pre-set image quality requirements as the second image quality requirement.

[0105] S20. Obtain the image parameters of the reference image.

[0106] Exemplarily, the reference image has corresponding image parameters, such as brightness parameters, color parameters, encoding parameters, resolution parameters, data volume parameters, etc. The image parameters can be obtained through the analysis of the reference image itself.

[0107] S30. Determine the image parameters of the reference image as the adjustment parameters.

[0108] Exemplarily, the adjustment parameters are used to adjust the first feature data and / or the second feature data, so as to ensure that the fused image obtained by fusing the first feature data and the second feature data meets the image quality requirement. After determining the image parameters of the reference image as the adjustment parameters, the reference image can be used as a reference when using the adjustment parameters to adjust the feature data corresponding to other source images, improving the quality of the fused image corresponding to other source images, and thus improving the image quality of the target images corresponding to each source image.

[0109] In an embodiment of the present application, as Figure 5 shown, the method further includes the following steps:

[0110] S40. Respond to a switching instruction for switching the source image.

[0111] Exemplarily, when an electronic device plays images from multiple data sources, it is necessary to switch multiple images. For example, during a live car race, it is necessary to switch images taken by cameras at multiple different positions. The user can actively operate or send a switching instruction to the electronic device (also a live device) through other devices. The electronic device responds to the switching instruction and starts the switching operation of multiple source images.

[0112] S50. Obtain the corresponding target image from the signal output path indicated by the switching instruction and output it.

[0113] Exemplarily, the switching instruction can indicate the target image that needs to be used currently and the target image that needs to be replaced. Each target image has a corresponding data source and a related signal output path. For example, the first data source has a first signal output path, and based on this first signal output path, the source image can be output to the electronic device. The second data source has a second signal output path, and based on this second signal output path, the source image can be output to the electronic device, etc. Therefore, the target image obtained from the signal output path indicated by the switching instruction is the image that the user currently needs to view, and thus this target image can be output for viewing. Of course, if the signal output path indicated by the updated switching instruction changes, the corresponding target image can be obtained from the signal output path indicated by the updated switching instruction and output, thereby realizing the switching operation of the target image.

[0114] Next, in combination with a specific embodiment, the process of an electronic device obtaining multiple source images and processing them based on its internal functional modules will be specifically described.

[0115] As Figure 6 shown, the electronic device has a media stream input adaptation interface (Signal Adapter Switcher), which can be respectively connected to a network stream, a serial digital video interface (12G / 6G / 3G SDI), and a high-definition multimedia interface (HDMI) to obtain the source images sent by the above multiple interfaces. The video decoding module (Video Decoder) in the electronic device can perform video decoding processing on the source images and input them into the video frame correction module (Video frame correction) of the electronic device. The video frame correction module includes a video color gamut format unit (YUV / RGB), a brightness unit (Brightness), a chroma unit (Chroma), an equalization unit (Equalization), and an RGB color gamut format unit. The video color gamut format unit is used to perform color gamut processing on the image data, the brightness unit is used to perform brightness processing on the image data, the chroma unit is used to perform chroma processing on the image data, and the equalization unit is used to adjust the parameter equalization degree of the image data. The intelligent model is connected to the video frame correction module and can control the video frame correction module, thereby realizing the processing of the source images. The video frame correction module is connected to the video encoding module (VideoEncoder), and the video encoding module compresses and encodes the corrected / repaired video frames and transmits them to the video output bus (Signal output bus). The media output bus is compatible with the network stream, SDI signal, and HDMI interface, and thus the target images are output through each interface.

[0116] An embodiment of the present application further provides an image processing device, which can be applied to an electronic device, such as Figure 7 as shown, including:

[0117] An acquisition module, configured to acquire source images respectively sent by at least two data sources, wherein the image qualities of at least two of the source images are different.

[0118] Exemplarily, the data source can be an image acquisition device for acquiring an image of a target object, and the target object is the object to be photographed, such as an object in various types of live broadcast scenes, a monitored object, a production product object, etc.

[0119] In one embodiment, at least two data sources can be data sources at different shooting angles, data sources shooting at different time periods, or data sources for different target objects. For example, in a live racing scene or a live sports event scene, multiple data sources can be set at different positions in the racing scene to capture different source images.

[0120] In one embodiment, multiple data sources respectively send their own source images. Since the device performance, shooting conditions, and shooting time of each data source may be different, the image qualities of the source images will be different. For example, the source image captured by the first data source and sent to the electronic device is the first image quality, and the source image captured by the second data source and sent to the electronic device is the second image quality, and the first image quality is higher than the second image quality.

[0121] The acquisition module acquires multiple source images from multiple data sources respectively, so that the electronic device can process each source image separately.

[0122] A decomposition module, configured to perform a decomposition operation on each of the source images respectively through an intelligent model to form multiple groups of illumination maps and reflection maps.

[0123] Exemplarily, the intelligent model can be a neural network model capable of processing images, such as a deep learning model. This intelligent model includes a convolutional neural network and an enhancement network, or is the convolutional neural network itself. A Convolutional Neural Network (CNN) is a deep learning model designed to process data with a grid structure (such as images and audio), including convolutional layers, pooling layers, and fully connected layers. Among them, the convolutional layer: performs a convolution operation by sliding a convolution kernel over the input data to extract local features of the data. The parameters in the convolution kernel are shared, which greatly reduces the number of model parameters, reduces the computational amount, and the risk of overfitting. The pooling layer: is usually used for downsampling, that is, compressing the feature map extracted by the convolutional layer. Common pooling methods include max pooling and average pooling. It can reduce the data dimension while retaining the main features, improving the robustness and computational efficiency of the model. The fully connected layer: After passing through multiple convolutional layers and pooling layers, the extracted features are integrated, and the features are mapped to the output space in a fully connected manner for the final classification or regression task. The enhancement network can enrich the training dataset of the CNN by generating high-quality images and further improve its performance.

[0124] In this embodiment, the decomposition module can use the intelligent model to perform a decomposition operation on each source image. Each source image can be decomposed into a corresponding illumination map and reflection map, thereby forming multiple groups of illumination maps and reflection maps. Among them, one source image corresponds to a group of illumination maps (illumination images) and reflection maps (reflection images). The illumination map is an image that can reflect the illumination situation of an object or a scene. There are obvious bright and dark parts in the formation of the illumination map. The direction, intensity, and color of the light will cause different light and shadow effects on the object surface, forming a rich sense of hierarchy and three-dimensionality. The reflection image is formed based on the principle of light reflection. When light irradiates the object surface, part of the light is absorbed and part of the light is reflected. The reflected light enters the human eye or the imaging device, and then the reflection image is formed.

[0125] In this embodiment, the decomposition module can use the intelligent model to perform a decomposition operation on the source image respectively based on the formation principles of the illumination map and the reflection map. The formed illumination map of the source image depicts the bright and dark parts, the direction, intensity, and color of the light in the source image. The formed reflection map of the source image depicts the light irradiation and reflection characteristics of the source image. The intelligent model decomposes each group of source images to form multiple groups of illumination maps and reflection maps. Each group of illumination maps and reflection maps corresponds to a source image.

[0126] An extraction module configured to perform feature extraction on the light map and the reflection map of each group respectively, generate corresponding first feature data and second feature data, and adjust the first feature data and / or the second feature data based on adjustment parameters, where the adjustment parameters are preset parameters or parameters determined based on at least one of the source images.

[0127] Exemplarily, the extraction module performs feature extraction on the light map and the reflection map of each group. Specifically, within the same group, the extraction module uses an intelligent model to perform feature extraction on the light map to generate corresponding first feature data, so that the first feature data can characterize the features of the current light map; and performs feature extraction on the reflection map to generate corresponding second feature data, so that the second feature data can characterize the features of the current reflection map. In another group, the extraction module uses the intelligent model to perform feature extraction on the light map again to generate corresponding first feature data, and performs feature extraction on the reflection map to generate corresponding second feature data.

[0128] In one embodiment, performing feature extraction on the light map includes: color feature extraction, texture feature extraction, edge feature extraction, and deep learning-based feature extraction. For example, on the one hand, in the RGB color space of the light map, histogram statistics are respectively performed on the red, green, and blue channels to obtain three histograms, which are combined to describe the color features of the light map. On the other hand, by calculating the occurrence frequency of different gray-level pixel pairs in the light map, a gray-level co-occurrence matrix is generated, and multiple texture features, such as contrast, correlation, energy, and entropy, can be extracted from this matrix.

[0129] Performing feature extraction on the reflection map includes: geometric feature extraction, texture feature extraction, color feature extraction, and deep learning-based feature extraction. For example, on the one hand, the edges and contours of objects in the reflection image are important geometric features. Edge detection algorithms, such as the Canny operator and the Sobel operator, can be used to extract the edges of objects in the reflection image. By connecting these edge points, the contours of the objects can be obtained. On the other hand, by calculating the occurrence frequency of different gray-level pixel pairs in a specific direction and distance, a gray-level co-occurrence matrix is generated. Features such as contrast, correlation, energy, and entropy can be extracted from the gray-level co-occurrence matrix, and these features can reflect information such as the roughness and texture direction of the reflection surface.

[0130] Adjustment parameters are used to adjust the first feature data and / or the second feature data, thereby ensuring that the fused image obtained by fusing the first feature data and the second feature data meets the image quality requirements. On the one hand, the adjustment parameters are preset parameters. Specifically, the adjustment parameters can be preset based on historical data or empirical data, so that when adjusting each group of the first feature data and the second feature data, the adjustment parameters are used for adjustment. On the other hand, the adjustment parameters are determined based on at least one source image. For example, they can be determined according to one or more source images among multiple source images. For example, they can be determined according to a source image with the best image quality. Thus, after using the adjustment parameters to adjust the first feature data and the second feature data, a target image with a higher image quality than the original image quality of other source images can be obtained after processing. In this way, each image obtained by the electronic device from multiple data sources can reach the required image quality, and the image quality of each image can also be kept consistent. For example, during a live broadcast, when switching the pictures of multiple data sources at the live broadcast site, the image quality of each obtained target image can be kept consistent.

[0131] A fusion module, configured to fuse each group of the first feature data and the second feature data respectively through the intelligent model to generate a fused feature map.

[0132] Exemplarily, fusing different types of feature data can make full use of the advantages of each feature, improving the data representation ability and the performance of the model. In this embodiment, the intelligent model is used to fuse each group of the first feature data and the second feature data after parameter adjustment respectively to generate corresponding fused feature maps, and these fused feature maps can more fully represent the features of the corresponding source images. The specific fusion method can be realized by various fusion means. For example, the fusion module can perform addition operation, weighted operation, splicing operation, and / or multiplication operation, etc. on each data unit in the first feature data and the corresponding data unit in the second feature data to obtain a fused data unit, and then obtain a fused feature map based on the fused data unit.

[0133] In one embodiment, since the image quality of each source image is different, the image quality corresponding to the fused feature maps obtained by the fusion module fusing each group of the first feature data and the second feature data is also different.

[0134] A construction module, configured to construct a corresponding target image based on the fused feature map.

[0135] Exemplarily, when reconstructing an image based on the fused feature map, the building block can use the generator in the intelligent model to take the fused feature map as input, and by continuously adjusting the parameters, make the generated image as close as possible to the real image. At the same time, the discriminator evaluates the generated image and feeds back to the generator to optimize the generation process. The fused feature map can be normalized to make its numerical range within a specific interval. The fused feature map can also be feature-adjusted, such as adding additional layers to further extract or transform features, or performing operations such as cropping and splicing on the fused feature map. The decoder can be used again to decode the features to achieve image reconstruction, and thus obtain the target image.

[0136] In one embodiment of the present application, the decomposition module is further configured to:

[0137] Input the source image into the convolutional neural network, so that the convolutional neural network outputs the corresponding illumination image and reflection image according to the learned features and their mapping relationship with the relevant images, where the loss function of the convolutional neural network is determined based on the reflectance consistency of the reflection map and the smoothness of the illumination map.

[0138] In one embodiment of the present application, the fusion module is further configured to:

[0139] Perform a preset operation on the first element in the first feature data and the corresponding second element in the second feature data to obtain the corresponding fused element, where the preset operation includes at least one of the following: addition operation, weighted operation, splicing operation, and multiplication operation;

[0140] Generate the fused feature map based on the fused element.

[0141] In one embodiment of the present application, the extraction module is further configured to:

[0142] Adjust all the first feature data and / or the second feature data corresponding to the source images based on the adjustment parameters; or,

[0143] Adjust the first feature data and / or the second feature data corresponding to the source images that do not meet the first image quality requirement based on the adjustment parameters, where the adjustment parameters include at least one of the following: brightness parameter, chroma parameter, and coding parameter.

[0144] In one embodiment of the present application, the image processing device further includes an evaluation module, and the evaluation module is configured to:

[0145] Evaluate the image quality of the target image based on the evaluation index;

[0146] Adjust the parameters of the intelligent model based on the evaluation results to optimize the intelligent model.

[0147] In one embodiment of the present application, the image processing device further includes a training module, and the training module is configured to:

[0148] Determine a training data set based on the irradiation map training data and the reflection map training data;

[0149] Based on the training data set, train the feature extraction strategy and the fusion strategy of the intelligent model.

[0150] In one embodiment of the present application, the extraction module is further configured to:

[0151] Based on the second image quality requirement, select one of the multiple source images as the reference image;

[0152] Obtain the image parameters of the reference image;

[0153] Determine the image parameters of the reference image as the adjustment parameters.

[0154] In one embodiment of the present application, the image processing device further includes a switching module, and the switching module is configured to:

[0155] Respond to a switching instruction for switching the source image;

[0156] Obtain the corresponding target image from the signal output path indicated by the switching instruction and output it.

[0157] An embodiment of the present application further provides an electronic device, including a processor and a memory, where an executable program is stored in the memory, and the memory executes the executable program to perform the steps of the method described above.

[0158] The above embodiments are only exemplary embodiments of the present application and are not used to limit the present application. The protection scope of the present application is defined by the claims. Those skilled in the art can make various modifications or equivalent replacements within the essence and protection scope of the present application, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present application.

Claims

1. An image processing method, comprising: Obtaining source images respectively sent by at least two data sources, wherein the image qualities of at least two of the source images are different; Performing a decomposition operation on each of the source images respectively through an intelligent model to form multiple groups of illumination maps and reflection maps; Performing feature extraction on each group of the illumination maps and the reflection maps respectively to generate corresponding first feature data and second feature data, and adjusting the first feature data and / or the second feature data based on adjustment parameters, wherein the adjustment parameters are preset parameters or parameters determined based on at least one of the source images; Fusing each group of the first feature data and the second feature data respectively through the intelligent model to generate a fused feature map; Constructing a corresponding target image based on the fused feature map.

2. The method according to claim 1, wherein the intelligent model has a convolutional neural network, and the performing a decomposition operation on each of the source images respectively through the intelligent model to form multiple groups of illumination maps and reflection maps comprises: Inputting the source image into the convolutional neural network, so that the convolutional neural network outputs corresponding illumination images and reflection images according to the learned features and their mapping relationships with related images, wherein the loss function of the convolutional neural network is determined based on the reflectance consistency of the reflection map and the smoothness of the illumination map.

3. The method according to claim 1, wherein the fusing each group of the first feature data and the second feature data respectively through the intelligent model to generate a fused feature map comprises: Performing a preset operation on a first element in the first feature data and a second element in the corresponding second feature data to obtain a corresponding fused element, wherein the preset operation includes at least one of the following: addition operation, weighted operation, splicing operation, and multiplication operation; Generating the fused feature map based on the fused element.

4. The method according to claim 1, wherein the adjusting the first feature data and / or the second feature data based on adjustment parameters comprises: Adjusting the first feature data and / or the second feature data corresponding to all the source images based on the adjustment parameters; Or, Adjusting the first feature data and / or the second feature data corresponding to the source images that do not meet the first image quality requirement based on the adjustment parameters, wherein the adjustment parameters include at least one of the following: brightness parameter, chroma parameter, and coding parameter.

5. The method according to claim 1, after constructing a corresponding target image based on the adjusted fused feature map, the method further comprises: Evaluating the image quality of the target image based on an evaluation index; Adjusting the parameters of the intelligent model based on the evaluation result to optimize the intelligent model.

6. The method according to claim 1, the method further comprises: Determining a training data set based on illumination map training data and reflection map training data; Training the feature extraction strategy and the fusion strategy of the intelligent model based on the training data set.

7. The method according to claim 1, the method further comprises: Select one of the multiple source images as the reference image based on the second image quality requirement; Obtain the image parameters of the reference image; Determine the image parameters of the reference image as the adjustment parameters.

8. The method according to claim 1, wherein the method further comprises: Respond to a switching instruction for switching the source image; Obtain the corresponding target image from the signal output path indicated by the switching instruction and output it.

9. An image processing apparatus, comprising: An acquisition module configured to acquire source images respectively sent by at least two data sources, wherein the image qualities of at least two of the source images are different; A decomposition module configured to perform a decomposition operation on each of the source images respectively through an intelligent model to form multiple groups of illumination maps and reflection maps; An extraction module configured to perform feature extraction on each group of the illumination maps and the reflection maps respectively to generate corresponding first feature data and second feature data, and adjust the first feature data and / or the second feature data based on adjustment parameters, wherein the adjustment parameters are preset parameters or parameters determined based on at least one of the source images; A fusion module configured to perform fusion on each group of the first feature data and the second feature data respectively through the intelligent model to generate a fusion feature map; A construction module configured to construct a corresponding target image based on the fusion feature map.

10. An electronic device, comprising a processor and a memory, wherein the memory stores an executable program, and the memory executes the executable program to perform the steps of the method according to claim 1.