A method, apparatus, electronic device, and storage medium for processing an image
By deploying an image acquisition module on the vehicle, calculating image evaluation scores using global and local feature vectors, selecting target images and performing image conversion, the problem that the existing technology cannot take into account the difficulty of image acquisition and shooting timing, and efficient image processing and batch filter addition are achieved.
Patent Information
- Application Number
- CN202210045010.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-14
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-01-14
AI Technical Summary
The existing image acquisition technology cannot take into account both the difficulty of image acquisition and the timing of shooting, which leads to difficulties for users to take high-quality and interesting photos on the road.
The image set is acquired by the image acquisition module deployed on the vehicle, and the global feature vector and local feature vector of the candidate image are determined, the image evaluation score is calculated, the target image is selected and the image conversion model is imported, and the processing image corresponding to the preset filter is output.
It realizes the rapid filtering of target images with high shooting quality from large batches of candidate images, and batch filter addition is realized through the image conversion model, which improves image processing efficiency and user experience.
Smart Images

Figure CN115115848B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of data processing, and particularly relates to an image processing method, apparatus, electronic device, and storage medium. Background Art
[0002] With the continuous improvement of people's living standards, photography has become an important means for users to record their daily lives. Users can take pictures of the scenery along the way through different devices such as cameras and mobile phones to leave good memories. Especially during the journey, users can, while the vehicle is moving, through the in-vehicle... Therefore, how to take high-quality and interesting photos has become one of the key concerns of users.
[0003] In the existing image acquisition technology, after a user obtains an image through a device such as a camera, the user can manually select a suitable image from the captured images and add corresponding filters to the image. During the journey, since the number of pictures obtained by the image acquisition device mounted on the vehicle is large, it takes a lot of time for the user to select a suitable image from the huge number of images. Or when using a handheld device (such as a mobile phone or a camera) to take pictures, the user also needs to manually press the shutter, and it is easy to miss the shooting opportunity. Thus, it can be seen that the existing image acquisition technology cannot take into account both the image acquisition difficulty and the shooting opportunity. Summary of the Invention
[0004] Embodiments of this application provide an image processing method, apparatus, electronic device, and storage medium, which can solve the problem that the existing image processing technology cannot take into account both the image acquisition difficulty and the shooting opportunity.
[0005] In a first aspect, embodiments of this application provide an image processing method, including:
[0006] Receiving an image set captured by an image acquisition module deployed on a vehicle; the image set includes a plurality of candidate images;
[0007] Respectively determining the global feature vector and the local feature vector of each candidate image, and calculating an image evaluation score of the candidate image according to the global feature vector and the local feature vector;
[0008] Selecting a target image to be converted from all the candidate images according to the image evaluation score;
[0009] Importing the target image into a preset image conversion model, and outputting a processed image corresponding to a preset filter.
[0010] In a possible implementation of the first aspect, separately determining the global feature vectors and local feature vectors of each of the candidate images, and calculating an image evaluation score of the candidate images according to the global feature vectors and the local feature vectors includes:
[0011] Dividing the candidate image into a plurality of image blocks, and generating block vectors corresponding to the respective image blocks;
[0012] Separately calculating an attention score between any one block vector and each of the other block vectors;
[0013] Weighting each of the other block vectors according to the attention score to obtain a fused feature vector of the any one block vector; the fused feature vector is specifically:
[0014]
[0015] where f’ x,y is the fused feature vector of the any one block vector; f x,y is the any one block vector; is the i-th other block vector; is the attention score between the i-th other block vector and the any one block vector; N is the total number of the image blocks;
[0016] Calculating the local feature vector corresponding to the candidate image based on all the fused feature vectors.
[0017] In a possible implementation of the first aspect, separately determining the global feature vectors and local feature vectors of each of the candidate images, and calculating an image evaluation score of the candidate images according to the global feature vectors and the local feature vectors includes:
[0018] Converting the candidate image into a global image vector, and importing the global feature vector into a convolutional layer in a preset global feature extraction model to obtain a convolutional feature vector corresponding to the global image vector;
[0019] Importing the convolutional feature vector into a fully connected layer in the global feature extraction model to obtain the global feature vector;
[0020] Combining the global feature vector and the local feature vector to generate a combined feature vector corresponding to the candidate image;
[0021] Importing the combined feature vectors corresponding to all the candidate images into a classification scoring module to determine the image evaluation scores of the respective candidate images.
[0022] In a possible implementation of the first aspect, the image conversion model includes a generation network and a discriminative network;
[0023] Before importing the target image into a preset image conversion model to output a processed image corresponding to a preset filter, it further includes:
[0024] Output a to-be-verified image corresponding to the training image through the generation network;
[0025] Import the to-be-verified image and multiple sample images generated based on the preset filter into the discriminative network to generate classification results corresponding to the to-be-verified image and the multiple sample images;
[0026] If the category to which the to-be-verified image belongs is within the category corresponding to the preset filter, it is recognized that the generation network has been trained, so as to output the processed image of the target image through the generation network;
[0027] If the category to which the to-be-verified image belongs is outside the category corresponding to the preset filter, adjust the parameters in the generation network based on the difference threshold, and return to execute outputting the to-be-verified image corresponding to the training image through the generation network.
[0028] In a possible implementation of the first aspect, the importing the to-be-verified image and multiple sample images generated based on the preset filter into the discriminative network to generate classification results corresponding to the to-be-verified image and the multiple sample images includes:
[0029] Extract a first feature vector corresponding to the to-be-verified image, and respectively extract second feature vectors corresponding to each of the sample images;
[0030] Calculate the vector similarity between the first feature vector and all the second feature vectors;
[0031] Classify the to-be-verified image and all the sample images based on the vector similarity.
[0032] In a possible implementation of the first aspect, the respectively determining the global feature vector and the local feature vector of each candidate image, and calculating the image evaluation score of the candidate image according to the global feature vector and the local feature vector includes:
[0033] Determine the target position information where the vehicle is located when the candidate image is taken through a positioning module;
[0034] Determine the driving direction of the vehicle according to a plurality of associated positions corresponding to the target position information; the associated positions are other position information whose acquisition time difference from the position information is less than a preset time threshold;
[0035] Determine the shooting position weight corresponding to the candidate image according to the target position information and the driving direction;
[0036] Calculate the image evaluation score of the candidate image according to the global feature vector, the local feature vector and the shooting position weight.
[0037] In a possible implementation manner of the first aspect, the receiving the image set captured by the image acquisition module deployed on the vehicle includes:
[0038] The vehicle driving video obtained by the driving recorder deployed in the vehicle; the vehicle driving video includes multiple video images; each video image is a first candidate image;
[0039] In response to a shooting instruction initiated by the user, obtain a second candidate image through a camera device deployed inside the vehicle;
[0040] Generate the image set according to the first candidate image and the second candidate image.
[0041] In a second aspect, an embodiment of the present application provides an image processing device, including:
[0042] An image set acquisition unit, configured to receive an image set captured by an image acquisition module deployed on a vehicle; the image set includes a plurality of candidate images;
[0043] An image evaluation score calculation unit, configured to respectively determine the global feature vector and the local feature vector of each candidate image, and calculate the image evaluation score of the candidate image according to the global feature vector and the local feature vector;
[0044] A target image selection unit, configured to select a target image to be converted from all the candidate images according to the image evaluation score;
[0045] A filter adding unit, configured to import the target image into a preset image conversion model and output a processed image corresponding to a preset filter.
[0046] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the method according to any one of the first aspects when executing the computer program.
[0047] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the method according to any one of the above first aspects.
[0048] Fifthly, an embodiment of the present application provides a computer program product, which when running on a server, causes the server to execute the method according to any one of the above first aspects.
[0049] Sixthly, an embodiment of the present application provides a vehicle equipped with the electronic device of the third aspect. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method according to any one of the above first aspects.
[0050] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: After receiving the image set fed back by the image acquisition device deployed on the vehicle, the global feature vectors and local feature vectors of each candidate image in the image set can be determined respectively, so as to score the candidate images in the global dimension and the local dimension, and based on the global feature vectors and local feature vectors, determine the image evaluation score for evaluating the shooting quality of the candidate images, thereby realizing the quantization process of the quality of the shooting, so that the target images with better shooting quality can be extracted from a large number of candidate images according to the image evaluation score, and then through the image conversion model associated with the filter style to be added, output the processed image corresponding to the target image, achieving the purpose of batch filter addition. Compared with the existing image processing technology, the image set in the embodiments of the present application is obtained by the image acquisition module deployed on the vehicle, which can avoid missing the shooting opportunity, and the electronic device does not require the user to manually select the appropriate target image from a large number of image sets, but can screen the images by calculating the image evaluation scores of each candidate image, greatly reducing the time consumed by the user when selecting images and reducing the unnecessary operations of the user, thereby improving the efficiency of image processing; on the other hand, the electronic device can batch process the images according to the filters preset by the user, and can also avoid the user adding filters one by one, further improving the efficiency of image processing and greatly enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0052] Figure 1 It is a schematic structural diagram of an image processing system applied to a vehicle provided by an embodiment of the present application;
[0053] Figure 2 It is an interaction flowchart of an image processing method provided by an embodiment of the present application;
[0054] Figure 3 It is a schematic structural diagram of a scoring model provided by an embodiment of the present application;
[0055] Figure 4 It is a schematic diagram of filter addition provided by an embodiment of the present application;
[0056] Figure 5 It is an interaction flowchart of an image processing method S202 provided by an embodiment of the present application;
[0057] Figure 6 It is a schematic diagram of the generation of local feature vectors provided by an embodiment of the present application;
[0058] Figure 7 It is a schematic diagram of an implementation manner of S202 of an image processing method provided by an embodiment of the present application;
[0059] Figure 8 It is a schematic diagram of an implementation manner of S204 of an image processing method provided by an embodiment of the present application;
[0060] Figure 9 It is a schematic structural diagram of an image conversion model provided by an embodiment of the present application;
[0061] Figure 10 It is a schematic diagram of an implementation manner of S202 of an image processing method provided by an embodiment of the present application;
[0062] Figure 11 It is a schematic diagram of an implementation manner of S201 of an image processing method provided by an embodiment of the present application;
[0063] Figure 12 It is a schematic structural diagram of an image processing device provided by an embodiment of the present application;
[0064] Figure 13 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0065] In the following description, specific details such as specific system architectures, technologies, etc. are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.
[0066] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.
[0067] In addition, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0068] The image processing method provided by the embodiments of the present application can be applied to electronic devices such as in-vehicle terminals, smartphones, servers, tablet computers, laptop computers, and netbooks that can implement image processing. In particular, the electronic device can specifically be a central control terminal installed in a vehicle, and the central control terminal can establish a communication connection with an image acquisition module deployed on the vehicle to process the image fed back by the image acquisition module and then display the processed image through a display screen configured in the central control terminal.
[0069] Exemplarily, Figure 1 shows a schematic structural diagram of an image processing system applied to a vehicle provided by an embodiment of the present application. Refer to Figure 1 As shown, the image processing system includes a central control terminal 11 inside the vehicle. There are two image acquisition modules deployed inside the vehicle, namely a driving recorder 21 for acquiring images of the external environment of the vehicle and a camera device 22 for acquiring images of the internal scene of the vehicle. A wired connection can be established between the central control terminal 11 and the driving recorder and the camera device 22 through a serial port, or a wireless communication connection can be established through a wireless communication module, which is specifically set according to the actual situation. Among them, the user can complete multi-angle photographing operations inside the carriage through the camera device 22. If the central control terminal is configured with a Bluetooth communication module, a smartphone 23 held by the user can also be used as an extended image acquisition module in the image processing system. In this case, the smartphone 23 can send the acquired images to the central control terminal 11 for processing.
[0070] Please refer to Figure 2 , Figure 2 shows a schematic implementation diagram of an image processing method provided by an embodiment of the present application. The method includes the following steps:
[0071] In S201, an image set captured by an image acquisition module deployed on a vehicle is received; the image set includes a plurality of candidate images.
[0072] In this embodiment, an image acquisition module is deployed inside the vehicle. The image acquisition module can be devices such as a driving recorder and an in-vehicle camera. Through the image acquisition module, a large number of candidate images can be acquired during the driving of the vehicle, and all the acquired candidate images are encapsulated to obtain the above-mentioned image set. Among them, if the image acquisition module is specifically a driving recorder, the driving recorder can obtain video data during the driving of the vehicle. In this case, the electronic device can regard each frame of the video data as a candidate image.
[0073] In a possible implementation manner, the image acquisition module can be configured with an acquisition frequency, and the image acquisition module can obtain a plurality of candidate images based on the above-mentioned acquisition frequency during the driving of the vehicle, so as to generate the above-mentioned image set.
[0074] In a possible implementation manner, the electronic device can be configured with a plurality of key position areas suitable for taking pictures. When the electronic device detects that the current position reaches the above-mentioned key position areas, it can send an acquisition trigger instruction to the image acquisition module (during the process that the vehicle is located in the above-mentioned key position areas, the electronic device can send multiple acquisition trigger instructions to the image acquisition module). The image acquisition module can acquire candidate images in response to the above-mentioned acquisition trigger instruction, so as to obtain the above-mentioned image set.
[0075] In a possible implementation manner, the electronic device can be configured with a start condition for image processing. For example, it can detect the number of candidate images included in the acquired image set. If the number of candidate images reaches a preset start quantity threshold, the image processing process can be started; for another example, the electronic device can obtain the running state of the vehicle. If it detects that the engine of the vehicle is in the off state, that is, the user may reach a transfer station or the end point of the journey, at this time, the images taken during the journey can be processed, and the image processing process is started.
[0076] In a possible implementation manner, the electronic device can output an operation interface through a display module, and a start control for batch processing of images is configured in the operation interface. If the electronic device detects that the user clicks the above-mentioned start control, it recognizes that the user needs to perform batch image processing on the received image set, and then starts the image processing process.
[0077] In S202, the global feature vector and the local feature vector of each of the candidate images are respectively determined, and an image evaluation score of the candidate image is calculated according to the global feature vector and the local feature vector.
[0078] In this embodiment, the electronic device may import each candidate image in the image set into the scoring model. The scoring module specifically includes two branches, namely a global extraction branch for determining the global feature vector and a local extraction branch for determining the local feature vector. When the candidate image is imported into the above scoring model, it will first determine that the candidate image is processed through the two branches respectively to determine the corresponding global feature vector and local feature vector. Then, the electronic device may merge the above two feature vectors and calculate the image evaluation score corresponding to the candidate image based on the merged vector.
[0079] In this embodiment, the above global feature vector determines the features of the candidate image in the global dimension. For example, the shooting quality of the candidate image can be evaluated from multiple dimensions, such as the composition of the picture, the ratio of brightness and darkness, the overall color of the picture, etc. The above dimensions are all related to the overall picture. In this case, the global feature vector output by the global extraction branch in the scoring model can be used to represent the score within the above overall dimension. Correspondingly, the local feature vector determines the features of the candidate image in the local dimension. For example, when evaluating the shooting quality of an image, in addition to being related to global features such as composition, color, and the ratio of brightness and darkness, local details will also affect the evaluation of the image. For example, the degree of blurring of the background of the picture, whether the shooting subject is in focus accurately, whether the details of the shooting subject are complete and clear, etc. The above local features can be determined by the local feature vector output by the local extraction branch in the scoring model.
[0080] Among them, the extraction module of the above global feature vector and the extraction module of the local feature vector may be constructed based on a neural network, and the neural network includes but is not limited to: convolutional neural network, long short-term memory neural network, etc. In particular, the neural network is specifically a convolutional neural network, which specifically includes three layers, namely a convolutional layer conv, a batch normalization layer BN, and a rectified linear unit Relu layer.
[0081] Exemplarily, Figure 3 shows a schematic structural diagram of the scoring model provided by an embodiment of the present application. Refer to Figure 3 As shown, the scoring model may perform parallel processing on the input candidate image through two branches, namely a global feature vector extraction branch and a local feature vector extraction branch. After the electronic device determines the above two feature vectors of the candidate image, it may splice the above two feature vectors to obtain a merged feature vector corresponding to the candidate image, and import it into the fully connected layer of the scoring model, so as to output the image evaluation score corresponding to the candidate image.
[0082] In a possible implementation, the image evaluation score is specifically used to determine the aesthetic characteristics of the candidate image. If the image evaluation score is higher, the image can be considered more beautiful. Therefore, a suitable target image can be selected based on this evaluation score.
[0083] In S203, according to the image evaluation score, a target image to be converted is selected from all the candidate images.
[0084] In this embodiment, since the image evaluation score can determine the shooting quality of the candidate image, the higher the value of the image evaluation score, the higher the corresponding shooting quality, that is, the more aesthetic it is. The electronic device can screen the above candidate images according to the image evaluation score, so as to determine the target image that needs to add a filter.
[0085] In a possible implementation, the electronic device can be configured with a score threshold. The candidate images with an image evaluation score less than the above score threshold are identified as filtered images that do not need to be style-converted, while the candidate images with an image evaluation score greater than or equal to the score threshold are identified as target images and subsequent processing is performed. Of course, the electronic device can also be set with an image upper limit value. If the number of candidate images with an image scoring score greater than or equal to the score threshold is greater than the above image upper limit value, the first N candidate images with larger image evaluation scores can be selected as target images.
[0086] In a possible implementation, the electronic device can also calculate the similarity between candidate images with each image evaluation score greater than the score threshold. If the similarity between any two candidate images is greater than a preset similarity threshold, then the candidate image with a larger image evaluation score is selected as the target image, thereby reducing the repetition between the selected target images.
[0087] In a possible implementation, the electronic device can be configured with an image upper limit value. The electronic device can arrange each candidate image in descending order of the image evaluation score and select the first N candidate images as target images; where the value of N is the above image upper limit value.
[0088] In S204, the target image is imported into a preset image conversion model, and a processed image corresponding to the preset filter is output.
[0089] In this embodiment, the electronic device can be configured with an image conversion module, and this image conversion module can add a preset filter to the input image. Exemplarily, Figure 4 shows a schematic diagram of filter addition provided by an embodiment of the present application. See Figure 4As shown in (a) in [reference], the target image before the input image conversion module, that is, the original image, after being processed by the above-mentioned image conversion module, a preset filter (such as a cartoon-style filter) is added to this image to obtain a processed image, that is, Figure 4 the image shown in (b) in [reference]. Based on this, the electronic device can add filters to the screened target images respectively through the image conversion model to obtain processed images that match the styles of the preset filters.
[0090] In this embodiment, the above-mentioned preset filter can be manually configured by the user, that is, the user can select one or more preset filters from the filter library as the preset filters used in batch processing. Each preset filter can correspond to an image conversion algorithm, or based on the selected multiple preset filters, an image conversion algorithm that can output different preset filters is generated (that is, the image conversion algorithm contains multiple branches, and different branches are used to output processed images of one preset filter).
[0091] In a possible implementation manner, after the electronic device generates the processed images of each target image, an image selection interface can be generated. The image selection interface contains all the target images and the processed images output based on the target images. If there are multiple preset filters selected by the user, each target image can correspond to multiple processed images. The user can select appropriate images in this image selection interface, for example, by checking and other methods, and after the selection is completed, store the target images and / or processed images selected by the user, and delete the unselected target images and processed images.
[0092] As can be seen from the above, in the image processing method provided by the embodiment of the present application, after receiving the image set fed back by the image acquisition device deployed on the vehicle, the global feature vectors and local feature vectors of each candidate image in the image set can be determined respectively, so as to score the candidate images in the global dimension and the local dimension, and based on the global feature vectors and local feature vectors, determine the image evaluation score for evaluating the shooting quality of the candidate images, thereby realizing the quantization process of the pros and cons of the shooting quality. Therefore, the target images with better shooting quality can be extracted from a large number of candidate images according to the image evaluation score, and then, through the image conversion model associated with the filter style to be added, the processed image corresponding to the target image is output, achieving the purpose of batch filter addition. Compared with the existing image processing technology, the image set in the embodiment of the present application is obtained by the image acquisition module deployed on the vehicle, which can avoid missing the shooting opportunity, and the electronic device does not need the user to manually select the appropriate target image from a large number of image sets. Instead, it can screen images by calculating the image evaluation scores of each candidate image, greatly reducing the time consumed by the user when selecting images and reducing unnecessary operations of the user, thereby improving the efficiency of image processing. On the other hand, the electronic device can batch process images according to the filters preset by the user, and can also avoid the user adding filters one by one, further improving the efficiency of image processing and greatly enhancing the user experience.
[0093] Figure 5 FIG. shows a specific implementation flowchart of an image processing method provided by the second embodiment of the present invention. Refer to Figure 5 , relative to Figure 2 the embodiment described above, in the image processing method provided by this embodiment, S202 includes: S501 to S504, which are specifically described in detail as follows:
[0094] Furthermore, the step of respectively determining the global feature vectors and local feature vectors of each candidate image, and calculating the image evaluation score of the candidate image according to the global feature vectors and the local feature vectors includes:
[0095] In S501, the candidate image is divided into a plurality of image blocks, and block vectors corresponding to each image block are generated.
[0096] In this embodiment, when the electronic device extracts the local feature vector of the candidate image, it can divide the candidate image into blocks, divide the candidate image into N image blocks, and the size of each image block is the same. Exemplarily, Figure 6 FIG. shows a schematic diagram of the generation of the local feature vector provided by an embodiment of the present application. Refer to Figure 6As shown, the candidate image is a C*H*W image, where C represents the number of channels of the candidate image. If the candidate image is a grayscale image, the corresponding number of channels is 1. If the candidate image is a color image, the corresponding number of channels is 3. That is, the number of channels C can be set according to the actual situation. H represents the pixel height of the candidate image, and W represents the pixel width of the candidate image. The electronic device can divide the above image into N blocks. Each block is specifically an image block of size C*k*k. The pixel width and pixel height of each image block are both k. According to the pixel values of each pixel point in the image block, a block vector about the image block can be generated. The block vector is specifically a multi-dimensional vector of C*k*k.
[0097] In S502, the attention scores between any block vector and each other block vector are calculated respectively.
[0098] In this embodiment, since it is necessary to determine the local features of the candidate image, the context attention mechanism can be added to improve the feature fusion of the foreground and background in the global feature vector, making the local features more abundant, which is beneficial to the robustness of the subsequent training model. Based on this, the electronic device needs to calculate the attention scores of each image block relative to the foreground and the background. Take an image block as the foreground feature block, and the other image blocks except this foreground feature block are the background feature blocks relative to this foreground feature block. Then, calculate the attention scores between this foreground feature block and each background feature block.
[0099] In this embodiment, when the electronic device calculates the attention score between this block vector and other block vectors, a reshape operation can be performed on other block vectors, as Figure 6 shown, that is, convert multiple block vectors arranged in two dimensions into (N - 1)*C*k*k tensors, that is Then, calculate the vector similarity between the block vector and other block vectors, and use this vector similarity as the attention score between the two. Exemplarily, if the vector similarity is the cosine similarity, the attention score is specifically expressed as:
[0100]
[0101] Among them, is the attention score between the block vector f x,y and the i-th other block vector , and ||f x,y || is the vector norm of the block vector f x,y , and is the vector norm of the i-th other block vector.
[0102] In S503, each of the other block vectors is weighted according to the attention score to obtain a fused feature vector of any one of the block vectors; specifically, the fused feature vector is:
[0103]
[0104] where f’ x,y is the fused feature vector of any one of the block vectors; f x,y is any one of the block vectors; is the i-th other block vector; is the attention score between the i-th other block vector and any one of the block vectors; N is the total number of the image blocks.
[0105] In this embodiment, after the electronic device calculates the attention scores between the foreground feature block (i.e., any one of the block vectors above) and each background feature block (i.e., other block vectors except any one of the block vectors), it can calculate the fused feature vector that fuses the foreground and the background, that is, weights and superimposes each background feature block based on the attention score, so as to obtain the fused feature vector that fuses the foreground and the background as the fused feature vector corresponding to any one of the block vectors. The above steps are performed for all block vectors to obtain the fused feature vectors corresponding to each block vector.
[0106] In S504, a local feature vector corresponding to the candidate image is calculated based on all the fused feature vectors.
[0107] In this embodiment, after all block vectors are converted into corresponding fused feature vectors, the local features corresponding to each region can be determined. All the fused feature vectors are spliced and reconstructed into a fused image, and a corresponding local feature vector is determined based on the fused image, as Figure 6 shown, for each fused image.
[0108] In the embodiment of the present application, by introducing context attention features when calculating the local feature vector, the feature fusion of the foreground and the background can be improved, so that the local features are richer, which is beneficial to the training of the subsequent network layers and makes the model more robust.
[0109] Figure 7 shows a specific implementation flowchart of an image processing method S202 provided in the third embodiment of the present invention. Refer to Figure 7 , relative to Figure 2 the above embodiment, S202 in the image processing method provided in this embodiment includes: S701 to S704, which are specifically described as follows:
[0110] In S701, the candidate image is converted into a global image vector, and the global feature vector is imported into the convolutional layer of a preset global feature extraction model to obtain a convolutional feature vector corresponding to the global image vector.
[0111] In this embodiment, the electronic device can represent the candidate image in the form of a vector. For example, an image of C*H*W can be converted into a multi-dimensional vector of C*H*W according to the pixel values of each pixel. If the image set contains M candidate images, it is converted into a global feature vector of M*C*H*W. To extract the global features of the candidate image, the electronic device can import the global image vector into the convolutional layer to output the corresponding convolutional feature vector. The convolutional layer can include multiple convolutional kernels, and the global feature vector is subjected to feature extraction through multiple convolutional kernels to obtain the convolutional feature vector. Among them, the convolutional layer can be a convolutional layer with a Conv+BN+ReLU structure.
[0112] In S702, the convolutional feature vector is imported into the fully connected layer of the global feature extraction model to obtain the global feature vector.
[0113] In this embodiment, after the electronic device calculates the convolutional feature vector corresponding to the candidate image, it can import it into the fully connected layer to output a one-dimensional global feature vector. For example, if the candidate image is an image of 3*244*244, the obtained global feature vector can be a 256-bit feature sequence.
[0114] In S703, the global feature vector and the local feature vector are combined to generate a combined feature vector corresponding to the candidate image.
[0115] In this embodiment, the electronic device can combine the determined global feature vector and local feature vector to obtain the combined feature vector of the candidate image. For example, both the global feature vector and the local feature vector are 256-bit feature sequences, and the combined combined feature vector is specifically a 512-bit feature sequence.
[0116] In S704, all the combined feature vectors corresponding to the candidate images are imported into the classification scoring module to determine the image evaluation scores of each candidate image.
[0117] In this embodiment, after outputting the combined feature vectors corresponding to all candidate images, such as obtaining M*512 feature sequences, all the obtained combined feature vectors can be imported into the classification scoring module for classification scoring to determine the aesthetic degree of the image and calculate the corresponding image evaluation score.
[0118] Figure 8The figure shows a specific implementation flowchart of an image processing method provided by the fourth embodiment of the present invention. Refer to Figure 8 relative to Figure 2 the embodiment, an image processing method provided in this embodiment includes: S801 to S804, which are specifically described in detail as follows:
[0119] Furthermore, the image conversion model includes a generation network and a discriminant network;
[0120] Before importing the target image into a preset image conversion model to output a processed image corresponding to a preset filter, it further includes:
[0121] In S801, a to-be-verified image corresponding to the training image is output through the generation network.
[0122] In this embodiment, the electronic device can perform adversarial training and learning on the generation network through the discriminant network, so that the output effect of the generation network is consistent with the preset filter. Based on this, the electronic device can obtain one or more training images for training, and output a to-be-verified image corresponding to the training image through the generation network.
[0123] Exemplarily, Figure 9 The figure shows a schematic structural diagram of an image conversion model provided by an embodiment of the present application. Refer to Figure 9 As shown, the image conversion model includes two networks, namely a generation network for performing image conversion, and the other is a determination network for verifying the similarity between the generation network and the preset filter. Unsupervised learning is performed through the generation network and the determination network.
[0124] In S802, the to-be-verified image and a plurality of sample images generated based on the preset filter are imported into the discriminant network to generate classification results corresponding to the to-be-verified image and the plurality of sample images.
[0125] In this embodiment, the electronic device stores a plurality of sample images generated based on the preset filter. In order to determine whether the style of the to-be-verified image is similar to that of the sample images to determine whether the conversion is successful, the to-be-verified image and the sample images can be imported into the discriminant network to generate corresponding classification results. The classification results are the results output after classifying the to-be-verified image and the sample images.
[0126] Furthermore, as another embodiment of the present application, the above S802 may specifically include the following three steps:
[0127] Step 1: Extract a first feature vector corresponding to the to-be-verified image, and respectively extract second feature vectors corresponding to each of the sample images.
[0128] Step 2: Calculate the vector similarity between the first feature vector and all the second feature vectors with each other.
[0129] Step 3: Classify the image to be verified and all the sample images based on the vector similarity.
[0130] In this embodiment, the discrimination network includes a feature extraction layer and a fully connected layer. The feature extraction layer is used to extract features from the image to be verified and the sample images respectively, obtaining the first feature vector and the second feature vectors. Then, calculate the vector similarity between each pair of feature vectors, and classify all the images according to this vector similarity. Multiple images with a vector similarity less than a preset threshold are grouped into one category, and two images with a vector similarity greater than the preset threshold are classified into two different categories, so as to obtain the corresponding classification result.
[0131] In the embodiment of the present application, by converting the image into the corresponding feature vector and calculating the similarity between different feature vectors, the image can be classified, improving the accuracy of classification.
[0132] In S803, if the category where the image to be verified is located is within the category corresponding to the preset filter, it is recognized that the generation network has been trained, so as to output the processed image of the target image through the generation network.
[0133] In this embodiment, after classifying the image to be verified, the electronic device can identify the image category where the image to be verified is located. If the image to be verified is within the image category corresponding to the preset filter, it is recognized that the image to be verified is similar to the sample image corresponding to the preset filter. At this time, it can be recognized that the conversion effect is good and the generation network has been trained. For the subsequent received target image, the generation network can be used for processing. When generating the corresponding processed image of the target network subsequently, there is no need for the discrimination network to supervise, but it can be directly output through the generation network.
[0134] In S804, if the category where the image to be verified is located is outside the category corresponding to the preset filter, adjust the parameters in the generation network based on the difference threshold, and return to execute outputting the image to be verified corresponding to the training image through the generation network.
[0135] In this embodiment, if the image to be verified is outside the image category corresponding to the preset filter, it is recognized that the image to be verified is not similar to the sample image corresponding to the preset filter. At this time, it can be recognized that the conversion effect is poor and the generation network needs to be adjusted. Therefore, it is necessary to adjust the parameters in the generation network and reprocess the target image.
[0136] In the embodiments of the present application, by using a discrimination network to supervise the output effect of the Shengcheng network, the matching degree between image processing and a preset filter can be improved, and the effect of the processed image after conversion can be enhanced.
[0137] Figure 10 Fig. shows a specific implementation flowchart of an image processing method S202 provided in the fifth embodiment of the present invention. Refer to Figure 10 , relative to Figures 2 - 9 any one of the above embodiments, an image processing method S202 provided in this embodiment includes: S1001 to S1004, which are specifically described in detail as follows:
[0138] Further, the step of respectively determining the global feature vector and the local feature vector of each candidate image, and calculating the image evaluation score of the candidate image according to the global feature vector and the local feature vector includes:
[0139] In S1001, a positioning module is used to determine the target position information of the vehicle when the candidate image is captured.
[0140] In S1002, according to a plurality of associated positions corresponding to the target position information, the driving direction of the vehicle is determined; the associated positions are other position information whose acquisition time difference from the position information is less than a preset time threshold.
[0141] In S1003, according to the target position information and the driving direction, the shooting position weight corresponding to the candidate image is determined.
[0142] In this embodiment, the electronic device can use the positioning module to determine the corresponding target position information when capturing the candidate image, and according to the driving direction of the vehicle, in combination with the target position information, determine the shooting angle corresponding to the candidate image. Different shooting angles have different degrees of influence on the shooting effect. Therefore, the shooting position weight corresponding to the candidate image can be calculated by combining the above two features.
[0143] In a possible implementation manner, the electronic device may record the key positions of different scenic spots (commonly known as check-in points). By calculating whether the target position information is close to the position where the scenic spot is located, the position score corresponding to the target position information can be determined, and according to the driving direction of the vehicle, it can be determined whether the shooting angle of the image acquisition module can capture the key markers in the scenic spot, so as to obtain the angle score. Therefore, according to the position score and the angle score, the shooting position weight corresponding to the candidate image can be calculated.
[0144] In S1004, the image evaluation score of the candidate image is calculated according to the global feature vector, the local feature vector, and the shooting position weight.
[0145] In this embodiment, the electronic device can calculate a preliminary evaluation score corresponding to the candidate image based on the global feature vector and the local feature vector, and superimpose the calculated shooting position weight on the basis of the preliminary evaluation score, so as to calculate the above-mentioned image evaluation score.
[0146] In the embodiment of the present application, by considering the influence of the position of the candidate image on the shooting quality, since the user often passes by scenic spots when driving a vehicle, if the candidate image obtained by shooting contains a specific scenic spot, the feasibility of higher shooting quality is higher, so as to improve the accuracy of extracting the selected target image.
[0147] Figure 11 The specific implementation flowchart of a method S201 for processing an image provided in the sixth embodiment of the present invention is shown. Refer to Figure 11 relative to Figures 2 - 9 Any of the above embodiments, a method S201 for processing an image provided in this embodiment includes: S2011 to S2013, which are specifically described in detail as follows:
[0148] In S2011, a vehicle driving video obtained by a driving recorder deployed in the vehicle; the vehicle driving video includes multiple video images; each frame of the video image is a first candidate image.
[0149] In S2012, in response to a shooting instruction initiated by the user, a second candidate image is obtained by a camera device deployed inside the vehicle.
[0150] In S2013, an image set is generated according to the first candidate image and the second candidate image.
[0151] In this embodiment, there are two types of image acquisition devices in the vehicle, namely a driving recorder and a camera device. The driving recorder can shoot in the driving direction or in the direction of the vehicle tail, and is specifically set according to the actual situation. The vehicle driving video obtained by the driving recorder, and multiple frames of images in the video are used as the first candidate images; on the other hand, the user can control the shooting device to collect images by initiating a shooting instruction, and obtain the second candidate images, so that the image set includes the first candidate images of the external scene of the vehicle and the second candidate images of the internal scene of the vehicle.
[0152] In the embodiment of the present application, candidate images of different scenes can be obtained through different types of image acquisition devices, which can improve the diversity of the types of collected images, make it more convenient for users to shoot suitable images, and improve the flexibility of image acquisition.
[0153] Figure 12The structural block diagram of an image processing device provided by an embodiment of the present invention is shown. Each unit included in the image processing device is used to execute Figure 2 each step implemented by the encryption device in the corresponding embodiment. For details, please refer to Figure 2 and Figure 2 the relevant descriptions in the corresponding embodiment. For the sake of convenience of description, only the parts related to this embodiment are shown.
[0154] See Figure 12 , the image processing method device includes:
[0155] An image set acquisition unit 121, configured to receive an image set captured by an image acquisition module deployed on a vehicle; the image set includes a plurality of candidate images;
[0156] An image evaluation score calculation unit 122, configured to respectively determine the global feature vector and the local feature vector of each of the candidate images, and calculate the image evaluation score of the candidate image according to the global feature vector and the local feature vector;
[0157] A target image selection unit 123, configured to select a target image to be converted from all the candidate images according to the image evaluation score;
[0158] A filter addition unit 124, configured to import the target image into a preset image conversion model and output a processed image corresponding to a preset filter.
[0159] Optionally, the image evaluation score calculation unit 122 includes:
[0160] An image block division unit, including dividing the candidate image into a plurality of image blocks and generating block vectors corresponding to each of the image blocks;
[0161] An attention score calculation unit, including calculating the attention score between any one block vector and each of the other block vectors respectively;
[0162] A fused feature vector generation unit, including weighting each of the other block vectors according to the attention score to obtain a fused feature vector of any one block vector; the fused feature vector is specifically:
[0163]
[0164] where f’ x,y is the fused feature vector of any one block vector; f x,y is any one block vector; is the i-th other block vector; is the attention score between the i-th other block vector and any one of the block vectors; N is the total number of the image blocks;
[0165] The local feature vector generation unit includes calculating the local feature vector corresponding to the candidate image based on all the fused feature vectors.
[0166] Optionally, the image evaluation score calculation unit 122 includes:
[0167] The convolutional feature vector generation unit includes converting the candidate image into a global image vector and importing the global feature vector into a convolutional layer of a preset global feature extraction model to obtain a convolutional feature vector corresponding to the global image vector;
[0168] The global feature vector generation unit includes importing the convolutional feature vector into a fully connected layer of the global feature extraction model to obtain the global feature vector;
[0169] The merged feature vector generation unit includes merging the global feature vector and the local feature vector to generate a merged feature vector corresponding to the candidate image;
[0170] The merged feature vector import unit includes importing the merged feature vectors corresponding to all the candidate images into a classification scoring module to determine the image evaluation scores of each candidate image.
[0171] Optionally, the image conversion model includes a generation network and a discriminant network;
[0172] The processing device further includes:
[0173] The image to be verified generation unit includes outputting an image to be verified corresponding to the training image through the generation network;
[0174] The image classification unit includes importing the image to be verified and a plurality of sample images generated based on the preset filters into the discriminant network to generate classification results corresponding to the image to be verified and the plurality of sample images;
[0175] The first determination unit includes if the category of the image to be verified is within the category corresponding to the preset filter, then recognizing that the generation network has been trained, so as to output the processed image of the target image through the generation network;
[0176] The second determination unit includes if the category of the image to be verified is outside the category corresponding to the preset filter, then adjusting the parameters in the generation network based on the difference threshold, and returning to execute outputting the image to be verified corresponding to the training image through the generation network.
[0177] Optionally, the image classification unit includes:
[0178] A feature vector generation unit, including extracting a first feature vector corresponding to the image to be verified, and respectively extracting second feature vectors corresponding to each of the sample images;
[0179] A vector similarity calculation unit, including calculating the vector similarities between the first feature vector and all the second feature vectors;
[0180] A partitioning unit, including classifying the image to be verified and all the sample images based on the vector similarities.
[0181] Optionally, the image evaluation score calculation unit 122 includes:
[0182] A target position information determination unit, including determining the target position information of the vehicle when the candidate image is captured through a positioning module;
[0183] A driving direction determination unit, including determining the driving direction of the vehicle according to a plurality of associated positions corresponding to the target position information; the associated positions are other position information whose acquisition time difference from the position information is less than a preset time threshold;
[0184] A shooting position weight determination unit, including determining the shooting position weight corresponding to the candidate image according to the target position information and the driving direction;
[0185] A shooting position weight weighting unit, including calculating the image evaluation score of the candidate image according to the global feature vector, the local feature vector, and the shooting position weight.
[0186] Optionally, the image set acquisition unit 121 includes:
[0187] A first acquisition unit, including a vehicle driving video obtained through a driving recorder deployed in the vehicle; the vehicle driving video includes multiple video images; each frame of the video image is a first candidate image;
[0188] A second acquisition unit, including obtaining a second candidate image through a camera device deployed inside the vehicle in response to a shooting instruction initiated by the user;
[0189] An image encapsulation unit, including generating the image set according to the first candidate image and the second candidate image.
[0190] Therefore, the image processing device provided by the embodiments of the present invention can also, after receiving the image set fed back by the image acquisition device deployed on the vehicle, respectively determine the global feature vectors and local feature vectors of each candidate image in the image set, so as to score the candidate images in the global dimension and the local dimension, and determine an image evaluation score for evaluating the shooting quality of the candidate images based on the global feature vectors and the local feature vectors, thereby realizing the quantification process of the quality of the captured images. Thus, the target images with better shooting quality can be extracted from a large number of candidate images according to the image evaluation score, and then, through the image conversion model associated with the filter style to be added, the processed images corresponding to the target images are output, achieving the purpose of batch filter addition. Compared with the existing image processing technologies, the image set in the embodiments of the present application is obtained through the image acquisition module deployed on the vehicle, which can avoid missing the shooting opportunity. Moreover, the electronic device does not require the user to manually select the appropriate target images from a large number of image sets. Instead, it can screen the images by calculating the image evaluation scores of each candidate image, greatly reducing the time consumed by the user when selecting images and reducing unnecessary operations of the user, thereby improving the efficiency of image processing. On the other hand, the electronic device can batch process the images according to the filters preset by the user, which can also avoid the user having to add filters one by one, further improving the efficiency of image processing and greatly enhancing the user experience.
[0191] It should be understood that Figure 12 in the structural block diagram of the image processing method device shown, each module is used to execute Figures 2 to 11 the corresponding steps in the corresponding embodiments, and for Figures 2 to 11 the corresponding steps in the corresponding embodiments have been explained in detail in the above embodiments. For details, please refer to Figures 2 to 11 and Figures 2 to 11 the relevant descriptions in the corresponding embodiments, which will not be elaborated here.
[0192] Figure 13 is the structural block diagram of an electronic device provided by another embodiment of the present application. As Figure 13 shown, the electronic device 1300 in this embodiment includes: a processor 1310, a memory 1320, and a computer program 1330 stored in the memory 1320 and executable on the processor 1310, such as the program of the image processing method. When the processor 1310 executes the computer program 1330, it implements the steps in each of the above embodiments of the image processing method, such as Figure 2 S201 to S204 shown. Or, when the processor 1310 executes the computer program 1330, it implements the functions of each module in the above Figure 13 corresponding embodiments. For example, Figure 12For the functions of the units 121 to 124 shown, please refer specifically to Figure 12 the relevant descriptions in the corresponding embodiments.
[0193] Exemplarily, the computer program 1330 can be divided into one or more modules. One or more modules are stored in the memory 1320 and executed by the processor 1310 to complete this application. One or more modules can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 1330 in the electronic device 1300. For example, the computer program 1330 can be divided into respective unit modules, and the specific functions of each module are as above.
[0194] The electronic device 1300 may include, but is not limited to, a processor 1310 and a memory 1320. Those skilled in the art can understand that Figure 13 this is only an example of the electronic device 1300 and does not constitute a limitation on the electronic device 1300. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may also include input and output devices, network access devices, buses, etc.
[0195] The so-called processor 1310 may be a central processing unit, or may also be other general-purpose processors, digital signal processors, application-specific integrated circuits, off-the-shelf programmable gate arrays, or other programmable logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0196] The memory 1320 may be an internal storage unit of the electronic device 1300, such as the hard disk or memory of the electronic device 1300. The memory 1320 may also be an external storage device of the electronic device 1300, such as a plug-in hard disk, a smart memory card, a flash memory card, etc. equipped on the electronic device 1300. Further, the memory 1320 may also include both the internal storage unit and the external storage device of the electronic device 1300.
[0197] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for processing an image, characterized in that, Including: Receiving an image set captured by an image acquisition module deployed on a vehicle; the image set includes multiple candidate images; Respectively determining the global feature vector and the local feature vector of each of the candidate images, and calculating an image evaluation score of the candidate image according to the global feature vector and the local feature vector; Selecting a target image to be converted from all the candidate images according to the image evaluation score; Importing the target image into a preset image conversion model, and outputting a processed image corresponding to a preset filter; The respectively determining the global feature vector and the local feature vector of each of the candidate images, and calculating an image evaluation score of the candidate image according to the global feature vector and the local feature vector includes: Converting the candidate image into a global image vector, and importing the global feature vector into a convolutional layer in a preset global feature extraction model to obtain a convolutional feature vector corresponding to the global image vector; Importing the convolutional feature vector into a fully connected layer in the global feature extraction model to obtain the global feature vector; Merging the global feature vector and the local feature vector to generate a merged feature vector corresponding to the candidate image; Importing the merged feature vectors corresponding to all the candidate images into a classification and scoring module to determine the image evaluation scores of each of the candidate images.
2. The processing method according to claim 1, characterized in that, The respectively determining the global feature vector and the local feature vector of each of the candidate images, and calculating an image evaluation score of the candidate image according to the global feature vector and the local feature vector includes: Dividing the candidate image into multiple image blocks, and generating block vectors corresponding to the image blocks; Respectively calculating the attention scores between any one block vector and each of the other block vectors; Weighting each of the other block vectors according to the attention scores to obtain a fused feature vector of the any one block vector; the fused feature vector is specifically: Among them, f x ’ ,y is the fusion feature vector of any one of the block vectors; f x,y is any one of the block vectors; is the i-th other block vector; is the attention score between the i-th other block vector and any one of the block vectors; N is the total number of the image blocks; Calculating the local feature vector corresponding to the candidate image based on all the fused feature vectors.
3. The processing method according to claim 1, characterized in that The image conversion model includes a generation network and a discriminant network; Before the importing the target image into a preset image conversion model and outputting a processed image corresponding to a preset filter, it further includes: Outputting a to-be-verified image corresponding to a training image through the generation network; Importing the to-be-verified image and multiple sample images generated based on the preset filter into the discriminant network to generate classification results corresponding to the to-be-verified image and the multiple sample images; If the category where the to-be-verified image is located is within the category corresponding to the preset filter, it is recognized that the generation network has been trained, so as to output the processed image of the target image through the generation network; If the category where the to-be-verified image is located is outside the category corresponding to the preset filter, adjusting the parameters in the generation network based on a difference threshold, and returning to execute the outputting the to-be-verified image corresponding to the training image through the generation network.
4. The processing method according to claim 3, wherein, Importing the image to be verified and multiple sample images generated based on the preset filters into the discrimination network to generate classification results corresponding to the image to be verified and the multiple sample images includes: Extracting a first feature vector corresponding to the image to be verified and respectively extracting second feature vectors corresponding to each of the sample images; Calculating the vector similarities between the first feature vector and all the second feature vectors; Classifying the image to be verified and all the sample images based on the vector similarities and the difference threshold.
5. The processing method according to any one of claims 1-4, characterized in that, Respectively determining the global feature vector and the local feature vector of each candidate image and calculating the image evaluation score of the candidate image according to the global feature vector and the local feature vector includes: Determining, by a positioning module, target position information where the vehicle is located when the candidate image is captured; Determining the driving direction of the vehicle according to multiple associated positions corresponding to the target position information; the associated positions are other position information whose acquisition time difference from the position information is less than a preset time threshold; Determining the shooting position weight corresponding to the candidate image according to the target position information and the driving direction; Calculating the image evaluation score of the candidate image according to the global feature vector, the local feature vector, and the shooting position weight.
6. The processing method according to any one of claims 1-4, characterized in that Receiving the image set captured by the image acquisition module deployed on the vehicle includes: Obtaining the vehicle driving video through a driving recorder deployed inside the vehicle; the vehicle driving video includes multiple video images; each video image is a first candidate image; In response to a shooting instruction initiated by the user, obtaining a second candidate image through a camera device deployed inside the vehicle; Generating the image set according to the first candidate image and the second candidate image.
7. An image processing device, characterized in that, Includes: An image set acquisition unit, configured to receive the image set captured by the image acquisition module deployed on the vehicle; the image set includes multiple candidate images; An image evaluation score calculation unit, configured to respectively determine the global feature vector and the local feature vector of each candidate image and calculate the image evaluation score of the candidate image according to the global feature vector and the local feature vector; A target image selection unit, configured to select a target image to be converted from all the candidate images according to the image evaluation score; A filter adding unit, configured to import the target image into a preset image conversion model and output a processed image corresponding to the preset filter; The image evaluation score calculation unit includes: A convolutional feature vector generation unit, including converting the candidate image into a global image vector and importing the global feature vector into a convolutional layer of a preset global feature extraction model to obtain a convolutional feature vector corresponding to the global image vector; A global feature vector generation unit, including importing the convolutional feature vector into a fully connected layer of the global feature extraction model to obtain the global feature vector; The combined feature vector generation unit includes combining the global feature vector and the local feature vector to generate a combined feature vector corresponding to the candidate image; The combined feature vector import unit includes importing the combined feature vectors corresponding to all the candidate images into the classification scoring module to determine the image evaluation scores of each candidate image.
8. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Image processing method and device, terminal and storage medium
CN108234870A
Image scene recognition method and device based on artificial intelligence and electronic equipment
CN112699855A