Image processing method and device, equipment and storage medium
By filtering the visual feature vectors with high significance in the set of feature vectors of the target image, the contradiction between the calculation efficiency of the image processing model and the accuracy of the result is solved, and efficient image processing is achieved.
Patent Information
- Application Number
- CN202510192292.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-06
AI Technical Summary
During image processing, directly inputting a large number of feature vectors will affect the calculation efficiency, and feature vector compression often leads to loss of visual information, affecting the accuracy of image processing results.
By obtaining the set of feature vectors of the target image, the feature significance value of each visual feature vector is determined, and the target visual feature vector that meets the preset conditions is selected for input to the image processing model.
The calculation efficiency of the image processing model is improved and the impact of feature vector compression on the accuracy of image processing results is reduced.
Smart Images

Figure CN120107709A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and in particular to an image processing method, device, equipment and storage medium. Background Art
[0002] In the process of image processing, features are usually extracted from the image to obtain a large number of feature vectors that can be used to describe the key information of the image. If a large number of feature vectors are directly input into the image processing model for calculation, it will affect the calculation efficiency of the image processing model. Therefore, before inputting the image feature vector into the model, the image feature vector needs to be compressed, that is, the number of image feature vectors is reduced.
[0003] In addition, the accuracy of the output results is also one of the important indicators to measure the performance of the image processing model. However, in related technologies, when compressing the image feature vector, it usually leads to a large amount of visual information loss, which has a great impact on the accuracy of the subsequent image processing results.
[0004] Therefore, how to improve the computational efficiency of the image processing model while reducing the impact on the accuracy of the image processing results is a technical problem that needs to be solved urgently. Summary of the invention
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an image processing method, apparatus, device and storage medium.
[0006] In a first aspect, an embodiment of the present disclosure provides an image processing method, the method comprising:
[0007] Acquire a feature vector set corresponding to a target image; wherein the feature vector set includes a visual feature vector extracted from the target image;
[0008] Determine a feature saliency value corresponding to each visual feature vector in the feature vector set; wherein the feature saliency value is used to characterize the saliency of the visual feature described by the visual feature vector in the target image;
[0009] Based on the feature saliency value, a target visual feature vector is determined from the visual feature vectors in the feature vector set; wherein the saliency of the visual feature described by the target visual feature vector in the target image meets a preset condition, and the target visual feature vector is used to be input into an image processing model to generate a processing result of the target image.
[0010] In an optional implementation, determining the feature saliency value corresponding to each visual feature vector in the feature vector set includes:
[0011] Determine the L2 norm corresponding to each visual feature vector in the feature vector set;
[0012] Accordingly, determining a target visual feature vector from visual feature vectors in the feature vector set based on the feature saliency value includes:
[0013] Based on the L2 norms corresponding to the visual feature vectors, a target visual feature vector is determined from the visual feature vectors in the feature vector set.
[0014] In an optional implementation, determining the feature saliency value corresponding to each visual feature vector in the feature vector set includes:
[0015] Determine the image information entropy corresponding to each visual feature vector in the feature vector set;
[0016] Accordingly, determining a target visual feature vector from visual feature vectors in the feature vector set based on the feature saliency value includes:
[0017] Based on the image information entropy, each visual feature vector in the feature vector set is arranged in descending order, and the first M visual feature vectors are determined as target visual feature vectors; wherein M is a preset integer.
[0018] In an optional implementation, determining the target visual feature vector from the visual feature vectors in the feature vector set based on the L2 norms corresponding to the visual feature vectors respectively includes:
[0019] Based on the L2 norms corresponding to the visual feature vectors, each visual feature vector in the feature vector set is centralized to obtain a centralized L2 feature score corresponding to each visual feature vector;
[0020] Converting the centralized L2 feature scores corresponding to the visual feature vectors into probability expressions to obtain the selection probabilities corresponding to the visual feature vectors;
[0021] A visual feature vector in the feature vector set whose selection probability is greater than a preset probability threshold is determined as a target visual feature vector.
[0022] In an optional implementation, the centralizing each visual feature vector in the feature vector set based on the L2 norm corresponding to each visual feature vector to obtain a centralized L2 feature score corresponding to each visual feature vector includes:
[0023] Determine an average value of the L2 norm corresponding to the feature vector set; wherein the average value of the L2 norm is obtained by calculating the average value of the L2 norms corresponding to each visual feature vector in the feature vector set;
[0024] The difference between the L2 norm corresponding to the visual feature vector in the feature vector set and the L2 norm average value is determined as the centralized L2 feature score corresponding to the visual feature vector.
[0025] In an optional implementation manner, after determining the visual feature vector in the feature vector set whose selection probability is greater than a preset probability threshold as the target visual feature vector, the method further includes:
[0026] The target visual feature vector is weighted by taking the probability of being selected as a weight to obtain a weighted visual feature vector corresponding to the target visual feature vector; wherein the weighted visual feature vector is used to be input into an image processing model to generate a processing result of the target image.
[0027] In an optional implementation, determining the visual feature vector in the feature vector set whose selection probability is greater than a preset probability threshold as the target visual feature vector includes:
[0028] Based on the selection probability of each visual feature vector in the feature vector set and a preset probability threshold, a mask array is created; wherein the mask array includes elements with element values of 1 or 0, and the elements with element values of 1 have a corresponding relationship with visual feature vectors whose selection probability is greater than the preset probability threshold;
[0029] The visual feature vector corresponding to the element whose element value in the mask array is 1 is determined as the target visual feature vector.
[0030] In an optional embodiment, the method is applied to a visual feature compression model, which includes a preset loss function, and the preset loss function is used to constrain the average value of the selection probabilities corresponding to the visual feature vectors in the feature vector set.
[0031] In a second aspect, the present disclosure provides an image processing device, the device comprising:
[0032] An acquisition module, used to acquire a feature vector set corresponding to a target image; wherein the feature vector set includes a visual feature vector extracted from the target image;
[0033] A first determination module is used to determine a feature saliency value corresponding to each visual feature vector in the feature vector set; wherein the feature saliency value is used to represent the saliency of the visual feature described by the visual feature vector in the target image;
[0034] The second determination module is used to determine a target visual feature vector from the visual feature vectors in the feature vector set based on the feature significance value; wherein the significance of the visual feature described by the target visual feature vector in the target image meets a preset condition, and the target visual feature vector is used to be input into an image processing model to generate a processing result of the target image.
[0035] In a third aspect, an embodiment of the present disclosure further provides an electronic device, comprising: a processor; a memory for storing executable instructions of the processor; the processor is used to read the executable instructions from the memory and execute the instructions to implement an image processing method as provided in an embodiment of the present disclosure.
[0036] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the image processing method provided by the embodiment of the present disclosure.
[0037] In a fifth aspect, the present disclosure provides a computer program product, wherein the computer program product comprises a computer program / instructions, and the computer program / instructions implement the above method when executed by a processor.
[0038] Compared with the prior art, the technical solution provided by the embodiments of the present disclosure has the following advantages:
[0039] In the image processing method provided by the embodiments of the present disclosure, a feature vector set corresponding to the target image is first obtained, and the feature vector set includes visual feature vectors extracted from the target image; then, feature saliency values corresponding to each visual feature vector in the feature vector set are determined, and the feature saliency values are used to characterize the saliency of the visual features described by the visual feature vectors in the target image; then, based on the feature saliency values, a target visual feature vector is determined from the feature vector set, so that it can be subsequently input into an image processing model to generate a processing result of the target image, and the saliency of the visual feature vector described by the target visual feature vector in the target image meets a preset condition.
[0040] It can be seen that the image processing method provided by the embodiment of the present disclosure can screen out the target visual feature vector that meets the preset conditions from the feature vector set based on the feature significance value corresponding to each visual feature vector in the feature vector set, thereby realizing the compression of the visual feature vector, which is subsequently input into the image processing model for processing, and can improve the computational efficiency of the image processing model. In addition, since the visual features described by the compressed target visual feature vector have a high degree of feature significance in the target image, while the computational efficiency of the image processing model is improved, the influence of the visual feature vector compression on the accuracy of the image processing results is greatly reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. Throughout the accompanying drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the originals and elements are not necessarily drawn to scale.
[0042] Figure 1 A flowchart of an image processing method provided by an embodiment of the present disclosure;
[0043] Figure 2 An image processing flow chart provided for an embodiment of the present disclosure;
[0044] Figure 3 A schematic diagram of the structure of an image processing device provided by an embodiment of the present disclosure;
[0045] Figure 4 A schematic diagram of the structure of an image processing device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0046] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure.
[0047] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.
[0048] The term "including" and its variations used herein are open inclusions, i.e., "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0049] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0050] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0051] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.
[0052] In the process of image processing, features are usually extracted from the image to obtain a large number of feature vectors that can be used to describe the key information of the image. If a large number of feature vectors are directly input into the image processing model for calculation, it will affect the calculation efficiency of the image processing model. Therefore, before inputting the image feature vector into the model, the image feature vector needs to be compressed, that is, the number of image feature vectors is reduced.
[0053] In addition, the accuracy of the output results is also one of the important indicators to measure the performance of the image processing model. However, in related technologies, when compressing the image feature vector, it usually leads to a large amount of visual information loss, which has a great impact on the accuracy of the subsequent image processing results.
[0054] Therefore, how to improve the computational efficiency of the image processing model while reducing the impact on the accuracy of the image processing results is a technical problem that needs to be solved urgently.
[0055] To this end, an embodiment of the present disclosure provides an image processing method. Specifically, a feature vector set corresponding to a target image is first obtained, and the feature vector set includes visual feature vectors extracted from the target image; then, feature saliency values corresponding to each visual feature vector in the feature vector set are determined, and the feature saliency values are used to characterize the saliency of the visual features described by the visual feature vectors in the target image; then, based on the feature saliency values, a target visual feature vector is determined from the feature vector set, so that it can be subsequently input into an image processing model to generate a processing result of the target image, and the saliency of the visual feature vector described by the target visual feature vector in the target image meets a preset condition.
[0056] It can be seen that the image processing method provided by the embodiment of the present disclosure can screen out the target visual feature vector that meets the preset conditions from the feature vector set based on the feature significance value corresponding to each visual feature vector in the feature vector set, thereby realizing the compression of the visual feature vector, which is subsequently input into the image processing model for processing, and can improve the computational efficiency of the image processing model. In addition, since the visual features described by the compressed target visual feature vector have a high degree of feature significance in the target image, while the computational efficiency of the image processing model is improved, the influence of the visual feature vector compression on the accuracy of the image processing results is greatly reduced.
[0057] Based on this, an embodiment of the present disclosure provides an image processing method, which is introduced below in conjunction with a specific embodiment.
[0058] Figure 1 The following is a flow chart of an image processing method provided by an embodiment of the present disclosure. The method can be executed by an image processing device, wherein the device can be implemented by software and / or hardware and can generally be integrated in an electronic device. Figure 1 As shown, the method includes:
[0059] S101: Obtain a set of feature vectors corresponding to a target image.
[0060] The feature vector set includes visual feature vectors extracted from the target image.
[0061] In the embodiments of the present disclosure, a target image refers to an image used for analysis or information extraction in a computer vision task. The target image may be any type of image, such as a landscape image, an object image, etc. The target image may also be any one of a plurality of image frames obtained by extracting frames from a video, and a corresponding feature vector set is obtained for the target image.
[0062] The feature vector set includes multiple visual feature vectors extracted from the target image. In practical applications, the feature vector set can be represented in the form of a matrix, for example, each row in the matrix can represent a visual feature vector. Visual feature vectors, also known as visual tokens, are used to describe the key information corresponding to each data block or each image area in the target image. Visual feature vectors can be used to perform image recognition or target detection on the target image.
[0063] In the disclosed embodiments, a visual encoder may be used to extract a visual feature vector from a target image. Specifically, as a part of a deep learning model, the visual encoder may automatically learn and extract key information from a target image through a self-attention mechanism as a visual feature vector of the target image, and then store the extracted visual feature vector using a feature vector set.
[0064] S102: Determine the feature saliency value corresponding to each visual feature vector in the feature vector set.
[0065] The feature saliency value is used to represent the saliency of the visual feature described by the visual feature vector in the target image.
[0066] In the disclosed embodiment, after obtaining the feature vector set corresponding to the target image, the corresponding feature saliency value is determined for each visual feature vector in the feature vector set. The feature saliency value can be used to characterize the saliency of the visual feature described by the visual feature vector in the target image. For example, assuming that a certain visual feature is located in the main part of the target image, the saliency value of the corresponding visual feature vector will be higher, indicating that the visual feature is more prominent in the target image. When the target image is subsequently processed, the visual feature vector with higher feature saliency will be assigned a higher weight for image processing tasks such as image recognition or target detection.
[0067] In addition, feature saliency can also be used to characterize the amount of information contained in the visual feature described by the visual feature vector in the target image, wherein the amount of information can be used to describe the importance of the visual feature in subsequent image processing tasks. It can be understood that the higher the feature saliency corresponding to the visual feature vector, the more information contained in the visual feature described by the visual feature vector, and the higher its importance to the image processing task.
[0068] In an optional implementation, after obtaining a feature vector set corresponding to the target image, the L2 norm corresponding to each visual feature vector in the feature vector set can be used to characterize the prominence or importance of the visual features described by each visual feature vector in the target image. Alternatively, the image information entropy corresponding to each visual feature vector in the feature vector set can be used to characterize the amount of information contained in the visual features described by each visual feature vector. The specific method will be described in detail in subsequent embodiments.
[0069] In practical applications, after extracting the visual feature vector from the target image, the high-frequency information extracted from the visual feature vector can also be used to characterize the prominence of the visual features described by the visual feature vector in the target image. It can be understood that high-frequency information (such as edges and details) is usually related to the prominence of the visual features. The more high-frequency information in the visual feature vector, the more significant the visual features described by the visual feature vector in the target image. Therefore, the prominence of the visual features can be determined by extracting the high-frequency information in the visual feature vector.
[0070] S103: Determine a target visual feature vector from the visual feature vectors in the feature vector set based on the feature saliency value.
[0071] The prominence of the visual features described by the target visual feature vector in the target image meets a preset condition, and the target visual feature vector is used to be input into an image processing model to generate a processing result of the target image.
[0072] In the embodiments of the present disclosure, since the feature saliency value can be used to characterize the saliency of the visual feature described by the visual feature vector in the target image, after determining the feature saliency values corresponding to each visual feature vector in the feature vector set, the target visual feature vector can be determined from the visual feature vectors included in the feature vector set based on the feature saliency, so that the target visual feature vector can be subsequently input into the image processing model for processing.
[0073] Among them, the target visual feature vector refers to a part of visual feature vectors screened out from the feature vector set, and the degree of prominence of the visual features described by this part of visual feature vectors in the target image meets the preset conditions. It can be understood that the target visual feature vector can be used to characterize some of the more significant or important visual features in the target image.
[0074] Specifically, based on the feature saliency, the visual feature vector in the feature vector set whose saliency in the target image meets the preset conditions can be determined as the target visual feature vector, thereby achieving the effect of compressing the visual feature vector in the feature vector set, that is, filtering out the visual feature vectors that do not meet the preset conditions. Since the visual features described by the retained target visual features have a higher saliency in the target image, it is possible to effectively reduce the computational cost while reducing the visual feature vectors, thereby improving the processing efficiency of the image processing model.
[0075] In the embodiment of the present disclosure, after determining the target visual feature vector from the visual feature vectors included in the feature vector set based on the feature saliency value, the target visual feature vector can also be input into the image processing model, and after being processed by the image processing model, the processing result of the target image is output.
[0076] Among them, image processing models may include multimodal models, image recognition models, image classification models, etc. Different image processing models can be used to perform different image processing tasks. For example, multimodal models combine the capabilities of image vision processing and natural language processing, and can be used to generate image description languages or perform image-based question and answering tasks.
[0077] The disclosed embodiments provide an image processing method. Specifically, a feature vector set corresponding to a target image is first obtained, wherein the feature vector set includes visual feature vectors extracted from the target image; then, a feature saliency value corresponding to each visual feature vector in the feature vector set is determined, wherein the feature saliency value is used to characterize the saliency of the visual feature described by the visual feature vector in the target image; then, based on the feature saliency value, a target visual feature vector is determined from the feature vector set so that it can be subsequently input into an image processing model to generate a processing result of the target image, wherein the saliency of the visual feature vector described by the target visual feature vector in the target image meets a preset condition.
[0078] It can be seen that the image processing method provided by the embodiment of the present disclosure can screen out the target visual feature vector that meets the preset conditions from the feature vector set based on the feature significance value corresponding to each visual feature vector in the feature vector set, thereby realizing the compression of the visual feature vector, which is subsequently input into the image processing model for processing, and can improve the computational efficiency of the image processing model. In addition, since the visual features described by the compressed target visual feature vector have a high degree of feature significance in the target image, while the computational efficiency of the image processing model is improved, the influence of the visual feature vector compression on the accuracy of the image processing results is greatly reduced.
[0079] In practical applications, the L2 norm is a method for measuring the length or size of a vector, also known as the Euclidean norm. For an n-dimensional vector, the corresponding L2 norm is calculated by summing the squares of each vector and then taking the square root. In computer vision tasks, the L2 norm can be used to measure the prominence of the feature vector described by the visual feature vector in the target image.
[0080] Therefore, after obtaining the feature vector set corresponding to the target image, the embodiment of the present disclosure can also use the L2 norm corresponding to each visual feature vector in the feature vector set to determine the target visual feature vector.
[0081] Specifically, first obtain a feature vector set corresponding to the target image, then determine the L2 norm corresponding to each visual feature vector in the feature vector set; then, based on the L2 norm corresponding to each visual feature vector, determine the target visual feature vector from the visual feature vectors in the feature vector set.
[0082] In the process of image feature extraction, if the L2 norm corresponding to a visual feature vector is large, it means that the visual feature or image area described by the visual feature vector is more prominent in the target image, that is, it occupies a more prominent position in the feature space of the target image, which may have a more important impact on subsequent image processing tasks, such as image classification, target detection, etc. In the embodiment of the present disclosure, the L2 norm corresponding to the visual feature vector can be calculated using formula (1):
[0083]
[0084] Among them, v represents the visual feature vector, which is a three-dimensional tensor with a dimension of b×n×d. b,i,j It represents the value of the i-th visual feature vector in the b-th sample in the j-th dimension. score[ ] represents the L2 norm corresponding to the i-th visual feature vector v in the feature vector set. It is a two-dimensional tensor with a dimension of b×n. b represents the number of samples. The b-th sample will be used to represent the target image below. n represents the number of visual feature vectors in the feature vector set. d represents the dimension of the hidden state of the visual feature vector. i represents the index of the visual feature vector, ranging from 0 to n-1. j represents the index of dimension d in the visual feature vector v, ranging from 0 to d-1.
[0085] The processing process corresponding to the above formula (1) includes: first, square and sum all dimensions of the i-th visual feature vector to obtain the sum result corresponding to the i-th visual feature vector, and then take the square root of the sum result to obtain the L2 norm corresponding to the i-th visual feature vector, so as to improve the performance and computational efficiency of the image processing model.
[0086] In an optional implementation, in order to further improve the performance and computational efficiency of the image processing model, after determining the L2 norms corresponding to each visual feature vector in the feature vector set, each visual feature vector in the feature vector set can be centralized based on the L2 norms corresponding to each visual feature vector to obtain a centralized L2 feature score corresponding to each visual feature vector.
[0087] Among them, the centering process refers to first calculating the average value of the L2 norms corresponding to all visual feature vectors in the feature vector set to obtain the L2 norm average value, and then subtracting the L2 norm average value from the L2 norm corresponding to each visual feature vector. The purpose of the centering process is to adjust the size of each visual feature vector so that it is measured relative to the center (i.e., the average value) of the feature vector set (i.e., the feature space), thereby eliminating the scale differences between the visual feature vectors, so that each visual feature vector is represented by its deviation from the average value.
[0088] In the disclosed embodiment, the centralized L2 feature score corresponding to the visual feature vector can be calculated using formula (2):
[0089]
[0090] Among them, score[b,i] represents the L2 norm corresponding to the i-th visual feature vector, which is a two-dimensional tensor with a dimension of b×n. center Represents the centralized L2 feature score corresponding to the visual feature vector v, which is a two-dimensional tensor with a dimension of b×n. k represents the index corresponding to the visual feature vectors used for summation, ranging from 0 to n-1.
[0091] The centralized L2 feature score calculated by the above formula (2) represents the relative importance or significance of the i-th visual feature vector in the b-th sample in the feature vector set. If the centralized L2 feature score is a positive value, it means that the visual feature described by the visual feature vector has a higher significance or importance in the feature vector set.
[0092] In the disclosed embodiment, the calculation process of the L2 norm corresponding to each visual feature vector may specifically include: first, determining the average L2 norm corresponding to the feature vector set; and then determining the difference between the L2 norm corresponding to the visual feature vector in the feature vector set and the average L2 norm as the centralized L2 feature score corresponding to the visual feature vector.
[0093] The L2 norm average is calculated based on the average value of the L2 norms corresponding to each visual feature vector in the feature vector set. The specific calculation formula is:
[0094]
[0095] The above calculation formula (3) indicates that the average value of the L2 norm score[b, k] corresponding to each visual feature vector in the feature vector set is calculated to obtain the L2 norm average value corresponding to the feature vector set.
[0096] After calculating the L2 norm average, calculate the difference between the L2 norm corresponding to each visual feature vector in the feature vector set and the L2 norm average, that is, subtract the L2 norm average from the L2 norm score[b, i] corresponding to the i-th visual feature vector. Get the centralized L2 feature score corresponding to the i-th visual feature vector.
[0097] In the disclosed embodiment, after centralizing each visual feature vector in the feature vector set to obtain the centralized L2 feature score corresponding to each visual feature vector, the centralized L2 feature score corresponding to each visual feature vector is converted into a probability expression to obtain the selection probability corresponding to each visual feature vector.
[0098] In practical applications, probability can be used to quantify the possibility of an event. The probability of an event refers to the numerical value of the possibility of the event occurring, and the probability value is between 0 and 1. The selected probability in the embodiment of the present disclosure is used to describe the possibility that the visual feature vector is used in the subsequent image processing task, or its importance in the subsequent image processing task. The higher the selected probability corresponding to the visual feature vector, the higher the importance of the visual feature vector in the subsequent image processing task.
[0099] In an optional implementation, a logical function is often used to map any real number into the interval (1, 2) so that it can represent a probability value. Therefore, the logical function can be used to convert the centralized L2 feature scores corresponding to each visual feature vector in the feature vector set into a probability expression to obtain the selection probability corresponding to each visual feature vector.
[0100] In the embodiment of the present disclosure, the selection probability corresponding to each visual feature vector can be calculated using formula (4):
[0101]
[0102] Among them, Score center It represents the centralized L2 feature score corresponding to the i-th visual feature vector in the b-th sample. It is a two-dimensional tensor with dimension b×n. P represents the selection probability corresponding to the i-th visual feature vector. e represents the base of the natural logarithm, which is approximately equal to 2.71828.
[0103] The processing process corresponding to the above formula (4) includes: first, obtaining the centralized L2 feature score corresponding to each visual feature vector through centralization processing, and then converting the centralized L2 feature score through a logical function to obtain the selection probability corresponding to each visual feature vector. The larger the value of the selection probability, the more significant the visual feature described by the visual feature vector in the entire target image, or the more information it contains, the more likely it is to be selected by the image processing model for subsequent processing, such as image classification, target detection and other image tasks.
[0104] In practical applications, the visual feature compression model can also be used to learn and predict the selection probability corresponding to each visual feature vector, and then characterize the amount of information contained in the target image or the prominence of the visual feature described by the visual feature vector in the target image according to the selection probability.
[0105] In the disclosed embodiment, after obtaining the selection probabilities corresponding to the respective visual feature vectors, the visual feature vectors in the feature vector set whose selection probabilities are greater than a preset probability threshold are determined as target visual feature vectors.
[0106] The preset probability threshold can be any value between 0 and 1, and can be set according to actual needs. The preset probability threshold is set to select the retained visual feature threshold according to the selection probability corresponding to each visual feature vector, that is, only the visual feature vector with a probability greater than the preset probability threshold is selected for subsequent calculation.
[0107] It can be seen that after determining the L2 norms corresponding to each visual feature vector in the feature vector set, the embodiment of the present disclosure first performs centralization processing on each visual feature vector in the feature vector set based on the L2 norms corresponding to each visual feature vector, so as to obtain the centralized L2 feature scores corresponding to each visual feature vector; then, the centralized L2 feature scores corresponding to each visual feature vector are converted into selection probabilities; then, based on the selection probabilities corresponding to each visual feature vector, the visual feature vector in the feature vector set whose selection probability is greater than the preset probability threshold is determined as the target visual feature vector, so that the target visual feature vector can be subsequently input into the image processing model for processing, thereby improving the computational efficiency of the image processing model while retaining the visual features with higher significance in the target image.
[0108] An optional implementation method, in order to further improve the processing efficiency of the image processing model for the target image, after determining the visual feature vector in the feature vector set whose selection probability is greater than a preset probability threshold as the target visual feature vector, the selection probability of the target visual feature vector can also be used as a weight to perform weighted processing on the target visual feature vector to obtain a weighted visual feature vector corresponding to the target visual feature vector.
[0109] In the disclosed embodiment, the weighted visual feature vector can be used for subsequent input into an image processing model, and after being processed by the image processing model, the processing result of the target image is output.
[0110] Since the selection probability can be used to characterize the amount of information contained in the visual features described by the visual feature vector, or to characterize the importance of the visual features described by the visual feature vector in the subsequent calculation process, the selection probability is used as the weight coefficient corresponding to the target visual feature vector, and the target visual feature vector is weighted to obtain the weighted visual feature vector. The weighted visual feature vector can also be used as the input of the image processing model to further improve the accuracy of the image processing results.
[0111] For example, the weighted visual feature vector corresponding to the target visual feature vector can be calculated using formula (5):
[0112] v p [b,i,j]=v[b,i,j]×P[b,i](5);
[0113] Where v represents the i-th target visual feature vector, P represents the probability of being selected corresponding to the i-th target visual feature vector, V P represents the weighted visual feature vector corresponding to the i-th target visual feature vector v, b represents the number of samples, n represents the number of target visual feature vectors, and d represents the dimension of the hidden state of the target visual feature vector.
[0114] In another optional implementation, the centralized L2 feature scores corresponding to each visual feature vector in the feature vector set are converted into probability expressions. After obtaining the selection probabilities corresponding to each visual feature vector, a mask array can be added to the selection probabilities corresponding to the visual feature vectors, so as to filter and obtain the target visual feature vectors for subsequent calculations through the mask array.
[0115] In the disclosed embodiment, a mask array is first created based on the selection probability of each visual feature vector in the feature vector set and a preset probability threshold; then the visual feature vector corresponding to the element with an element value of 1 in the mask array is determined as the target visual feature vector for subsequent input into the image processing model to output the processing result of the target image.
[0116] In the embodiment of the present disclosure, the mask array (Mask) includes elements with element values of 1 or 0, wherein the element with an element value of 1 corresponds to a visual feature vector whose selection probability is greater than a preset probability threshold, and the mask array can be used to filter out visual feature vectors whose selection probability values are greater than the preset probability threshold from a feature vector set.
[0117] Exemplarily, assuming that the preset probability threshold is 0.5, the visual feature vectors in the feature vector set include x1, x2, and x3, wherein the selection probability P1 corresponding to the visual feature vector x1 is 0.2, the selection probability P2 corresponding to the visual feature vector x2 is 0.6, and the selection probability P3 corresponding to the visual feature vector x3 is 0.8. According to the selection probabilities of the above-mentioned visual feature vectors and the preset probability threshold, the created mask array M is {0, 1, 1}, 0 indicates that the selection probability corresponding to the visual feature vector x1 is less than 0.5, and 1 indicates that the selection probabilities corresponding to the visual feature vectors x2 and x3 are both greater than 0.5. Then, the visual feature vectors corresponding to the elements with element values of 1 in the mask array M, i.e., the visual feature vectors x2 and x3, are determined as the target visual feature vectors.
[0118] It can be seen that the embodiment of the present disclosure performs weighted processing on the target visual feature vector by taking the selection probability as the weight coefficient corresponding to the target visual feature vector to obtain a weighted visual feature vector. Since the weighted visual feature vector can better reflect the importance of the visual feature vector in the subsequent calculation process, when the weighted visual feature vector is subsequently input into the image processing model for processing, the performance of the image processing model can be further improved.
[0119] In practical applications, a visual feature compression model can be used to compress the visual feature vector of an input target image, wherein the input of the visual feature compression model is a set of feature vectors corresponding to the target image, and the output is a target visual feature vector.
[0120] In the disclosed embodiment, in order to further improve the degree of compression of the visual feature vector of the target image, a preset loss function can also be introduced into the visual feature compression model to guide the model to constrain the average value of the selection probabilities corresponding to the visual feature vectors in the feature vector set during the training process.
[0121] Among them, the preset loss function may include a regularization loss function. The regularization loss function is a technology used to prevent model overfitting. By adding a regularization term to the loss function, the parameters of the model are constrained. For example, the average value of the selection probabilities corresponding to the visual feature vectors in the feature vector set is constrained so that the model assigns more probabilities to visual features with higher significance.
[0122] Since the embodiment of the present disclosure utilizes a preset loss function to constrain the average value of the selection probabilities corresponding to the visual feature vectors in the feature vector set during the model training process, that is, the average value of the selection probability is reduced, when the model is subsequently used to compress the visual feature vector of the target image, a target visual feature vector with a higher degree of compression can be output.
[0123] In the embodiment of the present disclosure, formula (6) can be used to reduce the average value of the selection probabilities corresponding to the visual feature vectors in the feature vector set:
[0124]
[0125] Where P represents the probability of being selected corresponding to the i-th target visual feature vector, b represents the number of samples, and n represents the number of target visual feature vectors. sparsity Represents the average value of the selection probabilities corresponding to the visual feature vectors in the feature vector set.
[0126] The above sparsity loss formula (6) is used to calculate the average value of the selection probability of all visual feature vectors in the feature vector set. The smaller the average value, the more the model tends to select only a few important visual feature vectors, which helps to improve the generalization ability and computational efficiency of the model.
[0127] It can be seen that the embodiment of the present disclosure introduces a preset loss function into the visual feature compression model to guide the model to constrain the average values of the selection probabilities corresponding to the visual feature vectors in the feature vector set during the training process, so that when the model is subsequently used to compress the visual feature vector of the target image, a target visual feature vector with a higher degree of compression can be output.
[0128] like Figure 2 The figure is a schematic diagram of an image processing flow provided by an embodiment of the present disclosure.
[0129] Firstly, a visual feature vector is extracted from the target image to obtain a feature vector set corresponding to the target image; then the feature vector set corresponding to the target image is input into a visual feature compression model.
[0130] After obtaining the feature vector set corresponding to the target image, the visual feature compression model first determines the feature saliency value corresponding to each visual feature vector in the feature vector set. The feature saliency value is used to characterize the saliency of the visual feature described by the visual feature vector in the target image. Then, based on the feature saliency value, the target visual feature vector is determined from the feature vector set.
[0131] After the target visual feature vector is obtained by screening from the feature vector set, the target visual feature vector is input into the image processing model. After being processed by the image processing model, the processing result corresponding to the target image is input. For example, after performing image recognition on the target image using the image processing model, the image recognition result corresponding to the target image is input.
[0132] In practical applications, in order to evaluate the compression effect of the visual feature vector corresponding to the target image, the target image can also be visualized based on the target visual feature vector to obtain a processed image. The more prominent area in the processed image can be used to represent the position of the visual features described by the target visual feature vector in the target image, so that the user can evaluate the compression effect of the visual feature vector based on the processed image.
[0133] In an optional implementation, the target visual feature vector may be determined from the visual feature vectors in the feature vector set based on the image information entropy corresponding to the visual feature vector.
[0134] Among them, image information entropy is used to measure the amount of information and complexity of the visual feature vector in the target image. Based on the image information entropy, the amount of information contained in the visual features described by the visual feature vector can be determined. That is, the larger the image information entropy, the more information or the richer the information contained in the visual features described by the visual feature vector, and the more important it is to the subsequent image processing tasks.
[0135] In the disclosed embodiment, a feature vector set corresponding to a target image is first obtained, and then the image information entropy corresponding to each visual feature vector in the feature vector set is determined, and then the visual feature vectors in the feature vector set are arranged in descending order based on the image information entropy, and the first M visual feature vectors are determined as target visual feature vectors, where M is a preset integer.
[0136] In the disclosed embodiment, since image information entropy can be used to characterize the amount of information contained in the visual features described by the visual feature vectors, after determining the image information entropy corresponding to each visual feature vector, a preset integer number of visual feature vectors with a frontier image information entropy can be determined as the target visual feature vectors, that is, the front preset number of visual feature vectors containing a larger amount of information can be determined as the target visual feature vectors, so as to be subsequently input into the image processing model to generate a processing result corresponding to the target image.
[0137] Exemplarily, assuming that the feature vector set includes 1000 visual feature vectors, and the preset integer M is 100, after determining the image information entropy corresponding to each visual feature vector, and sorting the visual feature vectors in the feature vector set in descending order based on the image information entropy, the first 100 vectors in the sorted visual feature vectors can be determined as the target visual feature vectors, which are used as the input of the image processing model, and the processing result corresponding to the input target image is obtained after being processed by the image processing model.
[0138] In the disclosed embodiment, after obtaining the feature vector set corresponding to the target image, the image information entropy corresponding to each visual feature vector in the feature vector set is determined, and the target visual feature vector containing more information is screened out from the feature vector set based on the image information entropy, thereby achieving compression of the visual feature vector. Since the visual features described by the compressed target visual feature vector contain more information in the target image, the computational efficiency of the image processing model is improved while the influence of the visual feature vector compression on the accuracy of the image processing results is greatly reduced.
[0139] In order to implement the above embodiments, the present disclosure also proposes an image processing device. Figure 3 The structure diagram of an image processing device provided by an embodiment of the present disclosure is shown in FIG. 1 , which can be implemented by software and / or hardware and can generally be integrated into an electronic device. Figure 3 As shown, the device comprises:
[0140] An acquisition module 301 is used to acquire a feature vector set corresponding to a target image; wherein the feature vector set includes a visual feature vector extracted from the target image;
[0141] A first determination module 302 is used to determine a feature saliency value corresponding to each visual feature vector in the feature vector set; wherein the feature saliency value is used to represent the saliency of the visual feature described by the visual feature vector in the target image;
[0142] The second determination module 303 is used to determine a target visual feature vector from the visual feature vectors in the feature vector set based on the feature significance value; wherein the significance of the visual feature described by the target visual feature vector in the target image meets a preset condition, and the target visual feature vector is used to be input into an image processing model to generate a processing result of the target image.
[0143] In an optional implementation manner, the first determining module includes:
[0144] A first determination submodule is used to determine the L2 norm corresponding to each visual feature vector in the feature vector set;
[0145] Accordingly, the second determining module includes:
[0146] The second determination submodule is used to determine a target visual feature vector from the visual feature vectors in the feature vector set based on the L2 norms corresponding to the visual feature vectors.
[0147] In an optional implementation manner, the first determining module includes:
[0148] A second determination submodule is used to determine the image information entropy corresponding to each visual feature vector in the feature vector set;
[0149] Accordingly, the second determining module includes:
[0150] The third determination submodule is used to arrange the visual feature vectors in the feature vector set in descending order based on the image information entropy, and determine the first M visual feature vectors as target visual feature vectors; wherein M is a preset integer.
[0151] In an optional implementation manner, the second determining submodule is specifically configured to:
[0152] Based on the L2 norms corresponding to the visual feature vectors, each visual feature vector in the feature vector set is centralized to obtain a centralized L2 feature score corresponding to each visual feature vector;
[0153] Converting the centralized L2 feature scores corresponding to the visual feature vectors into probability expressions to obtain the selection probabilities corresponding to the visual feature vectors;
[0154] A visual feature vector in the feature vector set whose selection probability is greater than a preset probability threshold is determined as a target visual feature vector.
[0155] In an optional implementation manner, the second determining submodule is further used to:
[0156] Determine an average value of the L2 norm corresponding to the feature vector set; wherein the average value of the L2 norm is obtained by calculating an average value based on the L2 norms corresponding to each visual feature vector in the feature vector set;
[0157] The difference between the L2 norm corresponding to the visual feature vector in the feature vector set and the L2 norm average value is determined as the centralized L2 feature score corresponding to the visual feature vector.
[0158] In an optional implementation manner, the second determining submodule is further used to:
[0159] The target visual feature vector is weighted by taking the probability of being selected as a weight to obtain a weighted visual feature vector corresponding to the target visual feature vector; wherein the weighted visual feature vector is used to be input into an image processing model to generate a processing result of the target image.
[0160] In an optional implementation manner, the second determining submodule is further used to:
[0161] Based on the selection probability of each visual feature vector in the feature vector set and a preset probability threshold, a mask array is created; wherein the mask array includes elements with element values of 1 or 0, and the elements with element values of 1 have a corresponding relationship with the visual feature vectors whose selection probability is greater than the preset probability threshold;
[0162] The visual feature vector corresponding to the element whose element value in the mask array is 1 is determined as the target visual feature vector.
[0163] In an optional embodiment, the method is applied to a visual feature compression model, which includes a preset loss function, and the preset loss function is used to constrain the average value of the selection probabilities corresponding to the visual feature vectors in the feature vector set.
[0164] In the image processing device provided by the embodiment of the present disclosure, specifically, first a feature vector set corresponding to the target image is obtained, and the feature vector set includes the visual feature vector extracted from the target image; then, the feature saliency value corresponding to each visual feature vector in the feature vector set is determined, and the feature saliency value is used to characterize the saliency of the visual feature described by the visual feature vector in the target image; then, based on the feature saliency value, the target visual feature vector is determined from the feature vector set, so that it can be subsequently input into the image processing model to generate the processing result of the target image, and the saliency of the visual feature vector described by the target visual feature vector in the target image meets the preset conditions.
[0165] It can be seen that the image processing method provided by the embodiment of the present disclosure can screen out the target visual feature vector that meets the preset conditions from the feature vector set based on the feature significance value corresponding to each visual feature vector in the feature vector set, thereby realizing the compression of the visual feature vector, which is subsequently input into the image processing model for processing, and can improve the computational efficiency of the image processing model. In addition, since the visual features described by the compressed target visual feature vector have a high degree of feature significance in the target image, while the computational efficiency of the image processing model is improved, the influence of the visual feature vector compression on the accuracy of the image processing results is greatly reduced.
[0166] The image processing device provided in the embodiments of the present disclosure can execute the image processing method provided in any embodiment of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method.
[0167] In order to implement the above embodiments, the present disclosure further proposes a computer program product, including a computer program / instruction, which implements the image processing method in the above embodiments when executed by a processor.
[0168] In addition to the above-mentioned method and apparatus, the embodiments of the present disclosure further provide a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a terminal device, the terminal device implements the image processing method described in the embodiments of the present disclosure.
[0169] In addition, the present disclosure also provides an image processing device, see Figure 4 As shown, this may include:
[0170] Processor 401, memory 402, input device 403 and output device 404. The number of processors 401 in the image processing device can be one or more. Figure 4 In some embodiments of the present disclosure, the processor 401, the memory 402, the input device 403 and the output device 404 may be connected via a bus or other means, wherein: Figure 4 The example of connecting through bus is taken in the following.
[0171] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing of the image processing device by running the software programs and modules stored in the memory 402. The memory 402 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc. In addition, the memory 402 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. The input device 403 can be used to receive input digital or character information, and generate signal input related to user settings and function control of the image processing device.
[0172] Specifically in this embodiment, the processor 401 will load the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 will run the applications stored in the memory 402, thereby realizing various functions of the above-mentioned image processing device.
[0173] It should be noted that, in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0174] The above description is only a specific embodiment of the present disclosure, so that those skilled in the art can understand or implement the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to the embodiments described herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An image processing method, characterized in that: include: Acquire a feature vector set corresponding to a target image; wherein the feature vector set includes a visual feature vector extracted from the target image; Determine a feature saliency value corresponding to each visual feature vector in the feature vector set; wherein the feature saliency value is used to characterize the saliency of the visual feature described by the visual feature vector in the target image; Based on the feature saliency value, a target visual feature vector is determined from the visual feature vectors in the feature vector set; wherein the saliency of the visual feature described by the target visual feature vector in the target image meets a preset condition, and the target visual feature vector is used to be input into an image processing model to generate a processing result of the target image.
2. The method according to claim 1, characterized in that The determining of the feature saliency value corresponding to each visual feature vector in the feature vector set includes: Determine the L2 norm corresponding to each visual feature vector in the feature vector set; Accordingly, determining a target visual feature vector from visual feature vectors in the feature vector set based on the feature saliency value includes: Based on the L2 norms corresponding to the visual feature vectors, a target visual feature vector is determined from the visual feature vectors in the feature vector set.
3. The method according to claim 1, characterized in that The determining of the feature saliency value corresponding to each visual feature vector in the feature vector set includes: Determine the image information entropy corresponding to each visual feature vector in the feature vector set; Accordingly, determining a target visual feature vector from visual feature vectors in the feature vector set based on the feature saliency value includes: Based on the image information entropy, each visual feature vector in the feature vector set is arranged in descending order, and the first M visual feature vectors are determined as target visual feature vectors; wherein M is a preset integer.
4. The method according to claim 2, characterized in that: The determining of the target visual feature vector from the visual feature vectors in the feature vector set based on the L2 norms corresponding to the visual feature vectors respectively includes: Based on the L2 norms corresponding to the visual feature vectors, each visual feature vector in the feature vector set is centralized to obtain a centralized L2 feature score corresponding to each visual feature vector; Converting the centralized L2 feature scores corresponding to the visual feature vectors into probability expressions to obtain the selection probabilities corresponding to the visual feature vectors; A visual feature vector in the feature vector set whose selection probability is greater than a preset probability threshold is determined as a target visual feature vector.
5. The method according to claim 4, characterized in that The centralizing each visual feature vector in the feature vector set based on the L2 norm corresponding to each visual feature vector to obtain a centralized L2 feature score corresponding to each visual feature vector includes: Determine an average value of the L2 norm corresponding to the feature vector set; wherein the average value of the L2 norm is obtained by calculating an average value based on the L2 norms corresponding to each visual feature vector in the feature vector set; The difference between the L2 norm corresponding to the visual feature vector in the feature vector set and the L2 norm average value is determined as the centralized L2 feature score corresponding to the visual feature vector.
6. The method according to claim 4 or 5, characterized in that: After determining the visual feature vector in the feature vector set whose selection probability is greater than a preset probability threshold as the target visual feature vector, the method further includes: The target visual feature vector is weighted by taking the probability of being selected as a weight to obtain a weighted visual feature vector corresponding to the target visual feature vector; wherein the weighted visual feature vector is used to be input into an image processing model to generate a processing result of the target image.
7. The method according to claim 4 or 5, characterized in that: The step of determining the visual feature vector in the feature vector set whose selection probability is greater than a preset probability threshold as the target visual feature vector comprises: Based on the selection probability of each visual feature vector in the feature vector set and a preset probability threshold, a mask array is created; wherein the mask array includes elements with element values of 1 or 0, and the elements with element values of 1 have a corresponding relationship with the visual feature vectors whose selection probability is greater than the preset probability threshold; The visual feature vector corresponding to the element whose element value in the mask array is 1 is determined as the target visual feature vector.
8. The method according to claim 4 or 5, characterized in that: The method is applied to a visual feature compression model, wherein the visual feature compression model includes a preset loss function, and the preset loss function is used to constrain the average value of the selection probabilities respectively corresponding to the visual feature vectors in the feature vector set.
9. An image processing device, characterized in that: The device comprises: An acquisition module, used to acquire a feature vector set corresponding to a target image; wherein the feature vector set includes a visual feature vector extracted from the target image; A first determination module is used to determine a feature saliency value corresponding to each visual feature vector in the feature vector set; wherein the feature saliency value is used to represent the saliency of the visual feature described by the visual feature vector in the target image; The second determination module is used to determine a target visual feature vector from the visual feature vectors in the feature vector set based on the feature significance value; wherein the significance of the visual feature described by the target visual feature vector in the target image meets a preset condition, and the target visual feature vector is used to be input into an image processing model to generate a processing result of the target image.
10. An electronic device, characterized in that: The electronic device comprises: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the image processing method described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to execute the image processing method described in any one of claims 1 to 8.
12. A computer program product, characterized in that The computer program product comprises a computer program / instruction, and when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 8 is implemented.