Image processing method and device, electronic equipment and storage medium

CN118608892BActive Publication Date: 2026-09-15GUANGZHOU KETENG INFORMATION TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410695732.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-31
Publication Date
2026-09-15
Estimated Expiration
2044-05-31

AI Technical Summary

Benefits of technology

[0022] The technical solution of this invention involves clustering at least two sample images in a sample image set to obtain a feature map set to be processed. This set includes a first complete set and at least one first proper subset. The first proper subset includes at least one feature map to be processed, providing data for subsequently determining a target proper subset. For each first proper subset, the image features of the current proper subset and the feature maps to be processed in the first complete set are evaluated differentially to obtain a first evaluation result. Based on at least one first evaluation result, a target proper subset corresponding to the target evaluation result is determined, achieving the goal of determining a representative first proper subset from the feature map set to be processed. The target proper subset is used as a feature map set to be used, and random noise is added to at least one feature map in the set to be used to obtain a feature map set to be determined, enriching the image features of the feature map set to be used. The difference values ​​in image features between the feature map set to be determined and the feature map set to be used and the preset feature map are determined. Based on the difference values, a target feature map with image features different from the feature map set to be used and the preset feature map is obtained. This facilitates the training of the model based on the target features corresponding to the target feature map, thereby improving the generalization ability of the image recognition model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118608892B_ABST
    Figure CN118608892B_ABST
Patent Text Reader

Abstract

An image processing method and device, electronic equipment and storage medium are disclosed. At least two sample images in a sample image set are clustered to obtain a to-be-processed feature map set. Each first proper subset and image features corresponding to to-be-processed feature maps in a first full set are differentially evaluated to obtain a first evaluation result. Based on at least one first evaluation result, a target evaluation result and a to-be-used feature map set are determined. Random noise is added to at least one to-be-used feature map in the to-be-used feature map set to obtain a to-be-determined feature map set. A difference value between each to-be-determined feature map in the to-be-determined feature map set and a preset feature map set in image features is calculated to obtain a second evaluation result based on at least one difference value. Based on the second evaluation result, a target feature map is obtained. The model is trained according to target features of the target feature map, thereby improving the generalization ability of the image recognition model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to an image processing method, apparatus, electronic device, and storage medium. Background Technology

[0002] Machine learning-based image recognition models rely on the collection of training samples. However, the collection of training samples is often affected by specific collection conditions, environment, and collection costs, so the training samples may not cover all possible image features. Therefore, image recognition models trained on such samples often exhibit low generalization ability.

[0003] Currently, to improve the generalization ability of image recognition models, training samples are often subjected to simple transformations, such as rotation, scaling, flipping, and color adjustment, to increase the image features of the training samples. However, the image features obtained by the above methods are all based on simple transformations of a limited set of image features, and their improvement on the generalization ability of image recognition models is very limited. Summary of the Invention

[0004] This invention provides an image processing method, apparatus, electronic device, and storage medium, which enriches the image features of training samples, facilitates model training based on target features of target feature maps, and thereby improves the generalization ability of image recognition models.

[0005] According to one aspect of the present invention, an image processing method is provided, the method comprising:

[0006] By clustering at least two sample images in the sample image set, a feature map set to be processed is obtained. The feature map set to be processed includes a first complete set and at least one first proper subset, and the first proper subset includes at least one feature map to be processed.

[0007] For each first proper subset, the image features corresponding to the feature maps to be processed in the current first proper subset and the first full set are evaluated differentially to obtain the first evaluation result, wherein the first evaluation result is used to characterize the difference between the first proper subset and the first full set in terms of image features;

[0008] Based on at least one first evaluation result, a target evaluation result and a target true subset corresponding to the target evaluation result are determined, and the target true subset is used as a feature map set to be used, wherein the target true subset is a true subset of at least one first true subset, and the feature map set to be used includes at least one feature map to be used;

[0009] Random noise is added to at least one feature map in the feature map set to be used to obtain a feature map set to be determined, and the difference value of each feature map to be determined in the feature map set to be determined and a preset feature map set in terms of image features is calculated, so as to obtain a second evaluation result based on at least one difference value, wherein the feature map set to be determined includes at least one feature map to be determined, and the preset feature map set includes a feature map set to be used and at least one preset feature map.

[0010] Based on the second evaluation results, the target feature map corresponding to the feature map to be determined is obtained.

[0011] According to another aspect of the present invention, an image processing apparatus is provided, the apparatus comprising:

[0012] The image clustering processing module is used to obtain a set of feature maps to be processed by performing clustering processing on at least two sample images in the sample image set. The set of feature maps to be processed includes a first complete set and at least one first proper subset, and the first proper subset includes at least one feature map to be processed.

[0013] The first evaluation result determination module is used to perform differential evaluation processing on the image features corresponding to the feature maps to be processed in the first set and the first set for each first proper subset, so as to obtain the first evaluation result. The first evaluation result is used to characterize the difference between the first proper subset and the first set in terms of image features.

[0014] The feature map set determination module is used to determine the target evaluation result and the target true subset corresponding to the target evaluation result based on at least one first evaluation result, and to use the target true subset as the feature map set to be used, wherein the target true subset is a true subset of at least one first true subset, and the feature map set to be used includes at least one feature map to be used;

[0015] A noise addition module is used to add random noise to at least one feature map in the feature map set to be used, to obtain a feature map set to be determined, and to calculate the difference value of each feature map to be determined in the feature map set to be determined and a preset feature map set in terms of image features, so as to obtain a second evaluation result based on at least one difference value, wherein the feature map set to be determined includes at least one feature map to be determined, and the preset feature map set includes a feature map set to be used and at least one preset feature map;

[0016] The target feature map determination module is used to obtain the target feature map corresponding to the feature map to be determined based on the second evaluation result.

[0017] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0018] At least one processor; and

[0019] A memory that is communicatively connected to at least one processor; wherein,

[0020] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the image processing method of any embodiment of the present invention.

[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the image processing method of any embodiment of the present invention.

[0022] The technical solution of this invention involves clustering at least two sample images in a sample image set to obtain a feature map set to be processed. This set includes a first complete set and at least one first proper subset. The first proper subset includes at least one feature map to be processed, providing data for subsequently determining a target proper subset. For each first proper subset, the image features of the current proper subset and the feature maps to be processed in the first complete set are evaluated differentially to obtain a first evaluation result. Based on at least one first evaluation result, a target proper subset corresponding to the target evaluation result is determined, achieving the goal of determining a representative first proper subset from the feature map set to be processed. The target proper subset is used as a feature map set to be used, and random noise is added to at least one feature map in the set to be used to obtain a feature map set to be determined, enriching the image features of the feature map set to be used. The difference values ​​in image features between the feature map set to be determined and the feature map set to be used and the preset feature map are determined. Based on the difference values, a target feature map with image features different from the feature map set to be used and the preset feature map is obtained. This facilitates the training of the model based on the target features corresponding to the target feature map, thereby improving the generalization ability of the image recognition model.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of an image processing method provided in an embodiment of the present invention;

[0026] Figure 2 This is a flowchart of an image processing method provided in an embodiment of the present invention;

[0027] Figure 3 This is a schematic diagram of the structure of an image processing device provided in an embodiment of the present invention;

[0028] Figure 4 This is a schematic diagram of the structure of an electronic device that implements the image processing method of the present invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] Example 1

[0032] Figure 1 This is a flowchart of an image processing method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations where target feature maps with image features different from existing image features are obtained, thereby enriching the image features of training samples. This method can be executed by an image processing device, which can be implemented in hardware and / or software, and can be configured in electronic devices such as mobile phones, computers, or servers. Figure 1 As shown, the method includes:

[0033] S110. By performing clustering processing on at least two sample images in the sample image set, a feature map set to be processed is obtained, wherein the feature map set to be processed includes a first complete set and at least one first proper subset, and the first proper subset includes at least one feature map to be processed.

[0034] The sample image set is a pre-obtained training sample set used to train the image recognition model. The sample image set contains multiple sample images. Each sample image contains corresponding image features. For example, image features can be color features, texture features, shape features, and spatial relationship features, etc. This embodiment does not limit the specific image features. Clustering can be understood as grouping sample images together based on similar image features. Optionally, the sample images can be clustered using a preset clustering algorithm. For example, the preset clustering algorithm can be at least one of K-means clustering, hierarchical clustering, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), and spectral clustering. The feature map set to be processed can be a set of feature maps obtained from the sample image set whose image features are consistent with the image features of the cluster centers. To facilitate the subsequent determination of the feature map set that accurately represents the image features in the sample image set, the feature map set to be processed can be divided into a first complete set and at least one first proper subset. The first complete set includes all feature maps to be processed, and the first proper subset includes at least one feature map to be processed. Therefore, based on the differences between the first proper subset and the first universal set, a representative set of feature maps is determined.

[0035] Specifically, in practical applications, image recognition models need to accurately identify image features within images. Therefore, it's necessary to improve the generalization ability of image recognition models so that they can recognize image features that have never been processed before. This can be achieved by determining the sample image features corresponding to at least two sample images in the acquired sample image set, and then clustering these features to obtain cluster centers. Based on the sample image set and the cluster centers, a set of feature maps to be processed that match the sample image features of the cluster centers is determined. For example, if the sample image set includes 10,000 sample images, the sample image features corresponding to these 10,000 images can be clustered to obtain 100 cluster centers. The sample image features corresponding to each cluster center can represent the features of other sample images in that cluster. Based on the sample image set and the cluster centers, feature maps that match the sample image features of the cluster centers are determined. These 100 feature maps are then used as the set of feature maps to be processed. To facilitate the subsequent determination of a feature set that accurately represents the features of all sample images in the sample image set, the feature set to be processed can be divided into a first complete set and at least one first proper subset, providing data support for the subsequent determination of a representative feature set.

[0036] Optionally, a feature map set to be processed is obtained by clustering at least two sample images in the sample image set, including: determining at least two sample image features corresponding to the sample images in the sample image set; inputting the at least two sample image features into a preset clustering algorithm to obtain at least two cluster centers corresponding to the sample image set; and obtaining a feature map set to be processed based on the sample image set and the cluster centers, wherein the image features corresponding to the feature map to be processed in the feature map set to be processed are consistent with the sample image features corresponding to the cluster centers.

[0037] Here, sample image features can be understood as the image features within a sample image. The pre-defined clustering algorithm can be an algorithm used to cluster these sample image features. Cluster centers correspond to the sample image features within the sample image. The sample image features corresponding to the cluster centers can represent the other sample image features within that cluster.

[0038] Specifically, at least two sample image features corresponding to sample images in the sample image set are obtained, and these sample image features are converted into feature vectors. A preset clustering algorithm is used to cluster these at least two feature vectors to obtain at least two cluster centers. Feature maps to be processed that match the sample image features corresponding to the cluster centers are determined from the sample image set, thus obtaining a set of feature maps to be processed based on the at least two cluster centers.

[0039] S120. For each first proper subset, perform differential evaluation processing on the image features corresponding to the feature maps to be processed in the current first proper subset and the first full set to obtain the first evaluation result.

[0040] The first evaluation result is used to characterize the difference in image features between the first proper subset and the first complete set. Optionally, if the difference in image features between the first proper subset and the first complete set is small, it indicates that the first proper subset is representative and can represent the first complete set. For example, if the first proper subset contains two image features, and these two image features can represent all image features in the first complete set, that is, if the difference between these two image features and the image features in the first complete set is small, then they can represent the image features in the first complete set.

[0041] Specifically, for each first proper subset, the image features corresponding to the current first proper subset and the first complete set can be input into a first preset evaluation function to achieve differential evaluation processing of the image features between the current first proper subset and the first complete set, obtaining the corresponding first evaluation result. Optionally, the first preset evaluation function can be a function transformed based on the maximum mean discrepancy (MMD) calculation formula. Taking the first evaluation result as an example, the greater the difference in image features between the first proper subset and the first complete set, the smaller the corresponding evaluation value. When the difference in image features between the first proper subset and the first complete set is smaller, the corresponding evaluation value is larger. Therefore, the image features of the first proper subset with the largest evaluation value can represent the image features of the first complete set.

[0042] S130. Based on at least one first evaluation result, determine the target evaluation result and the target true subset corresponding to the target evaluation result, and use the target true subset as the feature map set to be used.

[0043] The target proper subset is a proper subset of at least one first proper subset, and the set of feature maps to be used includes at least one feature map to be used.

[0044] The target evaluation result can be an evaluation result that meets the actual needs, determined by comparing at least one first evaluation result. Optionally, taking the first evaluation result as the evaluation value, the target evaluation result can be the first evaluation result with the highest evaluation value. The target proper subset is the first proper subset corresponding to the target evaluation result.

[0045] Specifically, based on actual needs, a target evaluation result and a first proper subset corresponding to the target evaluation result are selected from at least one first evaluation result, i.e., the target proper subset. The target proper subset is used as the feature map set to be used. Based on this, a feature map set that can represent the image features in the sample image set is determined.

[0046] S140. Add random noise to at least one feature map in the feature map set to be used to obtain a feature map set to be determined, and calculate the difference value of each feature map to be determined in the image features between the feature map set to be determined and the preset feature map set, so as to obtain a second evaluation result based on at least one difference value.

[0047] The set of feature maps to be determined includes at least one feature map to be determined, and the preset feature map set includes a set of feature maps to be used and at least one preset feature map. Adding random noise can be done by adding or multiplying the noise matrix with the pixel matrix corresponding to each pixel in the feature map to be used. Optionally, the random noise can be Gaussian noise. Accordingly, the set of feature maps to be used after adding random noise is used as the set of feature maps to be determined. The second evaluation result can be the result obtained by summing at least one difference value. The second evaluation result can be used to determine feature maps whose image features differ from the set of feature maps to be determined and the preset feature map. The preset feature map can be understood as a feature map corresponding to pre-obtained, existing, or known image features.

[0048] Specifically, to achieve diversity in image features, random noise can be added to the feature maps to be used. A preset mean and the calculated channel standard deviation corresponding to the feature map set to be used are input into a preset Gaussian noise generation function to obtain a random noise matrix. This random noise matrix is ​​added to the pixel matrix corresponding to the feature map to be used, resulting in the feature map to be determined. By adding random noise to at least one feature map to be used, a feature map set to be determined is obtained. A second preset evaluation function calculates the difference in image features between each feature map to be determined in the feature map set and the preset feature map set. At least one difference value is selected according to actual needs and summed to obtain the corresponding second evaluation result.

[0049] Optionally, random noise is added to at least one feature map in the feature map set to be used to obtain a feature map set to be determined, including: calculating the channel standard deviation corresponding to at least one feature map to be used; inputting the channel standard deviation and the preset mean into a preset Gaussian noise generation function to obtain a random noise matrix that conforms to a Gaussian distribution; for the feature map set to be used, adding the random noise matrix to the pixel matrix corresponding to each pixel of the current feature map to be used to obtain a feature map set to be determined.

[0050] The channel standard deviation measures the dispersion of data distribution across different channels or mappings generated by different convolutional kernels in the feature map to be used. The channel standard deviation is calculated by taking the standard deviation of all pixel values ​​in each channel of the feature map to be used. For example, if the feature map set to be used is P', then the channel standard deviation of the feature map set P' can be expressed as:

[0051] σ(P')={p'1,p'2,...,p' N}

[0052] In the above formula, σ(P') is the channel standard deviation of the feature set P' to be used, and p'1, ​​p'2, ..., p' NThe channel standard deviation is denoted by , and N is the number of feature maps to be used. The preset mean can be a value defined according to actual needs. Optionally, the preset mean can be zero. The preset Gaussian noise generation function is a function that outputs a random noise matrix conforming to a Gaussian distribution based on the input preset mean and channel standard deviation. The random noise matrix can be understood as a matrix composed of random noise vectors. The random noise vectors can be independently sampled from a Gaussian distribution determined based on the preset mean and channel standard deviation. The pixel matrix can be a combination of pixels in the feature maps to be used. By adding the random noise matrix to the pixel matrix, random noise can be added to each pixel in the feature maps to be used. The feature map to be used after adding random noise is used as the feature map to be determined. By adding random noise to at least one feature map to be used, a set of feature maps to be determined is obtained.

[0053] Specifically, the channel standard deviation corresponding to each feature map in the feature map set to be used is calculated. The channel standard deviation and a preset mean are input into a preset Gaussian noise generation function to determine the random noise matrix corresponding to the current feature map to be used. The random noise vectors in the random noise matrix follow a Gaussian distribution. The random noise matrix is ​​added to the pixel matrix corresponding to the current feature map to be used, thereby adding random noise to the current feature map. The feature map to be used after adding random noise is used as the feature map to be determined. By performing the above random noise addition process on at least one feature map to be used, the corresponding feature map set to be determined is obtained.

[0054] S150. Based on the second evaluation result, the target feature map corresponding to the feature map to be determined is obtained.

[0055] The target feature map can be a feature map selected from the feature maps to be determined based on the second evaluation result. The image features of the target feature map are different from the image features in the preset feature map set.

[0056] Specifically, after obtaining the difference values ​​between each feature map to be determined and the preset feature map set, the difference values ​​can be summed, and the summed result is used as the second evaluation result. When the summed result is the largest, the feature map to be determined corresponding to the difference value in the second evaluation result at this time is taken as the target feature map. Based on this, target feature maps with different image features from the preset feature map set are obtained, enriching the image features of the training samples and facilitating subsequent training of the image recognition model based on the target features in the target feature map.

[0057] The technical solution of this embodiment obtains a set of feature maps to be processed by clustering at least two sample images in a sample image set. This set includes a first complete set and at least one first proper subset, with each proper subset containing at least one feature map to be processed, providing data for determining the target proper subset. For each first proper subset, the image features of the current proper subset and the feature maps to be processed in the first complete set are evaluated differentially to obtain a first evaluation result. Based on at least one first evaluation result, a target proper subset corresponding to the target evaluation result is determined, achieving the goal of determining a representative first proper subset from the set of feature maps to be processed. The target proper subset is used as a set of feature maps to be used, and random noise is added to at least one feature map in the set of feature maps to be used to obtain a set of feature maps to be determined, enriching the image features of the set of feature maps to be used. The difference values ​​in image features between the set of feature maps to be determined, the set of feature maps to be used, and the preset feature maps are determined. Based on these difference values, a target feature map with image features different from the set of feature maps to be used and the preset feature maps is obtained, facilitating model training based on the target features corresponding to the target feature map, thereby improving the generalization ability of the image recognition model.

[0058] Example 2

[0059] Figure 2 This is a flowchart of an image processing method provided in Embodiment 2 of the present invention. This embodiment is a preferred embodiment among the above embodiments. For specific implementation details, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here.

[0060] like Figure 2 As shown, the method includes:

[0061] S210. By performing clustering processing on at least two sample images in the sample image set, a feature map set to be processed is obtained, wherein the feature map set to be processed includes a first complete set and at least one first proper subset, and the first proper subset includes at least one feature map to be processed.

[0062] S220. Determine the first image feature set corresponding to each first proper subset, wherein the first image feature set contains at least one image feature corresponding to the feature map to be processed in the current first proper subset.

[0063] Specifically, for at least one first proper subset, the image features corresponding to each feature map to be processed in the current proper subset are determined, and the set of image features corresponding to the current proper subset is taken as the first image feature set, so as to subsequently measure the difference between the first image feature set and the first complete set in terms of image features.

[0064] S230. For each first image feature set, calculate the maximum mean difference between the current first image feature set and the image features corresponding to the first complete set, and determine the first evaluation result based on the maximum mean difference.

[0065] The maximum mean discrepancy (MMD) value can be calculated using the formula for maximum mean discrepancy (MMD). Optionally, the MMD formula can be expressed as:

[0066]

[0067] In the above formula, S represents the image feature set corresponding to the first complete set, P represents the first image feature set corresponding to the first proper subset, and... s i and s j p represents the image features corresponding to the first complete set. i and p j Let k represent the image features in the first image feature set, and k be the kernel function. Where k(x) i ,x j )=exp(-γ||x i -x j ||), where γ is a preset constant adjusted according to actual needs. Correspondingly, the function corresponding to the first evaluation result, i.e., the first preset evaluation function, can be derived from the above formula for calculating the maximum mean difference, as follows:

[0068]

[0069] Specifically, for each first image feature set, the current first image feature set and the image feature set corresponding to the first complete set are input into the maximum mean difference calculation formula to measure the difference between the two and obtain the maximum mean difference value. According to the formula corresponding to the first evaluation result, the smaller the corresponding first evaluation result is when the maximum mean difference value is the largest. To obtain the most representative first proper subset, the corresponding first evaluation result can be determined when the maximum mean difference value is the largest. At this point, the obtained first evaluation result is the target evaluation result.

[0070] S240. Based on at least one first evaluation result, determine the target evaluation result and the target true subset corresponding to the target evaluation result, and use the target true subset as the feature map set to be used.

[0071] The target proper subset is a proper subset of at least one first proper subset, and the set of feature maps to be used includes at least one feature map to be used.

[0072] For example, in conjunction with the above example, based on at least one first evaluation result output by the first preset evaluation function, the largest first evaluation result is taken as the target evaluation result. Thus, the target proper subset P' corresponding to the target evaluation result is determined. The corresponding formula can then be expressed as follows:

[0073]

[0074] In the above formula, m p This represents the maximum number of elements that the first image feature set can contain, and is a pre-defined constant value. By maximizing the first preset evaluation function, the corresponding target evaluation result can be obtained, thereby determining the target proper subset corresponding to the target evaluation result. The target proper subset is then used as the feature map set to be used.

[0075] S250. Add random noise to at least one feature map in the feature map set to be used to obtain a feature map set to be determined, and calculate the difference value of each feature map to be determined in the image features between the feature map set to be determined and the preset feature map set, so as to obtain a second evaluation result based on at least one difference value.

[0076] The feature map set to be determined includes at least one feature map to be determined, and the preset feature map set includes the feature map set to be used and at least one preset feature map.

[0077] For example, combining the above example, the channel standard deviation σ(P') corresponding to the feature map set to be used is determined, and the preset mean is determined to be 0. Then, a random noise vector is sampled from the Gaussian distribution N(0,λ·diag(σ(S))) to add random noise to the feature map to be used based on the random noise vector. Here, λ is a preset precision parameter in the Gaussian distribution formula. At least one feature map to be used after adding random noise is taken as the feature map set D to be determined. Further, the preset feature map set can be expressed as P”=P'∪V, and the difference value of the feature map set D to be determined and the preset feature map set P” in image features can be calculated according to the second preset evaluation function. The second preset evaluation function is expressed as follows:

[0078]

[0079] In the above formula, d i For image features in D, p j Let P be the image features. k is the kernel function. Where k(x) i ,x j )=exp(-γ||x i -x j||), where γ is a preset constant adjusted according to actual needs. After obtaining at least one difference value based on the second preset evaluation function, the difference values ​​can be selected according to actual needs and summed to obtain the second evaluation result. The summing formula can be expressed as:

[0080]

[0081] In the above formula, g(x) l The selected difference value is denoted as . The maximum second evaluation result can be obtained by maximizing the above formula, thus determining the feature map to be determined corresponding to the difference value in the maximum second evaluation result. This feature map to be determined is then used as the target feature map C.

[0082] S260. Based on the second evaluation result, the target feature map corresponding to the feature map to be determined is obtained.

[0083] S270. Apply the target features in the target feature map to the model training of the image recognition model.

[0084] In this context, target features can be understood as image features in the target feature map.

[0085] Specifically, after obtaining the target feature map, the target features corresponding to the target feature map can be injected into the corresponding feature map during the model training process of the image recognition model. The image recognition model can be trained based on the feature map with injected target features, thereby improving the generalization ability of the image recognition model.

[0086] Optionally, the target features in the target feature map are applied to the model training of the image recognition model, including: performing instance normalization on the corresponding feature map in the image recognition model to obtain the normalized feature map; calculating the target standard deviation and target mean corresponding to the target features, as well as the original standard deviation and original mean corresponding to the original features in the image recognition model; determining the objective function based on the target standard deviation, target mean, original standard deviation, and original mean; and injecting the target features into the normalized feature map based on the objective function to train the image recognition model based on the target features.

[0087] Feature mapping is used to extract image features corresponding to the input image in the image recognition model. Feature mapping is related to the convolutional kernels of the image recognition model; the number of convolutional kernels used in the convolutional layers of the image recognition model determines the number of feature maps. Instance normalization can be used to convert the feature maps in each channel into a standard form. The target standard deviation and target mean can be the standard deviation and mean calculated based on the target features. The original standard deviation and original mean can be understood as the standard deviation and mean obtained based on the original features in the image recognition model. The objective function can be expressed as:

[0088]

[0089] In the above formula, a represents the target standard deviation, b represents the target mean, Z is the original feature in the image recognition model, μ(Z) is the original mean, and σ(Z) is the original standard deviation.

[0090] Specifically, the feature maps in the image recognition model are normalized to obtain normalized feature maps. The target standard deviation and target mean of the target features, as well as the original standard deviation and original mean of the original features in the image recognition model, are determined. The target standard deviation, target mean, original standard deviation, and original mean are input into the above formula to obtain the objective function. Based on the objective function, the target features are injected into the normalized feature maps. That is, the weight coefficients of the image recognition model are updated. It should be noted that the target feature injection operation can be applied to multiple convolutional blocks of the image recognition model.

[0091] Optionally, the method further includes: iteratively processing the acquisition process of the target feature map to obtain a feature map to be updated when a preset number of iterations is reached, wherein the image features of the feature map to be updated are different from the image features of the target feature map; updating the feature map set to be processed based on the feature map to be updated, so as to acquire the target feature map based on the updated feature map set to be processed.

[0092] The preset number of iterations can be a pre-set number of loops in the target feature map acquisition process. The feature map to be updated can be a feature map whose image features differ from the target feature map and all previous feature map sets.

[0093] Specifically, after obtaining the target feature map, it is added to the feature map set to be processed, and the process of obtaining the target feature map is repeated, i.e., iterative processing of the target feature map acquisition process. When the number of iterations reaches a preset iteration threshold, a feature map that differs from the image features in the current feature map set to be processed is identified as a new feature map. The feature map set to be processed is then updated based on the new feature map. For example, updates can be performed in chronological order to ensure that the feature map set to be processed always contains the latest obtained new feature maps. This is just an example, and this embodiment does not limit the update method. Thus, the target feature map is acquired based on the updated feature map set to be processed. Based on this, by continuously iterating and updating the feature map set to be processed, the diversity of image features in the feature map set to be processed can be effectively expanded, providing data support for improving model accuracy.

[0094] The technical solution of this embodiment obtains a feature map set to be processed by clustering at least two sample images in a sample image set. The feature map set to be processed includes a first complete set and at least one first proper subset, providing data for subsequent determination of the target proper subset. A first image feature set corresponding to each first proper subset is determined, and the maximum mean difference between each first image feature set and the image features corresponding to the first complete set is calculated. Based on the maximum mean difference, a first evaluation result is determined. Thus, based on at least one first evaluation result, a target proper subset corresponding to the target evaluation result is determined, achieving the purpose of determining a representative first proper subset from the feature map set to be processed. The target proper subset is used as a feature map set to be used, and random noise is added to at least one feature map in the feature map set to be used to obtain a feature map set to be determined, enriching the image features of the feature map set to be used. The difference values ​​in image features between the feature map set to be determined, the feature map set to be used, and the preset feature map are determined, so that a target feature map with image features different from the feature map set to be used and the preset feature map is obtained based on the difference values. By applying the target features corresponding to the target feature map to the model training of the image recognition model, the accuracy of the image recognition model in processing the input image can be improved without changing the semantic information of the input image, thereby improving the generalization ability of the image recognition model.

[0095] Example 3

[0096] Figure 3 This is a schematic diagram of the structure of an image processing device provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes: an image clustering processing module 310, a first evaluation result determination module 320, a feature map set to be used determination module 330, a noise addition module 340, and a target feature map determination module 350.

[0097] Image clustering processing module 310 is used to obtain a feature map set to be processed by clustering at least two sample images in the sample image set, wherein the feature map set to be processed includes a first complete set and at least one first proper subset, and the first proper subset includes at least one feature map to be processed; first evaluation result determination module 320 is used to perform differential evaluation processing on the image features corresponding to the feature maps to be processed in the first complete set for each first proper subset, and obtain a first evaluation result, wherein the first evaluation result is used to characterize the difference between the first proper subset and the first complete set in image features; feature map set to be used determination module 330 is used to determine the target evaluation result and the target proper subset corresponding to the target evaluation result based on at least one first evaluation result, and to... The target proper subset is used as the feature map set to be used, wherein the target proper subset is a proper subset of at least one first proper subset, and the feature map set to be used includes at least one feature map to be used; the noise addition module 340 is used to add random noise to at least one feature map to be used in the feature map set to be used to obtain the feature map set to be determined, and calculate the difference value of each feature map to be determined in the feature map set to be determined and the preset feature map set in terms of image features, so as to obtain a second evaluation result based on at least one difference value, wherein the feature map set to be determined includes at least one feature map to be determined, and the preset feature map set includes the feature map set to be used and at least one preset feature map; the target feature map determination module 350 is used to obtain the target feature map corresponding to the feature map to be determined based on the second evaluation result.

[0098] The technical solution of this embodiment obtains a set of feature maps to be processed by clustering at least two sample images in a sample image set. This set includes a first complete set and at least one first proper subset, with each proper subset containing at least one feature map to be processed, providing data for determining the target proper subset. For each first proper subset, the image features of the current proper subset and the feature maps to be processed in the first complete set are evaluated differentially to obtain a first evaluation result. Based on at least one first evaluation result, a target proper subset corresponding to the target evaluation result is determined, achieving the goal of determining a representative first proper subset from the set of feature maps to be processed. The target proper subset is used as a set of feature maps to be used, and random noise is added to at least one feature map in the set of feature maps to be used to obtain a set of feature maps to be determined, enriching the image features of the set of feature maps to be used. The difference values ​​in image features between the set of feature maps to be determined, the set of feature maps to be used, and the preset feature maps are determined. Based on these difference values, a target feature map with image features different from the set of feature maps to be used and the preset feature maps is obtained, facilitating model training based on the target features corresponding to the target feature map, thereby improving the generalization ability of the image recognition model.

[0099] Based on the above embodiments, optionally, the image clustering processing module includes: a sample image feature determination unit, used to determine at least two sample image features corresponding to sample images in the sample image set; a cluster center determination unit, used to input the at least two sample image features into a preset clustering algorithm to obtain at least two cluster centers corresponding to the sample image set; and a feature map set determination unit, used to obtain a feature map set to be processed based on the sample image set and the cluster centers, wherein the image features corresponding to the feature maps to be processed in the feature map set to be processed are consistent with the sample image features corresponding to the cluster centers.

[0100] Optionally, the first evaluation result determination module includes: a first image feature set determination unit, configured to determine a first image feature set corresponding to each first proper subset, wherein the first image feature set contains at least one image feature corresponding to the feature map to be processed in the current first proper subset; and a first evaluation result determination unit, configured to calculate, for each first image feature set, the maximum mean difference between the current first image feature set and the image features corresponding to the first full set, and determine the first evaluation result based on the maximum mean difference.

[0101] Optionally, the noise addition module includes: a standard deviation determination unit, used to calculate the channel standard deviation corresponding to at least one feature map to be used; a noise matrix generation unit, used to input the channel standard deviation and a preset mean into a preset Gaussian noise generation function to obtain a random noise matrix that conforms to a Gaussian distribution; and a feature map set determination unit, used to add the random noise matrix to the pixel matrix corresponding to each pixel of the current feature map to be used for the feature map set to be used, to obtain the feature map set to be determined.

[0102] Optionally, the device further includes a target feature application module for applying target features from the target feature map to the model training of the image recognition model.

[0103] Optionally, the target feature application module includes: a feature mapping normalization processing unit, used to perform instance normalization processing on the corresponding feature mapping in the image recognition model to obtain the normalized feature mapping; a mean and standard deviation calculation unit, used to calculate the target standard deviation and target mean corresponding to the target feature, as well as the original standard deviation and original mean corresponding to the original feature in the image recognition model; an objective function determination unit, used to determine the objective function based on the target standard deviation, target mean, original standard deviation, and original mean; and a feature injection unit, used to inject the target feature into the normalized feature mapping based on the objective function, so as to train the image recognition model based on the target feature.

[0104] Optionally, the device further includes: a feature map set update module, which is used to iteratively process the acquisition process of the target feature map so as to obtain the feature map to be updated when a preset number of iterations is reached, wherein the image features of the feature map to be updated are different from the image features of the target feature map; and to update the feature map set to be processed based on the feature map to be updated so as to acquire the target feature map based on the updated feature map set to be processed.

[0105] The image processing apparatus provided in the embodiments of the present invention can execute the image processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0106] Example 4

[0107] Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0108] like Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0109] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0110] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as image processing methods.

[0111] In some embodiments, the image processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the image processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the image processing method by any other suitable means (e.g., by means of firmware).

[0112] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0113] Computer programs for implementing the image processing methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0114] Example 5

[0115] Embodiment 5 of the present invention also provides a computer-readable storage medium storing computer instructions for causing a processor to execute an image processing method, the method comprising:

[0116] Clustering is performed on at least two sample images in the sample image set to obtain a feature map set to be processed. The feature map set includes a first complete set and at least one first proper subset, with each first proper subset containing at least one feature map to be processed. For each first proper subset, the image features corresponding to the feature maps to be processed in the current first proper subset and the first complete set are subjected to differential evaluation processing to obtain a first evaluation result. This first evaluation result characterizes the difference in image features between the first proper subset and the first complete set. Based on at least one first evaluation result, a target evaluation result and a target proper subset corresponding to the target evaluation result are determined, and the target proper subset is used as the set to be used. A feature map set is provided, wherein the target proper subset is a proper subset of at least one first proper subset, and the feature map set to be used includes at least one feature map to be used; random noise is added to at least one feature map to be used in the feature map set to be used to obtain a feature map set to be determined, and the difference value of each feature map to be determined in the feature map set to be determined and a preset feature map set in terms of image features is calculated, so as to obtain a second evaluation result based on at least one difference value, wherein the feature map set to be determined includes at least one feature map to be determined, and the preset feature map set includes a feature map set to be used and at least one preset feature map; based on the second evaluation result, a target feature map corresponding to the feature map to be determined is obtained.

[0117] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0118] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0119] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0120] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0121] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0122] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. An image processing method, characterized in that, include: By clustering at least two sample images in the sample image set, a feature map set to be processed is obtained, wherein the feature map set to be processed includes a first complete set and at least one first proper subset, and the first proper subset includes at least one feature map to be processed. For each of the first proper subsets, the image features corresponding to the feature maps to be processed in the current first proper subset and the first full set are subjected to differential evaluation processing to obtain a first evaluation result, wherein the first evaluation result is used to characterize the difference between the first proper subset and the first full set in terms of image features. Based on at least one of the first evaluation results, a target evaluation result and a target true subset corresponding to the target evaluation result are determined, and the target true subset is used as a feature map set to be used, wherein the target true subset is a true subset of at least one of the first true subsets, and the feature map set to be used includes at least one feature map to be used; Random noise is added to at least one feature map in the feature map set to be used to obtain a feature map set to be determined, and the difference value of each feature map to be determined in the feature map set to be determined and a preset feature map set in terms of image features is calculated, so as to obtain a second evaluation result based on at least one of the difference values, wherein the feature map set to be determined includes at least one feature map to be determined, and the preset feature map set includes the feature map set to be used and at least one preset feature map; Based on the second evaluation result, a target feature map corresponding to the feature map to be determined is obtained; The step of performing differential evaluation processing on the image features corresponding to the current first proper subset and the feature maps to be processed in the first full set to obtain a first evaluation result includes: A first image feature set is determined for each first true subset, wherein the first image feature set contains at least one image feature corresponding to the feature map to be processed in the current first true subset; for each first image feature set, the maximum mean difference between the current first image feature set and the image features corresponding to the first full set is calculated, and a first evaluation result is determined based on the maximum mean difference. The method further includes: applying the target features in the target feature map to the model training of the image recognition model; The step of applying the target features in the target feature map to the model training of the image recognition model includes: The feature maps corresponding to the image recognition model are normalized to obtain normalized feature maps; the target standard deviation and target mean of the target feature, as well as the original standard deviation and original mean of the original feature in the image recognition model, are calculated; an objective function is determined based on the target standard deviation, the target mean, the original standard deviation, and the original mean; based on the objective function, the target feature is injected into the normalized feature map to train the image recognition model based on the target feature.

2. The method according to claim 1, characterized in that, The process of clustering at least two sample images in the sample image set to obtain a feature map set to be processed includes: Determine at least two sample image features corresponding to the sample images in the sample image set; At least two of the sample image features are input into a preset clustering algorithm to obtain at least two cluster centers corresponding to the sample image set; Based on the sample image set and the cluster centers, the feature map set to be processed is obtained, wherein the image features corresponding to the feature maps to be processed in the feature map set are consistent with the sample image features corresponding to the cluster centers.

3. The method according to claim 1, characterized in that, The step of adding random noise to at least one feature map in the set of feature maps to be used to obtain a set of feature maps to be determined includes: Calculate the channel standard deviation corresponding to at least one of the feature maps to be used; The channel standard deviation and the preset mean are input into a preset Gaussian noise generation function to obtain a random noise matrix that conforms to a Gaussian distribution; For the feature map set to be used, the random noise matrix is ​​added to the pixel matrix corresponding to each pixel of the current feature map to be used, to obtain the feature map set to be determined.

4. The method according to claim 1, characterized in that, Also includes: The process of acquiring the target feature map is iteratively processed so that when a preset number of iterations is reached, a feature map to be updated is obtained, wherein the image features of the feature map to be updated are different from the image features of the target feature map; The feature map set to be processed is updated based on the feature map to be updated, so as to obtain the target feature map based on the updated feature map set.

5. An image processing apparatus, characterized in that, include: An image clustering processing module is used to obtain a set of feature maps to be processed by performing clustering processing on at least two sample images in a sample image set, wherein the set of feature maps to be processed includes a first complete set and at least one first proper subset, and the first proper subset includes at least one feature map to be processed. The first evaluation result determination module is used to perform differential evaluation processing on the image features corresponding to the feature maps to be processed in the first set for each first proper subset, and obtain a first evaluation result, wherein the first evaluation result is used to characterize the difference between the first proper subset and the first set in terms of image features. The feature map set determination module is used to determine a target evaluation result and a target true subset corresponding to the target evaluation result based on at least one first evaluation result, and to use the target true subset as a feature map set to be used, wherein the target true subset is a true subset of at least one first true subset, and the feature map set to be used includes at least one feature map to be used; A noise addition module is used to add random noise to at least one feature map in the feature map set to be used, to obtain a feature map set to be determined, and to calculate the difference value of each feature map to be determined in the feature map set to be determined and a preset feature map set in terms of image features, so as to obtain a second evaluation result based on at least one of the difference values, wherein the feature map set to be determined includes at least one feature map to be determined, and the preset feature map set includes the feature map set to be used and at least one preset feature map; The target feature map determination module is used to obtain a target feature map corresponding to the feature map to be determined based on the second evaluation result; The first evaluation result determination module includes: a first image feature set determination unit, configured to determine a first image feature set corresponding to each first proper subset, wherein the first image feature set contains at least one image feature corresponding to the feature map to be processed in the current first proper subset; and a first evaluation result determination unit, configured to calculate, for each first image feature set, the maximum mean difference between the current first image feature set and the image features corresponding to the first complete set, and determine a first evaluation result based on the maximum mean difference. The device further includes: a target feature application module, used to apply the target features in the target feature map to the model training of the image recognition model; The target feature application module includes: a feature mapping normalization processing unit, used to perform instance normalization processing on the corresponding feature mapping in the image recognition model to obtain a normalized feature mapping; a mean and standard deviation calculation unit, used to calculate the target standard deviation and target mean corresponding to the target feature, and the original standard deviation and original mean corresponding to the original feature in the image recognition model; a target function determination unit, used to determine a target function based on the target standard deviation, the target mean, the original standard deviation, and the original mean; and a feature injection unit, used to inject the target feature into the normalized feature mapping based on the target function, so as to train the image recognition model based on the target feature.

6. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the image processing method according to any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the image processing method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Classification model training method, clustering method and electronic equipment

    CN113918714A

  • Semi-supervised image classification method and semi-supervised image classification system

    CN116894985A