Multi-scale feature information processing and codebook generation method and device, equipment and medium
Through multi-scale feature information processing and codebook generation methods, the problem of local feature matching in image retrieval is solved, stable matching and efficient calculation are achieved, and fast retrieval of large-scale image data is suitable.
Patent Information
- Application Number
- CN202510342499.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is sensitive to scale changes in image retrieval, resulting in unstable matching effect. The calculation cost is too high when improving matching effect through multi-codebook schemes, making it difficult to take into account both accuracy and efficiency.
By extracting the multi-scale feature information of the target object, pooling process to generate feature maps of different scales, segmenting and generating feature subsets, performing aggregation processing and quantizing the aggregated feature set to generate a multi-scale feature codebook.
The stable matching of the target objects at different scales is achieved, the high calculation cost and long processing time of the multi-codebook scheme are reduced, the matching efficiency is improved, and it is suitable for rapid retrieval and feature comparison scenarios of large-scale image data.
Smart Images

Figure CN120196783A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment and storage medium for processing multi-scale feature information and generating a codebook. Background Art
[0002] In the field of image retrieval, in order to better accommodate the overall texture and local detail information of picture content, a combination of global features and local features is usually adopted for matching. Global features are mainly used for rough screening of targets during the first retrieval, while local features are used for fine comparison to improve the accuracy of recognition. During the local feature matching process, a codebook constructed based on local features is usually introduced to complete the recognition of target objects through the calculation of similarity between feature blocks. However, existing local feature matching methods perform poorly when faced with scale changes of target objects, resulting in difficulty in matching images of the same target at different scales, thus seriously affecting the retrieval effect. If a multi-codebook scheme is adopted to solve the scale problem, the computational amount and processing time will be significantly increased. Therefore, there are problems of low efficiency and low accuracy in local feature matching in the prior art for image retrieval.
[0003] In the field of medical and health, image retrieval technology is widely used in scenarios such as comparison of pathological images, identification of lesion areas, and duplication checking of medical data. Due to different magnification factors, lighting conditions, and shooting angles in the shooting of pathological section images, the scale representation forms of lesion areas change, and it is difficult for the prior art to accurately match these changing lesion areas through a local feature codebook with a fixed scale. For example, in the automatic identification of cancer lesions, it is difficult to match lesion images at different magnification factors, resulting in unstable recognition results. At the same time, in order to solve this problem, some technical solutions attempt to construct multiple codebooks to cover different scale features, but the introduction of multiple codebooks increases the computational cost and it is difficult to meet the real-time retrieval requirements of large-scale pathological image data. Therefore, existing lesion recognition methods are insufficient when faced with image scale changes, and their application effects in the field of medical imaging are limited.
[0004] In the financial field, image retrieval technology is applied to scenarios such as identity verification, fraud detection, and duplicate checking of loan documents. Taking biometrics as an example, identity verification systems usually perform matching by extracting local features of face images. However, due to the influence of shooting environment, angle, and expression changes, the scale representation of users' facial features will deviate, resulting in the failure of local feature matching and affecting the accuracy of verification results. Existing technologies often solve this problem through the matching of local feature codebooks. However, due to the sensitivity of the codebooks to scale changes, images taken at different distances and angles are difficult to match, especially in face recognition applications across devices and scenarios. To solve this problem, some technical solutions introduce a multi-codebook strategy to cover different facial feature scales, but the use of multi-codebooks increases the computational cost and is difficult to meet the requirements of real-time verification in the financial field. Therefore, the matching accuracy and efficiency of existing face recognition technologies in dealing with scale changes still need to be improved.
[0005] Generally speaking, the application of image retrieval technology in the fields of medical health and finance both face common problems. Existing technologies lack robustness to the scale changes of target objects during the local feature matching process, resulting in difficulty in matching image features of different scales; although the multi-codebook strategy can improve the matching effect, it also greatly increases the computational cost and is difficult to meet the rapid retrieval needs of large-scale data. Therefore, how to reduce the computational cost while improving the accuracy of image retrieval, especially to achieve scale invariance of local feature matching, has become a key challenge in the application of image retrieval technology in different fields. Summary of the Invention
[0006] The main purpose of the present invention is to provide a multi-scale feature information processing and codebook generation method, device, equipment, and storage medium, aiming to solve the technical problems that existing technologies are sensitive to scale changes during local feature matching, the matching effect is unstable, and the computational cost is too high when improving the matching effect through the multi-codebook scheme, making it difficult to balance accuracy and efficiency.
[0007] To achieve the above purpose, the present invention provides a multi-scale feature information processing and codebook generation method, including:
[0008] Obtain an image containing a target object;
[0009] Extract multi-scale feature information of the target object from the image;
[0010] Perform pooling processing on the multi-scale feature information according to a preset plurality of scales to generate pooling feature maps of different scales respectively;
[0011] Perform a segmentation operation on each pooling feature map to generate at least one feature subset;
[0012] Aggregate each feature subset to merge the scattered feature data in each feature subset into the corresponding aggregated feature set;
[0013] Perform quantization processing on the aggregated feature set to generate a multi-scale feature codebook.
[0014] Furthermore, to achieve the above object, the present invention provides a multi-scale feature information processing and codebook generation device, including:
[0015] An image acquisition module, configured to acquire an image containing a target object;
[0016] A multi-scale feature extraction module, configured to extract multi-scale feature information of the target object from the image;
[0017] A multi-scale pooling processing module, configured to perform pooling processing on the multi-scale feature information according to a plurality of preset scales to respectively generate pooling feature maps of different scales;
[0018] A feature segmentation module, configured to perform a segmentation operation on each pooling feature map to generate at least one feature subset;
[0019] A feature aggregation module, configured to aggregate each feature subset to merge the scattered feature data in each feature subset into the corresponding aggregated feature set;
[0020] A feature quantization and codebook generation module, configured to perform quantization processing on the aggregated feature set to generate a multi-scale feature codebook.
[0021] Furthermore, to achieve the above object, the present invention further provides a computer device, the computer device includes a memory, a processor, and a multi-scale feature information processing and codebook generation program stored on the memory and executable on the processor, and when the multi-scale feature information processing and codebook generation program is executed by the processor, it implements the steps of the multi-scale feature information processing and codebook generation method as described above.
[0022] Furthermore, to achieve the above object, the present invention further provides a computer-readable storage medium, on which a multi-scale feature information processing and codebook generation program is stored, and when the multi-scale feature information processing and codebook generation program is executed by a processor, it implements the steps of the multi-scale feature information processing and codebook generation method as described above.
[0023] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical and health. It discloses a method for processing multi-scale feature information and generating a codebook, including: obtaining an image containing a target object, extracting the multi-scale feature information of the target object, performing pooling processing on the multi-scale feature information according to a preset plurality of scales to generate pooling feature maps of different scales, performing a segmentation operation on each pooling feature map to generate at least one feature subset, performing an aggregation process on each feature subset to merge the scattered feature data into an aggregated feature set, and performing quantization processing on the aggregated feature set to generate a multi-scale feature codebook. Through the extraction of multi-scale feature information and unified quantization processing, the present invention realizes stable matching of the target object at different scales, solves the problem that local feature matching in the prior art is sensitive to scale changes; at the same time, by performing quantization processing on the aggregated feature set, the feature data is mapped to discrete codebook index values, avoiding the high computational cost and long processing time of the multi-codebook scheme; while ensuring the matching accuracy, it significantly reduces the computational cost and improves the matching efficiency, and is applicable to the scenarios of fast retrieval and feature comparison of large-scale image data. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0025] Figure 1 is a schematic diagram of an application environment of the method for processing multi-scale feature information and generating a codebook in an embodiment of the present invention;
[0026] Figure 2 is a schematic flowchart of an embodiment of the method for processing multi-scale feature information and generating a codebook of the present invention;
[0027] Figure 3 is a schematic diagram of functional modules of a preferred embodiment of the device for processing multi-scale feature information and generating a codebook of the present invention;
[0028] Figure 4 is a schematic diagram of a structure of a computer device in an embodiment of the present invention;
[0029] Figure 5 is another schematic diagram of a structure of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0031] The method for processing multi-scale feature information and generating a codebook provided by the embodiments of the present invention can be applied in, for example Figure 1In the application environment, the client communicates with the server through a network. The server can obtain an image containing a target object through the client, extract multi-scale feature information of the target object, perform pooling processing on the multi-scale feature information according to a plurality of preset scales, generate pooling feature maps of different scales, perform a segmentation operation on each pooling feature map to generate at least one feature subset, perform an aggregation process on each feature subset, merge the scattered feature data into an aggregated feature set, and perform quantization processing on the aggregated feature set to generate a multi-scale feature codebook. Through the extraction of multi-scale feature information and unified quantization processing, the present invention realizes stable matching of the target object at different scales, and solves the problem that local feature matching in the prior art is sensitive to scale changes. At the same time, by performing quantization processing on the aggregated feature set, the feature data is mapped to discrete codebook index values, avoiding the high computational cost and long processing time of the multi-codebook scheme. While ensuring the matching accuracy, the computational cost is significantly reduced, and the matching efficiency is improved, which is applicable to the scenarios of rapid retrieval and feature comparison of large-scale image data. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0032] Please refer to Figure 2 , Figure 2 FIG. is a schematic flowchart of an embodiment of a method for processing multi-scale feature information and generating a codebook provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0033] As Figure 2 shown, the method for processing multi-scale feature information and generating a codebook proposed by the present invention includes the following steps:
[0034] S10, obtaining an image containing a target object;
[0035] In this embodiment, an image of the target scene is obtained through a high-precision image acquisition device to ensure that the target object is clearly recorded in the image. The acquisition device may include a static image acquisition device (such as an industrial camera) or a dynamic image acquisition device (such as a smart image sensor). During the acquisition process, the resolution, shooting angle, lighting conditions, etc. of the device should be adjusted according to the target scene to meet the subsequent processing requirements.
[0036] Adjust the parameters of the image acquisition device, such as resolution, aperture, shutter speed, etc., to obtain clear and moderately contrasted images. Trigger the image acquisition device to collect an image containing the target object, and save the image data to the storage unit. Save the acquired image data in a standard format for easy parsing and loading in subsequent processing steps.
[0037] The acquired image data is input into the processing system and undergoes standard preprocessing operations. The standardization process includes resizing the image (e.g., uniformly adjusting to a fixed resolution), noise removal, and color space conversion (e.g., grayscale processing or standard color gamut adjustment) to ensure that the image data meets the input requirements of subsequent processing modules.
[0038] Load the image data into the processing unit and perform resizing operations, such as uniformly adjusting to a resolution suitable for subsequent processing. Apply a noise reduction algorithm to smooth the image data and remove the interference of environmental noise. Complete the standardization conversion of the color space to make the image data have a consistent color channel distribution.
[0039] Use image processing methods or intelligent target recognition techniques to detect the target object area in the image and generate corresponding target bounding box information. The target detection technology can reason based on specific rules or target features obtained from a trained model to ensure high accuracy of the target area.
[0040] Load the target detection model or specific target area rules in the processing unit. According to the preset rules or feature information, identify the target object area in the image, generate a bounding box, and label the coordinate positions of the target area. Extract the target area image to facilitate subsequent processing steps to focus on the feature analysis of the target object.
[0041] To further optimize the recognition effect of the target object, use a background separation method to segment the target area image and remove background interference. The extracted target area after segmentation can be further cropped to center it and adjust the boundaries to optimize the subsequent feature extraction and matching accuracy.
[0042] Use a segmentation method to distinguish the target area from the background area, such as segmenting based on specific rules or the characteristics of the target object. Crop and adjust the segmented target area image to ensure that the target object is located at the center of the image and is displayed within a fixed boundary range. Generate the adjusted target area image and output the standardized image data for feature extraction.
[0043] Example: In the field of medical and health, such as in the scenario of lesion area detection and recognition, in order to extract multi-scale features of the lesion area in a pathological section image, first collect the original image of the pathological section through a microscope device or a digital scanner. The acquired image may contain a large amount of background information and non-lesion areas, so the section image needs to be preprocessed.
[0044] In the collected pathological section images, target detection technology is used to locate potential lesion areas and generate bounding boxes for candidate regions. In the candidate regions, image segmentation technology is further applied to separate the lesion areas from the normal tissue background, ensuring the integrity of the lesion areas. At the same time, size adjustment operations are performed on the extracted lesion area images to make them meet the standardized resolution requirements, and through denoising and color channel standardization methods, the image quality and color consistency are improved. Finally, standardized lesion area images are obtained, providing high-quality input data for subsequent multi-scale feature extraction and analysis.
[0045] This process effectively reduces background interference, ensures that the target lesion areas can be extracted and processed with high precision, and at the same time avoids feature extraction errors caused by inconsistent image quality, improving the accuracy of lesion area recognition and comparison.
[0046] Through high-precision image acquisition devices and systematic image preprocessing techniques, it is ensured that the images of the target objects can be accurately acquired and extracted, and the quality of the input images is optimized through size adjustment, background separation, and boundary cropping. Through standardization and alignment operations of the target regions, the accuracy and matching performance of subsequent multi-scale feature extraction are effectively improved, avoiding feature analysis errors caused by background interference or poor input image quality.
[0047] S20, extract multi-scale feature information of the target object from the image;
[0048] In this embodiment, a feature extraction model is used to extract features from the preprocessed target image, initially obtaining a feature representation containing global and local information. The feature extraction model usually adopts a deep network structure and represents the image as a multi-dimensional feature vector. For example, when extracting deep features of an image, the feature dimension can be 128 dimensions, 256 dimensions, or 512 dimensions, which is selected according to the task complexity.
[0049] Load a pre-trained feature extraction model, such as a deep network structure suitable for texture feature extraction. Adjust the target image to a fixed size (e.g., 640×640 pixels) to ensure the standardization of the input data. Input the feature extraction model and extract feature representations layer by layer to generate a preliminary feature set containing multi-dimensional information, such as a feature map of **128×128×128 (H×W×C)**.
[0050] Design a multi-scale feature extraction strategy. By adjusting the network structure or adopting a multi-branch design, extract feature information of different scales from the same image. Multi-scale feature extraction can capture the hierarchical relationship between local features and global features in the image.
[0051] Construct multiple feature branches, and set different pooling kernel sizes and strides for each branch. For example, select pooling kernel sizes of 1×1, 2×2, and 3×3, with corresponding strides of 1, 2, and 3, to extract high-resolution, medium-resolution, and low-resolution feature information respectively.
[0052] Input the input image features into each branch respectively to extract the feature maps corresponding to the scales of each branch. For example, one branch may output a feature map with a size of 128×128×64 (high resolution), and another branch may output 32×32×64 (low resolution). Combine the feature maps output by each branch to form a multi-scale feature information set.
[0053] The multi-scale features may contain duplicate information or noise, so more refined feature representations need to be extracted through feature optimization and fusion operations.
[0054] Perform weighted summation on the feature maps of each branch according to the weights. For example, set the weight of the high-resolution feature to 0.5, the medium-resolution feature to 0.3, and the low-resolution feature to 0.2. Remove the noise information in the feature vector, for example, by setting a feature intensity threshold (such as feature values below 0.01 are excluded). Output the final optimized multi-scale feature map, for example, a unified feature map with a dimension of 128×128×64.
[0055] Example illustration: In the analysis of pathological section images, the lesion areas collected by microscopic equipment may have different magnifications and resolutions. To extract the multi-scale feature information of the lesion areas, first adjust the lesion image to 640×640 pixels and input it into the feature extraction module. Set three feature branches to extract global features (such as the shape contour of the lesion), medium-scale features (such as tissue texture distribution), and local features (such as cell structure details) respectively. Finally, fuse the features of each branch to generate a multi-dimensional feature map containing lesion features at different scales for subsequent comparison and classification.
[0056] Similarly, in the scenario of user authentication, the user's facial images captured by the camera may have scale differences in feature representation due to different shooting distances and angles. To address this issue, input the facial image into the multi-branch feature extraction module, set the high-resolution branch to extract the overall facial contour, the medium-resolution branch to extract key facial feature points (such as eyes and nose), and the low-resolution branch to capture local detail textures (such as skin texture). Through multi-scale feature fusion, generate a unified multi-scale facial feature map for identity comparison.
[0057] By constructing a multi-scale feature extraction mechanism, the global features and local detail features of the target object are comprehensively captured, enhancing the robustness to target objects of different scales. Through a branch structure or hierarchical feature extraction, the feature expression ability can be balanced between high resolution and low resolution. The optimized multi-scale feature information not only has the ability to describe the details of high-resolution features but also can reflect the overall pattern of low-resolution features, improving the feature matching performance of the target object under scale changes.
[0058] S30, perform pooling processing on the multi-scale feature information according to a plurality of preset scales, and generate pooling feature maps of different scales respectively;
[0059] In this embodiment, the preset parameters of multiple scales include the pooling kernel size and the pooling stride, and different parameters are used to extract the feature information of the target object at different resolutions. For example, the following parameters can be preset:
[0060] Pooling kernel size: 1×1, 2×2, 3×3, 5×5, etc.;
[0061] Pooling stride: 1, 2, 3, etc.
[0062] These parameters determine the feature resolutions of different scales, such that each scale corresponds to a specific feature map.
[0063] Set the parameter ranges of the pooling kernel and the stride in the algorithm initialization stage. For example, the pooling kernel sizes are 1×1, 2×2, 3×3, and the corresponding strides are 1 or 2. The parameters of each scale correspond to an independent feature extraction branch, and the specific feature resolution is determined by the parameters. For example, the 1×1 pooling kernel retains the detail information, and the 3×3 pooling kernel obtains the global information.
[0064] Perform dimensionality reduction processing on the input multi-scale feature information through a pooling operation to extract the feature information of the target object at different resolutions. The pooling can be performed in the following ways:
[0065] Average pooling: Calculate the average value of the feature values within the pooling area, which is suitable for smoothing and detail preservation;
[0066] Max pooling: Take the maximum value within the pooling area, which is suitable for extracting the significant points in the features.
[0067] Gradually apply the pooling operations with different pooling kernels to the input feature information. For example, the 1×1 pooling kernel generates a high-resolution feature map, and the 3×3 pooling kernel generates a low-resolution feature map. Use average pooling to retain the overall feature trend within the area, or use max pooling to extract the significant feature values. Generate pooling feature maps of multiple resolutions. For example, the size of the high-resolution feature map is 128×128×64, the medium-resolution feature map is 64×64×64, and the low-resolution feature map is 32×32×64.
[0068] After subjecting the multi-scale feature information to pooling operations at different scales, a series of pooled feature maps with specific resolutions are generated. These feature maps respectively represent the feature manifestation forms of the target object at different scales.
[0069] High-resolution feature map: Generated by a smaller pooling kernel (such as 1×1), it retains the detailed features of the target object, and the image resolution is relatively high, such as 128×128×64.
[0070] Medium-resolution feature map: Generated by a medium-sized pooling kernel (such as 2×2), it moderately retains the overall texture and local features of the target object, and the image resolution is, for example, 64×64×64.
[0071] Low-resolution feature map: Generated by a larger pooling kernel (such as 3×3), it captures the global information of the target object, and the image resolution is, for example, 32×32×64.
[0072] These pooled feature maps at different scales are stored in a specific data structure for subsequent feature segmentation and aggregation.
[0073] Example illustration: In pathological slice analysis, in order to extract the multi-scale feature information of the lesion area, first input the lesion area image into the pooling module. Set a 1×1 pooling kernel to extract high-resolution features (such as the clear contour of cell edges), a 2×2 pooling kernel to extract medium-resolution features (such as the overall pattern of cell distribution), and a 3×3 pooling kernel to extract low-resolution features (such as the overall shape of the lesion). Finally, feature maps with different resolutions are generated for subsequent lesion area segmentation and comparison.
[0074] Similarly, in the scenario of identity verification, input the standardized face image into the pooling module, and set pooling kernels of 1×1, 3×3, and 5×5 to extract high-resolution features (such as local facial texture), medium-resolution features (such as the layout of facial features), and low-resolution features (such as the overall facial contour) respectively. Through the fusion of pooled feature maps at different scales, multi-scale feature information is generated for subsequent identity matching operations.
[0075] By performing pooling operations on the multi-scale feature information according to a preset multiple scales to generate pooled feature maps at different scales, the feature manifestation forms of the target object globally and locally can be effectively retained. By adjusting the pooling kernel size and stride, feature maps with different resolutions are generated, enabling subsequent feature segmentation and aggregation processing to combine multi-scale features and enhance the feature matching robustness of the target object under different scale changes. At the same time, through the selection of various pooling methods, the expression ability of the feature maps can be optimized according to the task requirements.
[0076] S40. Perform a segmentation operation on each pooled feature map to generate at least one feature subset;
[0077] In this embodiment, in order to perform a segmentation operation on the pooled feature map, it is necessary to determine segmentation parameters according to a preset segmentation rule. The segmentation parameters include the size of the segmentation region, the grid division strategy, etc. For example, a grid size of 4×4, 8×8, or 16×16 can be selected, and the specific grid division strategy can be fixed-size division or adaptive division.
[0078] Preset the size of the segmentation region. For example, divide the pooled feature map into 16 feature regions according to a 4×4 grid. If an adaptive division strategy is adopted, the size of the segmentation region is dynamically adjusted according to the resolution of the feature map. Divide each pooled feature map according to the preset parameters to generate multiple small feature regions.
[0079] Divide the pooled feature map into multiple feature regions according to the segmentation parameters. The goal of grid division is to divide the feature map into equally sized region blocks, and each region block contains specific local feature information. This segmentation operation can retain the local features of the target object and facilitate subsequent feature aggregation.
[0080] Evenly divide the length and width of the feature map according to the size of the segmentation region. For example, divide a 128×128 feature map into 16×16 feature regions, and the size of each region is 8×8. If the size of the feature map cannot be evenly divided, a padding operation can be performed in the boundary region to make the size of the feature map match the segmentation parameters. The generated feature regions are saved in a specific data structure for subsequent processing.
[0081] For each divided feature region, extract a set of local feature vectors. Each set of local feature vectors contains all the feature data within the feature region and is used to represent the local feature pattern of the region.
[0082] Traverse each feature region, extract the feature values therein, and generate a set of feature vectors. Save the set of feature vectors of each feature region in a specific data structure, and assign a unique region identifier to each set. Use the extracted set of feature vectors as a feature subset for subsequent aggregation operations.
[0083] Convert the divided feature regions into feature subsets. Each feature subset corresponds to a set of feature vectors of a local feature region. According to the resolution of the feature map and the segmentation parameters, multiple feature subsets are finally generated, covering the local feature information of the target object.
[0084] Calculate the total number of segmented regions of the feature map. For example, divide a 128×128 feature map into an 8×8 grid, generating 256 feature regions, corresponding to 256 feature subsets. Store the set of feature vectors for each feature region as an independent feature subset. Output the segmented feature subsets to provide input data for subsequent feature aggregation and quantization processing.
[0085] Example illustration: In pathological slice analysis, the feature information of the lesion area may be distributed in different local regions. By dividing the pooled feature map of the lesion area into an 8×8 grid, extract the set of local feature vectors in each grid. For example, a 128×128×64 lesion feature map is divided into 256 feature subsets, and each feature subset corresponds to the lesion features of a local region. These feature subsets can be used for subsequent lesion recognition and comparison to improve the accuracy of lesion area recognition.
[0086] In the authentication scenario, perform a segmentation operation on the pooled feature map of the user's facial image to extract the feature information of the local facial regions. For example, divide the facial feature map into 16 feature subsets according to a 4×4 grid, and each feature subset corresponds to a set of local feature vectors of a facial region (such as eyes, nose, mouth, etc.). These feature subsets are used to compare the local features of the user's face to improve the robustness and accuracy of authentication.
[0087] By performing a segmentation operation on each pooled feature map and generating feature subsets, the local feature information of the target object can be effectively retained, and the feature matching accuracy of the target object can be improved. The choice of grid division strategy can be adjusted according to task requirements to adapt to target images with different resolutions and complexities. The generated feature subsets cover the global and local features of the target object, providing multi-level feature representations for subsequent feature aggregation and quantization processing.
[0088] S50, perform an aggregation process on each feature subset to merge the scattered feature data in each feature subset into the corresponding aggregated feature set;
[0089] In this embodiment, the scattered feature data in each feature subset is converted into a unified feature vector representation. The process of feature vectorization includes extracting and integrating the feature values of the data points within each feature region to convert them into a high-dimensional feature vector. For example, for a feature subset, a feature vector with a length of 128 dimensions or 256 dimensions is output.
[0090] Traverse the feature data in each feature subset, and integrate the feature values into a unified vector representation. Use normalization or standardization processing to limit the value of each dimension of the vector within a specific range (such as [0,1]) to eliminate the dimensional difference of the feature values. Output the set of feature vectors of each feature subset as the input for subsequent aggregation processing.
[0091] Extract the set of clustering centers from the pre-trained clustering model. The set of clustering centers is the center points of fixed feature patterns obtained through the training process and is used to represent the feature distributions of different classes. The dimension of each clustering center is consistent with the feature vector, usually 128 dimensions or 256 dimensions.
[0092] Load the pre-trained clustering model and extract the global set of clustering centers. The number of clustering centers in the set can be set according to the task requirements, such as 100, 500, or 1000 clustering centers. Each clustering center corresponds to a feature pattern and is used for subsequent feature matching and merging operations.
[0093] Match the feature vectors in each feature subset with the clustering centers in the set of clustering centers respectively, and calculate the distance from each feature vector to each clustering center. Commonly used distance calculation methods include Euclidean distance, cosine similarity, etc.
[0094] Traverse each feature vector, calculate the distance from this vector to each clustering center. Select the clustering center with the minimum distance as the target clustering center for this feature vector. Record the matching relationship between the feature vector and the target clustering center, providing the basic data for subsequent merging operations.
[0095] According to the principle of minimum distance, merge each feature vector into the target clustering center with the closest distance. The merging process aggregates the scattered feature vectors to the corresponding clustering centers according to similarity, forming a preliminary aggregated feature set.
[0096] For each target clustering center, collect all the feature vectors merged into this center. Calculate the central value (such as the average value or weighted average value) of the merged aggregated feature set as the new clustering center. Save the preliminary aggregated feature set in a standard data format for subsequent processing.
[0097] Perform feature optimization processing on the preliminary aggregated feature set, including operations such as removing redundant features, eliminating noise features, and merging similar features, to generate the final aggregated feature set.
[0098] Remove the redundant features in the aggregated feature set, such as merging duplicate feature vectors. Use a noise reduction algorithm to eliminate the abnormal feature values in the aggregated feature set. Merge similar feature vectors to further compress the dimension of the aggregated feature set and improve the compactness of feature representation.
[0099] Example illustration: In pathological image analysis, the characteristic information of the lesion area usually contains a large amount of scattered feature data. By performing aggregation processing on the feature subset of the lesion area, similar lesion features can be effectively merged. For example, the shape of the diseased cells represented by the feature vector is matched with the pre-trained clustering centers and merged into the corresponding lesion pattern centers, thus generating an aggregated lesion feature set. These aggregated feature sets can be used for the classification and comparison of lesions, improving the accuracy of diagnosis.
[0100] In the authentication scenario, the local feature information of the user's face may have different shooting angles and lighting variations. By performing aggregation processing on the face feature subset, the scattered face features of the user are merged into the clustering centers, such as feature patterns of eyes, nose, mouth, etc. The finally generated aggregated feature set is used for identity matching and verification, enhancing the robustness and precision of identity recognition.
[0101] By performing aggregation processing on each feature subset, the scattered feature data is merged into an aggregated feature set, which can effectively enhance the robustness and compactness of the features and reduce the interference of noise data. Using the pre-trained set of clustering centers for matching and merging operations makes the result of feature aggregation more representative and improves the efficiency of subsequent feature quantization and retrieval processes. At the same time, the optimized aggregated feature set has a higher feature expression ability and can more accurately reflect the diverse features of the target object.
[0102] S60, perform quantization processing on the aggregated feature set to generate a multi-scale feature codebook.
[0103] In this embodiment, the quantization processing needs to use the reference set of clustering centers as a basis to map each feature vector in the aggregated feature set to the nearest clustering center to generate discrete codebook index values.
[0104] Load the pre-trained clustering model and extract the set of clustering centers for quantization. Each clustering center represents a discretized feature pattern, and the feature dimension can be, for example, 128-dimensional, 256-dimensional or higher. According to the precision requirements of feature matching, set different numbers of clustering centers, such as 1024, 2048 or 4096 clustering centers.
[0105] For each feature vector in the aggregated feature set, calculate the distances from all clustering centers in the set of clustering centers respectively to determine the optimal quantization mapping for each feature vector.
[0106] Use Euclidean distance or cosine similarity to calculate the distance between the feature vector and the clustering center. Traverse each feature vector and calculate its distance from all clustering centers. Select the clustering center with the minimum distance as the target clustering center of the feature vector.
[0107] According to the distance matching results, a unique codebook index value is assigned to each feature vector. The codebook index value is a discrete identifier used to represent the cluster center to which the feature vector is merged.
[0108] For each feature vector, according to the principle of minimum distance, a codebook index value is assigned. The codebook index value is usually represented in binary or hexadecimal format, such as "0001", "0x1F", etc. The codebook index value of each feature vector is stored in the feature mapping table for subsequent encoding and storage.
[0109] The generated codebook index values are discretely encoded to convert the continuous feature vectors into discrete feature symbols, thereby achieving a compressed representation of the feature data.
[0110] Discretely encode the codebook index values, for example, using Huffman coding or arithmetic coding. The discretized codebook index values are stored in a fixed-length or variable-length format. The generated set of discrete feature symbols is the basic data of the multi-scale feature codebook.
[0111] Organize and store the set of discretized feature symbols to generate the final multi-scale feature codebook. The multi-scale feature codebook contains discretized feature information at different scales and can be used for the retrieval and matching of target objects.
[0112] Store the discrete feature symbols at different scales hierarchically or by feature type. According to the feature distribution of the target object, classify and manage the codebook by category or label. Output the generated multi-scale feature codebook and store it in the database for subsequent retrieval and matching operations.
[0113] Example illustration: In pathological image analysis, the lesion area usually requires multi-scale feature extraction and comparison to achieve accurate identification. Due to the differences in lesion morphology and size, traditional fixed-scale feature codebooks perform poorly when dealing with pathological sections at different magnifications. Therefore, the generation of a multi-scale feature codebook can effectively solve this problem.
[0114] In the specific implementation process, first input the pathological section image into the feature extraction module to extract the lesion feature information at different magnifications. The extracted feature information forms an aggregated feature set of the lesion after aggregation processing, such as cell edge contours, nuclear distribution patterns, tissue density, etc. Then, perform quantization processing on the aggregated feature set.
[0115] During the quantization process, each feature vector of the lesion is matched with a pre-trained set of clustering centers. For example, a feature vector may represent the shape features of a cell, and by comparing it with the clustering centers, the optimal clustering center for this feature vector is determined - such as "round cell pattern" or "oval cell pattern". Each feature vector is assigned a discretized codebook index value (e.g., "0001" represents the round cell pattern), generating a preliminary lesion feature codebook.
[0116] Finally, the discretized feature symbols are organized into a multi-scale lesion feature codebook. For example, for a lesion area containing different cell types, the generated codebook may contain the following:
[0117] Lesion feature index: "0001" represents round cells, "0010" represents oval cells, "0100" represents irregular cells;
[0118] Tissue density index: "1001" represents low-density tissue, "1010" represents medium-density tissue, "1100" represents high-density tissue.
[0119] The multi-scale lesion feature codebook can be stored in a database for subsequent retrieval and comparison of lesion areas. In practical applications, through the matching of the multi-scale feature codebook, the type and degree of lesions in the lesion area can be quickly identified. For example, when comparing newly collected pathological sections, the system can quickly determine which known lesion type the current lesion belongs to through codebook matching, thus assisting doctors in diagnosis.
[0120] Similarly, in the identity verification system of financial institutions, in order to identify the user's identity, the feature information of the user's facial image is usually extracted. However, due to different shooting angles, lighting conditions, and distances of the user's facial image, the scale of facial features may change. Therefore, using the traditional single-scale codebook matching method cannot handle these changing situations, resulting in verification failures or incorrect comparisons.
[0121] To improve the accuracy of identity verification, this problem can be solved by generating a multi-scale facial feature codebook.
[0122] In the specific implementation process, first, global features (such as face shape contour) and local features (such as details of eyes, nose, and mouth) are extracted from the user's standardized facial image. The extracted facial feature information forms an aggregated feature set after aggregation processing, such as eyebrow arc, eye spacing, nose width, etc. Subsequently, quantization processing is performed on the aggregated feature set.
[0123] During the quantization process, each facial feature vector is matched with a set of pre-trained clustering centers. For example, a feature vector may represent the shape of a user's nose. By comparing it with the clustering centers, the optimal clustering center for this feature vector is determined - such as the "high-bridge nose pattern" or the "flat-bridge nose pattern". Each feature vector is assigned a discretized codebook index value (e.g., "1001" represents the high-bridge nose pattern).
[0124] Finally, the discretized feature symbols are organized into a multi-scale facial feature codebook. For example, for different parts of a user's facial features, the generated codebook may contain the following:
[0125] Nose feature index: "1001" represents a high-bridge nose, and "1010" represents a flat-bridge nose;
[0126] Eye feature index: "0001" represents single eyelids, and "0010" represents double eyelids;
[0127] Mouth feature index: "0101" represents thin lips, and "0110" represents thick lips.
[0128] In practical applications, the bank's identity verification system can quickly determine whether the user's identity matches by comparing the multi-scale facial feature codebook. For example, when a user uses a mobile phone to take a selfie for identity verification, the system can ignore the feature changes caused by angle and distance variations and identify the user's main facial features through codebook matching, thus completing the identity verification.
[0129] By performing quantization processing on the aggregated feature set and generating a multi-scale feature codebook, it is possible to effectively convert continuous feature data into discretized feature symbols, thereby significantly reducing the resource consumption of storage and computing. The quantization processing ensures a compact representation of the feature data, improving the efficiency and accuracy of feature matching. The finally generated multi-scale feature codebook can be compatible with the feature of the target object with different scale changes, greatly enhancing the robustness of target retrieval and matching.
[0130] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as medical and health. It discloses a multi-scale feature information processing and codebook generation method, including: obtaining an image containing a target object, extracting multi-scale feature information of the target object, performing pooling processing on the multi-scale feature information according to a preset plurality of scales to generate pooling feature maps of different scales, performing a segmentation operation on each pooling feature map to generate at least one feature subset, performing an aggregation process on each feature subset to merge scattered feature data into an aggregated feature set, and performing quantization processing on the aggregated feature set to generate a multi-scale feature codebook. Through the extraction of multi-scale feature information and unified quantization processing, the present invention realizes stable matching of the target object at different scales, solves the problem that local feature matching in the prior art is sensitive to scale changes; at the same time, maps the feature data to discrete codebook index values, avoiding the high computational cost and long processing time of the multi-codebook scheme.
[0131] In one embodiment, before the above S20, it further includes:
[0132] S201, performing target detection processing on the image based on a pre-trained target detection model to obtain candidate regions of the target object;
[0133] S202, applying a target segmentation method in the candidate regions to separate the target object from the background region and generate a target object region image;
[0134] S203, performing cropping and boundary adjustment on the target object region image to obtain a target object image;
[0135] S204, performing size adjustment processing on the target object image;
[0136] S205, performing color channel normalization processing on the size-adjusted target object image to obtain a preprocessing image for extracting multi-scale feature information.
[0137] In this embodiment, target detection is to identify the position of the target object in the image through a pre-trained target detection model and generate candidate regions. The target detection model can be trained based on different feature categories, such as identifying breeding targets such as pigs and cows, or identifying human organs and lesion regions in medical images. The candidate regions usually represent the position and boundary of the target object in the form of a rectangular box.
[0138] Select a suitable target detection model, and the model type can be selected according to specific applications, such as a convolutional neural network (CNN) or a region proposal network (RPN); input the image, and the model outputs the rectangular box coordinates containing the position of the target object; for each rectangular box, record the candidate region coordinates and confidence score of the target object.
[0139] Object segmentation is to perform fine segmentation on the target objects in the candidate regions, distinguishing the boundaries of the target objects from the background regions. The segmentation can be binary segmentation (target and background) or multi-class segmentation (segmentation of different parts of the target). The goal of this step is to remove background interference and only retain the image part of the target objects.
[0140] Use image segmentation algorithms such as Graph Cut or Conditional Random Field (CRF); perform pixel-level classification on the target candidate regions, label the target pixels as foreground and the background pixels as background; output the regional image of the target objects, which only contains the pixel information of the target objects.
[0141] Cropping and boundary adjustment are used to standardize the size of the image region of the target objects for subsequent feature extraction operations. The cropping process can remove the redundant blank regions, and the boundary adjustment can ensure that the target objects are in the center of the image, preventing image offset from affecting the feature extraction effect.
[0142] Determine the bounding box of the target objects, crop the image to the size of the bounding box of the target objects; perform boundary adjustment on the target object region to make the target objects centered in the image; output the cropped image of the target objects, ensuring that the target objects are in the standardized image region.
[0143] Size adjustment is used to convert the target object image into a standard input size to fit the subsequent feature extraction module. Common standard sizes are 224×224, 640×640, etc. The size adjustment can be achieved by interpolation methods, including nearest neighbor interpolation, bilinear interpolation, and cubic spline interpolation.
[0144] Select a standard input size, such as 640×640 pixels; perform scaling or stretching operations on the target object image to adjust the image to the standard size; ensure that the size adjustment process does not affect the proportion and detail features of the target objects.
[0145] Color channel normalization is used to normalize the color values of the target object image to reduce the impact of illumination and color differences on feature extraction. The normalization process usually includes normalizing the mean and standard deviation of the RGB channels.
[0146] Calculate the mean and standard deviation of each color channel of the target object image; subtract the mean from each pixel value and then divide by the standard deviation to achieve color channel normalization; output the normalized target object image as the input image of the feature extraction module.
[0147] In this embodiment, through a series of preprocessing operations such as object detection, segmentation, cropping, size adjustment, and color normalization, standardized image data of the target object can be effectively extracted. The standardized image data has a unified input size and color characteristics, which can significantly improve the accuracy of subsequent feature extraction and comparison. Especially in the fields of healthcare and finance, it helps to reduce the inconsistencies caused by image resolution, shooting conditions, and lighting changes, and improve the robustness of retrieval and recognition.
[0148] In one embodiment, the above S30 includes:
[0149] S301, establishing multiple scale branches, each scale branch corresponding to a different pooling kernel size or sampling stride;
[0150] S302, inputting the multi-scale feature information into each scale branch respectively;
[0151] S303, performing pooling processing on the multi-scale feature information based on the pooling kernel size or sampling stride of each scale branch, and respectively generating pooling feature maps corresponding to different scales.
[0152] In this embodiment, the multiple scale branches are used to extract the feature information of the target object at different resolutions. The pooling kernel size and sampling stride of each scale branch are different, so that feature representations of different scales can be extracted.
[0153] Pooling kernel size: represents the size of the pooling window, such as 1×1, 2×2, 3×3, etc.
[0154] Sampling stride: represents the stride of the pooling window moving on the feature map, such as 1, 2, 3, etc.
[0155] Preset the pooling kernel sizes and sampling strides of multiple scales, such as 1×1, 2×2, and 3×3 pooling kernels. Establish corresponding scale branches, and each scale branch has independent pooling parameters. Each scale branch is responsible for extracting the feature information of the target object at a specific resolution.
[0156] The multi-scale feature information refers to the feature map extracted from the target object image, which is input into different scale branches for pooling processing. The multi-scale feature information is respectively input into each scale branch to ensure that each branch independently extracts features under different pooling parameters.
[0157] Initialize the parameters of each pooling branch, including the pooling kernel size and sampling stride. Input the multi-scale feature information into each branch in turn, and perform pooling processing according to their respective pooling parameters. Ensure that the input data is independently processed in each branch to avoid feature interference between different branches.
[0158] Pooling processing is an operation for dimensionality reduction and feature compression of the feature map, aiming to extract feature representations at different scales. Pooling processing can adopt the methods of max pooling or average pooling to generate pooling feature maps at different scales.
[0159] Max pooling: Take the maximum value of the feature values within each pooling window to retain the significant feature points in the feature map.
[0160] Average pooling: Take the average value of the feature values within each pooling window to smooth the feature changes in the feature map.
[0161] After performing pooling processing, pooling feature maps at different scales are output, such as high-resolution feature maps, medium-resolution feature maps, and low-resolution feature maps. Ensure that the pooling feature maps have consistent feature representations at different scales, facilitating subsequent feature segmentation and aggregation processing.
[0162] In this embodiment, by performing pooling processing on multi-scale feature information, pooling feature maps at different resolutions can be generated, providing multi-level feature information for subsequent feature extraction, segmentation, and aggregation. The design of different scale branches enables the system to be compatible with the global and local features of the target object, thereby improving the robustness and accuracy of feature matching.
[0163] In one embodiment, the above S40 includes:
[0164] S401, determining the segmentation parameters for each pooling feature map, where the segmentation parameters include a preset grid division strategy;
[0165] S402, based on the segmentation parameters, dividing each pooling feature map into multiple pooling feature regions according to the preset grid division strategy;
[0166] S403, performing a feature extraction operation on each pooling feature region to extract a subset feature vector set in each pooling feature region, and taking the subset feature vector set in each feature region as a feature subset.
[0167] In this embodiment, the determination of the segmentation parameters is the basis for the segmentation operation of each pooling feature map. The segmentation parameters usually include the grid size, the boundary strategy of the segmentation region, etc. The preset grid division strategy refers to dividing the pooling feature map into several grid regions of fixed size or adaptive size.
[0168] Grid size selection: Preset the size of the divided region of the grid. For example, divide a 128×128 pooling feature map into 8×8 or 16×16 grids.
[0169] Segmentation strategy setting: Select a fixed grid division or adaptive grid division strategy according to the task requirements.
[0170] Boundary processing: For feature maps that cannot be evenly divided, boundary padding or cropping is used to ensure that all feature regions are of the same size.
[0171] The splitting operation divides the pooled feature map into multiple feature regions according to the splitting parameters. Each feature region represents a local region of the image and contains a certain number of feature values.
[0172] According to the preset grid size, the length and width of the pooled feature map are divided into several feature regions respectively. Ensure that the number of pixels in each feature region is the same to avoid uneven feature region sizes. For the feature map after boundary padding, it is divided into feature regions in a neat grid division manner, and each feature region is marked as an independent feature region block.
[0173] The feature extraction operation is a process of extracting a set of local feature vectors from each pooled feature region. The subset of feature vectors in each feature region contains all the feature point values in that region and is used to represent the feature information of that region.
[0174] Traverse all the pixel points of each pooled feature region, extract the feature value of each pixel point as a feature vector. Combine all the feature vectors in a feature region together to form the subset of feature vectors of that region. Generate a unique subset identifier for each feature region, and store the feature vector set in association with the corresponding feature subset. The output feature subset contains the detailed feature representations of each local feature region in the pooled feature map, providing input data for subsequent aggregation and quantization operations.
[0175] In this embodiment, by performing a splitting operation on the pooled feature map, the global feature information of the image can be split into multiple local feature regions. The feature subset of each feature region can effectively retain the local feature detail information of the target object, providing a more refined feature representation for subsequent feature aggregation and quantization processing. This splitting method can improve the robustness of feature matching, reduce recognition errors caused by scale changes or image noise, and enhance the expression ability and recognition effect of image features.
[0176] In one embodiment, the above S50 includes:
[0177] S501, perform vectorization processing on the scattered feature data in each feature subset to generate a corresponding set of feature vectors;
[0178] S502, extract a set of global clustering centers from a pre-trained clustering model, and the set of global clustering centers contains multiple clustering centers;
[0179] S503, match each feature vector in each set of feature vectors with all the clustering centers in the set of clustering centers to determine the distance from each feature vector to each clustering center;
[0180] S504. Merge each feature vector into the closest clustering center according to the minimum distance principle to generate a preliminary aggregated feature set.
[0181] S505. Perform feature optimization processing on the preliminary aggregated feature set to generate a final aggregated feature set.
[0182] In this embodiment, the vectorization process is to convert the scattered feature data in the feature subset into high-dimensional feature vectors. Each feature vector represents the numerical features of a feature region, and the dimension of the vector usually depends on the number of features extracted. Through vectorization processing, unstructured feature data can be converted into a vector form that is convenient for calculation and comparison.
[0183] Traverse each feature subset, extract the pixel values or feature values in each subset; combine these feature values into a high-dimensional vector, such as a 128-dimensional or 256-dimensional feature vector; perform normalization or standardization processing to convert the feature values into a unified numerical range and reduce the influence of dimensional differences.
[0184] The set of clustering centers is a set of central points of feature patterns generated by a clustering algorithm (such as K-means). The global set of clustering centers represents the feature distribution of the entire data set, and each clustering center corresponds to a common feature pattern. The purpose of extracting the global set of clustering centers is to provide a reference benchmark for the matching of feature vectors.
[0185] Load a pre-trained clustering model and extract a set containing multiple clustering centers; each clustering center is a fixed high-dimensional vector representing the central position of a certain feature pattern; the number of clustering centers can be set according to the actual application scenario, such as 1024 clustering centers or 2048 clustering centers.
[0186] The matching process is to calculate the distance between the feature vector and the clustering center to determine the clustering center closest to the feature vector. Commonly used distance metrics include Euclidean distance and cosine similarity. The smaller the distance, the higher the matching degree between the feature vector and the clustering center.
[0187] Traverse each feature vector and calculate the distance from this vector to each clustering center; use the Euclidean distance formula to calculate the distance:
[0188]
[0189] where x represents the feature vector, c represents the clustering center, and n represents the dimension of the feature vector.
[0190] Record the distances from each feature vector to each clustering center to provide basic data for subsequent merging operations.
[0191] The merging process is to allocate the feature vectors to the corresponding cluster centers according to the principle of minimum distance, thereby generating a preliminary aggregated feature set. The preliminary aggregated feature set represents the aggregation result of the scattered features in the feature subset.
[0192] For each feature vector, select the cluster center with the closest distance to it as the merging target; gather all the feature vectors merged into the same cluster center together to generate a preliminary aggregated feature set; record the number of feature vectors and the eigenvalue distribution of each aggregated feature set.
[0193] Feature optimization processing is a process of further optimizing the preliminary aggregated feature set, including operations such as removing redundant features, eliminating noise features, and merging similar features. The optimized aggregated feature set is more representative and compact.
[0194] Perform a deduplication operation on each aggregated feature set to remove duplicate feature vectors; use a noise reduction algorithm to eliminate abnormal eigenvalues and reduce the interference of noise data; perform a merging operation on similar feature vectors to reduce the dimension of the aggregated feature set and enhance the compactness of the feature representation; output the final aggregated feature set as the input data for subsequent quantization processing.
[0195] In this embodiment, through the aggregation processing of the feature subset, the scattered feature data can be effectively merged into a more representative aggregated feature set. This aggregation processing method can enhance the compactness and robustness of the feature data, reduce the influence of noise data, and thus improve the accuracy of subsequent feature matching and recognition. The finally generated aggregated feature set has better discrimination ability and expression ability in the retrieval and matching process and is applicable to feature comparison tasks in various scenarios.
[0196] In one embodiment, the above S60 includes:
[0197] S601, extract a set of cluster centers from the clustering model, and the set of cluster centers contains multiple cluster centers;
[0198] S602, respectively determine the distances between each feature vector in the aggregated feature set and each cluster center in the set of cluster centers;
[0199] S603, according to the principle of minimum distance, determine the target cluster center corresponding to each feature vector and assign a codebook index value to each feature vector;
[0200] S604, establish a mapping relationship between the codebook index value and the corresponding feature vector to generate a preliminary quantization feature set;
[0201] S605, perform discretization processing on the preliminary quantization feature set to generate discrete codebook dictionary index values;
[0202] S606, encode the discrete codebook dictionary index values to generate the multi-scale feature codebook.
[0203] In this embodiment, the set of clustering centers is a set of reference points extracted from a pre-trained clustering model, and each clustering center represents the center point of a feature pattern. In the quantization process, these clustering centers are used to map continuous feature vectors to discrete index values.
[0204] Load a pre-trained clustering model, such as a K-means clustering model. Extract the set of clustering centers generated in the clustering model. Each clustering center is a high-dimensional feature vector. The number of clustering centers depends on the configuration of the clustering model, such as 1024 clustering centers or 2048 clustering centers. Each clustering center represents a common feature pattern, such as edge features, texture features, etc.
[0205] Distance calculation is the process of matching a feature vector with a clustering center. By calculating the distance between each feature vector and all clustering centers, it is possible to determine which clustering center the feature vector matches most closely.
[0206] Traverse each feature vector and calculate the distance to each clustering center in the set of clustering centers one by one. Use Euclidean distance or cosine similarity to calculate the distance between the feature vector and the clustering center, and record the distance from each feature vector to each clustering center to provide a basis for subsequent index assignment.
[0207] The principle of the minimum distance means that each feature vector is assigned to the clustering center with the closest distance. Through this mapping method, a unique codebook index value can be assigned to each feature vector, representing its corresponding clustering center.
[0208] Traverse each feature vector, and based on the previously calculated distance data, select the clustering center with the closest distance as the target clustering center. Assign a codebook index value to each feature vector, corresponding to the number of the target clustering center. For example, if the distance between feature vector A and clustering center 1 is the smallest, then assign the codebook index value "0001".
[0209] Establish a mapping relationship between each codebook index value and the feature vector, and a preliminary quantization feature set can be generated. The preliminary quantization feature set is an intermediate result of discretizing continuous feature data.
[0210] Pair each feature vector with its assigned codebook index value to form a mapping relationship. The preliminary quantization feature set contains all feature vectors and their corresponding index values, for example:
[0211] Feature vector A -> Index value 0001;
[0212] Feature vector B -> Index value 0010;
[0213] Feature vector C -> Index value 0101.
[0214] The discretization process further converts the codebook index values in the preliminary quantization feature set into discrete feature symbols. The discretized index values are usually represented as fixed-length binary codes or hexadecimal codes.
[0215] Convert each codebook index value into a discrete symbol according to a preset coding rule. For example, convert the index value 0001 to "A" and the index value 0010 to "B". The discretized codebook dictionary index values can use compression techniques such as Huffman coding or arithmetic coding to further reduce the data storage space.
[0216] For example:
[0217] Index value 0001 -> Discrete symbol A;
[0218] Index value 0010 -> Discrete symbol B;
[0219] Index value 0101 -> Discrete symbol C.
[0220] The coding process organizes the discretized codebook index values into a final multi-scale feature codebook. The multi-scale feature codebook can be represented as hierarchical encoded data for easy and quick retrieval and comparison. Group and encode the discrete codebook index values according to the feature hierarchy or feature category. Perform coding processing on each discrete symbol to generate the final multi-scale feature codebook. For example: Multi-scale feature codebook = {A, B, C, D,...}
[0221] Example illustration: In the field of medical and health, lesion recognition is one of the core tasks in medical image analysis. Taking lung CT images as an example, doctors need to identify the lesion areas from the images and compare them with historical case data to assist in diagnosis and treatment. The morphology, density, boundary features, etc. of different lesions may vary greatly, and the lesion areas in the images may show different scale changes due to factors such as imaging equipment, shooting angles, and lesion sizes. Therefore, for lesion recognition and matching, a codebook that can be compatible with multi-scale features needs to be constructed to improve the accuracy of retrieval and comparison.
[0222] First, the system extracts an aggregated feature set from the images of the lesion areas, such as the edge contours, density distributions, and texture information of the lesions. After vectorization processing of these feature data, a set of feature vectors is generated. Next, the system extracts a set of clustering centers from a pre-trained clustering model. These clustering centers represent different lesion feature patterns, such as circular lesions, lesions with irregular boundaries, high-density lesions, etc. Each clustering center corresponds to a specific feature category.
[0223] The system determines the closest cluster center by calculating the distance from each feature vector to the cluster center. According to the principle of minimum distance, each feature vector is assigned to the corresponding cluster center, and a preliminary set of quantized features is generated. Subsequently, these preliminary sets of quantized features are discretized, converting the feature data into discrete codebook index values, such as using letters or numbers to represent different lesion feature categories.
[0224] In the final encoding process, the system organizes the discrete codebook index values into a multi-scale feature codebook. This codebook contains multiple levels of feature patterns for representing lesion features at different scales. The multi-scale feature codebook can be used to quickly retrieve similar lesion images in the historical case database, providing a reference for doctors in diagnosis.
[0225] For example, when a doctor uploads a new lung CT image, the system can extract the feature vectors of the lesion area from the image and compare these features with the multi-scale feature codebook. By matching the edge shape, texture distribution, and density information of the lesion, the system can quickly screen out the records in the historical case database that are most similar to the current lesion, providing a reference for the possible lesion type and the trend of disease development for the doctor.
[0226] In this application, the generation of the multi-scale feature codebook can effectively solve the problem that traditional codebooks are sensitive to scale changes, reduce the error of lesion feature matching, and improve the accuracy of lesion retrieval. At the same time, the hierarchically stored codebook index values help to improve the retrieval efficiency of the system, reduce the consumption of computing resources, and ensure that the system has an efficient response ability in large-scale medical data retrieval.
[0227] In this embodiment, through the quantization process of the aggregated feature set, continuous feature vector data can be converted into discretized feature symbols, thereby significantly reducing the data storage space and computing resource consumption. The finally generated multi-scale feature codebook can effectively be compatible with the features of target objects at different scales, improving the robustness and accuracy of feature retrieval and matching. This quantization processing method is applicable to feature comparison tasks in multiple scenarios, such as image retrieval, identity verification, lesion recognition, and other fields.
[0228] In one embodiment, after S60 as described above, it further includes:
[0229] S701, annotate the multi-scale feature codebook to generate annotation information including the category information and label information of the target object;
[0230] S702, generate a unique index identifier for each discretized feature vector in the multi-scale feature codebook according to the annotation information, and establish a corresponding index mapping relationship;
[0231] S703: Store the annotated multi-scale feature codebook and the index mapping relationship according to a preset level.
[0232] In this embodiment, annotation is to give actual semantic information to the discretized feature vectors in the multi-scale feature codebook. Category information can indicate the type of target object, such as animal categories such as pigs and cows; label information can indicate the specific attributes or status of the feature vector, such as body shape, age, health status, etc.
[0233] Category information labeling: Determine the category to which the target object belongs based on the feature pattern corresponding to the feature codebook index value. For example, label a specific feature codebook index value as "cow" or "pig".
[0234] Label information annotation: Give the feature vector a more detailed description based on business needs, such as labels such as "adult pig" and "lesion area".
[0235] Annotation format: Annotation information can be stored in the form of key-value pairs, and each discretized feature vector corresponds to one or more annotation information items.
[0236] The index identifier is a coded value that uniquely identifies each discretized feature vector, which is convenient for subsequent retrieval and management. The index mapping relationship is a relational table that stores the codebook index value and the annotation information in correspondence.
[0237] Generate index identifier: Generate a unique identifier based on the codebook index value of the discretized feature vector, usually in the form of encoding, such as UUID, hash value or digital number.
[0238] Establish index mapping relationship: Establish mapping relationship between each index identifier and discretized feature vector and its annotation information to facilitate fast retrieval.
[0239] Hierarchical storage is to manage the multi-scale feature codebook and index mapping relationship in a hierarchical structure to support retrieval and matching operations at different levels. The preset levels can be divided by categories, tags, feature patterns, etc.
[0240] Hierarchical structure design: Design a hierarchical storage structure based on actual business needs, such as hierarchical management by category and tag.
[0241] Storage method: Distributed database, file system or memory storage can be used to ensure retrieval efficiency.
[0242] Example hierarchy:
[0243] The first level: divided by category, such as "pig" and "cow";
[0244] The second level: divided by labels, such as "adult", "juvenile", "lesion area";
[0245] The third level: Divide according to feature patterns, such as "edge features" and "texture features".
[0246] In this embodiment, through the annotation and hierarchical storage of the multi-scale feature codebook, the discretized feature vectors can be effectively associated with the actual business data, improving the accuracy of feature retrieval and matching. The annotation information enables the feature codebook to have stronger semantic expression ability, while the hierarchical storage structure ensures the efficiency and flexibility of data management, can be quickly adapted to different application scenarios, reduce the time cost of data retrieval, and improve the overall performance of the system.
[0247] In one embodiment, a multi-scale feature information processing and codebook generation device is provided, and the multi-scale feature information processing and codebook generation device corresponds one-to-one with the multi-scale feature information processing and codebook generation method in the above embodiment. Refer to Figure 3 , Figure 3 is a schematic diagram of the functional modules of a preferred embodiment of the multi-scale feature information processing and codebook generation device of the present invention. Image acquisition module 10, multi-scale feature extraction module 20, multi-scale pooling processing module 30, feature segmentation module 40, feature aggregation module 50, and feature quantization and codebook generation module 60. The detailed description of each functional module is as follows:
[0248] The image acquisition module 10 is used to acquire an image containing a target object;
[0249] The multi-scale feature extraction module 20 is used to extract multi-scale feature information of the target object from the image;
[0250] The multi-scale pooling processing module 30 is used to perform pooling processing on the multi-scale feature information according to a preset plurality of scales, and generate pooled feature maps of different scales respectively;
[0251] The feature segmentation module 40 is used to perform a segmentation operation on each pooled feature map to generate at least one feature subset;
[0252] The feature aggregation module 50 is used to perform aggregation processing on each feature subset to merge the scattered feature data in each feature subset into a corresponding aggregated feature set;
[0253] The feature quantization and codebook generation module 60 is used to perform quantization processing on the aggregated feature set to generate a multi-scale feature codebook.
[0254] In one embodiment, the multi-scale feature extraction module 20 is specifically used for:
[0255] Performing target detection processing on the image based on a pre-trained target detection model to obtain candidate regions of the target object;
[0256] Apply the target segmentation method in the candidate region to separate the target object from the background region and generate an image of the target object region;
[0257] Perform cropping and boundary adjustment on the target object region image to obtain a target object image;
[0258] Perform size adjustment processing on the target object image;
[0259] Perform color channel normalization processing on the target object image after size adjustment processing to obtain a preprocessed image for extracting multi-scale feature information.
[0260] In one embodiment, the multi-scale pooling processing module 30 is specifically configured to:
[0261] Establish multiple scale branches, each scale branch corresponding to a different pooling kernel size or sampling step;
[0262] Input the multi-scale feature information into each scale branch respectively;
[0263] Perform pooling processing on the multi-scale feature information based on the pooling kernel size or sampling step of each scale branch, and generate pooling feature maps corresponding to different scales respectively.
[0264] In one embodiment, the feature segmentation module 40 is specifically configured to:
[0265] Determine the segmentation parameters of each pooling feature map, where the segmentation parameters include a preset grid division strategy;
[0266] Based on the segmentation parameters, divide each pooling feature map into multiple pooling feature regions according to the preset grid division strategy;
[0267] Perform feature extraction operations on each pooling feature region to extract a subset feature vector set in each pooling feature region, and use the subset feature vector set in each feature region as a feature subset.
[0268] In one embodiment, the feature aggregation module 50 is specifically configured to:
[0269] Perform vectorization processing on the scattered feature data in each feature subset to generate a corresponding feature vector set;
[0270] Extract a global clustering center set from a pre-trained clustering model, where the global clustering center set contains multiple clustering centers;
[0271] Match the feature vectors in each feature vector set with all the clustering centers in the clustering center set respectively to determine the distances from each feature vector to each clustering center;
[0272] According to the minimum distance principle, each feature vector is merged into the nearest clustering center to generate a preliminary aggregated feature set;
[0273] Perform feature optimization processing on the preliminary aggregated feature set to generate a final aggregated feature set.
[0274] In one embodiment, the feature quantization and codebook generation module 60 is specifically configured to:
[0275] Extract a set of clustering centers from the clustering model, where the set of clustering centers includes multiple clustering centers;
[0276] For each feature vector in the aggregated feature set, determine the distances from each clustering center in the set of clustering centers;
[0277] According to the minimum distance principle, determine the target clustering center corresponding to each feature vector, and assign a codebook index value to each feature vector;
[0278] Establish a mapping relationship between the codebook index value and the corresponding feature vector to generate a preliminary quantized feature set;
[0279] Perform discretization processing on the preliminary quantized feature set to generate discrete codebook dictionary index values;
[0280] Perform encoding processing on the discrete codebook dictionary index values to generate the multi-scale feature codebook.
[0281] In one embodiment, the feature quantization and codebook generation module 60 is specifically configured to:
[0282] Annotate the multi-scale feature codebook to generate annotation information including the category information and label information of the target object;
[0283] Generate a unique index identifier for each discretized feature vector in the multi-scale feature codebook according to the annotation information, and establish a corresponding index mapping relationship;
[0284] Store the annotated multi-scale feature codebook and the index mapping relationship at a preset level.
[0285] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media, and internal memory. The non-volatile storage media stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage media. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a multi-scale feature information processing and codebook generation method.
[0286] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5 shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile storage media and internal memory. The non-volatile storage media stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage media. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a multi-scale feature information processing and codebook generation method.
[0287] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:
[0288] Obtain an image containing a target object;
[0289] Extract multi-scale feature information of the target object from the image;
[0290] Perform pooling processing on the multi-scale feature information according to a preset plurality of scales, and generate pooling feature maps of different scales respectively;
[0291] Perform a segmentation operation on each pooling feature map to generate at least one feature subset;
[0292] Perform an aggregation process on each feature subset to merge the scattered feature data in each feature subset into a corresponding aggregated feature set;
[0293] Perform quantization processing on the aggregated feature set to generate a multi-scale feature codebook.
[0294] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0295] Obtain an image including a target object;
[0296] Extract multi-scale feature information of the target object from the image;
[0297] Perform pooling processing on the multi-scale feature information according to a plurality of preset scales to generate pooling feature maps of different scales respectively;
[0298] Perform a segmentation operation on each pooling feature map to generate at least one feature subset;
[0299] Perform an aggregation process on each feature subset to merge the scattered feature data in each feature subset into a corresponding aggregated feature set;
[0300] Perform quantization processing on the aggregated feature set to generate a multi-scale feature codebook.
[0301] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can implement, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.
[0302] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0303] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0304] It should be noted that if there are software tools or components of other companies in the embodiments of the present application, they are only used for example introduction and do not represent actual use. The above-mentioned embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A multi-scale feature information processing and codebook generation method, characterized in that: The following steps are involved: Get an image containing the target object; Extracting multi-scale feature information of the target object from the image; Performing pooling processing on the multi-scale feature information according to a plurality of preset scales to generate pooling feature maps of different scales respectively; Perform a segmentation operation on each pooled feature map to generate at least one feature subset; Aggregate each feature subset to merge the scattered feature data in each feature subset into a corresponding aggregate feature set; The aggregated feature set is quantized to generate a multi-scale feature codebook.
2. The multi-scale feature information processing and codebook generation method according to claim 1, characterized in that: Before extracting the multi-scale feature information of the target object from the image, the method further includes: Performing target detection processing on the image based on a pre-trained target detection model to obtain a candidate region of the target object; Applying a target segmentation method in the candidate area to separate the target object from the background area and generate a target object area image; Performing cropping and boundary adjustment on the target object area image to obtain a target object image; Performing a resizing process on the target object image; A color channel normalization process is performed on the target object image after the size adjustment process to obtain a preprocessed image for extracting multi-scale feature information.
3. The multi-scale feature information processing and codebook generation method according to claim 1, characterized in that: Performing pooling processing on the multi-scale feature information according to the preset multiple scales to generate pooling feature maps of different scales respectively, including: Establish multiple scale branches, each scale branch corresponds to a different pooling kernel size or sampling step size; Inputting the multi-scale feature information into each scale branch respectively; Based on the pooling kernel size or sampling step size of each scale branch, pooling processing is performed on the multi-scale feature information to generate pooling feature maps corresponding to different scales.
4. The multi-scale feature information processing and codebook generation method according to claim 1, characterized in that: Perform a segmentation operation on each pooled feature map to generate at least one feature subset, including: Determining segmentation parameters of each pooling feature map, wherein the segmentation parameters include a preset grid division strategy; Based on the segmentation parameters, each pooling feature map is divided into a plurality of pooling feature regions according to a preset grid division strategy; A feature extraction operation is performed on each pooled feature region to extract a set of subset feature vectors in each pooled feature region, and the set of subset feature vectors in each feature region is used as a feature subset.
5. The multi-scale feature information processing and codebook generation method according to claim 1, characterized in that: Aggregation processing is performed on each feature subset to merge the scattered feature data in each feature subset into a corresponding aggregate feature set, including: Perform vectorization processing on the scattered feature data in each feature subset to generate a corresponding feature vector set; Extracting a global cluster center set from a pre-trained clustering model, wherein the global cluster center set includes multiple cluster centers; Matching the feature vectors in each feature vector set with all cluster centers in the cluster center set respectively, and determining the distance from each feature vector to each cluster center; According to the minimum distance principle, each feature vector is merged into the nearest cluster center to generate a preliminary aggregated feature set; Perform feature optimization processing on the preliminary aggregated feature set to generate a final aggregated feature set.
6. The multi-scale feature information processing and codebook generation method according to claim 1, characterized in that: The aggregated feature set is quantized to generate a multi-scale feature codebook, including: Extracting a cluster center set from the clustering model, wherein the cluster center set includes a plurality of cluster centers; Determining the distance between each feature vector in the aggregate feature set and each cluster center in the cluster center set; According to the minimum distance principle, the target cluster center corresponding to each feature vector is determined, and a codebook index value is assigned to each feature vector; Establishing a mapping relationship between the codebook index value and the corresponding feature vector to generate a preliminary quantized feature set; Discretizing the preliminary quantized feature set to generate discrete codebook dictionary index values; The discrete codebook dictionary index values are encoded to generate the multi-scale feature codebook.
7. The multi-scale feature information processing and codebook generation method according to claim 1, characterized in that: After quantizing the aggregated feature set to generate a multi-scale feature codebook, the method further includes: Annotating the multi-scale feature codebook to generate annotation information including category information and label information of the target object; Generating a unique index identifier for each discretized feature vector in the multi-scale feature codebook according to the annotation information, and establishing a corresponding index mapping relationship; The annotated multi-scale feature codebook and the index mapping relationship are stored according to a preset level.
8. A multi-scale feature information processing and codebook generation device, characterized in that: The multi-scale feature information processing and codebook generation device comprises: An image acquisition module, used for acquiring an image containing a target object; A multi-scale feature extraction module, used to extract multi-scale feature information of the target object from the image; A multi-scale pooling processing module, used to perform pooling processing on the multi-scale feature information according to a plurality of preset scales, and generate pooling feature maps of different scales respectively; A feature segmentation module, used to perform a segmentation operation on each pooled feature map to generate at least one feature subset; A feature aggregation module is used to perform aggregation processing on each feature subset to merge the scattered feature data in each feature subset into a corresponding aggregate feature set; The feature quantization and codebook generation module is used to quantize the aggregated feature set and generate a multi-scale feature codebook.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and a multi-scale feature information processing and codebook generation program stored in the memory and run on the processor. When the multi-scale feature information processing and codebook generation program is executed by the processor, the steps of the multi-scale feature information processing and codebook generation method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores a multi-scale feature information processing and codebook generation program, which, when executed by the processor, implements the steps of the multi-scale feature information processing and codebook generation method according to any one of claims 1 to 7.
Citation Information
Cited By
Gas pipeline leakage detection method, device, equipment and program product
CN120747047A