Image feature extraction method and device, equipment and medium
Through the self-supervised mapping learning mechanism, the feature vectors of viewing maps are generated and compared during the image feature extraction process, and the problem of relying on manual annotation and negative sample selection in the prior art is solved, and efficient image feature extraction and model generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510141899.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art relies on manual annotation data and negative sample selection in the process of image feature extraction, resulting in high cost of model training and the inability to accurately extract key features of similar samples, affecting the recognition effect and generalization ability.
By obtaining image data, the original view and enhanced view diagram are generated, the feature vectors are generated, and the differences between feature map vectors are compared through the self-supervised mapping learning mechanism, and the parameters of the feature encoding unit and mapping unit are adjusted to extract the feature representation of the target image.
Without relying on manual annotation data and negative sample selection, key features are extracted directly through image data, reducing model training costs, and improving the accuracy and generalization capabilities of the model in image recognition tasks.
Smart Images

Figure CN119992116A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence technology and medical health, and in particular to an image feature extraction method, device, equipment and storage medium. Background Art
[0002] Image feature coding methods are widely used in the medical and health fields and the financial field. In the medical and health field, medical image analysis relies on the accurate extraction of image features by the model for scenarios such as lesion identification and disease prediction. However, traditional image feature coding methods usually rely on a large amount of manually annotated data to train the model. This annotation process is not only time-consuming, labor-intensive and costly, but also difficult to ensure the consistency of the annotations. Especially when the lesion features are complex and diverse, manual annotation is difficult to fully cover, thus affecting the recognition effect of the model. In addition, existing contrastive learning methods often use positive and negative sample pairs in medical images to improve the model's recognition ability, but improper selection of negative samples may cause the model to ignore subtle lesion features, affecting the accuracy of early diagnosis.
[0003] In the financial field, image feature encoding methods are often used in automated business scenarios such as identity authentication, bill recognition, and transaction voucher review. Traditional methods also rely on manually labeled data to train models, but due to the diverse formats of financial documents, it is difficult to cover all types of documents, resulting in insufficient model generalization capabilities. At the same time, the differences in documents such as financial bills and contracts are often reflected in detailed features. If the negative samples of traditional contrastive learning methods are not properly selected, these key details may be ignored, thereby affecting the recognition accuracy of scenarios such as forged bill detection and invoice verification.
[0004] In summary, the existing image feature encoding methods in the fields of healthcare and finance all face common problems such as high manual labeling costs, difficulty in selecting negative samples, and insufficient extraction of key features of similar samples. These problems directly affect the recognition effect and practical application value of the model, making it difficult to meet the needs of high-precision image recognition. Summary of the invention
[0005] The main purpose of the present invention is to provide an image feature extraction method, device, equipment and storage medium, aiming to solve the technical problems that the existing technology relies on manually labeled data and negative sample selection in the process of image feature extraction, resulting in high model training costs and inability to accurately extract key features of similar samples, affecting recognition effects and generalization capabilities.
[0006] To achieve the above object, the present invention provides an image feature extraction method, comprising:
[0007] Acquire image data, and generate an original viewing angle map and an enhanced viewing angle map according to the image data;
[0008] The original view image is encoded by a first feature encoding unit to generate an original feature vector;
[0009] The enhanced viewing angle map is encoded by a second feature encoding unit to generate an enhanced feature vector;
[0010] Inputting the original feature vector and the enhanced feature vector into a first mapping unit to generate an original mapping vector and an enhanced mapping vector respectively;
[0011] Inputting the original mapping vector into a second mapping unit to generate an original feature mapping vector;
[0012] Determine a loss value based on the original feature mapping vector and the enhanced mapping vector;
[0013] Adjusting parameters of the first mapping unit and parameters of the first feature encoding unit according to the loss value;
[0014] adjusting parameters of the second feature encoding unit based on the adjusted parameters of the first mapping unit and the parameters of the first feature encoding unit;
[0015] A feature representation of the target image is extracted based on the adjusted second feature encoding unit.
[0016] Furthermore, to achieve the above object, the present invention provides an image feature extraction device, comprising:
[0017] A data preprocessing module, used to obtain image data and generate an original viewing angle map and an enhanced viewing angle map according to the image data;
[0018] An original encoding module, used for encoding the original view image through a first feature encoding unit to generate an original feature vector;
[0019] an enhanced coding module, configured to perform coding processing on the enhanced viewing angle map through a second feature coding unit to generate an enhanced feature vector;
[0020] A first feature mapping module, used for inputting the original feature vector and the enhanced feature vector into a first mapping unit to generate an original mapping vector and an enhanced mapping vector respectively;
[0021] A second feature mapping module, used for inputting the original mapping vector into a second mapping unit to generate an original feature mapping vector;
[0022] A loss determination module, configured to determine a loss value based on the original feature mapping vector and the enhanced mapping vector;
[0023] a parameter adjustment module, used for adjusting the parameters of the first mapping unit and the parameters of the first feature encoding unit according to the loss value;
[0024] a model updating module, configured to adjust parameters of the second feature encoding unit based on the adjusted parameters of the first mapping unit and the parameters of the first feature encoding unit;
[0025] A feature extraction module is used to extract a feature representation of a target image based on the adjusted second feature encoding unit.
[0026] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and an image feature extraction program stored in the memory and executable on the processor, wherein the image feature extraction program implements the steps of the image feature extraction method described above when executed by the processor.
[0027] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which an image feature extraction program is stored, and when the image feature extraction program is executed by a processor, the steps of the image feature extraction method as described above are implemented.
[0028] Beneficial effects: The present invention relates to the fields of artificial intelligence technology and medical health, and discloses an image feature extraction method, including: acquiring image data and generating an original viewing angle map and an enhanced viewing angle map, encoding the original viewing angle map and the enhanced viewing angle map respectively, generating an original feature vector and an enhanced feature vector, and then generating a feature mapping vector. By comparing the difference between the feature mapping vectors, the loss value is calculated, the parameters of the feature coding unit are adjusted, and the feature representation of the target image is extracted based on the adjusted feature coding unit. The present invention introduces a self-supervised mapping learning mechanism in the image feature extraction process, without relying on manually labeled data and negative sample selection, and directly generates the original viewing angle map and the enhanced viewing angle map through the image data itself, thereby effectively extracting the key features of the same type of samples; by adjusting the parameters of the feature coding unit and the mapping unit, the accuracy and generalization ability of the model in the image recognition task are improved; the model training cost can be reduced, and the dependence on manual intervention can be reduced, and it is suitable for high-precision image analysis scenarios such as medical health and financial bill recognition, and the automation level of image recognition and business processing efficiency are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:
[0030] Figure 1 A schematic diagram of an application environment of an image feature extraction method in an embodiment of the present invention;
[0031] Figure 2 A schematic diagram of a flow chart of an embodiment of an image feature extraction method of the present invention;
[0032] Figure 3 A schematic diagram of functional modules of a preferred embodiment of the image feature extraction device of the present invention;
[0033] Figure 4 A schematic diagram of the structure of a computer device in one embodiment of the present invention;
[0034] Figure 5 FIG. 4 is another schematic diagram of the structure of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0035] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.
[0036] The image feature extraction method provided by the embodiment of the present invention can be applied in the following aspects: Figure 1 In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can obtain image data through the user terminal and generate an original view map and an enhanced view map, encode the original view map and the enhanced view map respectively, generate an original feature vector and an enhanced feature vector, and then generate a feature mapping vector. By comparing the difference between the feature mapping vectors, the loss value is calculated, the parameters of the feature coding unit are adjusted, and the feature representation of the target image is extracted based on the adjusted feature coding unit. The present invention introduces a self-supervised mapping learning mechanism in the image feature extraction process, without relying on manual annotation data and negative sample selection, and directly generates the original view map and the enhanced view map through the image data itself, thereby effectively extracting the key features of the same sample; by adjusting the parameters of the feature coding unit and the mapping unit, the accuracy and generalization ability of the model in the image recognition task are improved; the model training cost can be reduced, and the dependence on manual intervention is reduced, which is suitable for high-precision image analysis scenarios such as medical health and financial bill recognition, and the automation level of image recognition and business processing efficiency are improved. Among them, the user terminal can be but not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server can be implemented with an independent server or a server cluster composed of multiple servers. The present invention is described in detail below through specific embodiments.
[0037] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of the image feature extraction method provided by the present invention. It should be noted that although the logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0038] like Figure 2As shown, the image feature extraction method proposed by the present invention includes the following steps:
[0039] S10, acquiring image data, and generating an original viewing angle map and an enhanced viewing angle map according to the image data;
[0040] In this embodiment, acquiring image data refers to collecting image data from a data source, and the data type may be a static image, a continuous frame image, a scanned file, etc. These image data are usually stored in digital formats, including JPEG, PNG, TIFF, etc., to ensure that the image quality meets the requirements of subsequent processing.
[0041] Ways to obtain image data may include:
[0042] Local image acquisition: directly read image files from local storage devices. Image reading can be achieved through file system interfaces, image processing libraries, etc.
[0043] Online image acquisition: Obtain image data from online data sources through APIs or web crawlers. For example, extract medical imaging data from hospital information systems or obtain scanned bills from financial business systems.
[0044] Real-time image acquisition: Real-time image data acquisition through cameras, scanners and other devices, suitable for image processing in dynamic scenes, such as real-time identity authentication or medical image acquisition.
[0045] The original view map can refer to the image data without any processing. The original view map is used as a positive sample for model training. By directly using the original image, it can be ensured that the model extracts features from the most initial image information. In certain scenarios, the original image can also be preprocessed (such as denoising, graying) to improve the recognition effect of the model, but the basic structure and content of the original image cannot be changed.
[0046] The enhanced view map is a picture generated by performing a series of data augmentation operations on the original picture. These operations are designed to generate new picture samples so that the model can learn the image features from different viewpoints and improve the robustness and generalization ability of the model.
[0047] Data augmentation operations include:
[0048] Rotation processing: randomly rotate the image by a certain angle to generate viewing angles of different angles.
[0049] Implementation method: Call the rotation function in the image processing library (such as OpenCV, PIL).
[0050] Cropping: Randomly crop part of the image to generate perspective images of different sizes.
[0051] Implementation method: Specify the coordinates and size of the cropping area, and perform the cropping operation through the image processing library.
[0052] Flip processing: Flip the image horizontally or vertically to generate a symmetrical perspective image.
[0053] Implementation method: Call the flip function of the image processing library to perform a mirror transformation on the image matrix.
[0054] Add Gaussian noise: Add random noise to the image to make the image show different texture characteristics.
[0055] Implementation method: Generate a Gaussian distributed noise matrix and superimpose the noise matrix on the image data.
[0056] Example: In the field of healthcare, there are many types of medical image data, including X-rays, CT scan images, pathological slice images, magnetic resonance imaging (MRI), etc. These images are used to assist doctors in identifying the location of lesions, diagnosing disease types, and predicting the progression of the disease. However, due to differences in patient body structure, scanning angles, lesion characteristics, etc., the same lesion may show different morphological characteristics from different perspectives. Traditional image recognition models rely on manually annotated data and it is difficult to fully cover these changes.
[0057] By obtaining the original medical image as the original view map, and performing data enhancement operations such as rotation, cropping, flipping, and adding noise on the image, an enhanced view map is generated. For example, for a lung CT scan image:
[0058] Rotation processing: simulates scanning perspectives at different angles to help the model learn the characteristic changes of lesions at different angles.
[0059] Cropping: Cropping is performed locally on the image so that the model can identify the local features of the lesion area.
[0060] Flip processing: simulate the feature changes when the image is flipped left and right or up and down to avoid misjudgment of the model due to changes in viewing angle.
[0061] Add noise: simulate the artifacts and noise that may appear in the actual scanning process to improve the model's robustness to image noise.
[0062] By enhancing the image data, the original view map and the enhanced view map are generated, thereby providing the model with richer training samples and avoiding reliance on manually labeled data. There is no need to select negative samples, and the recognition ability of the model can be improved only by learning positive sample pairs.
[0063] S20, encoding the original view image by a first feature encoding unit to generate an original feature vector;
[0064] In this embodiment, the original view map refers to an image sample obtained from the image data, which can be directly used for subsequent feature encoding processing. Before the image is input into the model, the original view map can be subjected to necessary preprocessing operations, such as image resizing, color channel conversion, normalization, etc. The purpose of these preprocessing operations is to standardize the image format, improve the compatibility and recognition effect of the model, but will not substantially change the core feature information of the image.
[0065] The size of the original view image can be adjusted to fit the model's input requirements, such as adjusting the image to a fixed size. In addition, the color format of the image can be converted, the RGB channel can be converted to a grayscale image, or the pixel values can be normalized to make the model more stable during training. If there is noise in the image, noise reduction can also be performed to improve the model's recognition effect on the image.
[0066] Encoding is a key step in converting the original view map into a feature vector. The first feature encoding unit usually uses a deep learning model such as a convolutional neural network, which can automatically extract feature information at different levels from the image. The initial feature extraction may include low-level features such as edges, contours, and textures, and as the number of network layers increases, more advanced semantic information can be captured.
[0067] The first feature encoding unit first performs a convolution operation on the input image, scanning various areas of the image to extract edge and texture information. Next, the feature map is reduced in dimension through a pooling operation, retaining the most important information while reducing computational complexity. In order to enhance the expressiveness of the model, after each layer of feature extraction, an activation function is applied to perform a nonlinear transformation on the eigenvalues. In addition, normalization can ensure that the eigenvalues are within a reasonable range, thereby avoiding the problem of gradient explosion or gradient vanishing.
[0068] The generation of the feature vector is the final output of the encoding process, usually represented by a fixed-length vector. This vector contains the key feature information of the image and serves as the input of the subsequent mapping unit.
[0069] Example description: In medical image analysis, pathological sections, CT scan images, etc. can be input into the model as original viewport images. For example, in the analysis of lung CT scan images, the first feature encoding unit is used to extract features such as image edges, lesion areas, and density distribution. The convolution operation is used to extract the contour information of the lesion, while the pooling operation can reduce noise interference and retain key features. The shape and location information of the lesion area are represented by feature vectors to assist doctors in diagnosing the disease.
[0070] In practical applications, the CT scan image can be normalized first to adjust the pixel value to between 0 and 1. Then, the features are extracted layer by layer through the convolutional network to finally generate a feature vector. This feature vector can be used by doctors or AI diagnostic systems to identify and classify different types of lesion areas.
[0071] In addition, in the automatic bill recognition scenario in the financial field, bill images such as invoices and contracts usually contain a large amount of text and table information. Through the first feature encoding unit, the bill image can be converted into a feature vector to extract key field information (such as invoice number, invoice date, etc.). The encoding process of the bill image can include preprocessing steps such as image cropping and noise removal. Subsequently, the text position and format layout of the bill are identified through convolution operations. The output of the feature vector contains key information such as invoice number and amount, providing data support for subsequent automated review.
[0072] By inputting the original view map into the first feature encoding unit and generating a feature vector through encoding processing, key feature information can be automatically extracted from the image data. This feature extraction method reduces the reliance on manual annotation and reduces the cost of model training. At the same time, through the multi-level feature extraction process, the model can accurately identify information such as edges, textures, and semantics in the image, thereby improving the accuracy and generalization ability of image recognition.
[0073] S30, encoding the enhanced viewing angle map by a second feature encoding unit to generate an enhanced feature vector;
[0074] In this embodiment, the enhanced view map is generated by performing data enhancement operations such as rotating, cropping, flipping, and adding noise to the original view map. This image processing method helps the model to recognize images under different view angles, different lighting conditions, and different background interferences, thereby improving the robustness of the model. Using the enhanced view map as input allows the model to learn the key features of the image under different changing conditions and improve the generalization ability of the model.
[0075] The second feature encoding unit is used to extract features from the enhanced view map and generate an enhanced feature vector. This unit can be an independent encoding module, usually including a convolutional neural network or other deep learning structure, for extracting high-level feature information from the enhanced view map.
[0076] The enhanced view map is preprocessed to meet the input requirements of the second feature encoding unit. For example, the image size, color format or normalization processing is adjusted. The enhanced view map is passed to the second feature encoding unit, and steps such as convolution, pooling and activation function processing are performed to extract the feature information of the image layer by layer. Through the encoding process, a fixed-length enhanced feature vector is finally output, which contains the feature representation of the image under the enhanced view, which is used for subsequent mapping unit processing.
[0077] The encoding process is the process of extracting features from the enhanced view map, which is similar to the processing of the original view map. However, the enhanced image samples enable the model to learn more feature representations under changing conditions. The second feature encoding unit extracts low-level features such as edges and textures of the image, as well as high-level information such as object shape and semantic features through multi-layer convolution and pooling operations.
[0078] This processing method can effectively improve the generalization ability of the model, so that it can still accurately identify target features when the image samples are rotated, flipped, or interfered by noise. In addition, through normalization processing and the introduction of activation functions, the numerical range of the enhanced feature vector is standardized, further improving the training stability of the model.
[0079] The convolution operation extracts the edge and texture features of the image; the pooling operation reduces the dimension of the feature map while retaining the most important information; the activation function processing introduces nonlinear transformation to enhance the recognition ability of the model; the normalization processing ensures that the numerical range of the feature vector is within a reasonable range to avoid gradient explosion or disappearance.
[0080] Example description: In medical image analysis, in order to improve the model's ability to identify lesion areas, data enhancement can be performed on CT scan images, MRI images, etc. For example, for the same lung CT image, multiple enhanced view images can be generated through operations such as rotation, cropping, and flipping. These enhanced view images are input into the second feature encoding unit to extract the feature information of the lesion area at different perspectives and generate enhanced feature vectors. These enhanced feature vectors can reflect the shape changes and density distribution of the lesion area, helping doctors to identify and diagnose diseases more accurately.
[0081] By encoding the enhanced view map through the second feature encoding unit, the model can learn the feature performance of the image under different changing conditions under the condition of data enhancement, and improve the robustness and generalization ability of the model. Compared with the feature extraction method that only uses the original view map, the processing of the enhanced view map can help the model better cope with image rotation, flipping, noise interference, etc., and avoid the model from being overly dependent on the features of specific image samples.
[0082] S40, inputting the original feature vector and the enhanced feature vector into a first mapping unit to generate an original mapping vector and an enhanced mapping vector respectively;
[0083] In this embodiment, the original feature vector and the enhanced feature vector are feature representations extracted from the image by the feature encoding unit. In order to further extract the high-order feature information of the image, these feature vectors need to be input into the first mapping unit. The core function of the first mapping unit is to convert the input feature vector into a feature representation in a new mapping space, thereby capturing the correlation between the features.
[0084] The process of inputting the first mapping unit includes data transmission and processing mechanisms. The first mapping unit is usually composed of a set of fully connected layers (Fully Connected Layer) or multi-layer perceptrons (MLP), which can perform feature transformation operations on the input feature vector to achieve the conversion from the original feature space to the mapping space.
[0085] The original feature vector and the enhanced feature vector are used as input data and passed to the input layer of the first mapping unit respectively. The feature vector is normalized at the input stage to ensure that the input data is within a reasonable numerical range and improve the training stability of the model. The feature vector is passed step by step through the neural network layer, and linear transformation and nonlinear activation operations are performed to convert the input feature vector into a mapping vector.
[0086] The original mapping vector and the enhanced mapping vector are new feature representations generated by transforming the original feature vector and the enhanced feature vector through the first mapping unit. The mapping vector expresses the high-order features of the image in the mapping space and can capture the semantic information and feature associations in the input image.
[0087] The first mapping unit performs a series of feature transformation operations to achieve the conversion from low-dimensional feature vectors to high-dimensional mapping vectors. In this process, the model can learn the differences and similarities between input data, thereby providing support for subsequent feature comparison and loss calculation.
[0088] The input feature vector is linearly transformed through the fully connected layer of the first mapping unit to map the low-dimensional feature vector to a high-dimensional space. A nonlinear activation function (such as ReLU or Sigmoid) is introduced after each linear transformation to enhance the model's ability to express complex features. The original mapping vector and the enhanced mapping vector are generated respectively, and these two mapping vectors will be used for subsequent contrastive learning and loss value calculation.
[0089] Example description: In a medical image analysis scenario, the original feature vector and the enhanced feature vector can represent the basic features of the lesion area and the changing features at different viewing angles, respectively. For example, in lung CT image analysis, the original feature vector may represent the edge features of the lesion, while the enhanced feature vector represents the performance of the lesion at different rotation angles. By converting these feature vectors into mapping vectors through the first mapping unit, the changing patterns of the lesions at different viewing angles can be captured.
[0090] By inputting the original feature vector and the enhanced feature vector into the first mapping unit and generating the original mapping vector and the enhanced mapping vector, the model can capture the high-order features of the input image in the mapping space. This feature representation method improves the feature expression ability of the model, especially in complex scenarios such as image perspective changes and background noise interference, which can enhance the model's ability to recognize key features.
[0091] S50, inputting the original mapping vector into a second mapping unit to generate an original feature mapping vector;
[0092] In this embodiment, the original mapping vector is a high-dimensional feature representation generated after the original feature vector is processed by the first mapping unit. In order to further extract deep image feature information, the original mapping vector is input into the second mapping unit. The second mapping unit is used to perform dimensionality reduction, feature transformation and nonlinear processing on the original mapping vector to generate a more representative original feature mapping vector.
[0093] The second mapping unit usually includes a fully connected layer, an activation function, and a dimensionality reduction processing module, which compresses the high-dimensional feature vector into a feature mapping vector of fixed length through multi-layer transformation. This process can effectively remove redundant information while retaining the core features of the image and improving the feature expression ability of the model.
[0094] Before inputting the original mapping vector, normalization is performed to ensure that the input data is within a reasonable range. The original mapping vector is passed to the fully connected layer of the second mapping unit to perform a linear transformation. The high-dimensional mapping vector is compressed to a fixed length through dimensionality reduction operations to reduce computational complexity.
[0095] The original feature map vector is the output result after dimensionality reduction, linear transformation and nonlinear processing of the original map vector. The feature map vector can express the core feature information of the image in a more concise way, providing data support for subsequent loss calculation and parameter adjustment.
[0096] In the process of generating the original feature mapping vector, the second mapping unit introduces nonlinear transformation through activation function processing, so that the model can learn the complex feature relationship of the image. At the same time, dimensionality reduction processing can effectively remove redundant features and improve the computational efficiency and generalization ability of the model.
[0097] The input raw mapping vector is linearly transformed through the fully connected layer. An activation function is introduced after each linear transformation to enhance the nonlinear expression ability of the model. The generated raw feature mapping vector is a fixed-length feature representation that contains the core feature information of the image.
[0098] Example description: In medical image analysis, the original mapping vector can represent the basic features of the lesion area. In order to further extract the key features of the lesion area, the original mapping vector is input into the second mapping unit to generate the original feature mapping vector. For example, when analyzing lung CT images, the texture features, density changes, and shape features of the lesion area are extracted by the second mapping unit to generate a simplified but more diagnostically valuable feature representation.
[0099] In financial bill recognition, the original mapping vector can represent the basic field information of the bill image. In order to improve the model's ability to recognize different types of bills, the original mapping vector needs to be input into the second mapping unit to generate the original feature mapping vector. For example, when recognizing invoice images, the second mapping unit extracts the feature representation of key fields such as invoice number, invoice date, amount, etc., to provide support for subsequent automated review.
[0100] By inputting the original mapping vector into the second mapping unit and generating the original feature mapping vector, the deep feature information of the image can be further extracted. This process effectively removes redundant information through dimensionality reduction, linear transformation and nonlinear processing, while retaining the core features of the image. Compared with directly using the feature representation of the original mapping vector, the original feature mapping vector is more representative and can improve the recognition accuracy and generalization ability of the model.
[0101] S60, determining a loss value based on the original feature mapping vector and the enhanced mapping vector;
[0102] In this embodiment, the original feature mapping vector and the enhanced mapping vector represent the feature representation of the image at different viewing angles. In order to measure the similarity or difference between them, the model needs to compare the two feature vectors. The higher the similarity, the more effective the enhanced viewing angle map is in retaining the core feature information of the original viewing angle map, thus proving the stability and consistency of the model in feature extraction.
[0103] The comparison process usually uses similarity measurement or difference measurement methods. Common similarity measurement methods include cosine similarity, Euclidean distance, etc., while the difference measurement can be achieved by calculating the error between feature vectors. This comparison process is the basis for the subsequent calculation of the loss value.
[0104] The similarity between the original feature map vector and the enhanced map vector is determined by calculating the cosine similarity or inner product between the two vectors. If the model uses the difference comparison method, the Euclidean distance or Manhattan distance between the two feature vectors can be calculated as their difference measure.
[0105] The loss value is a key indicator used to measure the error between the predicted result and the actual result during the model training process. In this step, the loss value reflects the degree of difference between the original feature mapping vector and the enhanced mapping vector. If the loss value is small, it means that the model can maintain the consistency of feature extraction under data enhancement and has good robustness; if the loss value is large, it means that the model's image enhancement processing is not stable enough, and the model parameters need to be adjusted to optimize the feature extraction capability.
[0106] The process of calculating the loss value is usually based on a preset loss function, such as mean square error (MSE), cross entropy loss, or contrast loss. The loss value calculated by the loss function is back-propagated during the training process to adjust the parameters of the model.
[0107] Select a loss function suitable for feature contrast learning, such as contrast loss or mean square error loss function. Take the difference between the original feature mapping vector and the enhanced mapping vector as input and calculate the loss value. The loss value is used as the optimization target of the model to adjust the parameters of the feature encoding unit and mapping unit of the model.
[0108] Example description: In medical image analysis, the original feature mapping vector can represent the key features of the lesion area, while the enhanced mapping vector represents the characteristic performance of the lesion under different viewing angles or different image processing conditions. By comparing the original feature mapping vector with the enhanced mapping vector, the consistency of the model's recognition of lesion features under different viewing angles can be evaluated. For example, in the analysis of lung CT scan images, the loss value can be calculated to determine whether the model can consistently identify the lung nodule area in scanned images at different angles.
[0109] In financial bill recognition, the original feature mapping vector can represent the basic field information of the bill image, while the enhanced mapping vector represents the feature changes of the bill under different scanning conditions. By comparing these two feature mapping vectors, the model's recognition consistency of the key fields of the bill under different scanning conditions can be evaluated. For example, in automated invoice review, the loss value can be calculated to determine whether the model can accurately identify key fields such as invoice number and amount.
[0110] By comparing the original feature map vector with the enhanced map vector and calculating the loss value, the consistency and stability of the model in the feature extraction process can be effectively evaluated. The loss value calculation result can guide the optimization process of the model and help the model maintain the ability to accurately recognize image features under different viewing angles, different lighting conditions and different background noise.
[0111] S70, adjusting parameters of the first mapping unit and parameters of the first feature encoding unit according to the loss value;
[0112] In this embodiment, during the training process of the neural network model, the loss value is used to measure the error between the model output result and the target result. By adjusting the parameters of the model according to the loss value, the feature extraction and mapping capabilities of the model can be optimized. Specifically, the larger the loss value, the larger the prediction error of the model, and the more parameters need to be adjusted; the smaller the loss value, the closer the prediction result of the model is to the target result, and the adjustment range is relatively small.
[0113] In this process, the model adjusts the parameters of the first mapping unit and the first feature encoding unit according to the loss value. The parameters of the first mapping unit are mainly used to control the mapping process of the feature vector, while the parameters of the first feature encoding unit are used to control the extraction process of the image features. By synchronously adjusting the parameters of these two units, the consistency of feature extraction of the model on images with different viewpoints can be improved.
[0114] The difference between the original feature mapping vector and the enhanced mapping vector is input into a preset loss function to calculate the loss value. According to the loss value, the parameter gradient of the first mapping unit and the first feature encoding unit is calculated by the back propagation algorithm. According to the parameter gradient and the preset learning rate, the parameters of the first mapping unit and the first feature encoding unit are updated.
[0115] Backpropagation is the core mechanism of deep learning model training. The parameter gradient of each layer of the network is calculated through the loss value, and then transmitted back to each unit of the model to adjust the parameters and reduce the prediction error. In this step, the backpropagation algorithm calculates the parameter gradient of the first mapping unit and the parameter gradient of the first feature encoding unit according to the loss value, and then uses these gradient values to update the parameters.
[0116] The parameter gradients of the first mapping unit and the first feature encoding unit are calculated through the back propagation of the loss value. According to the preset learning rate, the amplitude of parameter adjustment is controlled to avoid the parameter adjustment being too fast or too slow. The adjusted parameters are applied to the first mapping unit and the first feature encoding unit to optimize the feature extraction and mapping capabilities of the model.
[0117] Example description: In medical imaging diagnosis scenarios, feature extraction and parameter adjustment are important steps to improve the accuracy of disease recognition. Assume that the model is used to analyze lung CT scan images to identify features such as the shape, location, and density of lung lesions. In order to adapt to different scanning conditions (such as different angles, light intensities, or imaging devices), the model needs to have robust feature extraction capabilities.
[0118] In specific applications, the model takes the lung CT image as the original view map input and generates enhanced view maps of different view angles through data enhancement, such as rotating, flipping or adding noise to the image. Next, the model extracts the feature vectors of each view map through the feature encoding unit and inputs these feature vectors into the mapping unit to generate a mapping vector.
[0119] In order to ensure the consistency of the model's recognition of the lesion area under different viewing angles, the system compares the original feature mapping vector with the enhanced mapping vector, calculates the difference and generates a loss value. If the model's recognition results under different viewing angles are very different, the loss value will be higher, and the model will adjust the parameters of the first mapping unit and the first feature encoding unit through the back propagation algorithm. Finally, after multiple rounds of training, the model can accurately identify the characteristics of the lung lesion area under different scanning conditions.
[0120] By adjusting the parameters of the first mapping unit and the first feature encoding unit according to the loss value, the prediction error of the model can be effectively reduced and the feature extraction ability of the model can be improved. By synchronously adjusting the parameters of these two units, the model can better learn the feature performance of the image at different perspectives, thereby improving the understanding ability and generalization performance of image data.
[0121] S80, adjusting parameters of the second feature encoding unit based on the adjusted parameters of the first mapping unit and the parameters of the first feature encoding unit;
[0122] In this embodiment, the parameters of the first mapping unit and the first feature encoding unit are gradually adjusted during the model training process through the back propagation of the loss value. The update of these parameters is to optimize the feature extraction and mapping capabilities of the model on images of different viewing angles. The adjusted parameters include the learning results of the model on the current training data, which can effectively improve the robustness and recognition accuracy of the model.
[0123] The results of these parameter adjustments not only affect the performance of the first mapping unit and the first feature encoding unit, but also affect the subsequent second feature encoding unit. In order to ensure the overall consistency of the model and the coherence of feature extraction, it is necessary to use the adjustment results of the first mapping unit and the first feature encoding unit to synchronously update the parameters of the second feature encoding unit.
[0124] The parameters of the first mapping unit and the first feature encoding unit are calculated and updated by the back propagation algorithm. The adjusted parameters are passed to the second feature encoding unit for subsequent parameter synchronization. It is ensured that the parameter update process of the second feature encoding unit is consistent with the adjustment results of the first two units.
[0125] The parameter adjustment of the second feature encoding unit is a synchronous update based on the parameter changes of the first two units. This synchronous adjustment mechanism usually adopts the method of momentum update or weighted average update to ensure that the second feature encoding unit can inherit the learning results of the first mapping unit and the first feature encoding unit, further improving the feature extraction ability of the model.
[0126] The core idea of momentum update is to take into account the parameter change trend of the previous round of training and smooth the parameter adjustment at a certain ratio to avoid excessive parameter update process, which may lead to model instability.
[0127] The update coefficient is calculated according to the preset momentum formula, which is usually determined by the model's learning rate and the number of batches. The momentum formula is as follows:
[0128] v t =β·v t-1 +(1-β)θ t
[0129] Where: v t represents the update coefficient (momentum update value) of the current step; β is the momentum coefficient, which usually takes a value between 0 and 1 (for example, 0.9) and is used to control the influence of historical gradients on the current update; v t-1 represents the momentum update value of the previous step; θ t Indicates the current model parameter value (eg, the parameter of the first mapping unit).
[0130] To calculate the momentum update coefficient, the ratio of the filter frequency to the sampling frequency needs to be taken into account. The calculation formula is as follows:
[0131]
[0132] Wherein: FiltFreq represents the filter frequency; SampleFreq represents the sampling frequency; e is the base of the natural logarithm, which is approximately equal to 2.718.
[0133] According to the ratio of the filtering frequency to the sampling frequency, the momentum coefficient β is calculated by a preset formula. According to the current parameter value θ t and the updated value v of the previous round t-1 , calculate the momentum update value v for the current step t . Using the calculated update value v tAdjust the model parameters to ensure that the model's feature extraction capabilities are optimized.
[0134] The adjustment parameters of the first mapping unit and the first feature encoding unit are weighted averaged according to the momentum coefficient, and the parameter value after the weighted average is assigned to the second feature encoding unit to achieve synchronous update of the parameters.
[0135] In the initial stage of sliding average, due to the small number of historical parameter values, the parameter update value may have a large deviation. In order to reduce the impact of this initial deviation, it is usually necessary to correct the momentum update value. The correction formula is as follows:
[0136]
[0137] Where: v biased is the updated value of the momentum after correction; v t is the current momentum update value; β is the momentum update coefficient; t is the current time step.
[0138] When the time step number t is larger, the correction factor 1-β t As it approaches 1, the difference between the updated values before and after the correction gradually decreases. This correction method can effectively avoid excessive deviation in parameter updates in the early stages of the model.
[0139] Get the current momentum update value v t and time step t; calculate the correction factor 1-β t ; Use the modified formula Calculate the corrected momentum update value; output the corrected momentum update value v biased .
[0140] By adjusting the parameters of the second feature encoding unit based on the parameters of the first mapping unit and the parameters of the first feature encoding unit, the feature extraction capability and robustness of the model can be effectively improved. By synchronously updating the parameters of the three units, the model can maintain the consistency and accuracy of image feature recognition under different viewing angles and different data enhancement conditions.
[0141] S90: Extract a feature representation of the target image based on the adjusted second feature encoding unit.
[0142] In this embodiment, the second feature coding unit has been parameter adjusted by momentum updating. The adjusted second feature coding unit has a stronger feature extraction capability and can maintain the consistency and robustness of feature extraction under different image enhancement perspectives. Based on this unit, the feature representation of the target image can be extracted to effectively capture the key information in the target image and provide support for subsequent tasks (such as classification, target detection, etc.). The second feature coding unit can adopt a convolutional neural network (CNN) structure, which is divided into multiple convolutional layers and pooling layers. When extracting feature representation, the unit scans various areas of the input image to extract basic features such as edges, textures, shapes, and more advanced semantic features.
[0143] The target image refers to the image that is input into the adjusted second feature encoding unit for feature extraction after the model is trained.
[0144] The target image can vary according to different application scenarios and is not limited to a specific type of image. It can be either an original unprocessed image or an image that has been preprocessed or enhanced.
[0145] For example, in medical image analysis, the target image can be a CT scan image, an X-ray, an MRI image, etc. The model extracts the feature representation of the target image and identifies the shape, position, density change and other information of the lesion area. The types of target images can include lung CT scan images (for identifying lung nodules), heart MRI images (for identifying abnormal heart structures) and bone X-rays (for identifying fractures or osteoporosis). In lung nodule recognition, the target image is a CT scan image of the patient. The model extracts image features through the second feature encoding unit, identifies the size, shape and density change of lung nodules, and assists doctors in diagnosis.
[0146] In financial bill recognition, the target image can be bill images such as invoices, receipts, contracts, and ID card photos. The model extracts the feature representation of the target image and identifies key fields in the bill, such as invoice number, amount, date, etc. The types of target images can include value-added tax invoices (for automated review), receipt or contract scans (for information extraction and verification), and ID card photos (for identity verification). In automated invoice review, the target image is a scanned copy of the invoice. The model extracts the feature representation of the invoice through the second feature encoding unit, identifies the invoice number, invoicing date, amount, and other information, and reduces the workload of manual review.
[0147] In security monitoring, the target image can be a real-time image or video frame captured by a camera. The model extracts the feature representation of the target image and identifies information such as people, vehicles, and objects in the scene. The types of target images can include face images (for identity authentication), vehicle images (for license plate recognition and vehicle tracking), and public place surveillance images (for behavior recognition and anomaly detection)
[0148] Input the target image data into the second feature encoding unit. Use the convolution kernel to scan various areas of the target image and extract features such as edges and textures. Use the pooling layer to reduce the dimension of the feature map, retain the most important feature information, and reduce redundant data. After the convolution and pooling operations, apply the activation function for nonlinear transformation to enhance the expression ability of the model. Output the extracted feature representation as a fixed-length vector or tensor for subsequent tasks.
[0149] The feature representation of the target image is the output result after the image is processed by the second feature encoding unit, usually a vector of fixed length. This vector contains the core feature information of the image, such as edges, textures, shapes, color distribution, etc. The extracted feature representation can be used for a variety of tasks, such as image classification, object detection, image retrieval, etc.
[0150] Compared with the original image data, feature representation can express the key information of the image in a more concise way, reduce computational complexity, and improve the reasoning efficiency of the model. At the same time, this feature representation is more robust to changes such as image rotation, flipping, and noise.
[0151] The feature information of the target image is converted into a high-dimensional feature vector. The extracted feature representation is normalized to ensure that the vector value is within a reasonable range and to avoid gradient explosion or disappearance. The normalized feature representation is used as the final output for subsequent model use.
[0152] By extracting the feature representation of the target image based on the adjusted second feature encoding unit, the model's feature extraction capability can be effectively improved, and the model's ability to understand and express images can be enhanced. Compared with directly processing the original image data, the feature representation is more concise and accurate, which can reduce the computational complexity and improve the model's reasoning efficiency.
[0153] The present invention relates to the fields of artificial intelligence technology and medical health, and discloses an image feature extraction method, including: acquiring image data and generating an original viewing angle map and an enhanced viewing angle map, encoding the original viewing angle map and the enhanced viewing angle map respectively, generating an original feature vector and an enhanced feature vector, and then generating a feature mapping vector. By comparing the difference between the feature mapping vectors, the loss value is calculated, the parameters of the feature coding unit are adjusted, and the feature representation of the target image is extracted based on the adjusted feature coding unit. The present invention introduces a self-supervised mapping learning mechanism in the image feature extraction process, without relying on manually labeled data and negative sample selection, and directly generates the original viewing angle map and the enhanced viewing angle map through the image data itself, thereby effectively extracting the key features of the same type of samples; by adjusting the parameters of the feature coding unit and the mapping unit, the accuracy and generalization ability of the model in image recognition tasks are improved.
[0154] In one embodiment, the above S10 includes:
[0155] S101, acquiring image data, and using the image data as an original viewing angle image;
[0156] S102, performing cropping processing on the image data to generate a cropped enhanced viewing angle image;
[0157] S103, or flipping the image data horizontally or vertically to generate a flipped enhanced viewing angle image;
[0158] S104, or adding Gaussian noise to the image data to generate an enhanced viewing angle map after noise enhancement.
[0159] In this embodiment, the image data is the raw data input by the model, which can usually be medical images, bill scans, surveillance images, or product images. In the initial stage of data processing, the acquired image data is directly processed as the original view map. The original view map does not require additional data enhancement operations and is used by the model to extract the basic features of the image as a positive sample for model training and comparative learning.
[0160] Image data can be captured by a camera, collected by a scanning device, or imported from an image database. Perform preprocessing operations such as format conversion and size adjustment on the image data to ensure that the image data meets the input requirements of the model. The preprocessed image data is directly used as the original view map for the feature encoding unit to extract feature representation.
[0161] Cropping is one of the common methods for image data enhancement. It generates images with different perspectives by cropping the edge areas of the original image. The cropped enhanced perspective map can help the model identify local features in the image and improve the robustness and generalization ability of the model in feature extraction.
[0162] Randomly select the edge area of the image and crop a certain percentage of pixels. The cropped image is used as an enhanced view map for the model's comparative learning process. Cropping is suitable for medical image analysis (such as cropping specific areas of CT images) and bill recognition (such as cropping key field areas of invoices).
[0163] Horizontal or vertical flipping is a simple but effective data augmentation method that can increase the model's ability to adapt to changes in image orientation. Flipping does not change the core feature information of the image, but it changes the perspective of the image, thereby helping the model learn feature representations in different directions.
[0164] Select flip direction: Randomly select horizontal flip or vertical flip. Perform the selected flip operation on the original image data to generate a new enhanced view. Flip processing is suitable for scenarios such as face recognition, medical image analysis (such as analysis of symmetrical structures), and product recognition.
[0165] Gaussian noise is a common random noise that simulates image interference that may occur in real scenes, such as light changes during shooting, equipment failures, etc. By adding Gaussian noise to the original image data, the model's robustness to image noise can be enhanced, improving the model's performance in actual application scenarios.
[0166] Generate a random noise matrix based on Gaussian distribution. Superimpose the generated noise matrix on the original image data to generate an enhanced view with noise. Adding noise processing is suitable for medical image analysis (simulating noise interference of imaging equipment) and security monitoring image analysis (processing low-light or blurred images).
[0167] This embodiment can effectively improve the model's ability to extract image features from different viewing angles and in different noise environments by cropping, flipping, and adding noise to the original image data, thereby enhancing the model's robustness and generalization ability and reducing recognition errors caused by data changes.
[0168] In one embodiment, the above S40 includes:
[0169] S401, inputting the original feature vector and the enhanced feature vector into a first mapping unit, and performing normalization processing on the original feature vector and the enhanced feature vector through the first mapping unit;
[0170] S402: Input the normalized original feature vector and enhanced feature vector into the fully connected layer of the first mapping unit, and perform linear transformation and activation function processing respectively to generate an original mapping vector and an enhanced mapping vector.
[0171] In this embodiment, the original feature vector and the enhanced feature vector are image feature representations extracted by the first feature encoding unit and the second feature encoding unit, respectively. In order to further process these feature representations, they need to be input into the first mapping unit for feature mapping operation.
[0172] The main function of the first mapping unit is to transform the dimension and abstract the features of the input feature vector, and generate a more expressive mapping vector through linear transformation and nonlinear activation. This step can help the model learn deeper feature relationships and improve the feature extraction effect of the model.
[0173] The original feature vector and the enhanced feature vector are simultaneously input into the first mapping unit. It is ensured that the input data format meets the input requirements of the first mapping unit, such as checking the vector dimension and data type.
[0174] Normalization is one of the key steps in feature mapping, which can standardize the input feature vector value to a specific range (such as between 0 and 1). The purpose of normalization is to reduce the scale difference of the feature value, avoid the problem of gradient explosion or gradient disappearance during the training process of the model, and improve the convergence speed and stability of the model.
[0175] Batch normalization is performed on the original feature vectors and enhanced feature vectors to make the feature values of each batch more evenly distributed. Normalization is suitable for various image feature extraction tasks, such as medical image analysis, financial bill recognition, etc.
[0176] The fully connected layer is the basic layer in the neural network, which is used to perform linear transformation on the input data and map the high-dimensional feature vector to a new feature space. The normalized feature vector is input to the fully connected layer, which can further perform linear combination and feature abstraction on the features. Matrix multiplication is performed on the normalized feature vector to generate a new feature representation. A bias term is added to each feature vector to enhance the expressiveness of the model.
[0177] After completing the linear transformation, the output result needs to be processed by an activation function to introduce nonlinear characteristics and improve the model's ability to express complex feature relationships. Activation functions are commonly used such as ReLU (Rectified Linear Unit) or Sigmoid.
[0178] Apply activation function to the linearly transformed feature vector to enhance the nonlinear expression ability of the model. After linear transformation and activation function processing, the original mapping vector and enhanced mapping vector are generated respectively for subsequent feature comparison and loss calculation.
[0179] This embodiment can improve the feature extraction capability and robustness of the model by inputting the original feature vector and the enhanced feature vector into the first mapping unit to generate the original mapping vector and the enhanced mapping vector. It can not only learn the basic features of the image, but also capture the deep feature relationship of the image, providing data support for subsequent loss calculation and parameter adjustment.
[0180] In one embodiment, the above S50 includes:
[0181] S501, inputting the original mapping vector into a second mapping unit, and performing dimensionality reduction processing on the original mapping vector through the second mapping unit;
[0182] S502, inputting the original mapping vector after dimension reduction processing into the fully connected layer of the second mapping unit, and performing linear transformation processing on the original mapping vector after dimension reduction;
[0183] S503, applying an activation function to the original mapping vector that has been processed by the linear transformation to generate the original feature mapping vector.
[0184] In this embodiment, the original mapping vector is a feature vector output by the first mapping unit, which may have a high dimension and contain a large amount of feature information. Before further feature processing, the original mapping vector is subjected to dimensionality reduction processing, which can reduce the dimension of the vector, remove redundant information, and retain the most representative features. This can not only reduce the computational complexity, but also improve the training efficiency and generalization ability of the model.
[0185] The original mapping vector is input to the input layer of the second mapping unit. Through dimensionality reduction methods such as principal component analysis (PCA) or linear discriminant analysis (LDA), the dimension of the feature vector is reduced and the main feature information is retained. A feature vector with lower dimension but higher information density is generated for subsequent processing.
[0186] The original mapping vector after dimensionality reduction needs to be further transformed by the fully connected layer to map the low-dimensional feature vector to the new feature space. The purpose of linear transformation is to recombine feature information to enhance the model's expressiveness and help the model learn more abstract feature relationships.
[0187] Perform matrix multiplication on the original mapping vector after dimensionality reduction to generate a new feature representation. Add a bias term to each feature vector to enhance the feature expression ability of the model.
[0188] Activation function processing is an important step in introducing nonlinear characteristics into neural networks. After performing linear transformation, if activation function processing is not applied, the model can only learn linear features and is difficult to adapt to complex feature relationships. By applying activation functions, the model can have stronger feature learning capabilities.
[0189] Common activation functions include ReLU, Sigmoid, Tanh, etc. Select a suitable activation function based on the actual needs of the model. Apply the activation function to the linearly transformed vectors one by one to generate nonlinear feature representations. The vector processed by the activation function is the original feature mapping vector for use in subsequent tasks.
[0190] When the original mapping vector is input into the second mapping unit, the dimensionality reduction process can be further refined, and the accuracy and information density of the feature mapping can be improved through multi-level dimensionality reduction and feature fusion operations. Each level of dimensionality reduction uses a different weight matrix for transformation, and feature fusion is performed between each level to retain more key features.
[0191] The original mapping vector is input into the second mapping unit; a multi-level dimensionality reduction process is performed in the second mapping unit, and a different weight matrix is used for each level of dimensionality reduction; a feature fusion operation is performed on the vector after each level of dimensionality reduction, and the vectors after dimensionality reduction are concatenated or weighted averaged; the fused vector is input into the fully connected layer of the second mapping unit, and a linear transformation is performed; an activation function is applied to the vector after the linear transformation to generate the original feature mapping vector.
[0192] This embodiment can effectively improve the feature extraction capability of the model by inputting the original mapping vector into the second mapping unit through dimensionality reduction, linear transformation and activation function processing. Through this feature mapping process, the model can learn more abstract and efficient feature representation, reduce computational complexity, and improve recognition accuracy and robustness.
[0193] In one embodiment, the above S60 includes:
[0194] S601, performing a similarity comparison between the original feature mapping vector and the enhanced mapping vector to determine a similarity score between the original feature mapping vector and the enhanced mapping vector;
[0195] S602, determining the difference between the original feature mapping vector and the enhanced mapping vector based on the similarity score;
[0196] S603: Based on the difference between the original feature mapping vector and the enhanced mapping vector, determine the loss value according to a preset loss function.
[0197] In this embodiment, the similarity comparison is to measure two feature mapping vectors to evaluate their similarity in the feature space. The similarity score is usually calculated using methods such as cosine similarity or Euclidean distance. By calculating the similarity score, it can be evaluated whether the feature extraction effect of the model on images with different viewing angles is consistent.
[0198] The similarity can be calculated by the following formula:
[0199]
[0200] Where: A represents the original feature mapping vector; B represents the enhanced mapping vector; ||A|| and ||B|| represent the modulus of the vector; A·B represents the dot product of two vectors.
[0201] Get the original feature map vector and the enhanced map vector; compare the two using methods such as cosine similarity or Euclidean distance; output the similarity score for subsequent difference calculation.
[0202] Difference is a complementary measure of similarity, which is used to measure the degree of difference between two feature mapping vectors. The smaller the difference, the closer the two vectors are, and the better the feature extraction effect of the model; the larger the difference, the greater the difference between the two vectors, and the need for further optimization of the model.
[0203] The difference can be calculated by the following formula:
[0204] Difference(A,B)=1-Cosine Similarity(A,B)
[0205] Calculate the difference through the similarity score; determine whether the feature extraction effect of the model needs to be optimized based on the difference; output the difference value for loss value calculation.
[0206] The loss value is a key indicator in the model training process, which is used to measure the error between the model's prediction results and the target results. The smaller the loss value, the better the feature extraction effect of the model. Otherwise, the model parameters need to be adjusted.
[0207] Commonly used loss functions include mean square error (MSE) or contrastive loss function (Contrastive Loss):
[0208] The formula for calculating the mean square error is:
[0209]
[0210] Where: d i represents the difference between each pair of feature vectors; n represents the number of samples.
[0211] The calculation formula of the contrast loss function is:
[0212] Contrastive Loss = y·d 2 +(1-y)·max(0,md) 2
[0213] Among them: y is the label, indicating whether the two vectors are a positive sample pair; d represents the difference; m is the maximum difference threshold.
[0214] In the process of lung nodule recognition, the model needs to compare the feature mapping vectors generated by the original view map of the lung CT scan image and the enhanced view map to measure the recognition consistency of the model under different enhancement conditions (such as contrast enhancement, noise simulation, etc.). However, relying solely on a single similarity (such as cosine similarity) may not be able to fully capture the subtle differences in the edges and textures of lung nodules.
[0215] In addition to calculating the cosine similarity, the Euclidean distance, Manhattan distance, etc. are also calculated to obtain multiple similarity measurement results.
[0216] In medical scenarios, some lung nodules are polygonal or irregular in shape. Different similarity metrics may have different sensitivities to nodule morphology. By fusing multiple metrics, the consistency of nodule regions can be more comprehensively measured. In practice, learnable weights can be assigned to different metrics, and a small amount of annotated nodule data can be used for fine-tuning, thereby automatically optimizing the fusion ratio of each similarity.
[0217] In order to better identify the benign and malignant nature of lung nodules, in addition to measuring the difference between the two mapping vectors, an auxiliary classification task can also be introduced: the mapping vector is input into a small classifier to determine the "benign / malignant" label, and the error of its prediction result is included in the comprehensive calculation of the difference.
[0218] The difference can be defined as:
[0219] D=α·(1-S f )+(1-α)·ClassificationError
[0220] Where S f It is the comprehensive score obtained by fusing multiple similarities, ClassificationError represents the prediction error of the classifier, and α is the weight controlling the ratio of the two.
[0221] In this way, the model maintains feature consistency while also taking into account the ability to distinguish the pathological classification of lung nodules.
[0222] The final loss function can be based on the traditional MSE or contrast loss, and superimpose the error term caused by this part of auxiliary classification:
[0223] FinalLoss=BaseLoss(D)+β·ClassificationLoss
[0224] Among them, BaseLoss (D) represents the base loss, which is usually used to measure the performance of the model on feature extraction tasks; D here represents the difference, which is used to evaluate the difference between the original feature mapping vector and the enhanced mapping vector. ClassificationLoss represents the classification loss, which is used to measure the performance of the model on auxiliary classification tasks (such as benign / malignant tumor judgment, lesion area identification, etc.). β represents the weight coefficient, which is used to control the proportion of base loss and classification loss in the total loss.
[0225] In medical image analysis, this multi-task fusion can help the model better balance the two goals of "enhancing perspective consistency" and "benign / malignant classification" when extracting lung nodule features.
[0226] This embodiment can effectively measure the feature extraction effect of the model by calculating the loss value based on the similarity and difference between the original feature mapping vector and the enhanced mapping vector. Adjusting the model parameters according to the loss value can improve the consistency of feature extraction of the model under different perspectives, reduce recognition errors, and improve the robustness and accuracy of the model.
[0227] In one embodiment, the above S70 includes:
[0228] S701, determining a parameter gradient of the first mapping unit and a parameter gradient of the first feature encoding unit based on the loss value;
[0229] S702, back-propagating the parameter gradient of the first mapping unit and the parameter gradient of the first feature encoding unit to a parameter updating module;
[0230] S703: determining a parameter adjustment value of the first mapping unit and a parameter adjustment value of the first feature encoding unit based on a preset learning rate according to the parameter gradient of the first mapping unit and the parameter gradient of the first feature encoding unit;
[0231] S704: Based on the parameter adjustment value of the first mapping unit and the parameter adjustment value of the first feature encoding unit, update the parameters of the first mapping unit and the parameters of the first feature encoding unit through the parameter updating module.
[0232] In this embodiment, in the deep learning model, the parameter gradient is the rate of change of the loss value to the model parameter, which is used to guide the update direction of the parameter. Through the back propagation algorithm, the gradient of each parameter is calculated according to the loss value to determine which parameters need to be increased and which need to be reduced, so that the prediction error of the model is gradually reduced.
[0233] Back propagation is performed on the parameters of the first mapping unit and the first feature encoding unit, and the gradient of each parameter with respect to the loss value is calculated; and the parameter gradients of the first mapping unit and the first feature encoding unit are output.
[0234] Backpropagation is the core algorithm of deep learning models. It guides the update of parameters in each layer by propagating parameter gradients from the output layer to the input layer layer by layer. The parameter update module is an independent module that receives the gradients of backpropagation and calculates the adjustment values of parameters. Obtain the parameter gradients of the first mapping unit and the first feature encoding unit; propagate these gradient values forward from the output layer to the parameter update module; the parameter update module prepares to adjust the parameters according to the gradient values.
[0235] The learning rate is a hyperparameter that controls the step size of each parameter update. If the learning rate is too high, the model may not converge; if it is too low, the model training speed will slow down. Get the parameter gradients of the first mapping unit and the first feature encoding unit; calculate the adjustment value of each parameter according to the preset learning rate; output the adjustment value for the next parameter update.
[0236] Parameter updating is a key step in deep learning models. The calculated parameter adjustment values are applied to the model's parameters so that the model's prediction error gradually decreases. Parameter updating is usually done using optimization algorithms (such as SGD, Adam, etc.).
[0237] The parameter updating module receives the parameter adjustment values of the first mapping unit and the first feature encoding unit; updates the corresponding model parameters according to the adjustment values; and applies the updated parameters to the next round of training to improve the feature extraction capability of the model.
[0238] This embodiment can effectively improve the consistency and robustness of the model's feature extraction by adjusting the parameters of the first mapping unit and the first feature encoding unit according to the loss value. The back propagation of the parameter gradient is combined with the parameter update based on the learning rate to make the model's feature extraction results under different viewing angles more accurate, thereby improving the model's training effect and recognition accuracy.
[0239] In one embodiment, the above S80 includes:
[0240] S801, determining a momentum update coefficient;
[0241] S802, performing weighted averaging on the adjusted parameters of the first mapping unit and the parameters of the first feature encoding unit according to the momentum update coefficient to obtain a weighted average parameter value;
[0242] S803: Update the parameters of the second feature encoding unit according to the weighted average parameter value, wherein the parameters of the second feature encoding unit are updated only based on the parameters of the first mapping unit and the parameters of the first feature encoding unit.
[0243] In this embodiment, the momentum update coefficient is a coefficient used to control the "historical inertia" in the model parameter update process, which can effectively smooth the parameter update process and avoid the oscillation problem of parameter adjustment being too large or too small. The calculation of the momentum update coefficient is usually based on a preset formula, and the commonly used formula is:
[0244]
[0245] Where: β represents the momentum update coefficient; FiltFreq represents the filter frequency; SampleFreq represents the sampling frequency.
[0246] The value of the momentum update coefficient is usually between 0 and 1. The closer the value is to 1, the more the model relies on historical parameter values, and the closer the value is to 0, the more the model tends to update the current gradient.
[0247] Weighted average is to calculate the combined value of two parameter sets by assigning different weights to them. During the momentum update process, the model combines the historical parameter values with the current gradient update values to prevent the model from excessively deviating in a certain direction due to a single update. Weighted average formula:
[0248] θ t+1 =β·θ t +(1-β)·θ t '
[0249] Where: θ t+1 represents the parameter value after weighted average; β represents the momentum update coefficient; θ t represents the current model parameter value; θ t ′ represents the parameter value of the current gradient update.
[0250] The parameter update of the second feature encoding unit only depends on the parameter adjustment results of the first mapping unit and the first feature encoding unit. This update strategy can ensure that the parameter change of the second feature encoding unit is based on the training results of the first stage, thereby maintaining the consistency of feature extraction of the entire model.
[0251] Obtaining weighted averaged parameter values; using these parameter values to replace current parameters of the second feature encoding unit; ensuring that parameter update of the second feature encoding unit does not depend on other external information, but is only based on parameter update results of the first mapping unit and the first feature encoding unit.
[0252] This embodiment can effectively improve the consistency and robustness of the model's feature extraction by adjusting the parameters of the second feature encoding unit based on the momentum update mechanism. The introduction of the momentum update mechanism makes the model's parameter update more stable, avoiding the parameter oscillation problem caused by fluctuations in single training data, thereby improving the model's convergence speed and generalization ability.
[0253] In one embodiment, an image feature extraction device is provided, and the image feature extraction device corresponds one-to-one to the image feature extraction method in the above embodiment. Figure 3 , Figure 3 This is a functional module diagram of a preferred embodiment of the image feature extraction device of the present invention. Data preprocessing module 10, original encoding module 20, enhanced encoding module 30, first feature mapping module 40, second feature mapping module 50, loss determination module 60, parameter adjustment module 70, model updating module 80 and feature extraction module 90. The functional modules are described in detail as follows:
[0254] A data preprocessing module 10 is used to obtain image data and generate an original viewing angle map and an enhanced viewing angle map according to the image data;
[0255] The original encoding module 20 is used to encode the original view image through a first feature encoding unit to generate an original feature vector;
[0256] An enhanced coding module 30, configured to perform coding processing on the enhanced viewing angle map through a second feature coding unit to generate an enhanced feature vector;
[0257] A first feature mapping module 40, configured to input the original feature vector and the enhanced feature vector into a first mapping unit to generate an original mapping vector and an enhanced mapping vector respectively;
[0258] A second feature mapping module 50, configured to input the original mapping vector into a second mapping unit to generate an original feature mapping vector;
[0259] A loss determination module 60, configured to determine a loss value based on the original feature mapping vector and the enhanced mapping vector;
[0260] A parameter adjustment module 70, configured to adjust parameters of the first mapping unit and parameters of the first feature encoding unit according to the loss value;
[0261] A model updating module 80, configured to adjust parameters of the second feature encoding unit based on the adjusted parameters of the first mapping unit and the parameters of the first feature encoding unit;
[0262] The feature extraction module 90 is used to extract the feature representation of the target image based on the adjusted second feature encoding unit.
[0263] In one embodiment, the data preprocessing module 10 is specifically used for:
[0264] Acquire image data, and use the image data as an original viewing angle image;
[0265] Performing cropping processing on the image data to generate a cropped enhanced viewing angle image;
[0266] Or flipping the image data horizontally or vertically to generate a flipped enhanced viewing angle image;
[0267] Alternatively, Gaussian noise is added to the image data to generate an enhanced viewing angle map after noise enhancement.
[0268] In one embodiment, the first feature mapping module 40 is specifically configured to:
[0269] Inputting the original feature vector and the enhanced feature vector into a first mapping unit, and performing normalization processing on the original feature vector and the enhanced feature vector by the first mapping unit;
[0270] The normalized original feature vector and enhanced feature vector are input into the fully connected layer of the first mapping unit, and are processed by linear transformation and activation function respectively to generate the original mapping vector and enhanced mapping vector.
[0271] In one embodiment, the second feature mapping module 50 is specifically configured to:
[0272] Inputting the original mapping vector into a second mapping unit, and performing dimensionality reduction processing on the original mapping vector through the second mapping unit;
[0273] Inputting the original mapping vector after dimension reduction processing into the fully connected layer of the second mapping unit, and performing linear transformation processing on the original mapping vector after dimension reduction;
[0274] Applying an activation function to the original mapping vector after the linear transformation process to generate the original feature mapping vector.
[0275] In one embodiment, the loss determination module 60 is specifically configured to:
[0276] Comparing the original feature mapping vector with the enhanced mapping vector for similarity, and determining a similarity score between the original feature mapping vector and the enhanced mapping vector;
[0277] Based on the similarity score, determining a difference between the original feature map vector and the enhanced map vector;
[0278] Based on the difference between the original feature mapping vector and the enhanced mapping vector, the loss value is determined according to a preset loss function.
[0279] In one embodiment, the parameter adjustment module 70 is specifically used to:
[0280] Based on the loss value, determining a parameter gradient of the first mapping unit and a parameter gradient of the first feature encoding unit;
[0281] Back-propagating the parameter gradient of the first mapping unit and the parameter gradient of the first feature encoding unit to a parameter updating module;
[0282] Determining a parameter adjustment value of the first mapping unit and a parameter adjustment value of the first feature encoding unit based on a preset learning rate according to the parameter gradient of the first mapping unit and the parameter gradient of the first feature encoding unit;
[0283] Based on the parameter adjustment value of the first mapping unit and the parameter adjustment value of the first feature encoding unit, the parameters of the first mapping unit and the parameters of the first feature encoding unit are updated by the parameter updating module.
[0284] In one embodiment, the model updating module 80 is specifically configured to:
[0285] Determine the momentum update coefficient;
[0286] Performing weighted averaging on the adjusted parameters of the first mapping unit and the parameters of the first feature encoding unit according to the momentum update coefficient to obtain a weighted average parameter value;
[0287] The parameters of the second feature encoding unit are updated according to the weighted average parameter value, and the parameters of the second feature encoding unit are updated only based on the parameters of the first mapping unit and the parameters of the first feature encoding unit.
[0288] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, a memory, a network interface and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of the server side of an image feature extraction method.
[0289] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a user side of an image feature extraction method
[0290] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the following steps are implemented:
[0291] Acquire image data, and generate an original viewing angle map and an enhanced viewing angle map according to the image data;
[0292] The original view image is encoded by a first feature encoding unit to generate an original feature vector;
[0293] The enhanced viewing angle map is encoded by a second feature encoding unit to generate an enhanced feature vector;
[0294] Inputting the original feature vector and the enhanced feature vector into a first mapping unit to generate an original mapping vector and an enhanced mapping vector respectively;
[0295] Inputting the original mapping vector into a second mapping unit to generate an original feature mapping vector;
[0296] Determine a loss value based on the original feature mapping vector and the enhanced mapping vector;
[0297] Adjusting parameters of the first mapping unit and parameters of the first feature encoding unit according to the loss value;
[0298] adjusting parameters of the second feature encoding unit based on the adjusted parameters of the first mapping unit and the parameters of the first feature encoding unit;
[0299] A feature representation of the target image is extracted based on the adjusted second feature encoding unit.
[0300] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, the following steps are implemented:
[0301] Acquire image data, and generate an original viewing angle map and an enhanced viewing angle map according to the image data;
[0302] The original view image is encoded by a first feature encoding unit to generate an original feature vector;
[0303] The enhanced viewing angle image is encoded by a second feature encoding unit to generate an enhanced feature vector;
[0304] Inputting the original feature vector and the enhanced feature vector into a first mapping unit to generate an original mapping vector and an enhanced mapping vector respectively;
[0305] Inputting the original mapping vector into a second mapping unit to generate an original feature mapping vector;
[0306] Determine a loss value based on the original feature mapping vector and the enhanced mapping vector;
[0307] adjusting parameters of the first mapping unit and parameters of the first feature encoding unit according to the loss value;
[0308] adjusting parameters of the second feature encoding unit based on the adjusted parameters of the first mapping unit and the parameters of the first feature encoding unit;
[0309] A feature representation of the target image is extracted based on the adjusted second feature encoding unit.
[0310] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can refer to the relevant descriptions on the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0311] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0312] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0313] It should be noted that if software tools or components other than those of the Company appear in the embodiments of the present application, they are only used for illustration and do not represent actual use. The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the above-mentioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the above-mentioned embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents; and these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for extracting image features, characterized in that: The following steps are involved: Acquire image data, and generate an original viewing angle map and an enhanced viewing angle map according to the image data; The original view image is encoded by a first feature encoding unit to generate an original feature vector; The enhanced viewing angle map is encoded by a second feature encoding unit to generate an enhanced feature vector; Inputting the original feature vector and the enhanced feature vector into a first mapping unit to generate an original mapping vector and an enhanced mapping vector respectively; Inputting the original mapping vector into a second mapping unit to generate an original feature mapping vector; Determine a loss value based on the original feature mapping vector and the enhanced mapping vector; adjusting parameters of the first mapping unit and parameters of the first feature encoding unit according to the loss value; adjusting parameters of the second feature encoding unit based on the adjusted parameters of the first mapping unit and the parameters of the first feature encoding unit; A feature representation of the target image is extracted based on the adjusted second feature encoding unit.
2. The image feature extraction method according to claim 1, characterized in that: Acquiring image data, and generating an original viewing angle map and an enhanced viewing angle map according to the image data, including: Acquire image data, and use the image data as an original viewing angle image; Performing cropping processing on the image data to generate a cropped enhanced viewing angle image; Or flipping the image data horizontally or vertically to generate a flipped enhanced viewing angle image; Alternatively, Gaussian noise is added to the image data to generate an enhanced viewing angle map after noise enhancement.
3. The image feature extraction method according to claim 1, characterized in that: Inputting the original feature vector and the enhanced feature vector into a first mapping unit to generate an original mapping vector and an enhanced mapping vector respectively, comprising: Inputting the original feature vector and the enhanced feature vector into a first mapping unit, and performing normalization processing on the original feature vector and the enhanced feature vector by the first mapping unit; The normalized original feature vector and enhanced feature vector are input into the fully connected layer of the first mapping unit, and are processed by linear transformation and activation function respectively to generate the original mapping vector and enhanced mapping vector.
4. The image feature extraction method according to claim 1, characterized in that: Inputting the original mapping vector into a second mapping unit to generate an original feature mapping vector includes: Inputting the original mapping vector into a second mapping unit, and performing dimensionality reduction processing on the original mapping vector through the second mapping unit; Inputting the original mapping vector after dimension reduction processing into the fully connected layer of the second mapping unit, and performing linear transformation processing on the original mapping vector after dimension reduction; Applying an activation function to the original mapping vector after the linear transformation process to generate the original feature mapping vector.
5. The image feature extraction method according to claim 1, characterized in that: Determining a loss value based on the original feature mapping vector and the enhanced mapping vector includes: Comparing the original feature mapping vector with the enhanced mapping vector for similarity, and determining a similarity score between the original feature mapping vector and the enhanced mapping vector; Based on the similarity score, determining a difference between the original feature map vector and the enhanced map vector; Based on the difference between the original feature mapping vector and the enhanced mapping vector, the loss value is determined according to a preset loss function.
6. The image feature extraction method according to claim 1, characterized in that: Adjusting parameters of the first mapping unit and parameters of the first feature encoding unit according to the loss value includes: Based on the loss value, determining a parameter gradient of the first mapping unit and a parameter gradient of the first feature encoding unit; Back-propagating the parameter gradient of the first mapping unit and the parameter gradient of the first feature encoding unit to a parameter updating module; Determining a parameter adjustment value of the first mapping unit and a parameter adjustment value of the first feature encoding unit based on a preset learning rate according to the parameter gradient of the first mapping unit and the parameter gradient of the first feature encoding unit; Based on the parameter adjustment value of the first mapping unit and the parameter adjustment value of the first feature encoding unit, the parameters of the first mapping unit and the parameters of the first feature encoding unit are updated by the parameter updating module.
7. The image feature extraction method according to claim 1, characterized in that: Adjusting the parameters of the second feature encoding unit based on the adjusted parameters of the first mapping unit and the parameters of the first feature encoding unit includes: Determine the momentum update coefficient; Performing weighted averaging on the adjusted parameters of the first mapping unit and the parameters of the first feature encoding unit according to the momentum update coefficient to obtain a weighted average parameter value; The parameters of the second feature encoding unit are updated according to the weighted average parameter value, and the parameters of the second feature encoding unit are updated only based on the parameters of the first mapping unit and the parameters of the first feature encoding unit.
8. An image feature extraction device, characterized in that: The image feature extraction device comprises: A data preprocessing module, used to obtain image data and generate an original viewing angle map and an enhanced viewing angle map according to the image data; An original encoding module, used for encoding the original view image through a first feature encoding unit to generate an original feature vector; an enhanced coding module, configured to perform coding processing on the enhanced viewing angle map through a second feature coding unit to generate an enhanced feature vector; A first feature mapping module, used for inputting the original feature vector and the enhanced feature vector into a first mapping unit to generate an original mapping vector and an enhanced mapping vector respectively; A second feature mapping module, used for inputting the original mapping vector into a second mapping unit to generate an original feature mapping vector; A loss determination module, configured to determine a loss value based on the original feature mapping vector and the enhanced mapping vector; a parameter adjustment module, used for adjusting the parameters of the first mapping unit and the parameters of the first feature encoding unit according to the loss value; a model updating module, configured to adjust parameters of the second feature encoding unit based on the adjusted parameters of the first mapping unit and the parameters of the first feature encoding unit; A feature extraction module is used to extract a feature representation of a target image based on the adjusted second feature encoding unit.
9. A computer device, characterized in that: The computer device includes a memory, a processor, and an image feature extraction program stored in the memory and executable on the processor. When the image feature extraction program is executed by the processor, the steps of the image feature extraction method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The storage medium stores an image feature extraction program, which, when executed by a processor, implements the steps of the image feature extraction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Aviation plunger pump oil distribution disc abrasion detection method and system based on confrontation self-supervision
CN116070126A
Hyperspectral classification method based on band erasure and comparative learning
CN116681946A
Electrocardio multi-mode contrast learning method for single-lead arrhythmia diagnosis
CN118333130A