Label position feature fusion method based on attention mechanism and coordinate transformation

By introducing the fusion method of label position feature of attention mechanism and coordinate transformation in the image recognition model, the problem of poor modeling of target spatial relationships in traditional models is solved, and the accuracy of recognition of target positions in complex scenarios is improved.

CN120164064AInactive Publication Date: 2025-06-17YANCHENG INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510216535.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional deep learning models do not consider image-label correlation constraints when classifying multi-label images, resulting in poor modeling capabilities for target spatial relationships, and there are target position noise problems in complex scenarios, which affects the recognition accuracy of image recognition models.

Method used

A method of fusion of label position features based on attention mechanism and coordinate transformation is proposed. By introducing affine transformation of target coordinates, the modeling ability of the model to model the target spatial relationships, and the label spatial position information is fused through attention mechanism to solve the target position noise problem.

Benefits of technology

The accuracy of the image recognition model in identifying target positions in complex scenes is improved, the model's modeling ability to target spatial relationships is enhanced, and the effect of multi-label image classification is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120164064A_ABST
    Figure CN120164064A_ABST
Patent Text Reader

Abstract

The invention discloses a label position feature fusion method based on an attention mechanism and coordinate transformation. The method comprises the following steps: constructing an image recognition model; determining target local features and coordinate information of the input sample image based on an image recognition model; generating a position weight based on the target local feature and the coordinate information of the sample image, and calculating a position result; constructing a coordinate transformation function; and determining label position feature fusion information according to the position result and the coordinate transformation function, and adding the label position feature fusion information to the image recognition model to obtain a target image recognition model. An attention mechanism fusing label space position information is provided, and the modeling capability of a model on a target space relation is enhanced through coordinate transformation. Affine transformation of target coordinates is introduced, spatial relation modeling is enhanced, the problem of target position noise in a complex scene is solved, and the recognition accuracy of an image recognition model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly relates to a method for fusing label position features based on an attention mechanism and coordinate transformation. Background Art

[0002] With the advancement of autonomous driving research, multi-object recognition in autonomous driving is an essential part (such as simultaneously detecting vehicles, pedestrians, traffic signs, and their relative positions). Traditional deep learning models usually do not consider the image-label correlation constraint when classifying multi-label images. The strategy of classifying only based on the image's own features greatly limits the model performance. In existing image recognition models, the ability to model the target spatial relationship is poor, and there is a problem of target position noise in complex scenarios, resulting in inaccurate recognition of the image recognition model. Summary of the Invention

[0003] The present invention aims to solve at least one of the technical problems in the above technologies to some extent. For this purpose, the object of the present invention is to propose a method for fusing label position features based on an attention mechanism and coordinate transformation, propose an attention mechanism that fuses the spatial position information of labels, and enhance the model's ability to model the target spatial relationship through coordinate transformation. Introduce the affine transformation of the target coordinates to enhance the spatial relationship modeling, solve the problem of target position noise in complex scenarios, and improve the recognition accuracy of the image recognition model.

[0004] To achieve the above object, an embodiment of the present invention proposes a method for fusing label position features based on an attention mechanism and coordinate transformation, including:

[0005] Construct an image recognition model;

[0006] Based on the image recognition model, determine the target local features and coordinate information of the input sample image;

[0007] Generate position weights based on the target local features and coordinate information of the sample image, and calculate the position result;

[0008] Construct a coordinate transformation function;

[0009] Determine the label position feature fusion information according to the position result and the coordinate transformation function, and add the label position feature fusion information to the image recognition model to obtain the target image recognition model.

[0010] According to some embodiments of the present invention, determining the target local features and coordinate information of the input sample image based on the image recognition model includes:

[0011] Based on the image recognition model, load the input sample image, preprocess the sample image, including graying, denoising, and enhancing the contrast, to obtain a preprocessed sample image;

[0012] Based on the local feature extraction algorithm SIFT, local feature points and corresponding feature descriptors are extracted from the preprocessed sample image. The feature descriptor contains local feature information of the area around the local feature point, including the gradient direction and the gradient magnitude value;

[0013] The feature descriptor is matched with a preset target feature descriptor library, and the position of the target in the preprocessed sample image is determined according to the matching result, and the coordinate information is determined.

[0014] According to some embodiments of the present invention, a position weight is generated based on the target local feature and coordinate information of the sample image, and a position result is calculated, including:

[0015]

[0016] where g S (m) is the position result calculated based on the position weight; l is the number of image labels included in the sample image; the icon label is determined according to the target local feature and coordinate information; Q U is a convolution operation with a 1*1 convolution kernel, used to perform a convolution operation to extract specific features; is the feature vector of image label n; q nm is the relationship weight between image label n and image label m, which reflects the association degree between different label features and affects the calculation of the position weight;

[0017]

[0018] is the position weight between image label n and image label m; is the feature weight between image label n and image label m; is the position weight between image label b and image label m; is the feature weight between image label l and image label m;

[0019]

[0020] where Q L and Q W are the weight parameter and bias parameter of the fully connected layer respectively, which play a role in adjusting and mapping features during the feature calculation process; Ψ is a dot product operation, used to calculate the dot product of two feature vectors and measure the similarity between them; is the feature vector of image label m; e is a constant, and its value range is (0,1), used to normalize the calculation result;

[0021]

[0022] σ H is a high-dimensional conversion function for the position coordinate information of the image label; Q H is the position weight influence coefficient, and its value range is (0, 1); is the position coordinate information feature of the image label n; is the position coordinate information feature of the image label m.

[0023] According to some embodiments of the present invention, constructing a coordinate transformation function includes:

[0024]

[0025] where Γ(a n , b n , q n , f n ) is the coordinate transformation function; a n , b n are the coordinate components of point n, that is, the abscissa and the ordinate; q n is the scale parameter of point n; f n is the proportion parameter of point n; a m , b m are the coordinate components of point m, that is, the abscissa and the ordinate; q m is the scale parameter of point m; f m is the proportion parameter of point m; T is the transpose operation.

[0026] According to some embodiments of the present invention, determining the label position feature fusion information according to the position result and the coordinate transformation function includes:

[0027]

[0028] where r is the label position feature fusion information; a is the first fusion parameter, and its value is (0, 1), which is used to adjust the contribution of the position result in the fusion process; b is the second fusion parameter, and its value is (0, 1), which is used to adjust the contribution of the position feature determined by the coordinate transformation function in the fusion process. The sum of the first fusion parameter and the second fusion parameter is 1; is the feature fusion operation, including addition and splicing; Conv is the convolution operation, which is used to perform the convolution operation on the position feature Γ2 determined by the coordinate transformation function.

[0029] According to some embodiments of the present invention, determining the target local feature and coordinate information of the input sample image based on the image recognition model includes:

[0030] Performing binarization processing on the input sample image based on the image recognition model to obtain a binarized image;

[0031] Perform a morphological dilation operation on the area where the pixels in the binary image are 0, extract the area where the pixels are 0, determine the mapping relationship between the binary image and the input sample image, and determine the content area corresponding to the area where the pixels are 0 in the input sample image;

[0032] Perform image segmentation on the content area to obtain multiple image blocks;

[0033] Determine the key points in each image block;

[0034] Match the key points with the preset key points, and according to the matching result, use the feature influence parameter corresponding to the preset key point as the feature influence parameter of the key point;

[0035] Sum up the feature influence parameters of each key point included in the image block, and use the sum value as the feature value of the image block, which is used as the target local feature of the input sample image;

[0036] Based on HOG feature extraction and SVM classifier, determine each target category included in the input sample image and the position parameter of each target category, and determine the coordinate information.

[0037] According to some embodiments of the present invention, before performing image segmentation on the content area, it further includes:

[0038] Obtain the R-channel value, G-channel value, and B-channel value of each pixel point in the content area, perform weighted calculation based on the preset weighted parameters, and determine the feature value of each pixel point;

[0039] According to the feature value of each pixel point and the preset contour database, determine the target contour, determine the segmentation line according to the target contour, and perform image segmentation on the content area based on the segmentation line.

[0040] According to some embodiments of the present invention, before determining the label position feature fusion information according to the position result and the coordinate transformation function, determine the fusion order of the label position features.

[0041] According to some embodiments of the present invention, determining the fusion order of the label position features includes:

[0042] Construct a fusion simulation model;

[0043] According to the target local feature of each image block and the position weight of each image block, determine the simulation information;

[0044] Input the simulation information into the fusion simulation model for simulation based on the coordinate transformation function, and determine the fusion time and fusion gain of the fusion label position features of each image block;

[0045] Calculate the ratio of the fusion gain to the fusion time, sort them from largest to smallest, and determine the fusion order of the label position features based on the sorting result.

[0046] The present invention proposes a method for fusing label position features based on an attention mechanism and coordinate transformation. Based on the attention mechanism that fuses the spatial position information of labels, the model's ability to model the target spatial relationship is enhanced through coordinate transformation. The affine transformation of the target coordinates is introduced to enhance the spatial relationship modeling, solve the problem of target position noise in complex scenes, and improve the recognition accuracy of the image recognition model.

[0047] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification, or be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the structures specifically pointed out in the written specification and the drawings.

[0048] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings

[0049] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention, and do not constitute a limitation to the present invention. In the drawings:

[0050] Figure 1 is a flowchart of a method for fusing label position features based on an attention mechanism and coordinate transformation according to an embodiment of the present invention. Detailed Embodiments

[0051] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0052] As Figure 1 shown, the embodiment of the present invention proposes a method for fusing label position features based on an attention mechanism and coordinate transformation, including steps S1 - S5:

[0053] S1. Construct an image recognition model;

[0054] S2. Determine the target local features and coordinate information of the input sample image based on the image recognition model;

[0055] S3. Generate position weights based on the target local features and coordinate information of the sample image, and calculate the position result;

[0056] S4. Construct a coordinate transformation function;

[0057] S5. Determine the fused information of label position features based on the position result and the coordinate transformation function, and add the fused information of label position features to the image recognition model to obtain the target image recognition model.

[0058] The working principle of the above technical solution: The framework of the image recognition model is such as a convolutional neural network (CNN). The image recognition model is used to process the input sample image to extract the local features of the target of interest. At the same time, record the coordinate information of these features, that is, their positions in the image. Using the attention mechanism, calculate the position weights according to the local features of the target and the coordinate information. The position weights reflect the importance of different position features and are used for subsequent feature fusion. The position result represents the feature fusion result about the position calculated based on the position weights subsequently. Use self-attention or cross-attention mechanisms to capture the dependencies between features. Introduce position encoding or position embedding to enhance the model's ability to understand coordinate information. Design a coordinate transformation function for transforming the feature coordinates to adapt to different task requirements. The coordinate transformation includes operations such as translation, rotation, and scaling. Combine the position weights and the features after coordinate transformation for feature fusion. The fused features are used for training or inference to obtain the final image recognition result. Integrate the fused information of label position features in the feature extraction layer of the image recognition model.

[0059] The beneficial effects of the above technical solution: Based on the attention mechanism that fuses the spatial position information of the label, enhance the model's ability to model the target spatial relationship through coordinate transformation. Introduce the affine transformation of the target coordinates to enhance the spatial relationship modeling, solve the problem of target position noise in complex scenarios, and improve the recognition accuracy of the image recognition model.

[0060] According to some embodiments of the present invention, determining the local features and coordinate information of the target of the input sample image based on the image recognition model includes:

[0061] Load the input sample image based on the image recognition model, and preprocess the sample image, including grayscale conversion, denoising, and contrast enhancement, to obtain the preprocessed sample image;

[0062] Extract local feature points and corresponding feature descriptors from the preprocessed sample image based on the local feature extraction algorithm SIFT. The feature descriptors contain the local feature information of the area around the local feature points, including the gradient direction and the gradient magnitude;

[0063] Match the feature descriptors with a preset target feature descriptor library, and determine the position of the target in the preprocessed sample image according to the matching result to determine the coordinate information.

[0064] Working principle of the above technical solution: Grayscale processing: Convert a color image into a grayscale image. A grayscale image only contains luminance information and no color information, which helps reduce the computational amount and simplify subsequent processing steps. Denoising: Remove noises in the image, such as spots, stripes, etc. Denoising can improve the image quality and make local feature extraction more accurate. Contrast enhancement: Adjust the contrast of the image to make the target features in the image more obvious. This helps better identify the target in subsequent steps. SIFT (Scale-Invariant Feature Transform) is an algorithm for extracting local features from an image. It can detect key points (i.e., local feature points) in the image and generate a feature descriptor for each key point. A feature descriptor is a vector that contains local feature information of the area around the key point. In the SIFT algorithm, the feature descriptor usually includes the gradient direction and gradient magnitude of the area around the key point. Match the extracted feature descriptors with a preset target feature descriptor library and find the most similar feature pairs. According to the matching result, the position of the target in the preprocessed sample image can be determined. Specifically, the position of the target can be located by finding the feature descriptor that best matches the preset target feature descriptor and then determining its coordinates in the image. Once the position of the target in the image is determined, the coordinate information of the target can be extracted.

[0065] Beneficial effects of the above technical solution: Accurately determine the local features and coordinate information of the target in the input sample image based on the image recognition model.

[0066] According to some embodiments of the present invention, generate a position weight based on the local features and coordinate information of the target in the sample image, and calculate a position result, including:

[0067]

[0068] where, g S (m) is the position result calculated based on the position weight; l is the number of image labels included in the sample image; the icon label is determined according to the local features and coordinate information of the target; Q U is a convolution operation with a 1*1 convolution kernel, used to perform a convolution operation to extract specific features; is the feature vector of image label n; q nm is the relationship weight between image label n and image label m, reflecting the correlation degree between different label features and affecting the calculation of the position weight;

[0069]

[0070] is the position weight between image label n and image label m; is the feature weight between image label n and image label m; is the position weight between image label b and image label m; is the feature weight between image label l and image label m;

[0071]

[0072] where Q L and Q W are the weight parameter and bias parameter of the fully connected layer respectively, which play a role in adjusting and mapping features during the feature calculation process; Ψ is the dot product operation, used to calculate the dot product of two feature vectors and measure the similarity between them; is the feature vector of image label m; e is a constant, and its value range is (0, 1), which is used to normalize the calculation result;

[0073]

[0074] σ H is the high-dimensional conversion function for realizing the position coordinate information of the image label; Q H is the position weight influence coefficient, and its value range is (0, 1); is the position coordinate information feature of image label n; is the position coordinate information feature of image label m.

[0075] The working principle and beneficial effects of the above technical solution: The objective function g S (m) is calculated based on the position weight. This result is obtained by weighted summation of the feature vectors of multiple image labels, and the weight is determined by q nm . Based on the target local features and coordinate information of the sample image, the position weight is accurately generated, and then the position result is accurately determined.

[0076] According to some embodiments of the present invention, a coordinate transformation function is constructed, including:

[0077]

[0078] where Γ(a n , b n , q n , f n ) is the coordinate transformation function; a n , b n are the coordinate components of point n, that is, the abscissa and ordinate; q n is the scale parameter of point n; f n is the ratio parameter of point n; a m , bm are the coordinate components of point m, i.e., the abscissa and ordinate; q m is the scale parameter of point m; f m is the proportionality parameter of point m; T is the transpose operation.

[0079] The working principle and beneficial effects of the above technical solution: Γ(a n , b n , q n , f n ) is a coordinate transformation function that transforms the coordinate components of a point, (a n , b n ), the scale parameter q n , and the proportionality parameter f n into a new four-dimensional vector. This transformation also involves the corresponding parameters of another point m. The first component represents the relative position of point n and the reference point m in the abscissa direction (after scale and logarithmic transformation); the second component represents the relative position of point n and the reference point m in the ordinate direction (after proportionality and logarithmic transformation). The third component represents the relative difference in the scale parameter between point n and the reference point m; the fourth component represents the relative difference in the proportionality parameter between point n and the reference point m. The above four components are combined into a four-dimensional vector and a transpose operation is performed. This facilitates the accurate construction of the coordinate transformation function.

[0080] According to some embodiments of the present invention, determining the label position feature fusion information based on the position result and the coordinate transformation function includes:

[0081]

[0082] where r is the label position feature fusion information; a is the first fusion parameter, taking values in (0, 1), and is used to adjust the contribution of the position result in the fusion process; b is the second fusion parameter, taking values in (0, 1), and is used to adjust the contribution of the position feature determined by the coordinate transformation function in the fusion process. The sum of the first fusion parameter and the second fusion parameter is 1; is the feature fusion operation, including addition and splicing; Conv is the convolution operation, which is used to perform a convolution operation on the position feature Γ2 determined based on the coordinate transformation function.

[0083] The working principle and beneficial effects of the above technical solution: accurately determining the label position feature fusion information based on the position result and the coordinate transformation function.

[0084] According to some embodiments of the present invention, determining the target local feature and coordinate information of the input sample image based on the image recognition model includes:

[0085] Perform binarization processing on the input sample image based on the image recognition model to obtain a binarized image;

[0086] Perform morphological dilation operation on the area where the pixel is 0 in the binarized image, extract the area where the pixel is 0, determine the mapping relationship between the binarized image and the input sample image, and determine the content area corresponding to the area where the pixel is 0 in the input sample image;

[0087] Perform image segmentation on the content area to obtain multiple image blocks;

[0088] Determine the key points in each image block;

[0089] Match the key points with the preset key points, and according to the matching result, use the feature influence parameter corresponding to the preset key point as the feature influence parameter of the key point;

[0090] Sum up the feature influence parameters of each key point included in the image block, and use the sum value as the feature value of the image block, which is used as the target local feature of the input sample image;

[0091] Based on HOG feature extraction and SVM classifier, determine each target category included in the input sample image and the position parameter of each target category, and determine the coordinate information.

[0092] Working principle of the above technical solution: Convert the input sample image into a binary image, that is, the pixel values in the image only contain 0 and 255. Perform morphological dilation on the areas where the pixels in the binary image are 0 to expand these areas and reduce noise. Extract the areas where the pixels are 0 after dilation, and these areas correspond to specific objects in the image. According to the mapping relationship between the binary image and the input sample image, find the corresponding content areas in the input sample image for the areas where the pixels in the binary image are 0. Perform image segmentation on the content areas, divide them into multiple image blocks, and determine key points in each image block. Key point detection methods such as SIFT (Scale-Invariant Feature Transform) or SURF (Speeded Up Robust Features). These key points may be corner points, edge points, or points with specific features. Match the key points in each image block with the preset key points. According to the matching results, assign the feature influence parameters corresponding to the preset key points to the matched key points. Sum the feature influence parameters of all the key points in the image block to obtain the feature value of the image block. These feature values will be used as the target local features of the input sample image. Use the HOG (Histogram of Oriented Gradients) feature extraction method to extract features from the input sample image. Use an SVM (Support Vector Machine) classifier to classify the extracted features to determine each target category included in the input sample image and the position parameters of each target category. Based on the HOG feature extraction and SVM classification optimization to improve the accuracy and generalization ability of classification.

[0093] Beneficial effects of the above technical solution: Improve the accuracy and efficiency of determining the target local features and coordinate information of the input sample image based on the image recognition model.

[0094] According to some embodiments of the present invention, before performing image segmentation on the content area, it further includes:

[0095] Obtain the R-channel value, G-channel value, and B-channel value of each pixel point in the content area, perform weighted calculation based on the preset weighted parameters to determine the feature value of each pixel point;

[0096] Determine the target contour according to the feature value of each pixel point and the preset contour database, determine the segmentation line according to the target contour, and perform image segmentation on the content area based on the segmentation line.

[0097] Working principle of the above technical solution: Obtain the R (red), G (green), and B (blue) channel values of each pixel in the content area. According to the preset weighting parameters, perform weighted summation on the values of each channel to obtain the feature value of each pixel. The selection of the weighting parameters can be adjusted according to the specific task and the characteristics of the dataset to highlight certain color features or suppress noise. For example, the weighting parameters are 0.3, 0.3, and 0.4 respectively. The sum of the weighting parameters is 1. Match the calculated feature value of each pixel with the preset contour database. The preset contour database contains a series of contour features of known objects, and these features can be based on attributes such as color, shape, and texture. Use unsupervised learning methods (such as clustering) to automatically discover new contour features and add them to the database. Through matching, the part in the content area that is most similar to the target contour can be determined. According to the determined target contour, the segmentation line can be further determined. The segmentation line can be the boundary line along the target contour or calculated based on a certain rule (such as minimum cut set, maximum flow, etc.). Use these segmentation lines to perform image segmentation on the content area to obtain multiple image blocks.

[0098] Beneficial effects of the above technical solution: Further improve the accuracy and efficiency of the image segmentation method based on color channel weighting and the preset contour database, and also help improve the performance and robustness of the entire image recognition model.

[0099] According to some embodiments of the present invention, before determining the label position feature fusion information based on the position result and the coordinate transformation function, determine the fusion order of the label position features.

[0100] Beneficial effects of the above technical solution: Through a reasonable fusion order and method, the performance and accuracy of the image recognition or processing task can be improved.

[0101] According to some embodiments of the present invention, determining the fusion order of the label position features includes:

[0102] Construct a fusion simulation model;

[0103] Determine the simulation information according to the target local feature of each image block and the position weights of each image block;

[0104] Input the simulation information into the fusion simulation model for simulation based on the coordinate transformation function to determine the fusion time and fusion benefit of the fusion label position features of each image block;

[0105] Calculate the ratio of the fusion benefit to the fusion time, sort from large to small according to the ratio, and determine the fusion order of the label position features based on the sorting result.

[0106] Working principle of the above technical solution: The fusion simulation model can receive input information (such as the target local features of image blocks, position weights, etc.) and output a fusion result based on this information. A coordinate transformation function is integrated into the simulation model. The coordinate transformation function is used to adjust or transform the position information of the image blocks so as to consider the spatial relationship during the fusion process. Feature extraction is performed on each image block to obtain its target local features. The position weights are calculated according to factors such as the position, size, and importance of the image block in the image. The position weights reflect the relative importance of different image blocks during the fusion process. The extracted target local features and the calculated position weights are used as input information and input into the fusion simulation model. The input information is simulated based on the coordinate transformation function. The simulation process involves multiple iterations to explore the influence of different fusion orders and parameter combinations on the fusion result. During the simulation process, the fusion time of the fusion label position features of each image block is recorded. At the same time, the quality or benefit of the fusion result is evaluated. For each image block, the ratio of its fusion benefit to the fusion time is calculated. This ratio reflects the efficiency of the benefit brought by fusing this image block relative to the time consumed. According to the calculated ratio, the image blocks are sorted from large to small. Image blocks with higher ratios mean that they should be given priority during the fusion process. According to the sorting result, the fusion order of the label position features is determined. The fusion order should be in descending order of the ratio to ensure the best fusion effect within limited resources or time.

[0107] Beneficial effects of the above technical solution: By constructing a fusion simulation model, determining simulation information, performing simulation, calculating ratios, and sorting, etc., the fusion order of the label position features can be determined. This method helps to optimize resource allocation and improve fusion efficiency during the fusion process.

[0108] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and modifications.

Claims

1. A label position feature fusion method based on attention mechanism and coordinate transformation, characterized in that: include: Build an image recognition model; Determine the target local features and coordinate information of the input sample image based on the image recognition model; Generate position weights based on the target local features and coordinate information of the sample image, and calculate the position results; Construct coordinate transformation function; The label position feature fusion information is determined according to the position result and the coordinate transformation function, and the label position feature fusion information is added to the image recognition model to obtain the target image recognition model.

2. The label position feature fusion method based on attention mechanism and coordinate transformation as claimed in claim 1, characterized in that: Determine the target local features and coordinate information of the input sample image based on the image recognition model, including: Based on the image recognition model, an input sample image is loaded, and the sample image is preprocessed, including graying, denoising, and contrast enhancement, to obtain a preprocessed sample image; Based on the local feature extraction algorithm SIFT, local feature points and corresponding feature descriptors are extracted from the preprocessed sample image, wherein the feature descriptors contain local feature information of the area around the local feature points, including gradient direction and gradient modulus; The feature descriptor is matched with a preset target feature descriptor library, and the position of the target in the preprocessed sample image is determined according to the matching result, and the coordinate information is determined.

3. The label position feature fusion method based on attention mechanism and coordinate transformation as claimed in claim 1, characterized in that: Generate position weights based on the target local features and coordinate information of the sample image, and calculate the position results, including: Among them, g S (m) is the position result calculated based on the position weight; l is the number of image labels included in the sample image; the icon label is determined based on the local features and coordinate information of the target; Q U is a convolution operation with a 1*1 convolution kernel, used for Perform convolution operations to extract specific features; is the feature vector of image label n; q nm is the relationship weight between image label n and image label m, which reflects the degree of association between different label features and affects the calculation of position weight; is the position weight between image label n and image label m; is the feature weight between image label n and image label m; is the position weight between image label b and image label m; is the feature weight between image label l and image label m; Among them, Q L and Q W are the weight parameters and bias parameters of the fully connected layer, which play the role of adjusting and mapping features in the feature calculation process; Ψ is the dot product operation, which is used to calculate the dot product of two feature vectors and measure the similarity between them; is the feature vector of the image label m; e is a constant with a value range of (0,1) and is used to normalize the calculation results; σ H To realize the high-dimensional conversion function of image label position coordinate information; Q H is the position weight influence coefficient, and its value range is (0,1); is the location coordinate information feature of the image label n; is the location coordinate information feature of the image label m.

4. The label position feature fusion method based on attention mechanism and coordinate transformation as claimed in claim 3, characterized in that: Construct coordinate transformation functions, including: Among them, Γ(a n ,b n ,q n ,f n ) is the coordinate transformation function; a n , b n is the coordinate component of point n, i.e. the horizontal coordinate and the vertical coordinate; q n is the scale parameter of point n; f n is the scale parameter of point n; a m , b m is the coordinate component of point m, i.e. the horizontal coordinate and the vertical coordinate; q m is the scale parameter of point m; f m is the scale parameter of point m; T is the transpose operation.

5. The label position feature fusion method based on attention mechanism and coordinate transformation as claimed in claim 4, characterized in that: Determine the tag position feature fusion information based on the position result and the coordinate transformation function, including: Among them, r is the label position feature fusion information; a is the first fusion parameter, the value is (0, 1), which is used to adjust the contribution of the position result in the fusion process; b is the second fusion parameter, the value is (0, 1), which is used to adjust the contribution of the position feature determined by the coordinate transformation function in the fusion process, and the sum of the first fusion parameter and the second fusion parameter is 1; is a feature fusion operation, including addition and concatenation; Conv is a convolution operation, which is used to perform a convolution operation on the position feature Γ2 determined based on the coordinate transformation function.

6. The label position feature fusion method based on attention mechanism and coordinate transformation as claimed in claim 1, characterized in that: Determine the target local features and coordinate information of the input sample image based on the image recognition model, including: Binarize the input sample image based on the image recognition model to obtain a binary image; Perform morphological dilation operation on the area with zero pixels in the binary image, extract the area with zero pixels, determine the mapping relationship between the binary image and the input sample image, and determine the content area corresponding to the area with zero pixels in the input sample image; Perform image segmentation on the content area to obtain multiple image blocks; Determine the key points in each image patch; Matching the key point with the preset key point, and according to the matching result, taking the feature influence parameter corresponding to the preset key point as the feature influence parameter of the key point; The characteristic influencing parameters of each key point included in the image block are summed up, and the sum is used as the characteristic value of the image block and as the target local feature of the input sample image; Based on HOG feature extraction and SVM classifier, various target categories included in the input sample image and the position parameters of each target category are determined to determine the coordinate information.

7. The label position feature fusion method based on attention mechanism and coordinate transformation as claimed in claim 6, characterized in that: Before image segmentation of the content area, it also includes: Obtain the R channel value, G channel value, and B channel value of each pixel in the content area, perform weighted calculation based on preset weighting parameters, and determine the characteristic value of each pixel; According to the characteristic value of each pixel point and the preset contour database, the target contour is determined, the segmentation line is determined according to the target contour, and the image is segmented for the content area based on the segmentation line.

8. The label position feature fusion method based on attention mechanism and coordinate transformation as claimed in claim 6, characterized in that: Before determining the tag position feature fusion information based on the position result and the coordinate transformation function, determine the fusion order of the tag position features.

9. The label position feature fusion method based on attention mechanism and coordinate transformation as claimed in claim 8, characterized in that: Determine the fusion order of label position features, including: Construct fusion simulation model; Determine simulation information according to the target local features of each image block and the position weight of each image block; Input the simulation information into the fusion simulation model to perform simulation based on the coordinate transformation function, and determine the fusion time and fusion benefit of each image block for fusion label position features; Calculate the ratio of fusion benefit to fusion time, sort them from large to small according to the ratio, and determine the fusion order of label position features based on the sorting result.