Model application method, device and equipment based on regional self-attention mechanism

By introducing a processing method based on the regional self-attention mechanism in the visual big model, using the masking mechanism and attention adjustment mechanism to process images, the data deviation problem of the visual big model when processing noise or irrelevant information images is solved, and higher image processing accuracy is achieved.

CN120125798APending Publication Date: 2025-06-10张雨廷
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510287763.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-12-24
Filing Date
2025-03-12
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing visual models are prone to data deviations when processing images with noise or irrelevant information.

Method used

The model application method based on the regional self-attention mechanism is adopted, and the original image is processed through the masking mechanism, a mask matrix adapted to the regional self-attention mechanism is generated, and the image is adjusted through the attention adjustment mechanism to generate attention feature vectors, and finally the processed image is obtained through feature post-processing.

Benefits of technology

This method can better extract image features, reduce noise or unrelated information interference, reduce data deviation, and improve image processing accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125798A_ABST
    Figure CN120125798A_ABST
Patent Text Reader

Abstract

The invention discloses a model application method, device and equipment based on a regional self-attention mechanism, and relates to the technical field of model application, and the method comprises the steps: obtaining an original image; processing the original image through a mask mechanism to obtain a mask matrix adaptive to a region self-attention mechanism; based on the original image and the mask matrix of the adaptive area self-attention mechanism, adjusting through an attention adjustment mechanism to generate an attention adjustment result; generating an attention feature vector based on the original image and the attention adjustment result; and performing feature post-processing on the attention feature vector to obtain a processed image. According to the invention, data deviation can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of model application technology, and in particular to a model application method, device and equipment based on regional self-attention mechanism. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology, image processing through artificial intelligence technology has become the mainstream trend in the field of computer vision, and visual big models represent the latest progress of artificial intelligence technology in the field of image processing.

[0003] However, current large visual models are prone to data bias when processing images with noise or irrelevant information. Summary of the invention

[0004] The main purpose of the present invention is to provide a model application method, device and equipment based on the regional self-attention mechanism, aiming to solve the problem that the current large visual model is prone to data deviation when processing images with noise or irrelevant information.

[0005] To achieve the above object, the present invention provides a model application method based on regional self-attention mechanism, the method comprising:

[0006] Get the original image;

[0007] Processing the original image through a mask mechanism to obtain a mask matrix of an adaptation region self-attention mechanism;

[0008] Based on the mask matrix of the original image and the self-attention mechanism of the adaptation region, an attention adjustment mechanism is used to adjust and generate an attention adjustment result;

[0009] Based on the original image and the attention adjustment result, generating an attention feature vector;

[0010] Perform feature post-processing on the attention feature vector to obtain a processed image.

[0011] Optionally, the step of processing the original image through a mask mechanism to obtain a mask matrix adapted to the regional self-attention mechanism includes:

[0012] Performing instance segmentation on the original image to obtain a target mask matrix;

[0013] Segmenting the target mask matrix based on a preset side length to obtain a plurality of segmented regions;

[0014] Counting the pixel quantities of the plurality of segmented areas to obtain a block mask corresponding to each of the segmented areas;

[0015] A mask matrix of the adaptation region self-attention mechanism is generated based on the block masks corresponding to each of the segmented regions.

[0016] Optionally, the step of processing the original image through a mask mechanism to obtain a mask matrix adapted to the regional self-attention mechanism further includes:

[0017] Set the target mask matrix to a two-dimensional array A, set the preset side length to size, and set the threshold to threshold.

[0018] Create two a*b all-zero matrices, namely Co matrix as shown in formula (1) and Cr matrix as shown in formula (2), where the value of a is A.shape(0) / size and the value of b is A.shape(1) / size.

[0019]

[0020] The Co matrix is ​​used to record the number of pixels in the plurality of segmented regions, where co ij is the value corresponding to the matrix element, co ij The calculation formula is (3):

[0021]

[0022] The Cr matrix is ​​used to mask irrelevant areas, where cr ij is the value corresponding to the matrix element, cr ij The calculation formula is (4):

[0023]

[0024] Among them, the Cr matrix is ​​the mask matrix of the adaptation area self-attention mechanism.

[0025] Optionally, the mask matrix based on the original image and the adaptation region self-attention mechanism is adjusted by an attention adjustment mechanism to generate an attention adjustment result, comprising:

[0026] Performing image embedding processing on the original image to obtain an embedding vector X;

[0027] The embedding vector X and the trained weight matrix W q , W k , W v Perform matrix multiplication to obtain the query matrix Q, key matrix K and value matrix V. The formula is:

[0028] Q=X*W q ,

[0029] K=X*Wk ,

[0030] V=X*W v ;

[0031] Perform matrix multiplication of the query matrix Q and the transpose of the key matrix K to obtain the score matrix QK T , the formula is Q*K T =QK T ;

[0032] The score matrix QK is adjusted by the attention adjustment mechanism T and adjusting the mask matrix of the self-attention mechanism of the adaptation area to generate the attention adjustment result;

[0033] The step of generating an attention feature vector based on the original image and the attention adjustment result comprises:

[0034] The value matrix V is matrix-multiplied by the attention adjustment result to generate the attention feature vector Z'.

[0035] Optionally, the mask matrix of the adaptation region self-attention mechanism is set to Mask, and the score matrix QK is adjusted by the attention adjustment mechanism. T The step of adjusting the mask matrix of the self-attention mechanism of the adaptation region to generate the attention adjustment result includes:

[0036] The mask matrix Mask is converted into one dimension to obtain a one-dimensional mask matrix Mask_1D;

[0037] Transpose the one-dimensional mask matrix Mask_1D to obtain the transposed mask matrix (Mask_1D) T ;

[0038] The one-dimensional mask matrix Mask_1D and the transposed mask matrix (Mask_1D) are combined. T Perform matrix multiplication to obtain the mask autocorrelation matrix M, the formula is: Mask_1D*(Mask_1D) T =M;

[0039] The score matrix QK T Subtract the product of the mask autocorrelation matrix M and the preset parameter k to obtain the new QK T , if k>0, the background area is suppressed, if k<0, the foreground area is enhanced, the formula is: the new QK T =QK T ±k*M;

[0040] According to the new QK TDetermine attentional modulation outcomes.

[0041] Optionally, according to the new QK T Identify attention regulation outcomes, including:

[0042] The new QK T As QK i T , the attention adjustment result is determined by the following expression:

[0043]

[0044] Among them, DA i is the attention adjustment result of the i-th row, QK i T The result of multiplying the i-th row of the K matrix by the Q matrix. The result of multiplying all rows of the K matrix by the Q matrix is ​​the new QK T , is the scaling factor, k is the preset parameter, M is the mask autocorrelation matrix, and N is the number of rows of the K matrix.

[0045] Optionally, when background attention needs to be weakened, the value interval of the preset parameter k is (-∞, 0), and when foreground attention needs to be enhanced, the value interval of the preset parameter k is (0, +∞).

[0046] Optionally, the mask matrix based on the original image and the adaptation region self-attention mechanism is adjusted by an attention adjustment mechanism, and before the step of generating an attention adjustment result, the step includes:

[0047] Get the training dataset;

[0048] Based on the training data set, the pre-selected visual large model framework is trained to obtain the weight matrix W q , W k , W v .

[0049] The embodiment of the present invention further proposes a model application device based on a regional self-attention mechanism, the device comprising:

[0050] A data acquisition module, used for acquiring original images;

[0051] A processing module, used to process the original image through a mask mechanism to obtain a mask matrix adapted to the regional self-attention mechanism;

[0052] An attention adjustment module, configured to adjust the original image and the mask matrix of the self-attention mechanism of the adaptation region through the attention adjustment mechanism to generate an attention adjustment result; and generate an attention feature vector based on the original image and the attention adjustment result;

[0053] The feature post-processing module is used to perform feature post-processing on the attention feature vector to obtain a processed image.

[0054] An embodiment of the present invention also proposes a device, which includes a memory, a processor, and a model application based on a regional self-attention mechanism stored in the memory and executable on the processor. When the model application based on the regional self-attention mechanism is executed by the processor, the model application method based on the regional self-attention mechanism as described above is implemented.

[0055] An embodiment of the present invention also proposes a computer-readable storage medium, on which a model application based on a regional self-attention mechanism is stored. When the model application based on the regional self-attention mechanism is executed by a processor, the model application method based on the regional self-attention mechanism as described above is implemented.

[0056] The model application method, device and equipment based on the regional self-attention mechanism proposed in the embodiment of the present invention obtains an original image; processes the original image through a mask mechanism to obtain a mask matrix adapted to the regional self-attention mechanism; adjusts the original image and the mask matrix adapted to the regional self-attention mechanism through an attention adjustment mechanism to generate an attention adjustment result; generates an attention feature vector based on the original image and the attention adjustment result; and performs feature post-processing on the attention feature vector to obtain a processed image. The embodiment of the present invention processes the original image through a mask mechanism and an attention adjustment mechanism to generate an attention adjustment result, which can better extract image features and reduce the interference of noise or irrelevant information, thereby reducing data deviation and improving image processing accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a schematic diagram of the functional modules of the equipment to which the model application device based on the regional self-attention mechanism belongs in the present invention;

[0058] Figure 2 It is a flowchart of a first exemplary embodiment of a model application method based on a regional self-attention mechanism of the present invention;

[0059] Figure 3 A flow chart of the regional self-attention mechanism of the model application method based on the regional self-attention mechanism of the present invention.

[0060] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0061] It should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0062] The main solution of the embodiment of the present invention is as follows: obtaining an original image; processing the original image through a masking mechanism to obtain a masking matrix adapted to the regional self-attention mechanism; based on the original image and the masking matrix of the regional self-attention mechanism, adjusting through an attention adjustment mechanism to generate an attention adjustment result; based on the original image and the attention adjustment result, generating an attention feature vector; performing post-processing on the attention feature vector to obtain a processed image. By processing the original image through the masking mechanism and the attention adjustment mechanism, the embodiment of the present invention generates an attention adjustment result, which can better extract image features and reduce the interference of noise or irrelevant information, thereby reducing data deviation and improving the accuracy of image processing.

[0063] The embodiment of the present invention takes into account that: current large vision models are prone to data deviation when processing images with noise or irrelevant information.

[0064] Therefore, the embodiment of the present invention proposes a solution: obtaining an original image; processing the original image through a masking mechanism to obtain a masking matrix adapted to the regional self-attention mechanism; based on the original image and the masking matrix of the regional self-attention mechanism, adjusting through an attention adjustment mechanism to generate an attention adjustment result; based on the original image and the attention adjustment result, generating an attention feature vector; performing post-processing on the attention feature vector to obtain a processed image. By processing the original image through the masking mechanism and the attention adjustment mechanism, the embodiment of the present invention generates an attention adjustment result, which can better extract image features and reduce the interference of noise or irrelevant information, thereby reducing data deviation and improving the accuracy of image processing.

[0065] Specifically, referring to Figure 1 , Figure 1 is a schematic diagram of the functional modules of the device to which the model application device based on the regional self-attention mechanism of the present invention belongs. The model application device based on the regional self-attention mechanism can be a device independent of the device and capable of data processing, and it can be carried on the device in the form of hardware or software. The device can be a smart mobile terminal with data processing functions such as a mobile phone or a tablet computer, and can also be a fixed device or a server with data processing functions, etc.

[0066] In this embodiment, the device to which the model application device based on the regional self-attention mechanism belongs at least includes an output module 110, a processor 120, a memory 130, and a communication module 140.

[0067] The operating system and a model application based on the regional self-attention mechanism are stored in the memory 130; the output module 110 can be a display screen, etc. The communication module 140 can include a WIFI module, a Bluetooth module, etc., and communicates with external devices or servers through the communication module 140.

[0068] Among them, when the model application based on the regional self-attention mechanism in the memory 130 is executed by the processor, the following steps are implemented:

[0069] Obtain the original image;

[0070] Process the original image through a masking mechanism to obtain a mask matrix adapted to the regional self-attention mechanism;

[0071] Based on the original image and the mask matrix adapted to the regional self-attention mechanism, adjust through an attention adjustment mechanism to generate an attention adjustment result;

[0072] Based on the original image and the attention adjustment result, generate an attention feature vector;

[0073] Perform post-processing on the attention feature vector to obtain a processed image.

[0074] Furthermore, when the model application based on the regional self-attention mechanism in the memory 130 is executed by the processor, the following steps are also implemented:

[0075] Perform instance segmentation on the original image to obtain a target mask matrix;

[0076] Based on a preset side length, divide the target mask matrix to obtain a number of divided regions;

[0077] Count the pixel amounts of the several divided regions to obtain a block mask corresponding to each divided region;

[0078] Generate the mask matrix adapted to the regional self-attention mechanism based on the block masks corresponding to each divided region.

[0079] Furthermore, when the model application based on the regional self-attention mechanism in the memory 130 is executed by the processor, the following steps are also implemented:

[0080] Set the target mask matrix as a two-dimensional array A, set the preset side length as size, and set the threshold as threshold,

[0081] Establish two all-zero matrices of a*b, namely the Co matrix as shown in formula (1) and the Cr matrix as shown in formula (2), where the value of a is A.shape(0) / size and the value of b is A.shape(1) / size,

[0082]

[0083] Among them, the Co matrix is used to record the number of pixels in the several segmented regions, where co ij is the value corresponding to the matrix element, co ij The calculation formula of is (3):

[0084]

[0085] The Cr matrix is used to mask the irrelevant regions, where cr ij is the value corresponding to the matrix element, cr ij The calculation formula of is (4):

[0086]

[0087] Among them, the Cr matrix is the mask matrix of the self-attention mechanism for the adaptation region.

[0088] Furthermore, when the model application program based on the region self-attention mechanism in the memory 130 is executed by the processor, the following steps are further implemented:

[0089] Perform image embedding processing on the original image to obtain an embedding vector X;

[0090] Multiply the embedding vector X with the trained weight matrices W q 、W k 、W v to perform matrix multiplication to obtain a query matrix Q, a key matrix K, and a value matrix V. The formula is:

[0091] Q = X * W q ,

[0092] K = X * W k ,

[0093] V = X * W v ;

[0094] Multiply the query matrix Q with the transpose of the key matrix K to obtain a score matrix QK T , and the formula is Q * K T = QK T ;

[0095] Adjust the score matrix QK T and the mask matrix of the self-attention mechanism for the adaptation region through the attention adjustment mechanism to generate the attention adjustment result;

[0096] The step of generating the attention feature vector based on the original image and the attention adjustment result includes:

[0097] Multiply the value matrix V by the attention adjustment result to generate the attention feature vector Z'.

[0098] Further, when the model application program based on the regional self-attention mechanism in the memory 130 is executed by the processor, the following steps are also implemented:

[0099] Flatten the mask matrix Mask to obtain the one-dimensional mask matrix Mask_1D;

[0100] Transpose the one-dimensional mask matrix Mask_1D to obtain the transposed mask matrix (Mask_1D) T ;

[0101] Multiply the one-dimensional mask matrix Mask_1D by the transposed mask matrix (Mask_1D) T to obtain the mask autocorrelation matrix M, and the formula is: Mask_1D * (Mask_1D) T = M;

[0102] Subtract the product of the mask autocorrelation matrix M and the preset parameter k from the score matrix QK T to obtain the new QK T , if k > 0, suppress the background region, if k < 0, enhance the foreground region, and the formula is: the new QK T = QK T ± k * M;

[0103] Determine the attention adjustment result according to the new QK T .

[0104] Further, determining the attention adjustment result according to the new QK T includes:

[0105] Use the new QK T as QK i T , and determine the attention adjustment result through the following expression:

[0106]

[0107] where DA i is the attention adjustment result of the i-th row, QK i T is the result of multiplying the i-th row of the K matrix by the Q matrix, and the result of multiplying all rows of the K matrix by the Q matrix is the new QK T , is the scaling factor, k is a preset parameter, M is the masked autocorrelation matrix, and N is the number of rows of the K matrix.

[0108] Further, when the model application program based on the regional self-attention mechanism in the memory 130 is executed by the processor, the following steps are also implemented:

[0109] Obtain a training data set;

[0110] Train a pre-selected vision large model framework based on the training data set to obtain the weight matrices W q , W k , W v .

[0111] In this embodiment, through the above solution, specifically, an original image is obtained; the original image is processed through a masking mechanism to obtain a masking matrix adapted to the regional self-attention mechanism; based on the original image and the masking matrix of the adapted regional self-attention mechanism, adjustment is performed through an attention adjustment mechanism to generate an attention adjustment result; based on the original image and the attention adjustment result, an attention feature vector is generated; and post-processing is performed on the attention feature vector to obtain a processed image. In the embodiment of the present invention, the original image is processed through a masking mechanism and an attention adjustment mechanism to generate an attention adjustment result, which can better extract image features and reduce the interference of noise or irrelevant information, thereby reducing data deviation and improving the accuracy of image processing.

[0112] Based on the above device architecture but not limited to the above architecture, the model application method based on the regional self-attention mechanism of the present invention is proposed.

[0113] The execution subject of the method in this embodiment can be a model application device based on the regional self-attention mechanism. The model application device based on the regional self-attention mechanism can be a device independent of the device and capable of data processing, and can be carried on the device in the form of hardware or software.

[0114] Refer to Figure 2 , Figure 2 which is a schematic flowchart of the first exemplary embodiment of the model application method based on the regional self-attention mechanism of the present invention. The model application method based on the regional self-attention mechanism includes:

[0115] Step S10, obtain an original image;

[0116] It should be noted that the regional self-attention mechanism proposed by the model application method based on the regional self-attention mechanism of the present invention is a part of the vision large model and improves the content in the vision large model.

[0117] Among them, the large vision model can include convolutional neural networks (CNNs), generative adversarial networks (GANs), and so on.

[0118] Furthermore, the large vision model can also be a vision model introducing the self-attention mechanism, such as the CoAtNet collaborative attention network model, the vision model based on Transformer, etc.

[0119] The following will elaborate on the processing steps of the large vision model in combination with the regional self-attention mechanism:

[0120] (1) Extract features from the original image through the regional self-attention mechanism to obtain the attention feature vector.

[0121] (2) Pass the attention feature vector to the multi-layer fully connected linear layer to obtain the attention feature vector processed by the multi-layer fully connected linear layer. Among them, each neuron in the fully connected linear layer is connected to all neurons in the previous layer, and they are used to learn the complex relationships between features.

[0122] (3) Perform residual connection processing on each layer to obtain the attention feature vector after residual connection processing.

[0123] (4) Input the attention feature vector after residual connection processing into the multi-layer perceptron module to obtain the output result.

[0124] For the processing step (1) of the above large vision model, as an implementation, the image features of the original image can be extracted through the regional self-attention mechanism to enhance the attention of the foreground, and at the same time, suppress the attention of the background (the part with weak correlation with the foreground) to reduce the interference of the background on feature extraction.

[0125] Furthermore, different weights can also be assigned to the information of the foreground and the background to guide the concentration of attention.

[0126] As another implementation, the specific features or specific objects of the original image can be extracted through the regional self-attention mechanism to enhance the attention of the specific features or specific objects, and at the same time, suppress the attention of other parts to reduce the interference of other parts on feature extraction.

[0127] Furthermore, different weights can also be assigned to the specific features and non-specific features to guide the concentration of attention.

[0128] As another implementation, different self-attention mechanisms can also be adopted according to different regions to more accurately extract the features of the foreground and the background.

[0129] For example, in the foreground region, a fine-grained regional self-attention mechanism can be adopted. Through this mechanism, complex details and structures in the foreground region can be captured. Specifically, local windows can be used to limit the calculation range of attention, and attention weights are calculated only between pixels within the same foreground region.

[0130] For the background region, a loose regional self-attention mechanism is adopted. Since the background region usually has a simpler structure and fewer details, a larger local window or a global self-attention mechanism (but only within the background region) can be used to capture the global features and context information of the background.

[0131] Step S20: Process the original image through a masking mechanism to obtain a masking matrix adapted to the regional self-attention mechanism.

[0132] It should be noted that through the masking mechanism, key information in the original image can be extracted, and a masking matrix adapted to the regional self-attention mechanism can be generated.

[0133] Step S50: Based on the original image and the masking matrix of the adapted regional self-attention mechanism, adjust through an attention adjustment mechanism to generate an attention adjustment result.

[0134] It should be noted that the attention adjustment mechanism assigns attention weights to different regions in the original image based on the original image and the masking matrix of the adapted regional self-attention mechanism. By adjusting the attention weights of different regions, the model can pay more attention to the key information in the image and ignore or reduce the influence of irrelevant information.

[0135] Step S60: Based on the original image and the attention adjustment result, generate an attention feature vector.

[0136] Step S70: Perform post-processing on the attention feature vector to obtain a processed image.

[0137] Among them, the post-processing of the attention feature vector includes at least one of passing through multiple fully connected linear layers, residual connections, and multi-layer perceptron modules.

[0138] In this embodiment, through the above solution, the original image is obtained; the original image is processed through a masking mechanism to obtain a masking matrix adapted to the regional self-attention mechanism; based on the original image and the masking matrix of the regional self-attention mechanism, adjustment is performed through an attention adjustment mechanism to generate an attention adjustment result; based on the original image and the attention adjustment result, an attention feature vector is generated; and the attention feature vector is post-processed to obtain a processed image. In the embodiment of the present invention, the original image is processed through a masking mechanism and an attention adjustment mechanism to generate an attention adjustment result, which can better extract image features and reduce the interference of noise or irrelevant information, thereby reducing data deviation and improving the accuracy of image processing. Moreover, in this embodiment, different weights are assigned to foreground information and background information through the regional self-attention mechanism, which can guide the concentration of attention. In addition, this embodiment solves the problem that the traditional self-attention mechanism does not distinguish between different regions in the picture and processes them randomly. Furthermore, in this embodiment, the accuracy in multiple fields such as object detection and image segmentation can be improved through the regional self-attention mechanism. And the regional self-attention mechanism has strong universality and can be used in visual models with the self-attention mechanism introduced. Moreover, in this embodiment, a masking matrix adapted to the regional self-attention mechanism is generated through the masking mechanism, so that key information in the original image can be extracted. And in this embodiment, the attention adjustment mechanism assigns attention weights to different regions in the original image, so that more attention can be paid to the key information in the image, and the influence of irrelevant information can be ignored or reduced.

[0139] Based on the above Figure 2 In the embodiment shown, step S20 of processing the original image through a masking mechanism to obtain a masking matrix adapted to the regional self-attention mechanism includes:

[0140] Step S21, performing instance segmentation on the original image to obtain a target masking matrix;

[0141] It should be noted that instance segmentation is an important task in computer vision, aiming to accurately segment and identify each object instance in an image or video.

[0142] Step S22, dividing the target masking matrix based on a preset side length to obtain a number of divided regions;

[0143] Among them, the target masking matrix will be further processed. By setting a fixed side length as a standard, the target masking matrix is cut into multiple small pieces or regions, so that subsequent processing can focus more on local features and reduce computational complexity, etc.

[0144] Step S23, counting the pixel amounts of the several divided regions to obtain a block mask corresponding to each divided region;

[0145] Among them, the pixel quantity of each segmentation region can be statistically analyzed to obtain a block mask corresponding to each segmentation region.

[0146] Specifically, statistically analyzing the pixel quantity of each segmentation region is to calculate the number of pixels belonging to different objects within each segmentation region. Thus, the object distribution within each segmentation region can be reflected. For example, a segmentation region may mostly belong to the background, while other parts densely contain pixels of a specific object. In this way, the complex pixel-level target mask matrix can be converted into a more easily processed block mask.

[0147] Step S24, generate a mask matrix for the self-attention mechanism of the adaptation region based on the block masks corresponding to the respective segmentation regions.

[0148] Among them, all the block masks are combined into an overall mask matrix.

[0149] For example, the target mask matrix can be divided into 16*16 regions. By statistically analyzing the pixel quantity of each region, a corresponding block mask is obtained for each segmentation region. Arranging all the region masks in order can obtain the mask matrix for the self-attention mechanism of the adaptation region.

[0150] Furthermore, the above process can also be explained with a specific example.

[0151] Let the target mask matrix be a two-dimensional array A, the side length of each segmentation region be size, and the threshold be threshold.

[0152] Two all-zero matrices of a*b are established, namely the Co matrix as shown in formula (1) and the Cr matrix as shown in formula (2), where the value of a is A.shape(0) / size and the value of b is A.shape(1) / size.

[0153]

[0154] Among them, the Co matrix is used to record the number of pixels within each region, and the co ij is the value corresponding to the matrix element, as shown in formula (3):

[0155]

[0156] The Cr matrix is used to mask irrelevant regions, and the cr ij matrix element corresponding value, as shown in formula (4):

[0157]

[0158] Input the obtained Cr matrix and the picture into the attention adjustment mechanism, and the attention of foreground information and background information can be adjusted.

[0159] In this embodiment, the original image is processed by a masking mechanism to obtain a masking matrix adapted to the regional self-attention mechanism. The masking mechanism is not only the basis of the regional self-attention mechanism but also connects instance segmentation and the regional self-attention mechanism. The original image is subjected to instance segmentation to obtain a target masking matrix, which contains foreground and background information. The masking mechanism can extract the information therein and convert it into a masking matrix adapted to the regional self-attention mechanism.

[0160] Based on the above Figure 2 shown embodiment, in step S50, based on the original image and the masking matrix adapted to the regional self-attention mechanism, adjustment is performed through the attention adjustment mechanism, and the generation of the attention adjustment result includes:

[0161] Step S51, perform image embedding processing on the original image to obtain an embedding vector X;

[0162] It should be noted that in the self-attention mechanism, the original image usually cannot be directly processed, but the image needs to be converted into an embedding vector that the model can understand. This step is usually achieved through a convolutional neural network (CNN) or other types of neural networks, which can convert the pixel information of the image into a high-level feature vector, that is, an embedding vector.

[0163] Among them, the embedding vector captures the key information in the image, such as shape, texture, and color, etc.

[0164] Step S52, multiply the embedding vector X with the trained weight matrices W q 、W k 、W v to obtain a query matrix Q, a key matrix K, and a value matrix V. The formula is:

[0165] Q = X * W q ,

[0166] K = X * W k ,

[0167] V = X * W v ;

[0168] It should be noted that before obtaining the original image, a training data set needs to be obtained and the pre-selected visual large model framework is trained based on the training data set to obtain a visual large model, and the visual large model includes the trained weight matrices.

[0169] Among them, in the self-attention mechanism, the embedding vector is further processed into three parts: Query, Key, and Value. Correspondingly, the trained weight matrix is also divided into the weight matrix corresponding to the query part, the weight matrix corresponding to the key part, and the weight matrix corresponding to the value part.

[0170] Then, by multiplying the embedding vector with three different trained weight matrices. That is, multiplying the query part of the embedding vector with the weight matrix corresponding to the query part to obtain the query matrix; multiplying the key part of the embedding vector with the weight matrix corresponding to the key part to obtain the key matrix; multiplying the value part of the embedding vector with the weight matrix corresponding to the value part to obtain the value matrix.

[0171] Among them, the query matrix is used to measure the relevance of other parts, the key matrix is used to match with the query matrix to determine the attention weights, and the value matrix contains the information that needs to be weighted.

[0172] Step S53, multiply the query matrix Q with the transpose of the key matrix K to obtain the score matrix QK T , the formula is Q * K T = QK T ;

[0173] It should be noted that by calculating the dot product between the query matrix and the transpose of the key matrix, the score matrix is obtained.

[0174] Among them, each element in the score matrix represents a similarity score between a query and all keys. These scores reflect the correlation between different positions and are the basis for calculating the attention weights later.

[0175] Step S54, adjust the score matrix QK T and the mask matrix of the adaptive region self-attention mechanism through the attention adjustment mechanism to generate the attention adjustment result;

[0176] Specifically, one-dimensionalize the mask matrix of the adaptive region self-attention mechanism to obtain the one-dimensionalized mask matrix; then, transpose the one-dimensionalized mask matrix to obtain the transposed mask matrix; then, multiply the one-dimensionalized mask matrix with the transposed mask matrix to obtain the mask autocorrelation matrix; finally, subtract the product of the mask autocorrelation matrix and the preset parameter from the score matrix to obtain the attention adjustment result.

[0177] The step S60 of generating the attention feature vector based on the original image and the attention adjustment result includes:

[0178] Step S61: Multiply the value matrix V with the attention adjustment result to generate the attention feature vector Z'.

[0179] It should be noted that the attention feature vector contains the information weighted by the attention weights and reflects the importance of different positions in the image. The attention feature vector will be used as the input for subsequent processing steps to achieve various computer vision tasks, such as classification, detection, and segmentation.

[0180] In this embodiment, a masking mechanism and an attention adjustment mechanism are introduced. Through the masking mechanism, the vision large model can focus on key regions, thereby extracting more effective feature representations and improving accuracy. Through the attention adjustment mechanism, the vision large model can adaptively adjust the attention weights according to the characteristics of the input image and the task requirements.

[0181] Based on the embodiment described in steps S51 to S54 above, in step S54, the masking matrix of the adapted region self-attention mechanism is set as Mask, and the score matrix QK is adjusted through the attention adjustment mechanism T and the masking matrix of the adapted region self-attention mechanism to generate the attention adjustment result, including:

[0182] Step S541: Flatten the masking matrix Mask to obtain the one-dimensional masking matrix Mask_1D;

[0183] Among them, this embodiment will elaborate in detail on how to perform attention adjustment on the foreground and background based on the masking matrix of the adapted region self-attention mechanism.

[0184] Among them, as an implementation manner, all elements of the masking matrix (two-dimensional matrix) can be arranged in a one-dimensional array in a certain order (such as row-major or column-major) to obtain the one-dimensional masking matrix.

[0185] As another implementation manner, the masking matrix can also be flattened through built-in functions or methods in the deep learning framework.

[0186] Step S542: Transpose the one-dimensional masking matrix Mask_1D to obtain the transposed masking matrix (Mask_1D) T ;

[0187] Step S543: Multiply the one-dimensional masking matrix Mask_1D with the transposed masking matrix (Mask_1D) T to obtain the masking self-correlation matrix M, and the formula is: Mask_1D * (Mask_1D) T = M;

[0188] It should be noted that multiplying the one-dimensionalized mask matrix by its transposed mask matrix calculates the autocorrelation of the mask matrix, that is, the correlation between different positions in the mask.

[0189] Among them, the obtained mask autocorrelation matrix reflects the relative importance or relevance between different regions in the mask.

[0190] Step S544, subtract the product of the mask autocorrelation matrix M and the preset parameter k from the score matrix QK T to obtain a new QK T . If k > 0, suppress the background region; if k < 0, enhance the foreground region. The formula is: the new QK T = QK T ± k * M;

[0191] Determine the attention adjustment result according to the new QK T . Among them, subtracting the product of the score matrix (representing the similarity between the query matrix and the key matrix) and the mask autocorrelation matrix and the preset parameter is performed.

[0192] Among them, this preset parameter is used to control the influence degree of the mask autocorrelation matrix on the final attention adjustment result. Through the subtraction operation, the autocorrelation information in the mask matrix can be fused into the attention calculation, thereby realizing the adjustment of the attention weight.

[0193] Among them, determining the attention adjustment result according to the new QK T includes:

[0194] Take the new QK T as QK i T , and determine the attention adjustment result through the following expression:

[0195]

[0196] Among them, DA i is the attention adjustment result of the i-th row, QK i T is the result of multiplying the i-th row of the K matrix by the Q matrix, and the results of multiplying all rows of the K matrix by the Q matrix are the new QK T , is the scaling factor, k is the preset parameter, M is the mask autocorrelation matrix, and N is the number of rows of the K matrix.

[0197] The finally obtained attention adjustment result will be used to calculate the attention feature vector in the subsequent steps.

[0198] Use a specific example to explain the above process. Let the elements of the mask matrix of the adaptive region self-attention mechanism be BM = (BM ij ), and the elements of the QK T matrix be the elements of the pre-attention (Pre-Attn) matrix A = (A ij ), the elements of the attention suppression matrix be ABM = (ABM ij ), and the coefficient be k.

[0199] From the Hadamard product, formula (5) can be obtained:

[0200] ABM ij = A ij × BM ij ··········· (5)

[0201] Among them, the mask matrix BM is a two-dimensional matrix, and its elements are BM ij , which is used to indicate which positions need to be suppressed or enhanced.

[0202] Among them, the pre-attention matrix A (or called QK T ) is a two-dimensional matrix, and its elements are A ij , which is the result of the transpose multiplication of the query matrix Q and the key matrix K.

[0203] Among them, the attention suppression matrix ABM is a matrix generated by the Hadamard product of the mask matrix BM and the pre-attention matrix A, and its elements are ABM ij .

[0204] In addition, the coefficient k is a coefficient used to adjust the suppression intensity.

[0205] Suppose attn is an attention matrix updated over time or iteration steps, and ABM and the coefficient k are used to adjust this matrix. Thus, ABM suppresses attn, as shown in formula (6):

[0206] attn n+1 = attn n + ABM · k ············ (6)

[0207] Among them, attn n is the attention matrix before update, and attn n+1 is the attention matrix after update.

[0208] In addition, it should be noted that in the regional self-attention mechanism, when it is necessary to weaken the background attention, the theoretical value range of the preset parameter (the preset parameter is set to k) is (-∞, 0). However, since the background information has an auxiliary effect on tasks such as object detection and image classification, if the background information is significantly weakened, the model accuracy will decrease accordingly. A large number of experiments show that the value range of the parameter k is approximately (0, -1); when it is necessary to enhance the foreground attention, the theoretical value range of the parameter k is (0, +∞), and similarly, it should not be too large.

[0209] Taking the weakening of background attention as an example, when it is necessary to weaken the background attention, different values of the parameter k result in different degrees of background weakening. Based on different k-value dimensions of the model application method of the regional self-attention mechanism, including: k = 0 (not weakened), k = -0.2, k = -0.4, k = -0.6, k = -0.8, the attention situations of five different k-value dimensions.

[0210] Different attention suppression intensities have different effects on background suppression. The strength of attention is determined by the k value. The larger the k value, the stronger the attention allocated to this area, and vice versa, the weaker the attention allocated to this area. It can be obtained that: the larger the k value, the lower the degree of background weakening in attention, and vice versa, the smaller the k value, the greater the suppression of the background in attention.

[0211] In this embodiment, by one-dimensionalizing and transposing the mask matrix adapted to the regional self-attention mechanism, the subsequent calculation process can be simplified. Moreover, by calculating the mask self-correlation matrix, the distribution of attention in the sequence can be more accurately controlled.

[0212] Based on the embodiments described in the above steps S51 to S54, before the step S50, based on the original image and the mask matrix adapted to the regional self-attention mechanism, and through the attention adjustment mechanism for adjustment to generate the attention adjustment result, it includes:

[0213] Step S30, obtaining a training data set;

[0214] Step S40, training a pre-selected large visual model framework based on the training data set to obtain the weight matrices W q 、W k 、W v .

[0215] Among them, the pre-selectable large visual model framework has been described in the Figure 2 corresponding embodiment.

[0216] It should be noted that in the regional self-attention mechanism, the embedding vector is multiplied by the trained weight matrix to obtain a query matrix, a key matrix, and a value matrix, where the weight matrix is trained during the training of the visual large model framework.

[0217] Specifically, the weight matrix is randomly initialized first.

[0218] Then, during the training process, the weight matrix gradually learns the features of the training dataset. These features may be low-level features such as edges, textures, and shapes in the image, or may be more high-level features.

[0219] Through training, the weight matrix is gradually adjusted to the optimal state, enabling the visual large model to accurately fit the training dataset.

[0220] In this embodiment, by obtaining the training dataset and training the pre-selected visual large model framework based on the training dataset to obtain the visual large model, the image processing accuracy of the visual large model can be improved.

[0221] It should be noted that the regional self-attention mechanism proposed by the present invention adds two additional mechanisms compared to the self-attention mechanism. The first is the masking mechanism, and the second is the attention adjustment mechanism.

[0222] Correspondingly, the regional self-attention mechanism in the model application method based on the regional self-attention mechanism of the present invention is introduced below. Refer to Figure 3 , Figure 3 which is the schematic flowchart of the regional self-attention mechanism of the model application method based on the regional self-attention mechanism of the present invention. Specifically, this process includes the following steps:

[0223] Step S200: Multiply the embedding vector X by the weight matrices W q , W k , W v respectively to obtain the query matrix Q, the key matrix K, and the value matrix V.

[0224] Step S201: Multiply the query matrix Q by the transpose of the key matrix K to obtain QK T , where QK T contains QK 1 T , QK 2 T , QK 3 T ……QK N T .

[0225] Step S202: Process the original image through the masking mechanism to obtain the mask matrix Mask adapted to the regional self-attention mechanism.

[0226] Step S203: Input QK T and the mask matrix Mask of the adaptive region self-attention mechanism into the attention adjustment mechanism for adjustment to obtain an attention adjustment result.

[0227] Step S204: Multiply the value matrix V by the attention adjustment result to obtain Z'.

[0228] Further, in the said step S203: Input QK T and the mask matrix Mask of the adaptive region self-attention mechanism into the attention adjustment mechanism for adjustment to obtain an attention adjustment result, which includes steps S2031 to S2034:

[0229] Step S2031: Flatten the mask matrix Mask to obtain a flattened mask matrix.

[0230] Step S2032: Transpose the flattened mask matrix Mask to obtain a transposed mask matrix.

[0231] Step S2033: Multiply the flattened mask matrix by the transposed mask matrix to obtain a mask autocorrelation matrix M.

[0232] Step S2034: Subtract the product of the mask autocorrelation matrix M and a preset parameter k from QK T to obtain a new QK T . If k > 0, suppress the background region; if k < 0, enhance the foreground region. The formula is: new QK T = QK T ± k * M;

[0233] Determine the attention adjustment result according to the new QK T .

[0234] Among them, determining the attention adjustment result according to the new QK T includes:

[0235] Take the new QK T as QK i T , and determine the attention adjustment result through the following expression:

[0236]

[0237] where DA i is the attention adjustment result of the i-th row, QK i T is the result of multiplying the i-th row of the K matrix by the Q matrix, and the results of multiplying all rows of the K matrix by the Q matrix, the new QKT , where is the scaling factor, k is a preset parameter, M is the masked self-correlation matrix, and N is the number of rows of the K matrix.

[0238] An embodiment of the present application also proposes an object detection method for a model application method based on a regional self-attention mechanism, including:

[0239] Jointly training a preset visual self-attention model, a preset encoder, a preset decoder, and a preset one-to-one set matching model to obtain a trained visual self-attention model, encoder, decoder, and one-to-one set matching model;

[0240] Inputting the image to be detected, the first image feature output by the visual self-attention model, the second image feature output by the encoder, and the third image feature output by the decoder into the trained visual self-attention model, encoder, decoder, and one-to-one set matching model respectively to obtain the target object in the image to be detected and the position information of the target object, where the image obtained by the model application method based on the regional self-attention mechanism is used as the first image feature output by the visual self-attention model.

[0241] Further, jointly training the visual self-attention model, the encoder, the decoder, and the one-to-one set matching model to obtain a trained visual self-attention model, encoder, decoder, and one-to-one set matching model, including:

[0242] Inputting a preset image into the preset visual self-attention model to obtain a fourth image feature output from the preset visual self-attention model;

[0243] Inputting the fourth image feature into the preset encoder to obtain a fifth image feature output from the encoder;

[0244] Inputting the fifth image feature into a preset multi-scale feature extraction model to obtain a sixth image feature output;

[0245] Inputting the sixth image feature into each preset auxiliary head model to obtain a seventh image feature respectively output by each preset auxiliary head model;

[0246] Inputting the position encoding output by the preset multi-scale feature extraction model and the seventh image feature into the preset decoder to obtain an eighth image feature output from the decoder;

[0247] Inputting the eighth image feature into the preset one-to-one set matching model to obtain a first prediction set of the target object of the preset image;

[0248] Taking the seventh image feature respectively output by each preset auxiliary head model as a second prediction set, and constructing a loss function according to the first prediction set, the second prediction set, the first preset actual set, and the second preset actual set;

[0249] Adjust the parameters of the visual self-attention model, encoder, multi-scale feature extraction model, each auxiliary head model, decoder, and one-to-one set matching model according to the loss function value to obtain the trained visual self-attention model, encoder, multi-scale feature extraction model, each auxiliary head model, decoder, and one-to-one set matching model.

[0250] Furthermore, adjust the parameters of the visual self-attention model, encoder, multi-scale feature extraction model, each auxiliary head model, decoder, and one-to-one set matching model according to the loss function value to obtain the trained visual self-attention model, encoder, multi-scale feature extraction model, each auxiliary head model, decoder, and one-to-one set matching model, including:

[0251] Adjust the parameters of each auxiliary head model according to the loss function value constructed by the second prediction set corresponding to each auxiliary head model and the second preset actual set;

[0252] Based on the adjusted parameters of each auxiliary head model, adjust the parameters of the visual self-attention model, encoder, multi-scale feature extraction model, decoder, and one-to-one set matching model according to the loss function value constructed by the first prediction set and the first preset actual set;

[0253] Based on the adjusted parameters of the visual self-attention model, encoder, multi-scale feature extraction model, decoder, one-to-one set matching model, and each auxiliary head model, determine a new first prediction set, and determine a new loss function value corresponding to the new first prediction set and the first preset actual set;

[0254] When the difference between the loss function values corresponding to the current training round and the previous training round is less than the preset error, use the adjusted parameters of the visual self-attention model, encoder, multi-scale feature extraction model, decoder, one-to-one set matching model, and each auxiliary head model as the parameters of the trained visual self-attention model, encoder, multi-scale feature extraction model, decoder, one-to-one set matching model, and each auxiliary head model;

[0255] When the difference between the loss function values corresponding to the current training round and the previous training round is greater than or equal to the preset error, re-execute the step of adjusting the parameters of each auxiliary head model according to the loss function value constructed by the second prediction set corresponding to each auxiliary head model and the second preset actual set.

[0256] Furthermore, input the fifth image feature into the preset multi-scale feature extraction model to obtain the output sixth image feature, including:

[0257] Sample the fifth image feature using convolution kernels of different scales of a preset multi-scale feature extraction model, where the size of the convolution kernel is H×W, it has C input channels, the convolution kernel is K, and the size of the convolution kernel is k h ×k w , C' output channels, the convolution stride is s, the padding size is p, the sixth image feature output from the preset multi-scale feature extraction model is O, with a size of H'×W', where,

[0258]

[0259] where, O c′ (i, j) is the value of the C'-th channel of the sixth image feature at position (i, j), K c′,c (m, n) represents the (m, n) element of the C'-th output channel and the C-th input channel of the convolution kernel, I c (i·s + m - p, j·s + n - p) represents the pixel value of the C-th channel of the input feature map in the area covered by the convolution operation, b c′ is the bias term of the C'-th output channel, and the sixth image feature is O = {O 1 , O 2 ,..., O n}.

[0260] Further, input the sixth image feature into each preset auxiliary head model through the following expression to obtain the seventh image feature respectively output by each preset auxiliary head model:

[0261] Q aux = W q (σ(FCN(Flatten(O))) + α·P) + W a A

[0262] where, Q aux is the positive query, O is the sixth image feature, σ is the activation function, Flatten is the flattening function, FCN is the fully connected linear function, α is the learnable scalar weight, P is the position encoding, A is the auxiliary information vector, W q and W a are both learnable vector matrices.

[0263] Further, input the position encoding and the seventh image feature output by the preset multi-scale feature extraction model into the preset decoder through the following expression to obtain the eighth image feature output from the decoder:

[0264]

[0265] Final = LayerNorm(FCN(Attention(p q ) + Qaux ))

[0266] Among them, M is the number of sampling points of deformable self-attention, ΔP is the offset, and p q is the query position, and P sample ={P m |m = 1, 2,..., M} is the set of sampling points, and K Pm is the feature at the corresponding position obtained from the feature map K pij , V Pm is the feature at the corresponding position obtained from the feature map V pij , w ij is the weight, Attention(p q ) is the attention of the query position p q , is the positive query of the query position p q , LayerNorm is the normalization function, and Final is the eighth image feature output from the decoder.

[0267] In addition, an embodiment of the present application also proposes a model application device based on a regional self-attention mechanism, and the model application device based on a regional self-attention mechanism includes:

[0268] A data acquisition module for acquiring an original image;

[0269] A processing module for processing the original image through a masking mechanism to obtain a masking matrix adapted to the regional self-attention mechanism;

[0270] An attention adjustment module for adjusting based on the original image and the masking matrix adapted to the regional self-attention mechanism through an attention adjustment mechanism to generate an attention adjustment result; generating an attention feature vector based on the original image and the attention adjustment result;

[0271] A feature post-processing module for performing feature post-processing on the attention feature vector to obtain a processed image.

[0272] For the principle and implementation process of implementing the model application based on the regional self-attention mechanism in this embodiment, please refer to the above embodiments, and details are not described herein again.

[0273] In addition, an embodiment of the present application also proposes a device, and the device includes a memory, a processor, and a model application program based on a regional self-attention mechanism stored on the memory and executable on the processor. When the model application program based on the regional self-attention mechanism is executed by the processor, the steps of the model application method based on the regional self-attention mechanism as described above are implemented.

[0274] Since all the technical solutions of the foregoing embodiments are adopted when the model application program based on the regional self-attention mechanism is executed by a processor, it has at least all the beneficial effects brought by all the technical solutions of the foregoing embodiments, and thus will not be described in detail herein.

[0275] In addition, an embodiment of the present application further provides a computer-readable storage medium, on which a model application program based on the regional self-attention mechanism is stored. When the model application program based on the regional self-attention mechanism is executed by a processor, the steps of the model application method based on the regional self-attention mechanism as described above are implemented.

[0276] Since all the technical solutions of the foregoing embodiments are adopted when the model application program based on the regional self-attention mechanism is executed by a processor, it has at least all the beneficial effects brought by all the technical solutions of the foregoing embodiments, and thus will not be described in detail herein.

[0277] In this embodiment, through the above solution, specifically, an original image is obtained; the original image is processed by a masking mechanism to obtain a masking matrix adapted to the regional self-attention mechanism; based on the original image and the masking matrix adapted to the regional self-attention mechanism, adjustment is performed through an attention adjustment mechanism to generate an attention adjustment result; based on the original image and the attention adjustment result, an attention feature vector is generated; and the attention feature vector is post-processed to obtain a processed image. In the embodiment of the present invention, the original image is processed by the masking mechanism and the attention adjustment mechanism to generate an attention adjustment result, which can better extract image features and reduce the interference of noise or irrelevant information, thereby reducing data deviation and improving the accuracy of image processing.

[0278] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or method including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or method. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or method including the element.

[0279] The serial numbers of the embodiments of the present invention above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0280] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of the present invention.

[0281] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A model application method based on regional self-attention mechanism, characterized in that: The method comprises the following steps: Get the original image; Processing the original image through a mask mechanism to obtain a mask matrix of an adaptation region self-attention mechanism; Based on the mask matrix of the original image and the self-attention mechanism of the adaptation region, an attention adjustment mechanism is used to adjust and generate an attention adjustment result; Based on the original image and the attention adjustment result, generating an attention feature vector; Perform feature post-processing on the attention feature vector to obtain a processed image.

2. The method according to claim 1, characterized in that The step of processing the original image through a mask mechanism to obtain a mask matrix adapted to the regional self-attention mechanism includes: Performing instance segmentation on the original image to obtain a target mask matrix; Segmenting the target mask matrix based on a preset side length to obtain a plurality of segmented regions; Counting the pixel quantities of the plurality of segmented areas to obtain a block mask corresponding to each of the segmented areas; A mask matrix of the adaptation region self-attention mechanism is generated based on the block masks corresponding to each of the segmented regions.

3. The method according to claim 2, characterized in that The step of processing the original image by a mask mechanism to obtain a mask matrix adapted to the regional self-attention mechanism also includes: Set the target mask matrix to a two-dimensional array A, set the preset side length to size, and set the threshold to threshold. Create two a*b all-zero matrices, namely Co matrix as shown in formula (1) and Cr matrix as shown in formula (2), where the value of a is A.shape(0) / size and the value of b is A.shape(1) / size. The Co matrix is ​​used to record the number of pixels in the plurality of segmented regions, where co ij is the value corresponding to the matrix element, co ij The calculation formula is (3): The Cr matrix is ​​used to mask irrelevant areas, where cr ij is the value corresponding to the matrix element, cr ij The calculation formula is (4): Among them, the Cr matrix is ​​the mask matrix of the adaptation area self-attention mechanism.

4. The method according to claim 1, characterized in that The step of adjusting the mask matrix based on the original image and the self-attention mechanism of the adaptation region by the attention adjustment mechanism to generate the attention adjustment result comprises: Performing image embedding processing on the original image to obtain an embedding vector X; The embedding vector X and the trained weight matrix W q , W k , W v Perform matrix multiplication to obtain the query matrix Q, key matrix K and value matrix V. The formula is: Q=X*W q , K=X*W k , V=X*W v ; Perform matrix multiplication of the query matrix Q and the transpose of the key matrix K to obtain the score matrix QK T , the formula is Q*K T =QK T ; The score matrix QK is adjusted by the attention adjustment mechanism T and adjusting the mask matrix of the self-attention mechanism of the adaptation area to generate the attention adjustment result; The step of generating an attention feature vector based on the original image and the attention adjustment result comprises: The value matrix V is matrix-multiplied by the attention adjustment result to generate the attention feature vector Z'.

5. The method according to claim 4, characterized in that The mask matrix of the adaptation region self-attention mechanism is set to Mask, and the score matrix QK is adjusted by the attention adjustment mechanism T The step of adjusting the mask matrix of the self-attention mechanism of the adaptation region to generate the attention adjustment result includes: The mask matrix Mask is converted into one dimension to obtain a one-dimensional mask matrix Mask_1D; Transpose the one-dimensional mask matrix Mask_1D to obtain the transposed mask matrix (Mask_1D) T ; The one-dimensional mask matrix Mask_1D and the transposed mask matrix (Mask_1D) are combined. T Perform matrix multiplication to obtain the mask autocorrelation matrix M, the formula is: Mask_1D*(Mask_1D) T =M; The score matrix QK T Subtract the product of the mask autocorrelation matrix M and the preset parameter k to obtain the new QK T , if k>0, the background area is suppressed, if k<0, the foreground area is enhanced, the formula is: the new QK T =QK T ±k*M; according to the new QK T Determine attentional modulation outcomes.

6. The method according to claim 5, characterized in that According to the new QK T Identify attention regulation outcomes, including: The new QK T As QK i T , the attention adjustment result is determined by the following expression: Among them, DA i is the attention adjustment result of the i-th row, QK i T The result of multiplying the i-th row of the K matrix by the Q matrix. The result of multiplying all rows of the K matrix by the Q matrix is ​​the new QK T , is the scaling factor, k is the preset parameter, M is the mask autocorrelation matrix, and N is the number of rows of the K matrix.

7. The method according to claim 5, characterized in that When the background attention needs to be weakened, the value interval of the preset parameter k is (-∞, 0), and when the foreground attention needs to be enhanced, the value interval of the preset parameter k is (0, +∞).

8. The method according to claim 4, characterized in that The mask matrix based on the original image and the self-attention mechanism of the adaptation region is adjusted by the attention adjustment mechanism to generate the attention adjustment result, before the step includes: Get the training dataset; Based on the training data set, the pre-selected visual large model framework is trained to obtain the weight matrix W q , W k , W v .

9. A model application device based on regional self-attention mechanism, characterized in that: The device comprises: A data acquisition module, used for acquiring original images; A processing module, used to process the original image through a mask mechanism to obtain a mask matrix adapted to the regional self-attention mechanism; An attention adjustment module, configured to adjust the original image and the mask matrix of the self-attention mechanism of the adaptation region through the attention adjustment mechanism to generate an attention adjustment result; and generate an attention feature vector based on the original image and the attention adjustment result; The feature post-processing module is used to perform feature post-processing on the attention feature vector to obtain a processed image.

10. A model application device based on regional self-attention mechanism, characterized in that: The model application device based on the regional self-attention mechanism includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, the model application method based on the regional self-attention mechanism as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Panoramic segmentation method based on global edge attention

    CN112802039A

  • Image shadow removal model and construction method, device and application thereof

    CN115375589A

  • Medical image restoration method and system based on regional self-attention mechanism

    CN118552446A

  • Video coding method and device using in-loop filter based on transformer

    WO2023158127A1