Large visual model-based high-resolution image processing method and related system

By chunking, layer normalizing and local self-attention calculation of high-resolution images, combining residual connections and fully connected neural networks, the problems of high computational complexity and information tomography in high-resolution image processing are solved, and efficient and accurate image processing is achieved.

CN120339566APending Publication Date: 2025-07-18CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510426200.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Existing visual models have high computational complexity when processing high-resolution images, fixed window division leads to information failure, low memory access efficiency, and affects image processing efficiency and accuracy.

Method used

The image is divided into blocks and transcoded and mapped, and then layer normalization is performed to divide it into small pieces according to the preset size for local self-attention calculation. The feature map is processed using residual connections and fully connected neural networks to capture the relationship between features and form a high-quality final feature map.

Benefits of technology

It reduces the computational complexity, alleviates the problem of information faults, improves the efficiency and accuracy of high-resolution image processing, and meets the needs of intelligent inspections in the fields of power inspections and other fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339566A_ABST
    Figure CN120339566A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of image processing, and discloses a high-resolution image processing method based on a visual large model and a related system, and the method comprises the steps: carrying out the partitioning of an image, transcoding and mapping (obtaining a first feature image), converting a whole image processing problem into the processing of a smaller local block, and obtaining a second feature image; and the calculation amount caused by directly carrying out global self-attention calculation on the high-resolution whole image is effectively reduced. Layer normalization processing is carried out on the first feature map and the subsequent features, it is ensured that the features of all the layers have the unified mean value and variance, gradient disappearance and internal covariant offset can be relieved, and then the training process is more stable. According to the invention, the second feature map is divided into a plurality of small blocks according to the preset size, and local self-attention calculation is carried out in each small block, so that details and semantic relationships of a local region can be focused, high calculation cost caused by global self-attention is reduced, and full expression of local information is maintained at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing, and particularly relates to a high-resolution image processing method and related system based on a vision large model. Background Art

[0002] With the large-scale application of power drones in the inspection of power lines and equipment, the requirements for tasks such as intelligent analysis of inspection images, target detection, and defect and hidden danger identification are continuously increasing. This requires the ability to quickly and accurately process high-resolution images and extract useful information from them to ensure inspection efficiency and safety.

[0003] In recent years, the emergence of vision large models, such as the vision model ViT (Vision Transformer) and the self-attention model Swin Transformer, has provided new ideas and methods for tasks such as image target detection and classification. These models not only improve the semantic understanding ability of the model through the self-attention mechanism and hierarchical feature extraction, but also provide the possibility for high-resolution image processing.

[0004] In addition to the power industry, vision large models also show great potential in other high-resolution image application fields (such as satellite images, medical images, etc.), so it is of great practical significance to optimize their performance and improve their adaptability.

[0005] ViT divides an image into image patches of a fixed size and processes the image patches through the self-attention mechanism. Although it works well on low-resolution (such as 224×224) images, when the image resolution increases, the computational complexity approximately increases to the fourth power of the image side length, resulting in slow speed and low efficiency when processing high-resolution images.

[0006] Swin Transformer adopts a sliding window self-attention mechanism and divides the image into fixed windows for local attention calculation. Although this method reduces the computational amount, the fixed window division will lead to insufficient information interaction at the window edges or corners, forming "window mutations" and "attention dead corners", which in turn affect the transmission and fusion of the overall semantic information.

[0007] Whether it is ViT or Swin Transformer, their self-attention modules usually have a large number of memory read and write operations (such as multiple accesses to query vector Q, key vector K, value vector V, attention score matrix S, probability matrix P), resulting in low GPU utilization and limited memory bandwidth, thus affecting the calculation speed.

[0008] The optimized self-attention algorithm, Flash Attention, optimizes memory read and write (I / O) operations through strategies such as online update and tiling measurement, effectively reducing the number of memory accesses and significantly improving the computational efficiency in natural language processing (NLP) tasks.

[0009] Since NLP data is usually a one-dimensional sequence, the application of Flash Attention in two-dimensional image data requires further adaptation and optimization to overcome the different characteristics of two-dimensional data in terms of spatial structure and distribution, so as to be better applied to object detection and feature extraction tasks of high-resolution images.

[0010] Therefore, current research faces problems such as how to effectively reduce the computational complexity, optimize the memory access efficiency, and overcome the information discontinuity caused by fixed window partitioning while ensuring rich detail information in high-resolution images. Solving these challenges is of great significance for improving the efficiency and accuracy of image processing in fields such as power inspection. Summary of the Invention

[0011] The purpose of the present invention is to overcome the problems of high computational complexity and information discontinuity of fixed windows when ensuring high-resolution images, and to provide a high-resolution image processing method and related system based on a large vision model.

[0012] To achieve the above purpose, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a high-resolution image processing method based on a large vision model, including the following steps: Divide the input image into blocks, and perform transcoding mapping on the divided image to obtain a first feature map; Perform layer normalization on the first feature map to obtain a second feature map; According to a preset size, divide the second feature map into several small blocks, and perform local self-attention calculation on each small block to obtain a third feature map; Perform residual connection between the first feature map and the third feature map to obtain a fourth feature map; Process each feature of the fourth feature map to capture the mutual relationship between each feature and form a fifth feature map; Use the fifth feature map as the final feature map to perform detection according to the preset required information, and obtain the detection result of the final feature map, thus completing the high-resolution image processing method based on a large vision model.

[0013] A further improvement of the present invention lies in that the specific method of dividing the input image into blocks and performing transcoding mapping on the divided image to obtain a first feature map is as follows: Divide the input image into several non-overlapping blocks; Perform a convolution operation on each block to convert each block into a corresponding block vector; Combine all the block vectors to obtain a first feature map.

[0014] A further improvement of the present invention lies in the specific method of performing layer normalization on the first feature map to obtain a second feature map as follows: Obtain the block vectors of the first feature map; Calculate the mean and variance of the block vectors, adjust the block vectors to a distribution with a mean of 0 and a variance of 1, and add a scaling parameter and an offset parameter to the adjusted block vectors to obtain a second feature map.

[0015] A further improvement of the present invention lies in the specific method of dividing the second feature map into several small blocks according to a preset size and performing local self-attention calculation on each small block to obtain a third feature map as follows: Divide the second feature map into several small blocks according to a preset size; Combine each small block with a preset learnable linear mapping matrix to obtain a query vector, a key vector, and a value vector; Calculate an attention score matrix based on the query vector and the key vector; Combine the attention score matrix with a relative position bias term to obtain an attention score; Perform a normalized probability distribution calculation on the attention score using a normalized exponential function to obtain a probability matrix; Combine the probability matrix with the value vector to obtain a third feature map.

[0016] A further improvement of the present invention lies in the specific method of processing the features of the fourth feature map to capture the mutual relationship between each feature and form a fifth feature map as follows: Obtain all the features of the fourth feature map, perform normalization processing on all the features so that all the features have a distribution with a mean of 0 and a variance of 1; Send the normalized features into a fully connected neural network to capture the non-linear relationship between the features; Form a fifth feature map according to the non-linear relationship between all the features.

[0017] A further improvement of the present invention lies in that after forming the fifth feature map, perform layer normalization on the fifth feature map again until a fifth feature map that can be used as the final feature map is obtained.

[0018] A further improvement of the present invention lies in using the fifth feature map as the final feature map to perform detection according to the preset required information. When obtaining the detection result of the final feature map, the detection result of the final feature map includes the central coordinates of the final feature map, the width and height of the final feature map, the detection confidence information, and the category information to which the detected final feature map belongs.

[0019] In a second aspect, the present invention provides a high-resolution image processing system based on a vision large model, including: A chunking module for chunking the input image, performing transcoding mapping on the chunked image, and obtaining a first feature map; A layer normalization processing module for performing layer normalization processing on the first feature map to obtain a second feature map; A local self-attention calculation module for dividing the second feature map into several small chunks according to a preset size, performing local self-attention calculation on each small chunk, and obtaining a third feature map; A residual connection module for performing residual connection between the first feature map and the third feature map to obtain a fourth feature map; A relationship acquisition module for processing each feature of the fourth feature map, capturing the mutual relationship between each feature, and forming a fifth feature map; A detection module for using the fifth feature map as the final feature map to perform detection according to the preset required information, obtaining the detection result of the final feature map, and completing the high-resolution image processing method based on the vision large model.

[0020] A further improvement of the present invention lies in that the function of the chunking module is implemented by the following method: Dividing the input image into several non-overlapping chunks; Performing a convolution operation on each chunk to convert each chunk into a corresponding chunk vector; Combining all the chunk vectors to obtain a first feature map.

[0021] A further improvement of the present invention lies in that the function of the layer normalization processing module is implemented by the following method: Obtaining the chunk vectors of the first feature map; Calculating the mean and variance of the chunk vectors, adjusting the chunk vectors to a distribution with a mean of 0 and a variance of 1, and adding a scaling parameter and an offset parameter to the adjusted chunk vectors to obtain a second feature map.

[0022] A further improvement of the present invention lies in that the function of the local self-attention calculation module is implemented by the following method: Dividing the second feature map into several small chunks according to a preset size; Combining each small chunk with a preset learnable linear mapping matrix to obtain a query vector, a key vector, and a value vector; Calculate the attention score matrix based on the query vector and the key vector; Combine the attention score matrix with the relative position bias term to obtain the attention score; Perform a normalized probability distribution calculation on the attention score using the normalized exponential function to obtain the probability matrix; Combine the probability matrix with the numerical vector to obtain the third feature map.

[0023] A further improvement of the present invention lies in that the specific method for performing local self-attention calculation on each small block is as follows: Take each small block as the query vector, initialize or read the maximum value of the query vector, and normalize the accumulator and the value accumulator; Traverse the neighborhoods of all query vectors to obtain the key vectors and value vectors of the query vectors; Calculate the dot product of the query vector and the key vector to obtain the attention score matrix of the query vector neighborhood; Update the maximum value of the query vector according to the attention score matrix; Convert the updated maximum value into an exponential form and accumulate it into the exponential sum of the query vector neighborhood; If the maximum value of the query vector after update is different from the maximum value of the query vector before update, then adjust the weights of the normalized accumulator and the value accumulator according to the difference between the maximum value of the query vector after update and the maximum value of the query vector before update; if the maximum value of the query vector after update is the same as the maximum value of the query vector before update, then stop traversing other neighborhoods; After traversing all neighborhoods of the query vectors or stopping traversing other neighborhoods, divide the exponential sum of the query vector neighborhood by the normalized accumulator to obtain the finally weighted output representation, and at the same time update the maximum value of the current query vector; Take the updated maximum value of the current query vector and the adjusted value accumulator as the calculation result of the local self-attention.

[0024] A further improvement of the present invention lies in that the function of the relationship acquisition module is implemented by the following method: Obtain all the features of the fourth feature map, perform normalization processing on all the features so that all the features have a distribution with a mean of 0 and a variance of 1; Send the normalized features into a fully connected neural network to capture the non-linear relationships between the features; According to the non-linear relationships between all the features, form the fifth feature map.

[0025] A further improvement of the present invention lies in that it further includes a loop module, which is used to re-perform layer normalization processing on the fifth feature map after the fifth feature map is formed until the fifth feature map that can be used as the final feature map is obtained.

[0026] In a third aspect, the present invention provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the high-resolution image processing method based on a vision large model are implemented.

[0027] In a fourth aspect, the present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the high-resolution image processing method based on a vision large model are implemented.

[0028] Compared with the prior art, the present invention has the following beneficial effects: In the present invention, by dividing the image into blocks and performing transcoding mapping (to obtain a first feature map), the problem of processing the entire image is transformed into the processing of smaller local blocks, effectively reducing the computational complexity brought by directly performing global self-attention calculation on the high-resolution entire image. The present invention performs layer normalization processing on the first feature map and subsequent features to ensure that each layer of features has a unified mean and variance, which helps to alleviate gradient disappearance and internal covariate shift, and thus makes the training process more stable. By dividing the second feature map into several small blocks according to a preset size and performing local self-attention calculation within each small block, the present invention can focus on the details and semantic relationships of the local area, reduce the high computational cost brought by global self-attention, and at the same time maintain the full expression of local information. The present invention uses residual connection to fuse the original first feature map and the third feature map obtained through local self-attention processing, which not only retains the underlying details of the original image but also introduces local context relationships, thereby improving the information discontinuity problem caused by fixed window partitioning and promoting the continuous transmission of information between different scales. After performing normalization and fully connected neural network processing on the fourth feature map again, the present invention captures the global non-linear relationships between the normalized features to form a fifth feature map that is richer and has global consistency, providing a high-quality feature representation for subsequent detection tasks. Finally, using the fifth feature map as the input, the present invention extracts information such as target location and category through a preset detection head, making the overall method not only capable of efficiently processing high-resolution images but also performing excellently in terms of information integrity and detection accuracy, meeting the requirements of efficient intelligent detection in practical applications (such as power inspection). In summary, while reducing the computational complexity, the method effectively compensates for the information discontinuity caused by fixed windows through means such as multi-layer normalization, local self-attention, and residual connection, retains local details while capturing global semantics, and thus significantly improves the effect and efficiency of high-resolution image processing. Description of the Drawings

[0029] Figure 1 is a flowchart of the present invention; Figure 2 is a system diagram of the present invention; Figure 3 Schematic diagram of visual local self-attention pattern Figure 4 System diagram of Example 9 Detailed implementation manners

[0030] To further understand the content of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments are only for explaining the present invention rather than limiting it.

[0031] Example 1 Refer to Figure 1 , a high-resolution image processing method based on a visual large model, comprising the following steps: S1. Divide the input image into blocks, and perform transcoding mapping on the divided image to obtain a first feature map.

[0032] S2. Perform layer normalization processing on the first feature map to obtain a second feature map.

[0033] S3. According to a preset size, divide the second feature map into several small blocks, and perform local self-attention calculation on each small block to obtain a third feature map.

[0034] S4. Perform residual connection on the first feature map and the third feature map to obtain a fourth feature map.

[0035] S5. Process the features of the fourth feature map to capture the mutual relationship between the features and form a fifth feature map.

[0036] S6. Use the fifth feature map as the final feature map to perform detection according to the preset required information, obtain the detection result of the final feature map, and complete the high-resolution image processing method based on the visual large model.

[0037] Example 2 Refer to Figure 2 , a high-resolution image processing system based on a visual large model, comprising: A block division module for dividing the input image into blocks and performing transcoding mapping on the divided image to obtain a first feature map; A layer normalization processing module for performing layer normalization processing on the first feature map to obtain a second feature map; A local self-attention calculation module for dividing the second feature map into several small blocks according to a preset size and performing local self-attention calculation on each small block to obtain a third feature map; A residual connection module for performing residual connection on the first feature map and the third feature map to obtain a fourth feature map; A relationship acquisition module, which is used to process each feature of the fourth feature map, capture the mutual relationships between the features, and form a fifth feature map; A detection module, which is used to detect the fifth feature map as the final feature map according to the preset required information, obtain the detection result of the final feature map, and complete the high-resolution image processing method based on the vision large model.

[0038] Preferably, it further includes a loop module, which is used to perform layer normalization on the fifth feature map again after the fifth feature map is formed until the fifth feature map that can be used as the final feature map is obtained.

[0039] Embodiment 3: Based on the above embodiment, this embodiment further limits S1 as follows: S11, segment the input image into several non-overlapping blocks.

[0040] S12, perform a convolution operation on each block to convert each block into a corresponding block vector.

[0041] S13, combine all the block vectors to obtain a first feature map.

[0042] The block operation in this embodiment can enable the model to focus on the local regions of the image. Each block can independently capture the detailed features within the region, which is beneficial to extracting rich local information. After the image is segmented into blocks, each block can be processed in parallel, reducing the computational amount and memory occupation during global processing, and is especially suitable for high-resolution images. Through convolution transcoding in this embodiment, each block is mapped to a high-dimensional vector space, which can better express the semantic and texture information of the block, thereby improving the effect of subsequent self-attention processing. Therefore, this embodiment can not only fully extract the details of each local region in the image, but also effectively reduce the computational burden, and at the same time lay a solid foundation for subsequent more complex feature fusion and object detection tasks.

[0043] Embodiment 4: Based on the above embodiment, this embodiment further limits S2 as follows: S21, obtain the block vectors of the first feature map.

[0044] S22, calculate the mean and variance of the block vectors, adjust the block vectors to a distribution with a mean of 0 and a variance of 1, and add a scaling parameter and an offset parameter to the adjusted block vectors to obtain a second feature map.

[0045] The block vectors in this embodiment are adjusted to a distribution with a mean of 0 and a variance of 1, which helps to eliminate internal covariate shift, making the input distribution of each layer more stable, thus accelerating the training convergence of the network. The standardization process gives different block vectors a unified scale, which helps with the weight learning and feature fusion in subsequent processing and avoids training difficulties caused by differences in numerical ranges. By adding learnable scaling parameters and offset parameters in this embodiment, the network can automatically adjust the standardized feature distribution according to specific tasks, thereby enhancing the model's expressive ability while maintaining stability. Therefore, this embodiment not only makes the training process more stable and efficient but also provides a more balanced and adaptable feature representation for subsequent feature modeling and object detection tasks.

[0046] Embodiment 5: Based on the above embodiment, this embodiment further limits S3 as follows: S31. Divide the second feature map into several small blocks according to a preset size.

[0047] S32. Combine each small block with a preset learnable linear mapping matrix to obtain query vector Q, key vector K, and value vector V.

[0048] S33. Multiply query vector Q and key vector K to calculate the attention score matrix S.

[0049] S34. Add the relative position bias term to the attention score matrix S, and combine an attention mask or add an attention dropout mechanism to obtain attention score S1.

[0050] S35. Use the normalized exponential function softmax to calculate the normalized probability distribution of attention score S1 to obtain probability matrix P.

[0051] S36. Multiply probability matrix P and value vector V to obtain the third feature map.

[0052] In this embodiment, the image blocks are divided into several small blocks according to a fixed size, and each small block queries key vector K and value vector V from the W1×W2 small blocks closest to it. Figure 3 A sample with W1 = W2 = 3 is given. Since query vector Q, key vector K, and value vector V are all 6-dimensional vectors, their size is B×H×ht×wt×S×D, where B is the number of samples, H is the number of self-attention heads, ht is the number of small blocks in the vertical direction, wt is the number of small blocks in the horizontal direction, S is the number of image blocks in each small block, and D is the dimension of a single self-attention head.

[0053] To simplify the algorithm, the case where the chunks of the query vector Q and the key vector K span multiple small chunks is not considered. Therefore, to improve the utilization rate of the cache, a small chunk size as large as possible should be selected, and the small chunk size should preferably be a multiple of 64. It is noted that when the number of image patches in a small chunk is 8×8, the receptive field of self-attention has reached 24×24. For a feature map with a resolution of 16, the side length of its receptive field is 384 pixels, which has exceeded the processing size of most pre-trained models. At this time, the number of image patches in each chunk of the query vector Q and the key vector K is 64, and the cache occupancy is 33KB, which can be satisfied by most GPU graphics cards.

[0054] In the local self-attention mode, for different query vectors Q, the ranges of the query key vector K and the value vector V that need to be accessed in the on-chip loop are different. Therefore, it is necessary to calculate the width and height dimensions of the coordinates of the local query key vector K, the value vector V, and the attention bias term on the chip, and calculate the deviation of the data address to be loaded according to the coordinates.

[0055] The specific steps of local self-attention calculation are as follows: Step a: Initialize or read the local maximum , initialize (the denominator of the normalized exponential function softmax) to 1, and initialize the numerical cumulative value to 0.

[0056] Step b: For each small chunk in a certain row and a certain column, read the corresponding query vector Q, and then perform the following loop calculation (double loop for rows and columns): Step b1: Find the nearby 9 small chunks centered on this small chunk, with coordinates (curw, curh) respectively. If it exceeds the boundary of the entire image, delete it, and perform the following loop calculation (double loop for curw and curh): Step b1(1): Read the key vector K and the value vector V; Step b1(2): Calculate the attention score matrix S obtained by multiplying the query vector Q and the key vector K of the small chunk with coordinates (curw, curh). If there is a bias term, add it: S = dot(Q, K) + bias where dot represents the dot product.

[0057] Step b1(3): Iterative calculation:

[0058] where, is the current calculated local maximum, is the local maximum, represents the maximum value in the first dimension of S, and scale is the scaling factor.

[0059]

[0060]

[0061]

[0062] Among them, is the exponentiated fraction, represents the sum of p in the first dimension.

[0063]

[0064]

[0065]

[0066]

[0067] Among them, is an intermediate quantity, is the exponential sum of the current adjacent small block, acc is the cumulative score weighted value, and dot represents the dot product.

[0068] If , end the two-layer double loop.

[0069] Step c: Calculate the module output:

[0070]

[0071] Step d: Save the local maximum and the cumulative score weighted value acc.

[0072] The forward propagation of the image processing method is completed, and the inference process of the model is optimized.

[0073] In this embodiment, after dividing the second feature map into several small pieces, self-attention is calculated within each small piece, enabling the model to focus on feature interactions within the local region, fully capturing local structure and semantic information. By generating query, key, and value vectors and calculating attention scores based on them, the model can dynamically adjust the focus of attention, assigning higher weights to important features, thereby improving the quality of feature representation. This embodiment combines the relative position bias term with the attention score, which helps incorporate spatial relationships, enhances the model's perception of local position information, and further improves the effect of tasks such as object localization. This embodiment uses the softmax function to convert the attention score into a probability distribution, ensuring the numerical stability of the calculation process, facilitating gradient propagation, and the convergence of network training. Therefore, this embodiment can not only efficiently capture semantic and structural information in the local region but also achieve feature enhancement and dynamic weighting through the self-attention mechanism, improving the overall performance of the model in processing high-resolution images.

[0074] Embodiment 6: Based on the above embodiment, this embodiment further defines S5 as follows: S51. Obtain all the features of the fourth feature map and perform normalization processing on all the features so that all the features have a distribution with a mean of 0 and a variance of 1.

[0075] S52. Send the normalized features into a fully connected neural network to capture the non-linear relationships between the features.

[0076] S53. According to the non-linear relationships between all the features, form the fifth feature map.

[0077] In this embodiment, after normalizing all the features in the fourth feature map, each feature has a distribution with a mean of 0 and a variance of 1, which helps alleviate the training instability problem caused by the difference in numerical scales of different features. The normalized features are sent into a fully connected neural network. Using the powerful non-linear modeling ability of this network, complex interaction relationships between the features can be captured, thereby enhancing the discriminative ability of feature representation. This embodiment further processes the normalized features through a fully connected layer, which can integrate key information from various regions, making the finally formed fifth feature map more globally consistent and semantically rich, providing higher-quality features for subsequent detection and other tasks. Therefore, this embodiment not only helps improve the numerical stability and fusion effect of the features but also captures richer global relationships through non-linear mapping, enabling the final fifth feature map to better support downstream tasks.

[0078] Embodiment 7: Based on the above embodiment, an additional process is added as follows: After forming the fifth feature map, the fifth feature map is re - normalized until a fifth feature map that can be used as the final feature map is obtained.

[0079] In this embodiment, after multiple non - linear transformations and fusions, the features may exhibit scale deviation or distribution shift. Re - normalization helps to stabilize the values, making subsequent calculations (such as attention calculation, fully - connected operations, etc.) more robust. Each normalization can reduce the impact of internal covariate shift, ensure that the feature distributions between layers are relatively consistent, and contribute to the gradient transmission and the convergence of model training.

[0080] Embodiment 8: Based on the above - mentioned embodiment, this embodiment further defines S6 as follows: When using the fifth feature map as the final feature map to perform detection according to the preset required information and obtaining the detection result of the final feature map, the detection result of the final feature map includes the central coordinates of the final feature map, the width and height of the final feature map, the detection confidence information, and the category information to which the detected final feature map belongs.

[0081] Embodiment 9: Please refer to Figure 4 As shown, the present invention also provides an electronic device 100 for a high - resolution image processing method based on a vision large - model; the electronic device 100 includes a memory 101, at least one processor 102, a computer program 103 stored in the memory 101 and executable on the at least one processor 102, and at least one communication bus 104.

[0082] The memory 101 can be used to store the computer program 103. The processor 102 realizes the steps of the high - resolution image processing method based on the vision large - model described in Embodiment 1 by running or executing the computer program stored in the memory 101 and calling the data stored in the memory 101. The memory 101 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device 100 (such as audio data, etc.). In addition, the memory 101 can include non - volatile memory, such as a hard disk, a memory, a plug - in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other non - volatile solid - state storage devices.

[0083] The at least one processor 102 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 102 may be a microprocessor or the processor 102 may also be any conventional processor, etc. The processor 102 is the control center of the electronic device 100, and connects various parts of the entire electronic device 100 through various interfaces and lines.

[0084] The memory 101 in the electronic device 100 stores a plurality of instructions to implement a high-resolution image processing method based on a vision large model. The processor 102 can execute the plurality of instructions to implement: Divide the input image into blocks, perform transcoding mapping on the divided image to obtain a first feature map; Perform layer normalization processing on the first feature map to obtain a second feature map; According to a preset size, divide the second feature map into several small blocks, perform local self-attention calculation on each small block to obtain a third feature map; Perform residual connection on the first feature map and the third feature map to obtain a fourth feature map; Perform normalization processing on each feature of the fourth feature map, capture the mutual relationship between the normalized features, and form a fifth feature map; Use the fifth feature map as the final feature map to perform detection according to the preset required information, obtain the detection result of the final feature map, and complete the high-resolution image processing method based on the vision large model.

[0085] Embodiment 10: If the modules / units integrated in the electronic device 100 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, and read-only memory (ROM, Read-Only Memory).

[0086] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, system, or computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0087] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the specified function in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0088] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified function in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0089] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the processes Figure 1 one process or a plurality of processes and / or blocks Figure 1 steps for the functions specified in one block or a plurality of blocks.

[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement without departing from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A high-resolution image processing method based on a large vision model, characterized in that It includes the following steps: Divide the input image into blocks, perform transcoding mapping on the divided image to obtain the first feature map; Perform layer normalization on the first feature map to obtain the second feature map; According to the preset size, divide the second feature map into several small blocks, perform local self-attention calculation on each small block to obtain the third feature map; Perform residual connection on the first feature map and the third feature map to obtain the fourth feature map; Process each feature of the fourth feature map to capture the mutual relationship between each feature, forming the fifth feature map; Use the fifth feature map as the final feature map to perform detection according to the preset required information, obtain the detection result of the final feature map, and complete the high-resolution image processing method based on the vision large model.

2. The high-resolution image processing method based on a visual large model according to claim 1, wherein The specific method for dividing the input image into blocks and performing transcoding mapping on the divided image to obtain the first feature map is as follows: Segment the input image into several non-overlapping blocks; Perform a convolution operation on each block to convert each block into a corresponding block vector; Combine all block vectors to obtain the first feature map.

3. The high-resolution image processing method based on a visual large model according to claim 1, wherein The specific method for performing layer normalization on the first feature map to obtain the second feature map is as follows: Obtain the block vectors of the first feature map; Calculate the mean and variance of the block vectors, adjust the block vectors to a distribution with a mean of 0 and a variance of 1, and add a scaling parameter and an offset parameter to the adjusted block vectors to obtain the second feature map.

4. The high-resolution image processing method based on a vision large model according to claim 1, wherein, The specific method for dividing the second feature map into several small blocks according to the preset size and performing local self-attention calculation on each small block to obtain the third feature map is as follows: Divide the second feature map into several small blocks according to the preset size; Combine each small block with a preset learnable linear mapping matrix to obtain a query vector, a key vector, and a value vector; Calculate the attention score matrix according to the query vector and the key vector; Combine the attention score matrix with the relative position bias term to obtain the attention score; Perform normalized probability distribution calculation on the attention score using the normalized exponential function to obtain the probability matrix; Combine the probability matrix with the value vector to obtain the third feature map.

5. The high-resolution image processing method based on a visual large model according to claim 4, wherein, The specific method for performing local self-attention calculation on each small block is as follows: Take each small block as the query vector, initialize or read the maximum value of the query vector, and normalize the accumulator and the value accumulator; Traverse the neighborhood of all query vectors to obtain the key vector and the value vector of the query vector; Calculate the dot product of the query vector and the key vector to obtain the attention score matrix of the query vector neighborhood; Update the maximum value of the query vector according to the attention score matrix; Convert the updated maximum value into an exponential form and accumulate it into the exponential sum of the query vector neighborhood; If the maximum value of the updated query vector is different from the maximum value of the query vector before update, adjust the weights of the normalized accumulator and the value accumulator according to the difference between the maximum value of the updated query vector and the maximum value of the query vector before update; if the maximum value of the updated query vector is the same as the maximum value of the query vector before update, stop traversing other neighborhoods; After traversing all the neighborhoods of the query vectors or stopping traversing other neighborhoods, the exponential sum of the query vector neighborhoods is divided by the normalized accumulator to obtain the finally weighted output representation, and at the same time, the maximum value of the current query vector is updated; The sum of the maximum value of the updated current query vector and the adjusted value accumulator is used as the calculation result of local self-attention.

6. The high-resolution image processing method based on a visual large model according to claim 1, wherein The specific method of processing each feature of the fourth feature map to capture the mutual relationship between each bit of features and form the fifth feature map is as follows: Obtain all the features of the fourth feature map, and perform normalization processing on all the features so that all the features have a distribution with a mean of 0 and a variance of 1; Send the normalized features into a fully connected neural network to capture the non-linear relationship between the features; According to the non-linear relationship between all the features, form the fifth feature map.

7. The high-resolution image processing method based on a visual large model according to claim 1, wherein After forming the fifth feature map, perform layer normalization processing on the fifth feature map again until the fifth feature map that can be used as the final feature map is obtained.

8. The high-resolution image processing method based on a visual large model according to claim 1, wherein Use the fifth feature map as the final feature map to perform detection according to the preset required information. When obtaining the detection result of the final feature map, the detection result of the final feature map includes the center coordinates of the final feature map, the width and height of the final feature map, the detection confidence information, and the category information to which the final feature map belongs.

9. A high-resolution image processing system based on a large vision model, characterized in that, Including: A block module for dividing the input image into blocks, performing transcoding mapping on the divided image, and obtaining the first feature map; A layer normalization processing module for performing layer normalization processing on the first feature map to obtain the second feature map; A local self-attention calculation module for dividing the second feature map into several small blocks according to the preset size, performing local self-attention calculation on each small block, and obtaining the third feature map; A residual connection module for performing residual connection on the first feature map and the third feature map to obtain the fourth feature map; A relationship acquisition module for processing each feature of the fourth feature map to capture the mutual relationship between each feature and form the fifth feature map; A detection module for using the fifth feature map as the final feature map to perform detection according to the preset required information, obtaining the detection result of the final feature map, and completing the high-resolution image processing method based on the vision large model.

10. The high-resolution image processing system based on a large vision model according to claim 9, wherein The function of the block module is realized by the following method: Divide the input image into several non-overlapping blocks; Perform a convolution operation on each block to convert each block into a corresponding block vector; Combine all the block vectors to obtain the first feature map.

11. The high-resolution image processing system based on a visual large model according to claim 9, wherein, The function of the layer normalization processing module is realized by the following method: Obtain the block vectors of the first feature map; Calculate the mean and variance of the block vectors, adjust the block vectors to a distribution with a mean of 0 and a variance of 1, and add a scaling parameter and an offset parameter to the adjusted block vectors to obtain the second feature map.

12. The high-resolution image processing system based on a large vision model according to claim 9, wherein, The function of the local self-attention calculation module is realized by the following method: Divide the second feature map into several small blocks according to the preset size; Combine each small block with a preset learnable linear mapping matrix to obtain a query vector, a key vector, and a value vector; Calculate the attention score matrix according to the query vector and the key vector; Combine the attention score matrix with the relative position bias term to obtain the attention score; The normalized exponential function is used to calculate the normalized probability distribution of the attention scores, obtaining a probability matrix; The probability matrix is combined with the numerical vector to obtain the third feature map.

13. The high-resolution image processing system based on a large vision model according to claim 9, wherein The function of the relationship acquisition module is implemented by the following method: All features of the fourth feature map are obtained and normalized so that all features have a distribution with a mean of 0 and a variance of 1; The normalized features are fed into a fully connected neural network to capture the non-linear relationships between the features; According to the non-linear relationships between all features, a fifth feature map is formed.

14. The high-resolution image processing system based on a visual large model according to claim 9, wherein It further includes a loop module for re-performing layer normalization on the fifth feature map after the fifth feature map is formed until a fifth feature map that can be used as the final feature map is obtained.

15. An electronic device, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the high-resolution image processing method based on the vision large model described in any one of claims 1 to 8.

16. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the high-resolution image processing method based on the vision large model described in any one of claims 1 to 8.