An Image Processing Method and System Based on a Large Artificial Intelligence Model

By constructing an undirected graph structure and dynamic noise intensity calibration mechanism, the problem of difficult to balance computing efficiency and feature integrity in the image processing system is solved, and efficient image processing effect and improvement of medical image segmentation accuracy is achieved.

CN119991438BActive Publication Date: 2025-07-18西安圣瞳科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510458171.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

Existing image processing systems are difficult to balance computing efficiency and feature integrity at the block processing level, and static fusion strategies cannot adapt to complex texture changes. There is feature space mismatch and insufficient noise constraints during cross-level feature distillation, resulting in limited model generalization capabilities.

Method used

An undirected graph structure is constructed using overlapping blocking strategy, a global topological descriptor is generated by calculating the node similarity weight, a noise vector is generated by combining noise intensity and structural constraints, and a feature distillation is performed using the student network and the teacher network, and the final distillation feature is generated by combining cosine similarity matching and attention mechanism, and priority allocation and calculation graph optimization are performed to finally output high-quality images.

Benefits of technology

The dual improvement of image processing efficiency and accuracy is achieved, the boundary effect of traditional blocking methods is alleviated, the generalization error is reduced, and the medical image segmentation accuracy comparable to FP32 is achieved with FP16 accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991438B_ABST
    Figure CN119991438B_ABST
Patent Text Reader

Abstract

The present invention discloses an image processing method and system based on an artificial intelligence large model, which relates to the field of artificial intelligence technology and includes: obtaining an original image and dividing it into overlapping blocks, and generating an image block feature set through a fusion strategy; constructing an undirected graph structure based on the image block feature set, calculating node similarity weights and generating a global topological descriptor; generating importance weights according to the dimension importance of the global topological descriptor, combining the noise intensity and structural constraints to generate a noise vector, and generating an enhancement vector and an offset matrix through a UNet encoder-decoder; using a student network and a teacher network to perform feature distillation on the enhancement vector, and generating a final distilled feature through cosine similarity matching and an attention mechanism. In the knowledge distillation link of the present invention, a dynamic noise intensity calibration mechanism is introduced, and precise constraint of the noise direction is achieved through variance parameter linkage control, so that the generalization error is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to an image processing method and system based on an artificial intelligence large model. Background Art

[0002] With the rapid development of deep learning technology in the field of computer vision, an image processing method based on a convolutional neural network (CNN) has formed a complete technical system. Traditional image block processing mostly adopts a fixed-size sliding window or non-overlapping block strategy, which can ensure the stability of local feature extraction, but there are problems such as significant boundary effects and loss of global structural information. In recent years, graph neural networks (GNNs) have shown unique advantages in feature fusion by constructing node-edge structures to represent the topological relationship of images. However, existing methods still have limitations in dynamic weight allocation and cross-level feature interaction.

[0003] Current image processing systems generally face multi-dimensional technical challenges: at the block processing level, traditional methods are difficult to balance computational efficiency and feature integrity; at the feature fusion level, static fusion strategies cannot adapt to complex texture changes; at the model optimization level, static computational graphs are difficult to meet the requirements of dynamic resource allocation. It is particularly worth noting that there are problems of feature space mismatch in the cross-level feature distillation process of existing technologies, and there is a lack of effective noise constraint mechanisms, resulting in limited model generalization ability. Summary of the Invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an image processing method based on an artificial intelligence large model to solve the problems of cross-level feature space mismatch and insufficient dynamic noise constraint in the multi-level feature distillation process.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, the present invention provides an image processing method based on an artificial intelligence large model, which includes: acquiring an original image and dividing it into overlapping blocks, and generating a set of image block features through a fusion strategy;

[0008] Constructing an undirected graph structure based on the set of image block features, calculating node similarity weights, and generating a global topological descriptor;

[0009] Generating importance weights according to the dimension importance of the global topological descriptor, generating a noise vector by combining noise intensity and structural constraints, and generating an enhancement vector and an offset matrix through a UNet encoder-decoder;

[0010] Feature distillation of the enhanced vector is performed using the student network and the teacher network, and the final distilled features are generated through cosine similarity matching and the attention mechanism;

[0011] Priority assignment and computational graph optimization are performed on the final distilled features and the offset matrix, and a low-precision inference result is obtained through block parallel computing;

[0012] The low-precision inference result is subjected to super-resolution reconstruction and color correction, and a high-quality image after real-time processing is output through asynchronous transmission technology.

[0013] As a preferred embodiment of the image processing method based on the artificial intelligence large model of the present invention, wherein: the steps of generating the image patch feature set include the following,

[0014] Obtain the size of the original image, divide the image into multiple overlapping small patches, and set the overlapping area;

[0015] The features of the overlapping area are merged through a fusion strategy of taking the average value and the maximum value to generate a fused image patch set;

[0016] Based on the fused image patch set, local texture and edge features are extracted, and through downsampling and multi-layer convolution processing, a feature set of the image patches is obtained.

[0017] As a preferred embodiment of the image processing method based on the artificial intelligence large model of the present invention, wherein: the steps of generating the global topological descriptor include the following,

[0018] Assign unique identifiers to the image patch feature set to form a node list, calculate the Euclidean distance similarity of node pairs and normalize it to a weight value;

[0019] Construct an undirected graph structure with nodes as vertices and weights as edges, label the node positions and edge semantic labels; encode the node features with the position and semantic information, update the node features through message passing and multiple rounds of aggregation iteration, and perform an average operation on the node features to generate a global topological descriptor.

[0020] As a preferred embodiment of the image processing method based on the artificial intelligence large model of the present invention, wherein: the steps of generating the enhanced vector and the offset matrix include the following,

[0021] Calculate the importance weight according to the volatility of the global topological descriptor dimension value, preset a dynamic coefficient to mix the initial noise and the importance weight according to the iteration stage, generate a dynamic noise intensity and convert it into a noise vector;

[0022] Input the noise vector into the UNet encoder-decoder to generate mean and variance parameters;

[0023] The perturbation term is obtained by scaling the noise vector with the variance parameter and superimposed on the mean to generate the enhanced vector;

[0024] The position index of the image patch is encoded into a dense vector and concatenated with the enhanced vector, and the position-aware feature is generated through a fully connected layer and the bias is eliminated to obtain the offset matrix.

[0025] As a preferred solution of the image processing method based on the large artificial intelligence model described in the present invention, wherein: the generation of the final distilled feature includes the following steps,

[0026] The student network receives the enhanced vector and generates local features, and the teacher network generates global features, and the feature alignment is performed through cosine similarity matching and contrast loss;

[0027] The global attention map of the teacher network guides the student network to focus on the key areas, and the final distilled feature is generated by combining the adversarial training of the generator and the discriminator.

[0028] As a preferred solution of the image processing method based on the large artificial intelligence model described in the present invention, wherein: the obtaining of the low-precision inference result includes the following steps,

[0029] Priorities are assigned to the final distilled features and the precision format is bound, the high-precision features are compressed to generate low-precision features, and the joint state vector is generated by combining the gradient complexity;

[0030] The joint state vector is passed through the policy network inference generation module to activate the probability;

[0031] Resources are allocated to the high-priority modules and the parameters are adjusted, the low-priority modules are deleted to generate optimized calculations, the low-precision features are scaled and corrected and anti-overflow processing is performed, and the block features are divided and independent video memories are allocated;

[0032] Based on the hardware utilization rate, the matrix operations are dynamically scheduled for block execution to generate the low-precision inference result.

[0033] As a preferred solution of the image processing method based on the large artificial intelligence model described in the present invention, wherein: the output of the high-quality image after real-time processing includes the following steps,

[0034] The low-precision inference result is expanded to the original size through interpolation, the details are reconstructed and repaired through the super-resolution model, the color space is converted and gamma correction is performed, and it is distributed to the multi-screen interface through asynchronous transmission technology, and the compression algorithm is used to adapt to the requirements of the output device to obtain the high-quality image after real-time processing.

[0035] In a second aspect, the present invention provides an image processing system based on a large artificial intelligence model, including,

[0036] The image block division module acquires the original image and divides it into overlapping blocks, and generates an image block feature set through a fusion strategy;

[0037] The topology construction module constructs an undirected graph structure based on the image block feature set, calculates the node similarity weights, and generates a global topology descriptor;

[0038] The noise enhancement module generates importance weights according to the dimension importance of the global topology descriptor, combines the noise intensity and the structure constraint to generate a noise vector, and generates an enhanced vector and an offset matrix through a UNet encoder-decoder;

[0039] The knowledge distillation module performs feature distillation on the enhanced vector using a student network and a teacher network, and generates a final distilled feature through cosine similarity matching and an attention mechanism;

[0040] The calculation optimization module performs priority assignment and computational graph optimization on the final distilled feature and the offset matrix, and obtains a low-precision inference result through block parallel calculation;

[0041] The image output module performs super-resolution reconstruction and color correction on the low-precision inference result, and outputs a high-quality image after real-time processing through an asynchronous transmission technology.

[0042] In a third aspect, the present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and: when the computer program is executed by the processor, any step of the image processing method based on an artificial intelligence large model as described in the first aspect of the present invention is implemented.

[0043] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and: when the computer program is executed by the processor, any step of the image processing method based on an artificial intelligence large model as described in the first aspect of the present invention is implemented.

[0044] The beneficial effects of the present invention are as follows: by constructing a multi-dimensional collaborative optimization framework, the double improvement of image processing efficiency and accuracy is realized; at the feature extraction level, an overlapping block strategy is adopted in combination with the construction of a dynamic graph structure, effectively alleviating the boundary effect of traditional block methods; in the knowledge distillation link, a dynamic noise intensity calibration mechanism is introduced, and the precise constraint of the noise direction is realized through the linkage control of variance parameters, reducing the generalization error; the computational graph optimization module adopts a resource pre-allocation strategy guided by activation probability, compresses the model calculation amount while maintaining the original accuracy; the hierarchical precision allocation scheme realizes the medical image segmentation accuracy equivalent to FP32 under the FP16 precision through the key layer feature protection mechanism, and has a significant improvement compared with traditional quantization methods. Description of the Drawings

[0045] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0046] Figure 1 This is the flowchart of image block division and feature extraction in this embodiment.

[0047] Figure 2 This is the schematic diagram of topology construction and noise enhancement in this embodiment.

[0048] Figure 3 This is the schematic diagram of feature distillation and calculation optimization in this embodiment.

[0049] Figure 4 This is the schematic diagram of image output and real-time rendering in this embodiment. Specific Embodiments

[0050] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings of the specification.

[0051] In the following description, many specific details are set forth to facilitate a thorough understanding of the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0052] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that mutually excludes other embodiments.

[0053] In this embodiment, referring to Figures 1 to 4 , this embodiment provides an image processing method based on an artificial intelligence large model, including the following steps:

[0054] S1. Obtain the size of the original image, divide the image into multiple N×N small blocks, and starting from the upper left corner, gradually move to cover the entire image area; to reduce boundary effects, set a 50% overlap between adjacent small blocks. For each overlapping area, extract features from adjacent small blocks and select an appropriate fusion strategy to merge the features of the overlapping area. Common fusion strategies include taking the average, maximum, weighted average, etc. For example, taking the average can smooth the change of feature values, while taking the maximum can retain the most significant features; according to the selected strategy, fuse the features of the overlapping area. For example, when choosing to take the average, add all the features of the overlapping area and divide by the number of features to obtain the fused feature vector; replace the features of the corresponding overlapping area in the original small block with the fused feature vector to obtain a set of overlapping image block partitions.

[0055] According to the specific image processing task, determine the required type of features. For example, if it is an image classification task, high-level semantic features need to be extracted; if it is an object detection task, edge and texture features need to be extracted; find a pre-trained CNN model (Convolutional Neural Network) suitable for the corresponding image processing task; convert the overlapping image block partitions into the format required by the CNN model, and normalize the pixel values of the image blocks to scale them to the standardized range; batch process the scaled multiple image blocks; the CNN model scans the batch-processed multiple image blocks through a sliding small window (such as 3×3 pixels) to extract local texture and edge features and generate a preliminary feature map; perform downsampling (such as 2×2 max pooling) on the preliminary feature map to retain significant features and reduce the computational amount; after multiple layers of convolution and pooling, obtain a set of features of the image partitions.

[0056] Assign a unique identifier to each image block in the set of features of the image partitions to form a node list, where each node corresponds to the features of an image block; traverse all node pairs, calculate the similarity of each pair of nodes through the Euclidean distance to measure the difference between the features of the image blocks; preset a constant to normalize the similarity result to obtain a weight value reflecting the degree of similarity. The smaller the weight, the more similar the nodes; use the nodes as vertices and the weights as edges, and select an adjacency matrix or adjacency list to represent the undirected graph structure; label the position of each node in the original image in the undirected graph structure (such as node 1 is located in the upper left corner of the image), and add semantic labels to the edges (such as strong texture similarity or weak color difference) to obtain an enhanced set of image block partitions.

[0057] Further explanation, when constructing an undirected graph, if the number of edges is too large, set a weight threshold and only retain the edges with weights less than the weight threshold to reduce complexity.

[0058] Use the original feature vector corresponding to each node in the enhanced image block set as the initial representation; encode the position coordinates of the nodes and the semantic labels of the edges into low-dimensional vectors, and concatenate them with the original feature vectors to obtain the initialized node feature matrix; each node sends its own feature to all neighbor nodes, and the message content is the node feature itself; after all nodes receive the messages from neighbor nodes, they merge the information through an aggregation function (such as summation, averaging) to generate the aggregated temporary feature; update the node feature according to the aggregated temporary feature, and perform multiple rounds of propagation on the updated node feature. In each round, the node resends the message based on the latest feature and aggregates the neighbor information; as the number of rounds increases, the node feature gradually fuses the information of farther neighbors, captures the global structure, and dynamically regulates the importance of different hierarchical features through a learnable weight matrix to obtain the final node feature matrix after multiple rounds of iteration; perform an average operation on the features of all nodes in the final node feature matrix to generate the global topology descriptor.

[0059] It should be noted that the global topology descriptor is a multi-dimensional vector (such as containing K dimensions), and each dimension reflects different attributes of the global structure of the image (such as the distribution density of object connected regions, the number of holes, etc.).

[0060] S2. For each dimension in the global topology descriptor, calculate the corresponding weight value. When the value of the corresponding dimension is smaller, the weight is higher, indicating that this topological feature is more stable and important; conversely, if the corresponding dimension threshold is set low and the fluctuation is small (such as the dimension value is always close to the preset dimension threshold), then suppress the weight through an exponential function; add up the weights of all dimensions, take the reciprocal of the sum as the normalization constant, and then scale the weights of each dimension proportionally to finally generate the importance weight.

[0061] Furthermore, the purpose of generating the importance weight is to ensure that the value of the importance weight is between 0 and 1, which is used to quantify the overall importance of the global topological structure. The closer the importance weight is to 1, the more critical the global structure is.

[0062] Preset a dynamically increasing coefficient from 0 to 1 to control the mixing ratio of noise and structural constraints, and preset an initial noise intensity as the initial noise level in the diffusion process to provide an exploratory noise source; according to the dynamic coefficient of the current iteration stage, fuse the initial noise and the structural importance weight proportionally, specifically as follows:

[0063] In the initial iteration stage, the noise is mainly at the initial value to ensure that the diffusion process has sufficient randomness to explore the feature space; in the later iteration stage, the importance weight is mainly dominant, constraining the noise direction to conform to the global topological law (such as suppressing irrelevant texture noise and strengthening object boundaries), and finally generating the dynamic noise intensity; converting the dynamic noise intensity into a noise vector that matches the image patches in the enhanced image patch set; inputting the converted noise vector and the enhanced image patch set into the encoder of UNet in parallel, and through multi-layer convolution and pooling operations, extracting the interaction pattern between features and noise, and the encoder outputs a low-dimensional intermediate representation; inputting the low-dimensional intermediate representation into the decoder of UNet to generate the mean parameter and the variance parameter.

[0064] Further explanation, after generating the variance parameter, calibration is required, which means forcing the variance parameter to be coordinated with the dynamic noise intensity through the loss function. If the dynamic noise intensity is high, the variance parameter tends to increase, allowing a larger noise perturbation. If the dynamic noise intensity is low, the variance parameter is restricted to a smaller range to ensure feature stability.

[0065] Scaling the noise vector by the variance parameter to obtain the noise perturbation term; superimposing the noise perturbation term on the mean parameter to generate the enhanced vector; the image patch position index includes the row index and column index of the image patch. Converting the row index and column index of the image patch into a dense position vector and splicing it with the corresponding enhanced vector and inputting it into the fully connected layer to generate the position-aware feature; performing matrix subtraction and scaling operations on the position-aware feature to eliminate the computational bias caused by noise or local perturbation to obtain the offset matrix.

[0066] S3. The student network receives the enhanced vector and generates local student features that match the output dimension of the teacher network through multi-layer convolutional and pooling operations. The teacher performs global average pooling on the enhanced vector to generate global teacher features. The local features generated by the student network are matched with the global features generated by the teacher network pixel by pixel to calculate the cosine similarity. A similarity threshold is set. In this embodiment, the similarity threshold is set to 0.85. If the similarity between the local student features and the global teacher features is lower than the similarity threshold, the similarity gap of the contrast loss is calculated and backpropagated to force the local features of the student network to be highly similar to the teacher network, obtaining the preliminary distilled features. The teacher network generates a teacher attention map during the global average pooling operation. The preliminary distilled features are input into the self-attention mechanism to generate a student attention map. The L2 norm difference between the teacher attention map and the student attention map is calculated to guide the student network to focus on the regions considered key by the teacher network (such as object boundaries or texture-dense areas). The preliminary distilled features and the student attention map are concatenated through a fully connected layer to obtain the student refined features. A generator is set up to map the student refined features to the feature space of the teacher network through the generator. A discriminator is set up to distinguish between the real teacher features and the student refined features mapped to the feature space of the teacher network by the generator. The student refined features output by the generator and the student refined features before mapping are concatenated to generate the final distilled features.

[0067] S4. The final distilled features are used as global feature candidates. The average value of each dimension in the final distilled features is calculated, and the core trend of each dimension is retained to generate a global feature vector. The gradient change intensity of each pixel point in each image patch in the enhanced image patch set is calculated (such as the change amplitude in the horizontal and vertical directions) to generate a two-dimensional gradient matrix. The images in the two-dimensional gradient matrix are divided into multiple small regions, and the gradient change intensity within each region is summed to obtain a local complexity value. The complexity values of all local regions are summed to generate a scalar local complexity, which reflects the overall detail complexity of the image. The global feature vector and the scalar local complexity are concatenated to form a joint state vector.

[0068] The joint state vector is divided into multiple sub-vectors according to a preset rule. The divided sub-vectors are input into a multi-layer perceptron for policy network inference, as follows:

[0069] The input layer receives the divided sub-vectors and performs a linear transformation by multiplying with a weight matrix through a fully connected layer to generate an intermediate representation. The ReLU function is applied to the intermediate representation to set negative values to zero and retain positive values. The intermediate layer further linearly transforms the output of the first layer, and a Dropout layer (random suppression) is inserted after the output of the intermediate layer to randomly ignore the values of some neurons to suppress overfitting and extract abstract features. The output layer maps the extracted abstract features to the module activation probability space to generate the initial value of the module activation probability.

[0070] The initial probability of each module is exponentiated to convert it into a non - negative value, and the non - negative values of all modules are added together and then compressed into the range of 0 to 1 to obtain the normalized activation probability; set an activation probability threshold, mark the modules with the normalized activation probability greater than or equal to the activation probability threshold as high - priority, and mark the unactivated modules as low - priority; further sort the high - priority modules according to the activation probability, the higher the probability, the higher the priority, pre - allocate resource requirements for each high - priority module to form a resource pre - allocation plan; perform a downgrading process on the low - priority modules; according to the processing results of high - priority and low - priority modules, delete all unactivated modules and cut off the input / output connections of unactivated modules, only retaining the backbone path; dynamically adjust the internal parameters of the module (such as convolution kernel size, pooling stride) according to the remaining module resources after deletion to generate an optimized computational graph configuration.

[0071] S5. Assign priorities according to the role of the final distilled features in the corresponding task. For example, assign them as key - layer features and secondary - layer features; according to the layering result, assign a target precision format to each layer. Key - layer features are bound to high - precision features (such as FP32), and secondary - layer features are bound to low - precision features (such as FP16 / BF16); perform compression conversion on the high - precision features to release the memory occupied by the original high - precision features, only retaining the low - precision version to obtain low - precision features; determine the maximum / minimum value range according to the numerical distribution of the low - precision features; linearly scale and correct the low - precision features through a scaling factor and compress them into the maximum / minimum value range; add an offset to the scaled low - precision features to avoid underflow (such as losing precision for negative values) to obtain stabilized features.

[0072] Further explanation, key - layer features are features that play a decisive role in the task goal, usually characterized by high sensitivity and global dependence. High sensitivity means that a tiny change in the value will lead to a significant deviation in the result (such as object boundary coordinates), and global dependence means that cross - regional correlation analysis is required (such as lesion areas in medical images); secondary - layer features are features that have less impact on the task goal or have a higher redundancy, usually characterized by locality and fault tolerance. Locality means that it only affects local details (such as background texture, light and shadow), and fault tolerance means that numerical errors have a limited impact on the overall result (such as hair details, cloth folds).

[0073] The relevance of priority assignment to the task specifically includes task - driven feature importance judgment and the trade - off between hardware resources and efficiency, as follows:

[0074] Task - driven feature importance judgment includes classification tasks, detection tasks, and segmentation tasks, etc.

[0075] Among them, the key layer features refer to the discriminative regions of object categories (such as the ears of a cat and the nose of a dog) in the classification task, and the secondary layer features refer to the background color distribution and the texture of non-target regions in the classification task.

[0076] The key layer features refer to the bounding box coordinates and the positions of key points (such as the facial feature points) in the detection task, and the secondary layer features refer to the slightly smooth regions on the object surface (such as the color gradient of a wall) in the detection task.

[0077] The key layer features refer to the boundary between the target and the background and the multi-organ junction regions in the segmentation task, and the secondary layer features refer to the uniform regions inside the organs (such as uniform liver tissue) in the segmentation task.

[0078] For the trade-off between hardware resources and efficiency, more computing resources need to be allocated in the high-precision layer to ensure the computational accuracy of key features. In the low-precision layer, the video memory occupancy is compressed to improve the computational throughput, but the precision loss needs to be avoided through PASA optimization (online pseudo-average shift).

[0079] Compare the stabilized features with the offset matrix region by region, and analyze the distribution differences between the two; according to the numerical range of the current stabilized features, scale or translate the intensity of the offset matrix, and combine with the adaptation range of the hardware to limit the adjustment amplitude of the offset matrix to avoid overcompensation, and obtain the dynamically adjusted offset matrix; combine the stabilized features with the adjusted offset matrix and correct them through element-wise operations to obtain anti-overflow features; according to the block strategy configured by the optimized computational graph, divide the anti-overflow features into multiple local blocks, mark the data dependencies between the blocks, and allocate independent video memory spaces for each block to obtain a set of block features; according to the types of the block features, select the optimal hardware operator, perform matrix operations block by block, and dynamically adjust the block scheduling order according to the hardware resource utilization rate (such as the GPU core idle rate) to maximize the throughput and obtain the low-precision inference result.

[0080] Further explanation, allocating independent video memory spaces for each block specifically includes the following steps

[0081] For the feature maps with uniform distribution (such as the organ regions in medical images), fixed-size blocks are used. For high-complexity regions (such as object edges and texture-dense areas), the block size is reduced to avoid exceeding the video memory limit for a single block. For non-integer multiple block regions, overlapping block or padding strategies (such as mirror padding) are adopted to ensure data integrity. By calculating the input-output relationships between blocks in the graph, strong and weak dependencies are established. Strong dependency means that the calculation of block B needs to wait for block A to complete (such as cascaded convolutional layers), which is marked as serial execution. Weak dependency means that block C and block D can be calculated in parallel (such as feature processing of different channels), which is marked as parallel execution. Before the start of the computing task, a continuous large area (video memory pool) is pre-allocated from the video memory. An independent physical address space is allocated from the video memory pool for each block, the video memory size of each block is calculated, and free blocks that meet the requirements are searched in the video memory pool.

[0082] S6. Align the feature dimension of the low-precision inference result with the original image size to generate a low-resolution feature map; expand the low-resolution feature map to the original image size through an interpolation algorithm; use a pre-trained super-resolution model to reconstruct the expanded low-resolution feature map to generate a high-resolution feature map, repair blurred edges and enhance detailed textures; convert the high-resolution feature map from grayscale or single-channel format to three-channel RGB format; perform gamma correction on the converted high-resolution feature map to adjust brightness and contrast to adapt to the visual characteristics of the human eye, and convert the adjusted high-resolution feature map from the device's native color space to the standard color space; quickly transmit the standardized color image to the video memory through asynchronous transmission technology to avoid the synchronization waiting delay between the central processing unit and the graphics processing unit; select a lossless / lossy compression algorithm to compress the standardized color image according to the output device requirements; distribute the compressed image data to the output interfaces of multiple display screens to obtain a high-quality image after real-time processing.

[0083] This embodiment also provides an image processing system based on an artificial intelligence large model, including:

[0084] An image block module that acquires the original image and divides it into overlapping blocks, and generates an image block feature set through a fusion strategy;

[0085] A topology construction module that constructs an undirected graph structure based on the image block feature set, calculates the node similarity weights, and generates a global topology descriptor;

[0086] A noise enhancement module that generates importance weights according to the dimension importance of the global topology descriptor, combines the noise intensity and structural constraints to generate a noise vector, and generates an enhancement vector and an offset matrix through a UNet encoder-decoder;

[0087] The knowledge distillation module uses the student network and the teacher network to perform feature distillation on the enhanced vector, and generates the final distilled features through cosine similarity matching and the attention mechanism;

[0088] The calculation optimization module performs priority allocation and computational graph optimization on the final distilled features and the offset matrix, and obtains a low-precision inference result through block parallel computing;

[0089] The image output module performs super-resolution reconstruction and color correction on the low-precision inference result, and outputs a high-quality image after real-time processing through asynchronous transmission technology.

[0090] This embodiment also provides a computer device, which is applicable to the case of an image processing method based on an artificial intelligence large model, and includes: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the image processing method based on an artificial intelligence large model proposed in the above embodiment.

[0091] This computer device may be a terminal, and this computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of this computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a carrier network, NFC (Near Field Communication), or other technologies. The display screen of this computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of this computer device may be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0092] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the image processing method based on an artificial intelligence large model as proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disc.

[0093] In summary, through the construction of a multi-dimensional collaborative optimization framework, the present invention realizes the double improvement of image processing efficiency and accuracy; at the feature extraction level, the overlapping block strategy is combined with the construction of a dynamic graph structure to effectively alleviate the boundary effect of the traditional block method; in the knowledge distillation link, a dynamic noise intensity calibration mechanism is introduced, and the precise constraint of the noise direction is realized through the linkage control of the variance parameter, so that the generalization error is reduced; the computational graph optimization module adopts a resource pre-allocation strategy guided by the activation probability to compress the model calculation amount while maintaining the original accuracy; the hierarchical accuracy allocation scheme realizes the medical image segmentation accuracy equivalent to FP32 under the FP16 accuracy through the key layer feature protection mechanism, which is significantly improved compared with the traditional quantization method.

[0094] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. An image processing method based on an artificial intelligence large model, characterized in that: Including, Obtain the original image and segment it into overlapping blocks, and generate a set of image block features through a fusion strategy; Construct an undirected graph structure based on the set of image block features, calculate the node similarity weights, label the node position coordinates and the semantic labels of the edges, obtain the enhanced image block set, and generate a global topological descriptor according to the enhanced image block set; Generate importance weights according to the dimension of the global topological descriptor, combine the noise intensity with the importance weights to convert them into a noise vector, and input the converted noise vector and the enhanced image block set into the UNet encoder-decoder in parallel to generate an enhanced noise vector and an offset matrix; Use the student network and the teacher network to perform feature distillation on the enhanced noise vector, and generate the final distilled features through cosine similarity matching and attention mechanism; Perform priority assignment and computational graph optimization configuration on the final distilled features and the offset matrix, and obtain a low-precision inference result through block parallel computing; Perform super-resolution reconstruction and color correction on the low-precision inference result, and output a high-quality image after real-time processing through asynchronous transmission technology.

2. The image processing method based on the large artificial intelligence model according to claim 1, wherein: The generation of the set of image block features includes the following steps, Obtain the size of the original image, divide the image into multiple overlapping small blocks, and set the overlapping area; Merge the features of the overlapping area through a fusion strategy of taking the average value and the maximum value to generate a set of fused image blocks; Extract local texture and edge features based on the set of fused image blocks, and obtain a set of features of the image blocks through downsampling and multi-layer convolution processing.

3. The image processing method based on an artificial intelligence large model according to claim 1, characterized in that: The generation of the global topological descriptor includes the following steps, Assign unique identifiers to the set of image block features to form a node list, calculate the Euclidean distance between node pairs and normalize it to a weight value; Construct an undirected graph structure with nodes as vertices and weights as edges, and label the node position coordinates and the semantic labels of the edges; Encode the position coordinates of the nodes and the semantic labels of the edges into low-dimensional vectors, update the node features through message passing and multiple rounds of aggregation iteration, and perform an average operation on the node features to generate a global topological descriptor.

4. The image processing method based on the artificial intelligence large model according to claim 1, wherein: The generation of the enhanced noise vector and the offset matrix includes the following steps, Calculate the importance weights according to the volatility of the dimension value of the global topological descriptor, preset a dynamic coefficient to mix the initial noise and the importance weights according to the iteration stage, generate a dynamic noise intensity and convert it into a noise vector; Input the converted noise vector and the enhanced image block set into the UNet encoder-decoder in parallel to generate mean and variance parameters; Scale the noise vector through the variance parameter to obtain a noise perturbation term, and superimpose it on the mean parameter to generate an enhanced noise vector; Encode the image block position index into a dense vector and splice it with the enhanced noise vector, generate a position-aware feature through a fully connected layer and eliminate the bias to obtain an offset matrix.

5. The image processing method based on an artificial intelligence large model according to claim 1, wherein: The generation of the final distilled features includes the following steps, The student network receives the enhanced noise vector and generates local features, the teacher network generates global features, and performs feature alignment through cosine similarity matching and contrast loss; The global attention map of the teacher network guides the student network to focus on key regions, and combines the adversarial training of the generator and the discriminator to generate the final distilled features.

6. The image processing method based on the artificial intelligence large model according to claim 1, wherein: The obtaining of the low-precision inference result includes the following steps: Assign priorities to the final distilled features and bind precision formats, compress the high-precision features to generate low-precision features, and generate a joint state vector by combining gradient complexity; Infer the activation probability of the joint state vector through the policy network inference generation module; Allocate resources to high-priority modules and adjust parameters, delete low-priority modules, generate an optimized computational graph configuration, perform scaling correction and anti-overflow processing on the low-precision features, divide the block features, and allocate independent video memory; Dynamically schedule block execution of matrix operations based on hardware utilization to generate low-precision inference results.

7. The image processing method based on the large artificial intelligence model according to claim 1, wherein: The output of the high-quality image after real-time processing includes the following steps: Expand the low-precision inference result to the original size through interpolation, reconstruct and repair details through a super-resolution model, convert the color space and perform gamma correction, distribute it to a multi-screen interface through asynchronous transmission technology, and adapt to the output device requirements using a compression algorithm to obtain a high-quality image after real-time processing.

8. An image processing system based on an artificial intelligence large model, based on the image processing method based on an artificial intelligence large model according to any one of claims 1 to 7, characterized in that: Including: An image block module that obtains an original image and divides it into overlapping blocks, and generates an image block feature set through a fusion strategy; A topology construction module that constructs an undirected graph structure based on the image block feature set, calculates the node similarity weights, labels the node position coordinates and the semantic labels of the edges, obtains an enhanced image block set, and generates a global topology descriptor according to the enhanced image block set; A noise enhancement module that generates importance weights according to the dimension of the global topology descriptor, combines the noise intensity and the importance weights to be converted into a noise vector, and inputs the converted noise vector and the enhanced image block set into the UNet encoder-decoder in parallel to generate an enhanced noise vector and an offset matrix; A knowledge distillation module that performs feature distillation on the enhanced noise vector using a student network and a teacher network, and generates a final distilled feature through cosine similarity matching and an attention mechanism; A calculation optimization module that performs priority allocation and computational graph optimization configuration on the final distilled feature and the offset matrix, and obtains a low-precision inference result through block parallel calculation; An image output module that performs super-resolution reconstruction and color correction on the low-precision inference result, and outputs a high-quality image after real-time processing through asynchronous transmission technology.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the image processing method based on the artificial intelligence large model according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the image processing method based on the artificial intelligence large model according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Vehicle re-identification method based on feature enhancement

    CN114005096A

  • Tumor analysis method based on pathological tissue image

    CN119480151A