Image processing method and system based on artificial intelligence large model

By building a multi-dimensional collaborative optimization framework, the problem of feature space mismatch and insufficient dynamic noise constraints during cross-level feature distillation in image processing is solved, and the image processing efficiency and accuracy is improved, the boundary effect is alleviated and generalization error is reduced.

CN119991438AActive Publication Date: 2025-05-13西安圣瞳科技有限公司

Patent Information

Application Number
CN202510458171.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
2045-04-14

AI Technical Summary

Technical Problem

In the image processing, the prior art has the problem of feature space mismatch and insufficient dynamic noise constraints during cross-level feature distillation, resulting in limited model generalization capabilities.

Method used

By acquiring the original image and dividing it into overlapping blocks, an image block feature set is generated, an undirected graph structure is constructed, node similarity weight is calculated, global topology descriptors are generated, noise vectors are generated based on noise intensity and structural constraints, enhancement vectors and offset matrices are generated through UNet encoder-decoder, feature distillation is used for student network and teacher network, calculation graphs are optimized to obtain low-precision inference results, and high-quality images are output through super-resolution reconstruction and color correction.

Benefits of technology

The dual improvement of image processing efficiency and accuracy is achieved, the boundary effect of traditional blocking methods is alleviated, the generalization error is reduced, and the medical image segmentation accuracy comparable to high precision is achieved at low accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991438A_ABST
    Figure CN119991438A_ABST
Patent Text Reader

Abstract

The invention discloses an image processing method and system based on an artificial intelligence large model, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining an original image, segmenting the original image into overlapped blocks, and generating an image block feature set through a fusion strategy; constructing an undirected graph structure based on the image block feature set, calculating a node similarity weight and generating a global topology descriptor; generating an importance weight according to the dimension importance of the global topology descriptor, generating a noise vector in combination with noise intensity and structural constraint, and generating an enhancement vector and an offset matrix through a UNet encoder-decoder; and performing feature distillation on the enhanced vector by using a student network and a teacher network, and generating a final distillation feature through cosine similarity matching and an attention mechanism. According to the method, a dynamic noise intensity calibration mechanism is introduced in a knowledge distillation link, and accurate constraint of a noise direction is realized through variance parameter linkage control, so that generalization errors are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an image processing method and system based on an artificial intelligence large model. Background Art

[0002] With the rapid development of deep learning technology in the field of computer vision, image processing methods based on convolutional neural networks (CNNs) have formed a complete technical system. Traditional image block processing mostly uses fixed-size sliding windows or non-overlapping block strategies. Although it can ensure the stability of local feature extraction, it has problems such as significant boundary effects and loss of global structural information. In recent years, graph neural networks (GNNs) have demonstrated unique advantages in feature fusion by constructing node-edge structures to represent image topological relationships. However, existing methods still have limitations in dynamic weight allocation and cross-level feature interaction.

[0003] Current image processing systems generally face multi-dimensional technical challenges: at the block processing level, traditional methods have difficulty balancing computational efficiency and feature integrity; at the feature fusion level, static fusion strategies cannot adapt to complex texture changes; at the model optimization level, static computational graphs have difficulty meeting dynamic resource allocation requirements. It is particularly noteworthy that existing technologies have feature space mismatch problems in the cross-level feature distillation process, and lack an effective noise constraint mechanism, which limits the model's generalization ability. Summary of the invention

[0004] In view of the above existing problems, the present invention is proposed.

[0005] Therefore, the present invention provides an image processing method based on an artificial intelligence large model to solve the problems of cross-level feature space mismatch and insufficient dynamic noise constraints in the multi-level feature distillation process.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides an image processing method based on an artificial intelligence large model, which comprises: obtaining an original image and dividing it into overlapping blocks, and generating an image block feature set through a fusion strategy; Building an undirected graph structure based on the image block feature set, calculating node similarity weights and generating a global topology descriptor; Generate importance weights according to the dimensional importance of the global topology descriptor, generate noise vectors by combining noise intensity and structural constraints, and generate enhancement vectors and offset matrices through a UNet encoder-decoder; The student network and the teacher network are used to perform feature distillation on the enhanced vector, and the final distilled features are generated through cosine similarity matching and attention mechanism; Prioritizing and optimizing the calculation graph of the final distilled features and the offset matrix, and obtaining a low-precision reasoning result through block parallel calculation; The low-precision inference results are subjected to super-resolution reconstruction and color correction, and high-quality images after real-time processing are output through asynchronous transmission technology.

[0007] As a preferred solution of the image processing method based on artificial intelligence large model described in the present invention, wherein: the generating of the image block feature set comprises the following steps: Get the size of the original image, divide the image into multiple overlapping blocks, and set the overlapping area; The features of the overlapping areas are merged by taking the average and maximum values ​​to generate a fused set of image blocks; Based on the fused image block set, local texture and edge features are extracted, and a feature set of the image blocks is obtained through downsampling and multi-layer convolution processing.

[0008] As a preferred solution of the image processing method based on artificial intelligence large model described in the present invention, wherein: the generation of global topology descriptor includes the following steps: Assign unique identifiers to the image block feature set to form a node list, calculate the Euclidean distance similarity of the node pairs and normalize it into a weight value; Construct an undirected graph structure with nodes as vertices and weights as edges, and mark the node positions and edge semantic labels; encode node features with position and semantic information, update node features through message passing and multiple rounds of aggregation iterations, average the node features, and generate a global topology descriptor.

[0009] As a preferred solution of the image processing method based on artificial intelligence large model described in the present invention, wherein: the generation of enhancement vector and offset matrix includes the following steps: The importance weight is calculated according to the volatility of the dimension value of the global topological descriptor, and the preset dynamic coefficient is used to mix the initial noise and the importance weight according to the iteration stage to generate the dynamic noise intensity and convert it into a noise vector; Input the noise vector into the UNet encoder-decoder to generate mean and variance parameters; The noise vector is scaled by the variance parameter to obtain the disturbance term, which is then added to the mean to generate the enhancement vector. The image block position index is encoded into a dense vector and concatenated with the enhancement vector. The fully connected layer generates position-aware features and eliminates the bias to obtain the offset matrix.

[0010] As a preferred solution of the image processing method based on artificial intelligence large model of the present invention, wherein: the generation of the final distillation feature includes the following steps: The student network receives the enhanced vector and generates local features, while the teacher network generates global features and performs feature alignment through cosine similarity matching and contrast loss; The global attention map of the teacher network guides the student network to focus on the key areas, and the adversarial training of the generator and the discriminator is combined to generate the final distilled features.

[0011] As a preferred solution of the image processing method based on artificial intelligence large model described in the present invention, wherein: the obtaining of low-precision reasoning results includes the following steps: Assign priorities to the final distilled features and bind the precision format, compress high-precision features to generate low-precision features, and combine the gradient complexity to generate a joint state vector; The joint state vector is passed through the policy network to generate the module activation probability; Allocate resources and adjust parameters for high-priority modules, delete low-priority modules to generate optimized calculations, perform scaling correction and anti-overflow processing on low-precision features, divide block features and allocate independent video memory; Dynamically schedule block-by-block matrix operations based on hardware utilization to generate low-precision inference results.

[0012] As a preferred solution of the image processing method based on artificial intelligence large model described in the present invention, wherein: the output of high-quality image after real-time processing includes the following steps: The low-precision inference results are expanded to the original size through interpolation, the details are restored through super-resolution model reconstruction, the color space is converted and gamma correction is performed, and they are distributed to multi-screen interfaces through asynchronous transmission technology. The compression algorithm is used to adapt to the output device requirements to obtain high-quality images after real-time processing.

[0013] In a second aspect, the present invention provides an image processing system based on an artificial intelligence large model, comprising: The image segmentation module obtains the original image and segments it into overlapping blocks, and generates an image segmentation feature set through a fusion strategy; A topology construction module, which constructs an undirected graph structure based on the image block feature set, calculates node similarity weights and generates a global topology descriptor; A noise enhancement module generates importance weights according to the dimensional importance of the global topology descriptor, generates a noise vector by combining the noise intensity and the structural constraint, and generates an enhancement vector and an offset matrix through a UNet encoder-decoder; The knowledge distillation module uses the student network and the teacher network to perform feature distillation on the enhanced vector, and generates the final distilled features through cosine similarity matching and attention mechanism; A calculation optimization module, which performs priority assignment and calculation graph optimization on the final distillation features and the offset matrix, and obtains low-precision reasoning results through block parallel calculation; The image output module performs super-resolution reconstruction and color correction on the low-precision inference results, and outputs high-quality images after real-time processing through asynchronous transmission technology.

[0014] In a third aspect, the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, any step of the image processing method based on the artificial intelligence large model as described in the first aspect of the present invention is implemented.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the image processing method based on the artificial intelligence large model as described in the first aspect of the present invention.

[0016] The beneficial effects of the present invention are as follows: by constructing a multi-dimensional collaborative optimization framework, the image processing efficiency and accuracy are both improved; at the feature extraction level, an overlapping block strategy is adopted in combination with a dynamic graph structure to effectively alleviate the boundary effect of the traditional block method; in the knowledge distillation link, a dynamic noise intensity calibration mechanism is introduced, and the precise constraint of the noise direction is achieved through the linkage control of variance parameters, so that the generalization error is reduced; the computational graph optimization module compresses the model calculation amount while maintaining the original accuracy through the resource pre-allocation strategy guided by the activation probability; the hierarchical precision allocation scheme achieves medical image segmentation accuracy equivalent to FP32 at FP16 accuracy through the key layer feature protection mechanism, which is a significant improvement over the traditional quantization method. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0018] Figure 1 This is a flow chart of image segmentation and feature extraction in this embodiment.

[0019] Figure 2 Schematic diagram of topology construction and noise enhancement in this embodiment.

[0020] Figure 3 Schematic diagram of feature distillation and computational optimization in this embodiment.

[0021] Figure 4 It is a schematic diagram of image output and real-time rendering in this embodiment. DETAILED DESCRIPTION

[0022] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0023] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.

[0025] In this embodiment, refer to Figure 1~Figure 4 , this embodiment provides an image processing method based on an artificial intelligence large model, comprising the following steps: S1. Get the size of the original image, divide the image into multiple N×N small blocks, start from the upper left corner, and gradually move to cover the entire image area; in order to reduce the boundary effect, set 50% overlap between adjacent small blocks, for each overlapping area, extract features from adjacent small blocks, and select a suitable fusion strategy to merge the features of the overlapping area. Common fusion strategies include taking the average, maximum value, weighted average, etc. For example, taking the average value can smooth the change of feature values, while taking the maximum value can retain the most significant features; according to the selected strategy, fuse the features of the overlapping area. For example, when choosing to take the average value, add all the features of the overlapping area and divide them by the number of features to obtain the fused feature vector; replace the features of the corresponding overlapping area in the original small block with the fused feature vector to obtain a set of overlapping image block partitions.

[0026] According to the specific image processing task, determine the required feature type. For example, if it is an image classification task, it is necessary to extract high-level semantic features; if it is an object detection task, it is necessary to extract edge and texture features; find a pre-trained CNN model (convolutional neural network) suitable for the corresponding image processing task; convert overlapping image blocks into the format required by the CNN model, and normalize the pixel values ​​of the image blocks and scale them to the standardized range; batch process the scaled multiple image blocks; the CNN model scans the batched image blocks through a small sliding window (such as 3×3 pixels), extracts local texture and edge features, and generates a preliminary feature map; downsamples the preliminary feature map (such as 2×2 maximum pooling) to retain significant features and reduce the amount of calculation; after multiple layers of convolution and pooling, a feature set of the image blocks is obtained.

[0027] Assign a unique identifier to each image block in the feature set of the image block to form a node list, in which each node corresponds to the feature of an image block; traverse all node pairs, calculate the similarity of each pair of nodes through Euclidean distance, and measure the difference between the features of the image blocks; normalize the similarity results by a preset constant to obtain a weight value reflecting the degree of similarity, and the smaller the weight, the more similar the nodes are; take the nodes as vertices and the weights as edges, and select an adjacency matrix or an adjacency list to represent the undirected graph structure; mark the position of each node in the undirected graph structure in the original image (such as node 1 is located in the upper left corner of the image), add semantic labels to the edges (such as strong texture similarity or weak color difference), and obtain an enhanced image block set.

[0028] To further illustrate, when constructing an undirected graph, if the number of edges is too large, a weight threshold is set to only retain edges with weights less than the weight threshold to reduce complexity.

[0029] The original feature vector corresponding to each node in the enhanced image block set is used as the initial representation; the node position coordinates and the semantic label of the edge are encoded into a low-dimensional vector, which is concatenated with the original feature vector to obtain the initialized node feature matrix; each node sends its own features to all neighboring nodes, and the message content is the node feature itself; after all nodes receive messages from neighboring nodes, they merge the information through aggregation functions (such as summation and averaging) to generate aggregated temporary features; the node features are updated according to the aggregated temporary features, and the updated node features are propagated for multiple rounds. In each round, the node resends the message based on the latest features and aggregates the neighbor information; as the number of rounds increases, the node features gradually merge the information of more distant neighbors, capture the global structure, and dynamically adjust the importance of features at different levels through a learnable weight matrix to obtain the final node feature matrix after multiple rounds of iterations; the features of all nodes in the final node feature matrix are averaged to generate a global topology descriptor.

[0030] It should be noted that the global topology descriptor is a multidimensional vector (eg, containing K dimensions), and each dimension reflects different properties of the global structure of the image (eg, the distribution density of connected areas of objects, the number of holes, etc.).

[0031] S2. For each dimension in the global topology descriptor, calculate the corresponding weight value. The smaller the value of the corresponding dimension, the higher the weight, indicating that the topological feature is more stable and important. On the contrary, if the corresponding dimension threshold is set low and the fluctuation is small (such as the dimension value is always close to the preset dimension threshold), the weight is suppressed by the exponential function. Add the weights of all dimensions, get the reciprocal of the sum as the normalization constant, and then scale the weights of each dimension proportionally to generate the importance weight.

[0032] To further explain, the purpose of generating importance weights is to ensure that the value of the importance weights is between 0 and 1, which is used to quantify the overall importance of the global topological structure. The closer the importance weight is to 1, the more critical the global structure is.

[0033] The dynamic coefficient that increases monotonically from 0 to 1 is preset to control the mixing ratio of noise and structural constraints. The initial noise intensity is preset as the initial noise level in the diffusion process to provide an exploratory noise source. According to the dynamic coefficient of the current iteration stage, the initial noise and structural importance weight are proportionally fused, as follows: In the initial iteration stage, the noise is mainly based on the initial value to ensure that the diffusion process has sufficient randomness to explore the feature space; in the later iteration stage, the importance weight is mainly used to constrain the noise direction to conform to the global topological law (such as suppressing irrelevant texture noise and strengthening object boundaries), and finally generate dynamic noise intensity; the dynamic noise intensity is converted into a noise vector that matches the image block in the enhanced image block set; the converted noise vector and the enhanced image block set are input into the encoder of UNet in parallel, and the interaction pattern between features and noise is extracted through multi-layer convolution and pooling operations, and the encoder outputs a low-dimensional intermediate representation; the low-dimensional intermediate representation is input into the decoder of UNet to generate mean parameters and variance parameters.

[0034] To further explain, after the variance parameters are generated, they need to be calibrated, which means that the variance parameters are forced to coordinate with the dynamic noise intensity through the loss function. If the dynamic noise intensity is high, the variance parameters tend to increase, allowing larger noise disturbances. If the dynamic noise intensity is low, the variance parameters are limited to a smaller range to ensure feature stability.

[0035] The noise vector is scaled by the variance parameter to obtain the noise disturbance term; the noise disturbance term is superimposed on the mean parameter to generate an enhancement vector; the image block position index includes the row index and column index of the image block, the row index and column index of the image block are converted into a dense position vector, and concatenated with the corresponding enhancement vector to input the fully connected layer to generate position-aware features; matrix subtraction and scaling operations are performed on the position-aware features to eliminate the calculation deviation caused by noise or local disturbances, and obtain the offset matrix.

[0036] S3. The student network receives the enhanced vector, and generates local student features that match the output dimension of the teacher network through multi-layer convolution and pooling operations; the teacher performs a global average pooling operation on the enhanced vector to generate a global teacher feature; the local features generated by the student network are matched with the global features generated by the teacher network on a pixel-by-pixel basis, and the cosine similarity is calculated; a similarity threshold is set. In this embodiment, the similarity threshold is set to 0.85. If the similarity between the local student features and the global teacher features is lower than the similarity threshold, the similarity gap of the contrast loss is calculated and back-propagated to force the local features of the student network to maintain a high similarity with the teacher network, and obtain preliminary distilled features; the teacher network is subjected to the global average pooling operation. In the process, a teacher attention map is generated, and the preliminary distilled features are input into the self-attention mechanism to generate a student attention map; the L2 norm difference between the teacher attention map and the student attention map is calculated to guide the student network to pay attention to the areas that the teacher network considers to be critical (such as object boundaries or texture-dense areas); the preliminary distilled features are concatenated with the student attention map through the fully connected layer to obtain the student refined features; a generator is set to map the student refined features to the feature space of the teacher network through the generator; a discriminator is set to distinguish between the real teacher features and the student refined features mapped by the generator to the feature space of the teacher network; the student refined features output by the generator are concatenated with the student refined features before mapping to generate the final distilled features.

[0037] S4. Take the final distilled feature as a global feature candidate, perform average calculation on each dimension in the final distilled feature, retain the core trend of each dimension, and generate a global feature vector; calculate the gradient change intensity of the pixel points of each image block in the enhanced image block set (such as the change amplitude in the horizontal and vertical directions) to generate a two-dimensional gradient matrix; divide the image in the two-dimensional gradient matrix into multiple small areas, sum the gradient change intensity in each area, and obtain the local complexity value; sum the complexity values ​​of all local areas to generate a scalar local complexity to reflect the overall detail complexity of the image; concatenate the global feature vector and the scalar local complexity to form a joint state vector.

[0038] The joint state vector is divided into multiple sub-vectors according to the preset rules; the sub-vectors after division are input into the multi-layer perceptron for policy network reasoning, as follows: The input layer receives the segmented sub-vectors, performs linear transformation by multiplying them with the weight matrix through the fully connected layer, and generates an intermediate representation; the ReLU function is applied to the intermediate representation to set negative values ​​to zero and retain positive values; the intermediate layer further performs linear transformation on the output of the first layer, and inserts a Dropout layer (random inhibition) after the output of the intermediate layer to randomly ignore the values ​​of some neurons, inhibit overfitting, and extract abstract features; the output layer maps the extracted abstract features to the module activation probability space to generate the initial value of the module activation probability.

[0039] The initial probability of each module is converted into a non-negative value through exponential operation, and the non-negative values ​​of all modules are added and compressed to the interval of 0 to 1 to obtain the normalized activation probability; the activation probability threshold is set, and the modules with normalized activation probability greater than or equal to the activation probability threshold are marked as high priority, and the inactivated modules are marked as low priority; the high-priority modules are further sorted according to the activation probability, and the higher the probability, the higher the priority, and resource requirements are pre-allocated for each high-priority module to form a resource pre-allocation plan; the low-priority modules are downgraded; according to the processing results of high priority and low priority, all inactivated modules are deleted, and the input / output connections of the inactivated modules are cut off, leaving only the trunk path; according to the remaining amount of module resources after deletion, the internal parameters of the module (such as convolution kernel size and pooling step size) are dynamically adjusted to generate an optimized calculation graph configuration.

[0040] S5. Assign priorities according to the role of the final distilled features in the corresponding tasks, for example, assign them to key layer features and secondary layer features; according to the stratification results, assign target precision formats to each layer, bind key layer features to high-precision features (such as FP32), and bind secondary layer features to low-precision features (such as FP16 / BF16); compress and convert high-precision features to release the memory occupied by the original high-precision features, retain only the low-precision version, and obtain low-precision features; determine the maximum / minimum value range based on the numerical distribution of low-precision features; linearly scale and correct the low-precision features through scaling factors and compress them within the maximum / minimum value range; add offsets to the scaled low-precision features to avoid underflow (such as loss of precision due to negative values) to obtain stabilized features.

[0041] To further explain, key layer features are features that play a decisive role in the task objectives, and usually have the characteristics of high sensitivity and global dependence. High sensitivity means that a small change in the value will lead to a significant deviation in the result (such as the boundary coordinates of an object), and global dependence means that cross-regional correlation analysis is required (such as the lesion area in a medical image); secondary layer features are features that have little impact on the task objectives or have high redundancy, and usually have the characteristics of locality and fault tolerance. Locality means that it only affects local details (such as background texture, light and shadow), and fault tolerance means that numerical errors have limited impact on the overall result (such as hair details, cloth wrinkles).

[0042] The correlation between priority allocation and tasks specifically includes task-driven feature importance judgment and the trade-off between hardware resources and efficiency, as follows: Task-driven feature importance judgment includes classification tasks, detection tasks, and segmentation tasks.

[0043] The key layer features in the classification task refer to the discriminative areas of object categories (such as cat ears and dog noses), and the secondary layer features in the classification task refer to the background color distribution and the texture of non-target areas.

[0044] In the detection task, key layer features refer to bounding box coordinates and key point positions (such as facial features), etc., while secondary layer features refer to small smooth areas on the surface of objects (such as wall color gradients) in the detection task.

[0045] In the segmentation task, the key layer features refer to the boundary between the target and the background and the boundary area of ​​multiple organs. The secondary layer features refer to the uniform area inside the organ (such as uniform liver tissue) in the segmentation task.

[0046] The trade-off between hardware resources and efficiency is that more computing resources need to be allocated in the high-precision layer to ensure the calculation accuracy of key features. In the low-precision layer, the computing throughput can be improved by compressing the video memory occupancy, but PASA optimization (online pseudo average shift) is required to avoid accuracy loss.

[0047] The stabilized features are compared with the offset matrix region by region to analyze the distribution differences between the two. According to the numerical range of the current stabilized features, the strength of the offset matrix is ​​scaled or translated. Combined with the hardware adaptation range, the adjustment range of the offset matrix is ​​limited to avoid overcompensation, and the dynamically adjusted offset matrix is ​​obtained. The stabilized features are combined with the adjusted offset matrix and corrected through element operations to obtain anti-overflow features. According to the block strategy of the optimized computational graph configuration, the anti-overflow features are divided into multiple local blocks, the data dependencies between blocks are marked, and independent video memory space is allocated to each block to obtain a set of block features. According to the type of block features, the optimal hardware operator is selected, matrix operations are performed block by block, and the block scheduling order is dynamically adjusted according to the hardware resource utilization (such as GPU core idle rate) to maximize throughput and obtain low-precision inference results.

[0048] Further explanation: allocating independent video memory space to each block specifically includes the following steps: For evenly distributed feature maps (such as organ areas in medical images), fixed-size blocks are used. For high-complexity areas (such as object edges and texture-dense areas), the block size is reduced to avoid single-block video memory overruns. For non-integer multiple block areas, overlapping blocks or filling strategies (such as mirror filling) are used to ensure data integrity. By calculating the input-output relationship between blocks in the graph, strong and weak dependencies are established. Strong dependency means that the calculation of block B needs to wait for block A to be completed (such as cascaded convolutional layers), which is marked as serial execution. Weak dependency means that blocks C and D can be calculated in parallel (such as feature processing of different channels), which is marked as parallel execution. Before the calculation task starts, large continuous blocks (video memory pool) are pre-marked from the video memory, and independent physical address space is allocated to each block from the video memory pool. The video memory size of each block is calculated, and free blocks that meet the requirements are searched in the video memory pool.

[0049] S6. Align the feature dimension of the low-precision inference result with the original image size to generate a low-resolution feature map; expand the low-resolution feature map to the original image size through an interpolation algorithm; use a pre-trained super-resolution model to reconstruct the expanded low-resolution feature map to generate a high-resolution feature map, repair blurred edges and enhance detail textures; convert the high-resolution feature map from grayscale or single-channel format to a three-channel RGB format; perform gamma correction on the converted high-resolution feature map, adjust the brightness and contrast to adapt to the visual characteristics of the human eye, and convert the adjusted high-resolution feature map from the device's native color space to a standard color space; quickly transfer the standardized color image to the video memory through asynchronous transmission technology to avoid synchronization delays between the central processing unit and the graphics processing unit; select a lossless / lossy compression algorithm to compress the standardized color image according to the output device requirements; distribute the compressed image data to multiple display output interfaces to obtain high-quality images after real-time processing.

[0050] This embodiment also provides an image processing system based on an artificial intelligence large model, including: The image segmentation module obtains the original image and segments it into overlapping blocks, and generates an image segmentation feature set through a fusion strategy; A topology construction module, which constructs an undirected graph structure based on the image block feature set, calculates node similarity weights and generates a global topology descriptor; A noise enhancement module generates importance weights according to the dimensional importance of the global topology descriptor, generates a noise vector by combining the noise intensity and the structural constraint, and generates an enhancement vector and an offset matrix through a UNet encoder-decoder; The knowledge distillation module uses the student network and the teacher network to perform feature distillation on the enhanced vector, and generates the final distilled features through cosine similarity matching and attention mechanism; A calculation optimization module, which performs priority assignment and calculation graph optimization on the final distillation features and the offset matrix, and obtains low-precision reasoning results through block parallel calculation; The image output module performs super-resolution reconstruction and color correction on the low-precision inference results, and outputs high-quality images after real-time processing through asynchronous transmission technology.

[0051] This embodiment also provides a computer device, which is suitable for the image processing method based on the artificial intelligence large model, including: a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute computer executable instructions to implement the image processing method based on the artificial intelligence large model proposed in the above embodiment.

[0052] The computer device may be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device may be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device may be a touch layer covered on the display screen, or a key, trackball or touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.

[0053] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the image processing method based on the artificial intelligence large model proposed in the above embodiment is implemented; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, referred to as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, referred to as EEPROM), erasable programmable read-only memory (Erasable Programmable Read Only Memory, referred to as EPROM), programmable read-only memory (Programmable Red-Only Memory, referred to as PROM), read-only memory (Read-Only Memory, referred to as ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0054] In summary, the present invention achieves a dual improvement in image processing efficiency and accuracy by constructing a multi-dimensional collaborative optimization framework; at the feature extraction level, an overlapping blocking strategy is adopted in combination with a dynamic graph structure to effectively alleviate the boundary effect of the traditional blocking method; in the knowledge distillation link, a dynamic noise intensity calibration mechanism is introduced, and the precise constraint of the noise direction is achieved through the linkage control of the variance parameters, so that the generalization error is reduced; the computational graph optimization module compresses the model calculation amount while maintaining the original accuracy through the resource pre-allocation strategy guided by the activation probability; the hierarchical precision allocation scheme achieves medical image segmentation accuracy equivalent to FP32 at FP16 accuracy through the key layer feature protection mechanism, which is a significant improvement over the traditional quantization method.

[0055] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. An image processing method based on an artificial intelligence large model, characterized in that: include: Get the original image and divide it into overlapping blocks, and generate a set of image block features through a fusion strategy; Building an undirected graph structure based on the image block feature set, calculating node similarity weights and generating a global topology descriptor; Generate importance weights according to the dimensions of the global topology descriptor, generate noise vectors by combining noise intensity and structural constraints, and generate enhancement vectors and offset matrices through a UNet encoder-decoder; The student network and the teacher network are used to perform feature distillation on the enhanced vector, and the final distilled features are generated through cosine similarity matching and attention mechanism; Prioritizing and optimizing the calculation graph of the final distilled features and the offset matrix, and obtaining a low-precision reasoning result through block parallel calculation; The low-precision inference results are subjected to super-resolution reconstruction and color correction, and high-quality images after real-time processing are output through asynchronous transmission technology.

2. The image processing method based on artificial intelligence large model as claimed in claim 1, characterized in that: The generating of the image block feature set comprises the following steps: Get the size of the original image, divide the image into multiple overlapping blocks, and set the overlapping area; The features of the overlapping areas are merged by taking the average and maximum values ​​to generate a fused set of image blocks; Based on the fused image block set, local texture and edge features are extracted, and a feature set of the image blocks is obtained through downsampling and multi-layer convolution processing.

3. The image processing method based on artificial intelligence large model as claimed in claim 1, characterized in that: The generating of the global topology descriptor comprises the following steps: Assign unique identifiers to the image block feature set to form a node list, calculate the Euclidean distance of the node pairs and normalize them into weight values; Construct an undirected graph structure with nodes as vertices and weights as edges, and mark the node position coordinates and edge semantic labels; The node location coordinates and edge semantic labels are encoded into low-dimensional vectors. The node features are updated through message passing and multiple rounds of aggregation iterations. The node features are averaged to generate a global topology descriptor.

4. The image processing method based on artificial intelligence large model as claimed in claim 1, characterized in that: The generating of the enhancement vector and the offset matrix comprises the following steps: The importance weight is calculated according to the volatility of the dimension value of the global topological descriptor, and the preset dynamic coefficient is used to mix the initial noise and the importance weight according to the iteration stage to generate the dynamic noise intensity and convert it into a noise vector; Input the noise vector into the UNet encoder-decoder to generate mean and variance parameters; The noise vector is scaled by the variance parameter to obtain the disturbance term, which is then added to the mean to generate the enhancement vector. The image block position index is encoded into a dense vector and concatenated with the enhancement vector. The fully connected layer generates position-aware features and eliminates the bias to obtain the offset matrix.

5. The image processing method based on artificial intelligence large model as claimed in claim 1, characterized in that: The generating of the final distillation characteristics comprises the following steps, The student network receives the enhanced vector and generates local features, while the teacher network generates global features and performs feature alignment through cosine similarity matching and contrast loss; The global attention map of the teacher network guides the student network to focus on the key areas, and the adversarial training of the generator and the discriminator is combined to generate the final distilled features.

6. The image processing method based on artificial intelligence large model as claimed in claim 1, characterized in that: The low-precision reasoning result is obtained by the following steps: Assign priorities to the final distilled features and bind the precision format, compress high-precision features to generate low-precision features, and combine the gradient complexity to generate a joint state vector; The joint state vector is passed through the policy network to generate the module activation probability; Allocate resources and adjust parameters for high-priority modules, delete low-priority modules to generate optimized calculations, perform scaling correction and anti-overflow processing on low-precision features, divide block features and allocate independent video memory; Dynamically schedule block-by-block matrix operations based on hardware utilization to generate low-precision inference results.

7. The image processing method based on artificial intelligence large model as claimed in claim 1, characterized in that: The output of the high-quality image after real-time processing comprises the following steps: The low-precision inference results are expanded to the original size through interpolation, the details are restored through super-resolution model reconstruction, the color space is converted and gamma correction is performed, and they are distributed to multi-screen interfaces through asynchronous transmission technology. The compression algorithm is used to adapt to the output device requirements to obtain high-quality images after real-time processing.

8. An image processing system based on an artificial intelligence big model, based on the image processing method based on an artificial intelligence big model according to any one of claims 1 to 7, characterized in that: include: The image segmentation module obtains the original image and segments it into overlapping blocks, and generates an image segmentation feature set through a fusion strategy; A topology construction module, which constructs an undirected graph structure based on the image block feature set, calculates node similarity weights and generates a global topology descriptor; A noise enhancement module generates importance weights according to the dimensional importance of the global topology descriptor, generates a noise vector by combining the noise intensity and the structural constraint, and generates an enhancement vector and an offset matrix through a UNet encoder-decoder; The knowledge distillation module uses the student network and the teacher network to perform feature distillation on the enhanced vector, and generates the final distilled features through cosine similarity matching and attention mechanism; A calculation optimization module, which performs priority assignment and calculation graph optimization on the final distillation features and the offset matrix, and obtains low-precision reasoning results through block parallel calculation; The image output module performs super-resolution reconstruction and color correction on the low-precision inference results, and outputs high-quality images after real-time processing through asynchronous transmission technology.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the image processing method based on the artificial intelligence large model described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image processing method based on the artificial intelligence large model described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Vehicle re-identification method based on feature enhancement

    CN114005096A

  • Bridge structure damage identification method and system based on machine vision

    CN118691917A

  • Tumor analysis method based on pathological tissue image

    CN119480151A

  • Method for reducing multi-modal characteristic quantity of large model based on multi-level coding

    CN119559477A

  • Image super-resolution method based on knowledge distillation compression model and device thereof

    US20240233077A1

Cited By

  • Game scene optimization method and system based on dynamic rendering

    CN120586387A

  • Image super-resolution reconstruction method for hypha topological structure

    CN120976025A

  • A mycelium topology-oriented image super-resolution reconstruction method

    CN120976025B