Feature reconstruction method based on collaborative memory

By introducing a collaborative memory module and optimizing the stage-by-stage encoding loss, the problem of insufficient style consistency and detail coordination in image generation in existing technologies is solved, and high-quality feature reconstruction and image generation are achieved.

CN121330463APending Publication Date: 2026-01-13CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511595837.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-04
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing technologies are insufficient in terms of overall style consistency, detail coordination, and semantic coherence during image generation. Traditional encoder-decoder structures are prone to losing important details when processing complex image features, and are unable to fully capture the multi-scale and cross-domain correlations between images, resulting in low quality of reconstructed images.

Method used

A collaborative memory module is introduced to transform encoded features into a set of memory features. Decoding is performed using collaborative memory vectors, and combined with stage-by-stage encoding loss optimization, an output image with a style highly consistent with the input image is generated.

Benefits of technology

It achieves higher-quality feature reconstruction, generates image content with a style highly consistent with the input image, improves the quality and generalization performance of image generation tasks, can capture multi-scale and multi-level information of images more precisely, and enhances the learning ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure SMS_1
    Figure SMS_1
Patent Text Reader

Abstract

The invention discloses a feature reconstruction method based on collaborative memory. The invention relates to the field of image feature reconstruction, and aims to solve the key problems of low reconstruction quality, inaccurate image style transmission and the like caused by insufficient feature representation and information bottleneck in the existing method. In the prior art, although a CNN model can automatically learn hierarchical feature representation with semantic information from data, generated images still have the phenomena of detail loss, inconsistent styles and the like. According to the method, the collaborative memory module is introduced, multiple mode features are jointly extracted through multiple parallel memory units, and finally the multiple mode features are fused into collaborative memory vectors containing rich memory information to be used for image generation. The features reconstructed by the collaborative memory module can retain abundant image mode information, thereby ensuring high accuracy of the output image in details, semantics and styles.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of image feature reconstruction, collaborative memory, artificial intelligence vision tasks and image generation, and specifically relates to a feature reconstruction method based on collaborative memory, which can effectively realize high-quality feature reconstruction. The features reconstructed using the collaborative memory mode can generate image content consistent with the style of the input image. BACKGROUND

[0002] In recent years, with the rapid development of deep learning technology, artificial intelligence has made revolutionary breakthroughs in the field of image processing. Especially in image feature learning, image generation, image repair, style transfer and other tasks, methods based on convolutional neural networks (CNN) have shown strong capabilities. Traditional image processing methods rely on hand-designed features, while deep learning can automatically learn hierarchical and semantically meaningful feature representations from data through multiple layers of convolution and nonlinear activation. Encoder-Decoder structures, such as U-Net, Autoencoder, etc., are widely used in image feature compression and reconstruction, and can effectively capture global and local information of images. Although existing technologies have made great progress in the field of image processing, there are still some problems to be solved in realizing high-quality feature reconstruction and generating image content consistent with the style of the input image: traditional encoder-decoder structures are prone to lose important detail information when processing complex image features, or have difficulty in fully capturing the multi-scale and cross-domain relevance between images. Especially when highly accurate feature reconstruction is required, information bottlenecks often result in low-quality reconstructed images, blurred details, and even artifacts. In image generation and style transfer tasks, how to accurately capture and transfer the style information of the source image and effectively fuse it with the target content is a great challenge. Existing methods may perform well in certain aspects of style, but still have deficiencies in overall style consistency, detail coordination and semantic coherence. Existing image processing models mostly use flat feature processing methods, lacking an effective memory mechanism to store and retrieve cross-scale, semantically rich feature information related to the current processing task. This non-memory or weak memory mode limits the model's learning and application of deep image laws.

[0003] The application provides a feature reconstruction method based on collaborative memory, aiming at overcoming the above problems in the prior art, realizing better feature reconstruction, and generating image content highly consistent with the style of an input image. The most core innovation of the application lies in the introduction of a collaborative memory module. The module can process image coding features output by an encoder in an innovative way to form a memory feature set. The memory feature set can more comprehensively and robustly capture multi-scale and multi-level information of an image and has certain memory capacity. This method provides more fine control for image generation tasks, more rich context information for image repair tasks, and more consistent style fusion capability for style transfer tasks. The introduction of collaborative memory is expected to enhance the learning capacity of the model, so that the model can more effectively learn and adapt to complex and variable image data, thereby improving the generalization performance. The application significantly improves the quality of image feature reconstruction and the style consistency of generated images by introducing a collaborative memory module and an innovative training mechanism, and brings new technical breakthroughs to the field of artificial intelligence vision tasks and image generation. SUMMARY

[0004] The application aims to provide a feature reconstruction method based on collaborative memory, aiming to solve the problems that the prior art still has deficiencies in the consistency of the overall style, the coordination of details, and the coherence of semantics when generating images. The application introduces a collaborative memory module, which can fuse multiple mode features, thereby more finely controlling the features, and the image decoded from the features can also maintain higher consistency with the original image.

[0005] The technical scheme adopted by the application to solve the above technical problems is:

[0006] 1. A feature reconstruction method based on collaborative memory, characterized in that the method comprises the following steps:

[0007] S1. Compressing an input image to generate coding features, comprising using an encoder composed of convolution to compress the features of the image, and generating corresponding coding features for input of a collaborative memory module;

[0008] S2. Processing the coding features to generate a collaborative memory vector, comprising using a collaborative memory module to process the coding features of the image, converting the coding features into a memory feature set, and fusing the memory feature set through convolution operation to obtain a collaborative memory vector;

[0009]

[0009] S3. Decoding the collaborative memory vector, comprising using a decoder composed of convolution to decode the collaborative memory vector, and obtaining an output image consistent in size with the input image after decoding;

[0010] S4. Encoding the output image, comprising using an encoder composed of convolution to encode the output image;

[0011] S5. Training optimization target, including input image and output image pixel reconstruction loss optimization and input image and output image stage-by-stage encoding loss optimization.

[0012] 2. The feature reconstruction method based on collaborative memory according to claim 1, wherein the specific process of S1 is:

[0013] This step aims to compress the input image in dimension, and compress the three-dimensional input image into encoded features through convolution operation. The encoded features generated in this stage are used as the input of the collaborative memory module in the next stage.

[0014] In this stage, an input image X1 with a dimension of CxHxW is input into the model. In this stage, the input image X1 is compressed into encoded features query by the encoder A. The encoder A is composed of four layers of 3x3 convolution layers, normalization layers and LeakyReLU activation functions. The encoding steps are as follows.

[0015]

[0016] In the encoding process, the dimension of X1_1 is , the dimension of X1_2 is , the dimension of X1_3 is , and the dimension of query is . The query is the final encoded feature in this stage, which compresses the convolution feature information of the original input image and is used for encoded feature decomposition and corresponding memory feature reading in the S2 stage.

[0017] 3. The feature reconstruction method based on collaborative memory according to claim 1, wherein the specific process of S2 is:

[0018] This step aims to map the encoded features query that retain the features of the original input image to multiple different memory feature spaces through the collaborative memory module, so as to generate a set of memory features. The fused memory features contain more rich information representation, thereby realizing more robust feature reconstruction.

[0019] In this stage, the encoded feature query with a dimension of is received. Then, the feature is decomposed in channel dimension, and K encoded feature strips Q are obtained after decomposition, wherein , . The collaborative memory module contains S memory units, denoted as MU, . Each memory unit contains P memory items, denoted as MI, The dimension of each memory item is P and the dimension of the encoding feature bar Q is also P 1 1. The collaborative memory module uses memory units MU to read the memory features MF, and each memory unit MU reads the corresponding memory feature MF, The collaborative memory module uses multiple memory units to read memory features in parallel. During the reading of memory features MF, the memory items MI in the memory units MU are kept updated.

[0020] Reading process: first take the first memory unit as an example to introduce the memory reading process. Before the collaborative memory module reads the memory features, the encoding feature query is decomposed into K encoding feature bars Q. Q and the memory items in the first memory unit perform cosine similarity calculation. Here, each in the encoding feature bar Q is sequentially compared with all memory items , and the calculation result of Q and is a P×K correlation matrix, which represents the spatial distance relationship between each encoding feature bar and the memory items in the first memory unit . Continue to apply the Softmax function to the P×K correlation matrix in the vertical direction for normalization, and after normalization, the corresponding vertical direction correlation probability is obtained.

[0021]

[0022] The vertical direction correlation probability is regarded as the weight W of the memory items in the first memory unit , and each corresponding weight w is multiplied by each memory item in the memory item one by one and then added, and in this way the memory feature is read. The reading method is to read each memory feature bar one by one for each encoding feature bar.

[0023]

[0024] After reading each memory feature bar , it is spliced along the channel dimension to obtain the memory feature , and the memory feature read by the collaborative memory module is consistent with the dimension of the encoding feature query, both of which are The above describes how the encoded query utilizes the first memory unit in the co-memory module. Read the first memory feature The detailed process is as follows: In the collaborative memory module, there are S memory units (MUs), corresponding to the generation of S memory features (MFs). The reading of each memory feature in the collaborative memory module is a parallel process, occurring concurrently with the reading of the first memory feature. Similarly, this can be extended to all memory units (MUs).

[0025]

[0026] The collaborative memory module generates S memory features (MFs) by reading S memory units (MUs). All memory features (MFs) are then concatenated along the channel dimension to obtain a concatenated feature with dimension 1. The fusion memory features can be represented as The fused features are then fused through convolution, resulting in a co-memory vector (cmv). The co-memory vector has a dimension of [missing information]. .

[0027]

[0028] Co-memory vectors (CMVs) contain multiple different memory pattern information. Compared with traditional convolution extraction, they can extract multiple pattern information of images more clearly, such as lighting, texture, and shape, so that a more realistic image representation can be obtained after decoding.

[0029] Update Process: The co-memory module contains S memory units MU, and each memory unit contains P memory items MI. During training of the co-memory module, after each read operation, all memory items in the co-memory module are updated synchronously to ensure that any deviations in the memory items within the memory units are corrected in a timely manner. Referring to the read operation, let's still start with the first memory unit. Let's take the first memory unit as an example to introduce the memory items. The update process.

[0030] The co-memory module calculates Q and Q during the reading process. We obtain a P×K correlation matrix, which is used in the update process. The difference is that a Softmax function is applied to this correlation matrix along the horizontal direction during the update process to obtain a normalized similarity matrix. .

[0031]

[0032] Since this matrix calculates each encoded feature bar With each memory item Based on the similarity relationship, an index set can be built. To represent each memory item The index set contains all coded feature bars with the highest correlation. During the index set construction process, there may be cases where the same memory item is similar to multiple coded feature bars; that is, a specific coded feature bar can appear in the similarity index sets of multiple memory items. Among them, After the index set is built, the index records with the lowest correlation are traversed and denoted as the weight item. The weighted item with the highest correlation is denoted as Furthermore, the normalized similarity matrix is ​​further normalized to obtain... The purpose of normalization is to eliminate the influence of the weights least relevant to the specified memory item, thereby avoiding interference from the updated memory item by the encoded feature strip.

[0033]

[0034] The similarity matrix between the normalized encoded feature bars and memory items sets the weight of the encoded feature bar least relevant to a given memory item to 0. Used as weights for the first memory unit Each memory entry is updated, thus eliminating the influence of this coded feature strip on the memory entries. After eliminating the least relevant coded feature strip, all other coded feature strips are used according to the calculated... As weights, the weights and each encoded feature bar are used. The method of multiplying and then adding is applied to the first memory unit. Each memory item Update.

[0035]

[0036] By updating each memory entry, each memory entry can continuously approach the aggregation center of the relevant encoded features, thereby obtaining more accurate image prototype pattern information.

[0037] This is the first memory unit. Middle memory item The update process. In the collaborative memory module, the memory items in each memory unit are updated in parallel, so this can be extended to all memory units (MUs).

[0038]

[0039] Through the memory retrieval and memory unit update processes within the collaborative memory module, this module can rapidly retrieve multiple memory features that model various patterns of the original image's encoded features. During the modeling process, the update process ensures that the retrieved patterns are not affected by extreme cases, allowing for the retrieval of more detailed information representations and thus achieving higher-quality feature reconstruction.

[0040] 4. The feature reconstruction method based on collaborative memory according to claim 1, wherein the specific process of step S3 is as follows:

[0041] This step aims to upsample the co-memory vector generated by the co-memory module to generate an output image X2 with the same dimensional size as the input image X1, so as to ensure that the detailed information of the input image is fully recovered.

[0042] In the decoding stage of the co-memory vector, a decoder consisting of a combination of four layers of deconvolution, batch normalization, and upsampling of activation functions is used. The dimension of the co-memory vector (cmv) is... .

[0043]

[0044] The collaborative memory vector cmv is upsampled to dimension after passing through the first deconvolution layer. intermediate decoding features , After the second deconvolution layer, it is upsampled to dimension [dimensionality]. intermediate decoding features , After the third deconvolution layer, it is upsampled to dimension [dimensionality]. intermediate decoding features , After the fourth deconvolution layer, it is upsampled into an output image X2 with dimensions C×H×W.

[0045] 5. The feature reconstruction method based on collaborative memory according to claim 1, wherein the specific process of S4 is as follows:

[0046] This step aims to encode the output image generated by S3. The purpose of re-encoding the output image is to complete the consistency check between the input and output images at the feature level, so as to ensure that the generated output image can better match the features of the original input image, thereby indirectly stimulating the output image to be more realistic.

[0047] The output image X2 is encoded using encoder B. Encoder B uses a similar combination of convolution, batch normalization, and LeakyReLU activation functions as encoder A. To better ensure consistency, the dimensions of each encoding stage are kept consistent with those of encoder A. The encoding process ends by compressing the image to a dimension of [dimensional value missing]. The eigenvector fv.

[0048]

[0049] The output image X2 is encoded into a dimension of after passing through the first convolutional layer. X2_1 has the same dimensions as X1_1; X2_1 is encoded into a form with dimensions after a second convolutional layer. X2_2 has the same dimensions as X1_2; X2_2 is encoded into a dimension after a third convolutional layer. X2_3 has the same dimensions as X1_3; X2_3 is encoded into a dimension after a fourth convolutional layer. The dimension of fv is the same as the dimension of the co-memory vector cmv.

[0050] After encoding the output image X2, encoding stage features with the same dimensions as the input image X1 are generated at each stage of the encoding process. Ideally, if the image details of the input image X1 and the output image X2 are consistent, the features acquired by the encoder with the same structure should also be consistent. Therefore, these features are used in stage S5 to optimize the consistency loss of the encoding features between the input and output images.

[0051] 6. The feature reconstruction method based on collaborative memory according to claim 1, wherein the specific process of step S5 is as follows:

[0052] This step aims to define the training objectives of the entire neural network architecture to ensure that the training of the neural network can be optimized as expected. This stage mainly involves optimizing the pixel reconstruction loss of the input and output images and optimizing the stage-by-stage encoding loss of the input and output images.

[0053] Optimization Objective 1: Optimization of Pixel Reconstruction Loss

[0054] Pixel reconstruction loss optimization This refers to using L2 distance to perform pixel-by-pixel loss calculations on the input image X1 and the output image X2 to ensure that the details on the image surface can be clearly restored.

[0055]

[0056] Optimization Objective 2: Optimization of Stage-by-Stage Encoding Loss

[0057] Stage-by-stage encoding loss optimization This refers to using L1 distance to perform phase-by-phase feature consistency optimization on X1_1, X1_2, X1_3 generated in encoder A and the co-memory vector cmv reconstructed by the co-memory module, and X2_1, X2_2, X2_3, fv generated in encoder B. This ensures that the features encoded by the input image X1 and the output image X2 are consistent during training, thereby optimizing the generation quality of the output image X2.

[0058]

[0059] Pixel reconstruction loss ensures that the model can generate an output image X2 that is consistent with the details and style of the input image X1, while staged encoding loss ensures that the model corrects the features of the input image X1 and the output image X2, thereby ensuring that the model can optimize in the direction of generating high-quality output images.

[0060] Compared with existing technologies, the beneficial effects of this invention are:

[0061] 1. Compared with existing technologies, the most significant advantage of this invention lies in its ability to achieve higher-quality feature reconstruction and generate image content that is highly consistent with the style of the input image. Traditional image processing methods, especially models based on simple encoder-decoder structures, often suffer from information loss or insufficient feature representation when processing complex images, resulting in reconstructed images that are blurry in detail or have an inconsistent overall style. This invention, by introducing a co-memory module, can more effectively capture and store multi-scale, deep semantic features of images, forming a co-memory vector containing rich memory information. Subsequently, combined with the training objective of stage-wise encoding loss, the model can not only faithfully restore the image content during the decoding process but also accurately imitate the style features of the input image, thereby generating a higher-quality output image that is visually more natural, more realistic, and more stylistically consistent.

[0062] 2. The innovation of this invention lies not only in the improved reconstruction quality, but also in its enhanced capabilities and superior generalization performance for image generation tasks. Through a collaborative memory mechanism, the model learns more refined and robust image representations, resulting in superior performance in advanced visual tasks such as style transfer, content generation, and image inpainting. In content generation, it produces new images that are more semantically logical and richer in detail. The design of the collaborative memory module helps the model establish a more robust internal knowledge representation, thus exhibiting stronger generalization ability and reducing the likelihood of overfitting when faced with unseen new data. This collaborative memory learning paradigm allows the model to move beyond simple pixel mapping and intelligently generate images based on a deep understanding and memory of image features, opening up new possibilities for applications in artificial intelligence visual tasks.

[0063] Figure 1 This is a schematic diagram of a feature reconstruction method based on collaborative memory.

[0064] Figure 2 This is an image visualization of the reconstructed features generated by a feature reconstruction method based on collaborative memory. Detailed Implementation

[0065] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent.

[0066] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0067] Figure 1 This is a schematic diagram of a feature reconstruction method based on collaborative memory. Figure 1 As shown, the model is first fed an input image X1 with dimensions C×H×W. In this stage, the input image X1 is compressed into the encoded feature query by encoder A. Encoder A consists of four cascaded 3×3 convolutional layers, normalization layers, and LeakyReLU activation functions, and its encoding steps are as follows.

[0068]

[0069] During the encoding process, the dimension of X1_1 is The dimension of X1_2 is The dimension of X1_3 is The dimensions of the query are The query is the final encoded feature of this stage, which compresses the convolutional feature information of the original input image and is used for the encoded feature decomposition and corresponding memory feature retrieval in the S2 stage.

[0070] After encoding is complete, the query output from the previous stage is input into the collaborative memory module. In this stage, the module first receives the dimension obtained from the encoding in the previous stage. The query is encoded as a feature. This feature is then decomposed along the channel dimension, resulting in K encoded feature bars Q, where... , The collaborative memory module contains S memory units, denoted as MU. Each memory unit contains P memory items, denoted as MI. The dimension of each memory item is the same as the Q dimension of the encoded feature bar. 1 1. The collaborative memory module uses memory units (MUs) to read memory features (MFs). Each memory unit (MU) reads the corresponding memory feature (MF). The collaborative memory module uses a parallel reading method across multiple memory units to read memory features. During the reading of memory features (MF), the memory entries (MI) in the memory unit (MU) are kept updated.

[0071] Reading process: First, starting with the first memory unit Let's take an example to illustrate the memory retrieval process. Before the collaborative memory module retrieves the memory features, the encoded feature query is decomposed into K encoded feature bars Q. Q is related to the first memory unit. Memory items in Perform cosine similarity calculation. Here, each feature bar Q is encoded. All are sequentially associated with all memory items. Perform similarity calculation, Q and The result is a P×K correlation matrix, which represents the relationship between each encoded feature bar and the first memory unit. Memory items in The spatial distance relationship is determined. Further, the correlation matrix of dimension P×K is normalized vertically using the Softmax function. After normalization, the corresponding vertical correlation probabilities are obtained. .

[0072]

[0073] Associate the probability in the vertical direction Considered the first memory unit Memory items in The weights W are then assigned to each corresponding weight w and the memory term. Each memory item in The memory features are retrieved by performing weighted multiplication and summing the results. The reading method involves reading each memory feature bar one by one for each encoded feature bar. .

[0074]

[0075] After reading each memory feature bar Then, the features are spliced ​​along the channel dimension to obtain the memory features. Memory features obtained through the collaborative memory module Consistent with the dimensions of the encoded feature query, all are... The above describes how the encoded query utilizes the first memory unit in the co-memory module. Read the first memory feature The detailed process is as follows: In the collaborative memory module, there are S memory units (MUs), corresponding to the generation of S memory features (MFs). The reading of each memory feature in the collaborative memory module is a parallel process, occurring concurrently with the reading of the first memory feature. Similarly, this can be extended to all memory units (MUs).

[0076]

[0077] The collaborative memory module generates S memory features (MFs) by reading S memory units (MUs). All memory features (MFs) are then concatenated along the channel dimension to obtain a concatenated feature with dimension 1. The fusion memory features can be represented as The fused features are then fused through convolution, resulting in a co-memory vector (cmv). The co-memory vector has a dimension of [missing information]. .

[0078]

[0079] Co-memory vectors (CMVs) contain multiple different memory pattern information. Compared with traditional convolution extraction, they can extract multiple pattern information of images more clearly, such as lighting, texture, and shape, so that a more realistic image representation can be obtained after decoding.

[0080] Update Process: The co-memory module contains S memory units MU, and each memory unit contains P memory items MI. During training of the co-memory module, after each read operation, all memory items in the co-memory module are updated synchronously to ensure that any deviations in the memory items within the memory units are corrected in a timely manner. Referring to the read operation, let's still start with the first memory unit. Let's take the first memory unit as an example to introduce the memory items. The update process.

[0081] The co-memory module calculates Q and Q during the reading process. We obtain a P×K correlation matrix, which is used in the update process. The difference is that a Softmax function is applied to this correlation matrix along the horizontal direction during the update process to obtain a normalized similarity matrix. .

[0082]

[0083] Since this matrix calculates each encoded feature bar With each memory item Based on the similarity relationship, an index set can be built. To represent each memory item The index set contains all coded feature bars with the highest correlation. During the index set construction process, there may be cases where the same memory item is similar to multiple coded feature bars; that is, a specific coded feature bar can appear in the similarity index sets of multiple memory items. Among them, After the index set is built, the index records with the lowest correlation are traversed and denoted as the weight item. The weighted item with the highest correlation is denoted as Furthermore, the normalized similarity matrix is ​​further normalized to obtain... The purpose of normalization is to eliminate the influence of the weights least relevant to the specified memory item, thereby avoiding interference from the updated memory item by the encoded feature strip.

[0084]

[0085] The similarity matrix between the normalized encoded feature bars and memory items sets the weight of the encoded feature bar least relevant to a given memory item to 0. Used as weights for the first memory unit Each memory entry is updated, thus eliminating the influence of this coded feature strip on the memory entries. After eliminating the least relevant coded feature strip, all other coded feature strips are used according to the calculated... As weights, the weights and each encoded feature bar are used. The method of multiplying and then adding is applied to the first memory unit. Each memory item Update.

[0086]

[0087] By updating each memory entry, each memory entry can continuously approach the aggregation center of the relevant encoded features, thereby obtaining more accurate image prototype pattern information.

[0088] This is the first memory unit. Middle memory item The update process. In the collaborative memory module, the memory items in each memory unit are updated in parallel, so this can be extended to all memory units (MUs).

[0089]

[0090] Through the memory retrieval and memory unit update processes within the collaborative memory module, this module can rapidly retrieve multiple memory features that model various patterns of the original image's encoded features. During the modeling process, the update process ensures that the retrieved patterns are not affected by extreme cases, allowing for the retrieval of more detailed information representations and thus achieving higher-quality feature reconstruction.

[0091] The co-memory module generates a co-memory vector (cmv), which serves as the input for the next stage. In the co-memory vector decoding stage, a decoder consisting of four layers of deconvolution, batch normalization, and upsampling of activation functions is used. The co-memory vector (cmv) has a dimension of [missing information]. .

[0092]

[0093] The collaborative memory vector cmv is upsampled to dimension after passing through the first deconvolution layer. intermediate decoding features , After the second deconvolution layer, it is upsampled to dimension [dimensionality]. intermediate decoding features , After the third deconvolution layer, it is upsampled to dimension [dimensionality]. intermediate decoding features , After the fourth deconvolution layer, it is upsampled into an output image X2 with dimensions C×H×W.

[0094] Furthermore, the output image X2 is encoded using encoder B. Encoder B employs a similar combination of convolution, batch normalization, and LeakyReLU activation functions as encoder A. To ensure better consistency checks, the dimensions of each encoding stage are kept consistent with those of encoder A. The encoding process ends with compressing the image to a dimension of [dimensional value missing]. The eigenvector fv.

[0095]

[0096] The output image X2 is encoded into a dimension of after passing through the first convolutional layer. X2_1 has the same dimensions as X1_1; X2_1 is encoded into a form with dimensions after a second convolutional layer. X2_2 has the same dimensions as X1_2; X2_2 is encoded into a dimension after a third convolutional layer. X2_3 has the same dimensions as X1_3; X2_3 is encoded into a dimension after a fourth convolutional layer. The dimension of fv is the same as the dimension of the co-memory vector cmv.

[0097] After encoding the output image X2, encoding stage features with the same dimensions as the input image X1 are generated at each stage of the encoding process. Ideally, if the image details of the input image X1 and the output image X2 are consistent, the features acquired by the encoder with the same structure should also be consistent. Therefore, these features are used in stage S5 to optimize the consistency loss of the encoding features between the input and output images.

[0098] Finally, there is the training objective of optimizing the entire network architecture.

[0099] Optimization Objective 1: Optimization of Pixel Reconstruction Loss

[0100] Pixel reconstruction loss optimization This refers to using L2 distance to perform pixel-by-pixel loss calculations on the input image X1 and the output image X2 to ensure that the details on the image surface can be clearly restored.

[0101]

[0102] Optimization Objective 2: Optimization of Stage-by-Stage Encoding Loss

[0103] Stage-by-stage encoding loss optimization This refers to using L1 distance to perform phase-by-phase feature consistency optimization on X1_1, X1_2, X1_3 generated in encoder A and the co-memory vector cmv reconstructed by the co-memory module, and X2_1, X2_2, X2_3, fv generated in encoder B. This ensures that the features encoded by the input image X1 and the output image X2 are consistent during training, thereby optimizing the generation quality of the output image X2.

[0104]

[0105] Pixel reconstruction loss ensures that the model can generate an output image X2 that is consistent with the details and style of the input image X1, while staged encoding loss ensures that the model corrects the features of the input image X1 and the output image X2, thereby ensuring that the model can optimize in the direction of generating high-quality output images.

[0106] Figure 2 This image visualization demonstrates the reconstructed features generated using a collaborative memory-based feature reconstruction method. It showcases the effects of the bottle bottom dataset, capsule dataset, nut dataset, and fabric dataset. The input image represents the content input to the model, while the output image represents the feature reconstruction results from the collaborative memory module. It can be seen that the model effectively reconstructs the texture, lighting, color, rotation angle, and other image contextual information for all datasets, achieving high-quality image generation.

[0107] This invention proposes a feature reconstruction method based on collaborative memory, aiming to significantly enhance the feature extraction and representation capabilities of artificial intelligence vision tasks and greatly improve the output quality and style consistency of image generation tasks. This technology effectively solves key problems commonly found in existing methods, such as insufficient feature representation, information bottlenecks leading to low-quality reconstructed images and loss of detail, as well as inaccurate image style transfer and poor semantic coherence in style transfer and content generation. The core of this technology lies in the introduction of an innovative collaborative memory module. This module does not simply perform linear or convolutional fusion of encoded features, but rather transforms image encoded features into a collaborative memory vector containing rich memory information through the following core mechanism: First, the encoded features are treated as input to a latent memory bank; then, the collaborative memory module uses multiple parallel memory units to perform memory transformation on the encoded features, forming various pattern features. It explicitly extracts, organizes, and integrates relevant, cross-scale, and representative feature fragments from memory to form a memory feature set; finally, this memory feature set is aggregated and refined through convolutional operations to generate a collaborative memory vector that can effectively guide the subsequent decoding process and contains target style information. This invention achieves high-quality feature reconstruction, ensuring high accuracy in detail, semantics, and style of the output image, and is capable of generating image content that is highly consistent with the input image in visual style. This powerful capability enables the technology to be widely applied in image generation, artificial intelligence vision tasks, image recognition, scene understanding, and any other field requiring deep feature learning, precise semantic representation, and stylized content creation.

[0108] Finally, the details of the above examples of the present invention are merely illustrative of the invention. Any modifications, improvements, and substitutions to the above embodiments by those skilled in the art should be included within the scope of protection of the claims of the present invention.

Claims

1. A feature reconstruction method based on collaborative memory, characterized in that, The method includes the following steps: S1. Compress the input image to generate coded features, including using an encoder composed of convolutions to compress the image features and generate corresponding coded features for input to the co-memory module; S2. Processing encoded features to generate co-memory vectors, including using a co-memory module to process image encoded features, converting encoded features into a set of memory features, and fusing the set of memory features through convolution operations to obtain a co-memory vector; S3. Decode the co-memory vector, including using a decoder composed of convolutions to decode the co-memory vector, and obtaining an output image with the same size as the input image after decoding; S4. Encode the output image, including encoding the output image using an encoder composed of convolutions; S5. Training optimization objectives include optimizing the pixel reconstruction loss of the input and output images and optimizing the stage-by-stage encoding loss of the input and output images.

2. The feature reconstruction method based on collaborative memory according to claim 1, characterized in that, The specific process of S1 is as follows: This step aims to compress the dimensions of the input image, compressing the three-dimensional input image into encoded features through convolution operations. The encoded features generated in this stage serve as the input to the collaborative memory module in the next stage. In this stage, an input image X1 with dimensions C×H×W is first input to the model. The input image X1 is then compressed into an encoded feature query by encoder A. Encoder A consists of four cascaded 3×3 convolutional layers, normalization layers, and LeakyReLU activation functions, and its encoding steps are as follows. During the encoding process, the dimension of X1_1 is The dimension of X1_2 is The dimension of X1_3 is The dimensions of the query are The query is the final encoded feature of this stage, which compresses the convolutional feature information of the original input image and is used for the encoded feature decomposition and corresponding memory feature retrieval in the S2 stage.

3. The feature reconstruction method based on collaborative memory according to claim 1, characterized in that, The specific process of S2 is as follows: This step aims to map the encoded feature query, which retains the original input image features, to multiple different memory feature spaces through a collaborative memory module, thereby generating a set of memory features. The fused memory features contain richer information representations, thus achieving more robust feature reconstruction. In this stage, the dimension obtained from the encoding in the previous stage is first received. The query is encoded as a feature. This feature is then decomposed along the channel dimension, resulting in K encoded feature bars Q, where... , The collaborative memory module contains S memory units, denoted as MU. Each memory unit contains P memory items, denoted as MI. The dimension of each memory item is the same as the Q dimension of the encoded feature bar. 1 1. The collaborative memory module uses memory units (MUs) to read memory features (MFs). Each memory unit (MU) reads the corresponding memory feature (MF). The collaborative memory module uses a parallel reading method across multiple memory units to read memory features. During the reading of memory features (MF), the memory entries (MI) in the memory unit (MU) are kept updated. Reading process: First, starting with the first memory unit Let's take an example to illustrate the memory retrieval process. Before the collaborative memory module retrieves the memory features, the encoded feature query is decomposed into K encoded feature bars Q. Q is related to the first memory unit. Memory items in Perform cosine similarity calculation. Here, each feature bar Q is encoded. All are sequentially associated with all memory items. Perform similarity calculation, Q and The result is a P×K correlation matrix, which represents the relationship between each encoded feature bar and the first memory unit. Memory items in The spatial distance relationship is determined. Further, the correlation matrix of dimension P×K is normalized vertically using the Softmax function. After normalization, the corresponding vertical correlation probabilities are obtained. . Associate the probability in the vertical direction Considered the first memory unit Memory items in The weights W are then assigned to each corresponding weight w and the memory term. Each memory item in The memory features are retrieved by performing weighted multiplication and summing the results. The reading method involves reading each memory feature bar one by one for each encoded feature bar. . After reading each memory feature bar Then, the features are spliced ​​along the channel dimension to obtain the memory features. Memory features obtained through the collaborative memory module Consistent with the dimensions of the encoded feature query, all are... The above describes how the encoded query utilizes the first memory unit in the co-memory module. Read the first memory feature The detailed process is as follows: In the collaborative memory module, there are S memory units (MUs), corresponding to the generation of S memory features (MFs). The reading of each memory feature in the collaborative memory module is a parallel process, occurring concurrently with the reading of the first memory feature. Similarly, this can be extended to all memory units (MUs). The collaborative memory module generates S memory features (MFs) by reading S memory units (MUs). All memory features (MFs) are then concatenated along the channel dimension to obtain a concatenated feature with dimension 1. The fusion memory features can be represented as The fused features are then fused through convolution, resulting in a co-memory vector (cmv). The co-memory vector has a dimension of [missing information]. . Co-memory vectors (CMVs) contain multiple different memory pattern information. Compared with traditional convolution extraction, they can extract multiple pattern information of images more clearly, such as lighting, texture, and shape, so that a more realistic image representation can be obtained after decoding. Update Process: The co-memory module contains S memory units MU, and each memory unit contains P memory items MI. During training of the co-memory module, after each read operation, all memory items in the co-memory module are updated synchronously to ensure that any deviations in the memory items within the memory units are corrected in a timely manner. Referring to the read operation, let's still start with the first memory unit. Let's take the first memory unit as an example to introduce the memory items. The update process. The co-memory module calculates Q and Q during the reading process. We obtain a P×K correlation matrix, which is used in the update process. The difference is that a Softmax function is applied to this correlation matrix along the horizontal direction during the update process to obtain a normalized similarity matrix. . Since this matrix calculates each encoded feature bar With each memory item Based on the similarity relationship, an index set can be built. To represent each memory item The index set contains all coded feature bars with the highest correlation. During the index set construction process, there may be cases where the same memory item is similar to multiple coded feature bars; that is, a specific coded feature bar can appear in the similarity index sets of multiple memory items. Among them, After the index set is built, the index records with the lowest correlation are traversed and denoted as the weight item. The weighted item with the highest correlation is denoted as Furthermore, the normalized similarity matrix is ​​further normalized to obtain... The purpose of normalization is to eliminate the influence of the weights least relevant to the specified memory item, thereby avoiding interference from the updated memory item by the encoded feature strip. The similarity matrix between the normalized encoded feature bars and memory items sets the weight of the encoded feature bar least relevant to a given memory item to 0. Used as weights for the first memory unit Each memory entry is updated, thus eliminating the influence of this coded feature strip on the memory entries. After eliminating the least relevant coded feature strip, all other coded feature strips are used according to the calculated... As weights, the weights and each encoded feature bar are used. The method of multiplying and then adding is applied to the first memory unit. Each memory item Update. By updating each memory entry, each memory entry can continuously approach the aggregation center of the relevant encoded features, thereby obtaining more accurate image prototype pattern information. This is the first memory unit. Middle memory item The update process. In the collaborative memory module, the memory items in each memory unit are updated in parallel, so this can be extended to all memory units (MUs). Through the memory retrieval and memory unit update process in the collaborative memory module, the module is able to quickly read multiple memory features that model various patterns of the original image encoded features. During the modeling process, the update process ensures that the read pattern is not affected by extreme cases, reads more detailed information representation, and thus achieves higher quality feature reconstruction.

4. The feature reconstruction method based on collaborative memory according to claim 1, characterized in that, The specific process of S3 is as follows: This step aims to upsample the co-memory vector generated by the co-memory module to generate an output image X2 with the same dimensional size as the input image X1, so as to ensure that the detailed information of the input image is fully recovered. In the decoding stage of the co-memory vector, a decoder consisting of a combination of four layers of deconvolution, batch normalization, and upsampling of activation functions is used. The dimension of the co-memory vector (cmv) is... . The collaborative memory vector cmv is upsampled to dimension after passing through the first deconvolution layer. intermediate decoding features , After the second deconvolution layer, it is upsampled to dimension [dimensionality]. intermediate decoding features , After the third deconvolution layer, it is upsampled to dimension [dimensionality]. intermediate decoding features , After the fourth deconvolution layer, it is upsampled into an output image X2 with dimensions C×H×W.

5. The feature reconstruction method based on collaborative memory according to claim 1, characterized in that, The specific process of S4 is as follows: This step aims to encode the output image generated by S3. The purpose of re-encoding the output image is to complete the consistency check between the input and output images at the feature level, so as to ensure that the generated output image can better match the features of the original input image, thereby indirectly stimulating the output image to be more realistic. The output image X2 is encoded using encoder B. Encoder B uses a similar combination of convolution, batch normalization, and LeakyReLU activation functions as encoder A. To better ensure consistency, the dimensions of each encoding stage are kept consistent with those of encoder A. The encoding process ends by compressing the image to a dimension of [dimensional value missing]. The eigenvector fv. The output image X2 is encoded into a dimension of after passing through the first convolutional layer. X2_1 has the same dimensions as X1_1; X2_1 is encoded into a form with dimensions after a second convolutional layer. X2_2 has the same dimensions as X1_2; X2_2 is encoded into a dimension after a third convolutional layer. X2_3 has the same dimensions as X1_3; X2_3 is encoded into a dimension after a fourth convolutional layer. The dimension of fv is the same as the dimension of the co-memory vector cmv. After encoding the output image X2, encoding stage features with the same dimensions as the input image X1 are generated at each stage of the encoding process. Ideally, if the image details of the input image X1 and the output image X2 are consistent, the features acquired by the encoder with the same structure should also be consistent. Therefore, these features are used in stage S5 to optimize the consistency loss of the encoding features between the input and output images.

6. The feature reconstruction method based on collaborative memory according to claim 1, characterized in that, The specific process of S5 is as follows: This step aims to define the training objectives of the entire neural network architecture to ensure that the training of the neural network can be optimized as expected. This stage mainly involves optimizing the pixel reconstruction loss of the input and output images and optimizing the stage-by-stage encoding loss of the input and output images. Optimization Objective 1: Optimization of Pixel Reconstruction Loss Pixel reconstruction loss optimization This refers to using L2 distance to perform pixel-by-pixel loss calculations on the input image X1 and the output image X2 to ensure that the details on the image surface can be clearly restored. Optimization Objective 2: Optimization of Stage-by-Stage Encoding Loss Stage-by-stage encoding loss optimization This refers to using L1 distance to perform phase-by-phase feature consistency optimization on X1_1, X1_2, X1_3 generated in encoder A and the co-memory vector cmv reconstructed by the co-memory module, and X2_1, X2_2, X2_3, fv generated in encoder B. This ensures that the features encoded by the input image X1 and the output image X2 are consistent during training, thereby optimizing the generation quality of the output image X2. Pixel reconstruction loss ensures that the model can generate an output image X2 that is consistent with the details and style of the input image X1, while staged encoding loss ensures that the model corrects the features of the input image X1 and the output image X2, thereby ensuring that the model can optimize in the direction of generating high-quality output images.