An image processing method of a contrast learning world model based on slot attention mechanism feature fusion

By fusing slot attention mechanism with contrastive learning world model in image processing, this paper solves the problem of difficulty in fusing slot attention mechanism with world model in existing technologies, and achieves efficient image feature extraction and behavior prediction.

CN121010862BActive Publication Date: 2026-02-03BEIJING NORMAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511114577.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2026-02-03
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing technologies cannot effectively integrate slot attention mechanisms with world models, resulting in low efficiency in image data analysis and an inability to perform effective feature extraction and real-time image understanding.

Method used

A contrastive learning world model based on slot attention mechanism feature fusion is adopted. It combines convolutional neural network, multilayer perceptron, adapter and graph neural network, and combines slot attention mechanism and multilayer perceptron output for feature fusion. The model is optimized through a three-stage training strategy.

Benefits of technology

It improves the quality of image understanding and the accuracy of behavior prediction, enhances the structural decoupling ability and interpretability of the model, and realizes effective feature extraction and long-term time series modeling of image data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010862B_ABST
    Figure CN121010862B_ABST
Patent Text Reader

Abstract

The application discloses an image processing method of a contrast learning world model based on a slot attention mechanism feature fusion, and comprises the following steps: acquiring image data, processing the image data through a world model to obtain an image data processing result, wherein the image data processing result comprises a target prediction result, a target state estimation result and a future behavior prediction result in the image; wherein the world model comprises a convolutional neural network, a multilayer perceptron, an adapter and a graph neural network connected in sequence, the convolutional neural network and the multilayer perceptron are connected in parallel with a slot attention mechanism, the slot attention mechanism and the output of the multilayer perceptron are taken as the input of the adapter, and the output of the slot attention mechanism and the adapter is taken as the input of the graph neural network; and the world model is trained through a stage-by-stage training strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and computer vision technology, and in particular relates to an image processing method based on a contrastive learning world model with feature fusion using slot attention mechanism. Background Technology

[0002] In the development of intelligent agent systems, world models, as a key structure for understanding environmental dynamics and making inferences and predictions, have received widespread attention. While traditional contrastive learning methods can learn high-level feature representations of images, they face bottlenecks in long-term modeling, feature interpretability, and training efficiency. Slot attention, as an object-centered mechanism, possesses excellent structural decoupling capabilities and interpretability, but faces integration difficulties when combined with existing world models. Current technologies cannot effectively analyze real-time acquired image data using world models, and thus cannot effectively extract features from the image data. Summary of the Invention

[0003] To address the aforementioned technical problems, this invention proposes an image processing method based on a contrastive learning world model with feature fusion using a slot attention mechanism, thereby resolving the issues present in the prior art.

[0004] To achieve the above objectives, this invention provides an image processing method based on a slot attention mechanism feature fusion contrastive learning world model, comprising:

[0005] Image data is acquired, and the image data is processed through a world model to obtain image data processing results, wherein the image data processing results include target prediction results, target state estimation results, and future behavior prediction results in the image;

[0006] The world model includes a convolutional neural network, a multilayer perceptron, an adapter, and a graph neural network connected in sequence. The convolutional neural network and the multilayer perceptron are connected in parallel with a slot attention mechanism. The output of the slot attention mechanism and the multilayer perceptron is used as the input of the adapter, and the output of the slot attention mechanism and the adapter is used as the input of the graph neural network.

[0007] The world model is trained using a phased training strategy.

[0008] Optionally, the process of processing the image data using a world model includes:

[0009] Feature extraction is performed on the image data using a convolutional neural network, and the extracted features are then encoded at their locations. The encoded data is then processed by a multilayer perceptron to obtain multilayer image features.

[0010] The image data is processed by a slot attention mechanism to obtain slot vectors. The latent objects in the slot vectors are then clustered and represented by the attention mechanism to obtain slot vector representations.

[0011] The slot vector representation and multi-level image features are connected and fused. The connection and fusion result is processed by spatial feature fusion through an adapter and then concatenated with the slot vector representation. The connection and fusion result is detected by a graph neural network to obtain a structured representation. The structured representation is then mapped to obtain the image data processing result.

[0012] Optionally, the world model adopts a two-stage training strategy, wherein the first stage is to train the convolutional neural network, multilayer perceptron, adapter and graph neural network, the second stage is to train the slot attention mechanism, and the third stage is to train the world model as a whole.

[0013] Optionally, the loss function used for training the slot attention mechanism is the mean squared error loss function.

[0014] Optionally, the loss function used to train convolutional neural networks, multilayer perceptrons, adapters, and graph neural networks may be a contrastive learning loss function.

[0015] Optionally, the process of training the slot attention mechanism includes:

[0016] The slot attention mechanism is trained for several rounds. For each round of training, the input image data is normalized, weighted, processed by GRU, and the slot vector is updated. Attention weights are calculated using the input image data and slot vectors, and weighted by the mean of the attention weights. The weighted result of the previous and current training rounds is input into the GRU for processing. Based on the processed slot vectors, a new slot vector is generated by a multilayer perceptron to update the slot vectors. The slot attention mechanism is then trained based on the updated slot vectors.

[0017] Optionally, the training process for convolutional neural networks, multilayer perceptrons, and graph neural networks includes:

[0018] The parameters of the frozen slot attention mechanism are set, backpropagation of the frozen slot attention mechanism in the world model is turned off, and only the contrastive learning loss function is used to train the convolutional neural network, multilayer perceptron, and graph neural network.

[0019] Optionally, the adapter's data processing procedure is as follows:

[0020] The output data of the slot attention mechanism and the multilayer perceptron are concatenated. The number of objects is adapted by the CNN output features through an adapter. The concatenated result and the output result of the slot attention mechanism are combined through a cross-scale attention mechanism to obtain the output data of the adapter.

[0021] On the other hand, the present invention provides an image processing system based on a contrastive learning world model with feature fusion using a slot attention mechanism, for performing the aforementioned method.

[0022] Compared with the prior art, the present invention has the following advantages and technical effects:

[0023] This invention provides a world modeling method that extracts multi-level features from input images using a convolutional neural network (CNN) combined with positional encoding and a multilayer perceptron. It initializes slot vectors using a slot attention mechanism and clusters and represents potential objects in the image. The multi-level features and slot representations are fused and input into a graph neural network (GNN) to achieve structured modeling of relationships between objects. A three-stage training strategy is employed: first, the CNN and GNN parts are compared and optimized; then, the slot module is trained using mean squared error (MSE) loss; and finally, spatial adaptation is used to improve the overall performance and training stability of the model. This application achieves unsupervised object modeling by introducing a slot mechanism and learning object relationships using graph structures, ultimately improving image understanding quality and behavior prediction accuracy. Attached Figure Description

[0024] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0025] Figure 1 This is the overall architecture of the world model in this embodiment of the invention. Detailed Implementation

[0026] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0027] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0028] This invention provides an image processing method for a contrastive learning world model of a fusion slot attention module. By constructing a contrastive learning world model of a fusion slot attention module, it is possible to further and effectively extract features, monitor behavior, and predict long-term effects from image data.

[0029] This invention aims to provide an image processing method for a contrastive learning world model that integrates a slot attention module. This method combines a slot attention mechanism with multi-layer feature representation and can be used for tasks such as agent environment modeling, behavior policy learning, and long-term prediction. It enhances the modeling capability of the world model through object-level feature representation. Specifically, by fusing the slot attention module with multi-layer features extracted by a convolutional neural network (CNN), it achieves better modeling and prediction of independent objects in an image.

[0030] The present invention has the following advantages:

[0031] 1. Slot attention provides object-level semantic modeling and interpretability; 2. Feature fusion mechanism enhances the generalization ability of model structure; 3. Two-stage training improves prediction accuracy; 4. It has a pluggable design and can be flexibly integrated into various deep learning frameworks.

[0032] A similar description is provided for the above technical solutions:

[0033] This technical solution provides an image processing method for a contrastive learning world model using a fusion slot attention module, such as... Figure 1 This includes the following steps:

[0034] Step 1: The input image first extracts spatial features through a convolutional neural network (CNN), and then passes through a converter adapter, which includes position encoding and a multilayer perceptron for further processing to form multi-layer image features.

[0035] The SlotAttention Encoder (SAE) receives image features, initializes a set of slot vectors, and uses an attention mechanism to cluster and represent potential objects in the image.

[0036] Multi-level features and slot representations are connected and fused using a spatial feature fusion method. An adapter module is added to ensure feature correspondence, and the features are passed into the GNN module to model the relationships between objects, forming a structured representation.

[0037] The model output is used for image prediction, state estimation, and future behavior simulation. The output data includes predicted target content, estimated target state, and future target behavior data.

[0038] For physics classroom teaching, the above-described solution of this invention can extract and define the bounding boxes of relevant targets (such as balls, cars, or other physical models) in physics images provided by teachers, estimate the target states, and determine the relevant motion trajectories of the targets in subsequent images. In the above, the input is the provided image of the physical model, and the output includes the bounding box content, state, and relevant motion trajectory of the targets in the image.

[0039] Step 2: The model adopts a three-stage training strategy. In the first stage, the CNN-GNN part is trained separately using the contrastive learning loss function, and the slot attention module is frozen to achieve initial convergence. In the second stage, the slot attention is trained separately using the MSE loss function. In the third stage, the overall training is performed end-to-end to achieve final adaptation.

[0040] Specifically, step one: preliminary training

[0041] For the CNN used in this invention, the following operations are performed, following the formula:

[0042]

[0043] in, The c-th output feature map m represents t The value of (channel) at spatial location (i,j), C in The input feature map has 10 channels, k represents the kernel size, m represents the number of channels, u represents the horizontal offset, v represents the vertical offset, and x represents the feature map (the input physical image s). t ), where w represents the weights in the convolutional neural network and b represents the bias term.

[0044] For the graph neural network used in this invention, the following operations are performed, following the formula:

[0045]

[0046] in, It is the representation of node i in the l-th layer; is the set of neighbors of node i; AGGREGATE aggregates neighbor information (average, summation, maximum value, attention weighting, etc.); UPDATE means updating with aggregated information and its current state, which represents the update operation of the graph neural network.

[0047] The first stage trains the CNN, Adapter, and GNN network components. At this point, the slot attention module is frozen, its backpropagation throughout the network is disabled, and only the contrastive learning loss function is used.

[0048]

[0049] Where d represents the square of the Euclidean distance, and z represents the prediction in the latent space. γ represents the actual value in the latent space, T represents the translation based on TransE, a represents the action, t represents the step size, and γ represents the hyperparameter, which is usually 1.

[0050] Step 2: Pre-trained slot attention module

[0051] To obtain a slot attention encoder (or slot attention module) that can be used for downstream tasks, an encoder-decoder for training needs to be constructed first using a training method similar to VAE. For each training round, the forward propagation includes normalization, weighting, GRU, and updating the slot vector information. The data processing procedure for the forward propagation, or slot attention model, is as follows: First, the attention needs to be calculated according to the following formula:

[0052]

[0053] Where attn represents the slot attention coefficient, D represents the dimension of the input vector, 1 / sqrt(D) forms a scaling factor, k(inputs) represents the linear transformation of the input, q(slots) represents the slot-based query, axis indicates that the softmax operation is normalized along the slot dimension, and slots represents the object slots.

[0054] For updates during time-slot iterations, to improve the stability of the attention mechanism, it is usually weighted according to the average attention value. The formula is as follows:

[0055] updates=WeightedMean(weights=attn+∈, values=v(inputs)) (5)

[0056] Where updates represents the weighted result, weights represents the weights, attn represents the slot attention coefficient, ∈ represents a very small positive number, values ​​represents the weighted value, v represents the learnable linear projection, and inputs represents the input image;

[0057] Use GRU to update the slot, input the weighted results of the previous and current rounds into GRU, and update the MLP results:

[0058] slots=GRU(state=slots_prev, inputs=updates) (6)

[0059] slots+=MLP(LyaerNorm(slots))) (7)

[0060] Where slots_prev represents the slot information in the previous iteration, state represents the state of the input GRU model, inputs represents the input data, updates represents the weighted result in formula (5), and slots+ represents the difference iteration of the slot information.

[0061] For backpropagation or reverse update, the slot attention module is pre-trained and then self-supervised using a separate training set; based on the MSE loss function:

[0062]

[0063] Among them, y i This represents the true value of the i-th sample. Let represent the predicted value of the i-th sample, and n represent the number of samples.

[0064] In the third stage, backpropagation is initiated for end-to-end adaptive training to enhance the network's adaptability. An attention-based spatial feature fusion method is employed, proposing an adapter based on multi-layer bilinear upsampling to adapt CNN features according to the number of objects. Furthermore, a cross-scale attention mechanism is used to achieve dynamic feature adaptation, combining CNN and object-centric attention with adaptive weights. The adaptive method formula in the adapter is as follows:

[0065]

[0066] Among them, Z i V is the output feature at position i. j It is the original input feature vector, where the original input feature vector V j W is the result of concatenating the Adapter and slot vector information. v It is the linear transformation weight matrix of the value vector, corresponding to Figure 1 W in V The value vector is obtained by linear mapping through slot vector information, e ij The unnormalized attention scores for query i and key j are represented by the following formula.

[0067]

[0068] Where C is the scaling factor, Q i K represents the query vector at position i. j W represents the key vector at position j. Q and W k These are learnable weight matrices used for projecting the query and key, respectively. The query vector and key vector are obtained through a linear mapping of the original input feature vectors, corresponding to... Figure 1 W q and Wk .

[0069] The proposed adapter follows the formula below

[0070]

[0071] Where x represents the position of the input feature in the original image (h) i ,w j The channel vector of ), where W represents the weight of the science department, b represents the bias, and w ij This indicates that the distance (h) is based on the target location (h′, w′). i ,w j The bilinear interpolation weights satisfy ∑ i,j w ij =1.

[0072] In backpropagation, the loss function of the first-stage contrastive learning is used for end-to-end parameter updates.

[0073] The world model in this invention has the following effects on data processing:

[0074] 1. In the world model, obtain the positional information of objects by using object-centric features to improve the completeness of the world model features;

[0075] 2. The fusion method proposed in this paper uses a spatial attention mechanism to enable the convolutional features and object-centered features to adapt to each other.

[0076] 3. The training method provided in this proposal can bring this difficult-to-converge network to convergence through a three-stage training approach.

[0077] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image processing method based on a contrastive learning world model with feature fusion using a slot attention mechanism, characterized in that, include: Image data is acquired, and the image data is processed through a world model to obtain image data processing results, wherein the image data processing results include target prediction results, target state estimation results, and future behavior prediction results in the image; The world model includes a convolutional neural network, a multilayer perceptron, an adapter, and a graph neural network connected in sequence. The convolutional neural network and the multilayer perceptron are connected in parallel with a slot attention mechanism. The output of the slot attention mechanism and the multilayer perceptron is used as the input of the adapter, and the output of the slot attention mechanism and the adapter is used as the input of the graph neural network. The world model is trained using a phased training strategy; The process of processing the image data using a world model includes: Feature extraction is performed on the image data using a convolutional neural network, and the extracted features are then encoded at their locations. The encoded data is then processed by a multilayer perceptron to obtain multilayer image features. The image data is processed by a slot attention mechanism to obtain slot vectors. The latent objects in the slot vectors are then clustered and represented by the attention mechanism to obtain slot vector representations. The slot vector representation and multi-level image features are connected and fused. The connection and fusion result is processed by spatial feature fusion through an adapter and then concatenated with the slot vector representation. The connection and fusion result is detected by a graph neural network to obtain a structured representation. The structured representation is then mapped to obtain the image data processing result. The world model adopts a three-stage training strategy. The first stage is to train the convolutional neural network, multilayer perceptron, adapter, and graph neural network. The second stage is to train the slot attention mechanism. The third stage is to train the world model as a whole.

2. The method according to claim 1, characterized in that, The loss function used for training the slot attention mechanism is the mean squared error loss function.

3. The method according to claim 1, characterized in that, The contrastive learning loss function is used to train convolutional neural networks, multilayer perceptrons, adapters, and graph neural networks.

4. The method according to claim 1, characterized in that, The process of training the slot attention mechanism includes The slot attention mechanism is trained for several rounds. For each round of training, the input image data is normalized, weighted, processed by GRU, and the slot vector is updated. Attention weights are calculated using the input image data and slot vectors, and weighted by the mean of the attention weights. The weighted result of the previous and current training rounds is input into the GRU for processing. Based on the processed slot vectors, a new slot vector is generated by a multilayer perceptron to update the slot vectors. The slot attention mechanism is then trained based on the updated slot vectors.

5. The method according to claim 1, characterized in that, The training process for convolutional neural networks, multilayer perceptrons, and graph neural networks includes: The parameters of the frozen slot attention mechanism are set, backpropagation of the frozen slot attention mechanism in the world model is turned off, and only the contrastive learning loss function is used to train the convolutional neural network, multilayer perceptron, and graph neural network.

6. The method according to claim 1, characterized in that, The adapter's data processing procedure is as follows: The output data of the slot attention mechanism and the multilayer perceptron are concatenated. The number of objects is adapted by the CNN output features through an adapter. The concatenated result and the output result of the slot attention mechanism are combined through a cross-scale attention mechanism to obtain the output data of the adapter.

7. An image processing system based on a contrastive learning world model with feature fusion using a slot attention mechanism, characterized in that, Used to perform the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Conditional object-centric slot attention learning for video and other sequential data

    CN115018067A

  • Multi-modal fine-grained sentiment analysis method based on momentum contrast learning

    CN117435732A