Picture searching and displaying method and device

By generating and fusing multimodal feature vectors, using graph neural networks and adversarial network technologies, efficient image search and display are achieved, solving the accuracy and efficiency of complex visual content searches in the existing technology, and users can quickly obtain high-quality pictures that conform to the design style.

CN120407838AInactive Publication Date: 2025-08-01BEIJING CHUANGZUOMEIHAO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510536962.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing image search technology lacks accuracy in the search requirements of complex visual content, low image generation and display efficiency, making it difficult for users to quickly obtain high-quality pictures.

Method used

By generating weighted fusion of text search feature vectors, image feature vectors and design situation feature vectors, the graph neural network model is used to retrieve potentially associated candidate images in the image database, and style transfers through adversarial network models, so that the color, texture and composition of the candidate images are consistent with the style of the user's current design canvas page, and are displayed and decomposed into editable objects.

Benefits of technology

It improves the speed of image generation and display, improves the user experience, and allows users to quickly obtain high-quality pictures that match the design style.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407838A_ABST
    Figure CN120407838A_ABST
Patent Text Reader

Abstract

The invention discloses a picture search and display method and device. The method comprises the following steps: generating a text search feature vector according to a search keyword; analyzing a representative image in the canvas page through an image encoder model and generating an image feature vector; generating a design situation feature vector according to the context information in the current design canvas page; performing weighted fusion on the text search feature vector, the image feature vector and the design situation feature vector to generate a multi-modal feature vector; retrieving the multi-modal feature vectors in an image database to obtain potentially associated candidate images; performing style migration on the candidate images through the adversarial network model to enable the color, texture and composition of the candidate images to be consistent with the style of the canvas page currently designed by the user; the candidate image is displayed and decomposed into a plurality of editable objects. According to the method and the device, the picture generation and display speed is further improved, the experience of the user in complex visual content search is improved, and the user can quickly obtain high-quality pictures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of software design, and particularly to a method and device for image search and display. Background Art

[0002] Existing image search technologies usually perform matching based on keywords or metadata. Although they can provide a certain degree of search accuracy, for complex visual content search requirements, the existing technologies are often inadequate. In addition, the efficiency of image generation and display in the existing technologies is relatively low, and it is difficult for users to quickly obtain the required high-quality images.

[0003] Some existing AI image search tools (such as image search in search engines, AI image generation technologies) can provide certain intelligent search functions, but there are still deficiencies in the real-time performance and accuracy of image generation and display.

[0004] To solve the above problems, the present invention develops a method for image search and display to improve the speed of image generation and display and improve the user experience. Summary of the Invention

[0005] In view of the above problems, a method and device for image search and display are proposed to improve the speed of image generation and display and improve the user experience.

[0006] According to one aspect of the present invention, a method for image search and display is provided, including:

[0007] Generating a text search feature vector according to a search keyword input by a user; analyzing a representative image in the canvas page through an image encoder model and generating an image feature vector; generating a design context feature vector according to context information in the current design canvas page, where the context information includes the color distribution, font type, element type, and design style of the canvas page;

[0008] Performing weighted fusion on the text search feature vector, the image feature vector, and the design context feature vector to generate a multi-modal feature vector;

[0009] Retrieving the multi-modal feature vector in an image database through a graph neural network model to obtain potentially associated candidate images, where the image database contains material images and their corresponding multi-modal feature vectors;

[0010] Performing style transfer on the candidate images through an adversarial network model to make the colors, textures, and compositions of the candidate images consistent with the style of the user's current design canvas page; displaying the candidate images and decomposing them into multiple editable objects, where the position, size, color, and transparency of each editable object can be adjusted separately.

[0011] In an alternative approach, the method further includes:

[0012] Generating a saliency map on the candidate image according to a saliency detection algorithm, where the saliency map indicates the visually most attention-grabbing regions in the image;

[0013] Based on the saliency map, using an image segmentation algorithm to segment the candidate image into multiple objects, where regions with high saliency are preferentially retained during the segmentation process;

[0014] Extracting the feature vectors of each object and matching them with the keywords input by the user;

[0015] Performing weighted fusion according to the matching results of each object with the keywords and the saliency scores of each object in the original image to generate a candidate image representation.

[0016] In an alternative approach, the further steps of analyzing the representative images in the canvas page through the image encoder model and generating image feature vectors include:

[0017] Calculating the semantic correlation degree between each image in the canvas page and the design theme through a lightweight attention network, and screening the top 30% of the images with the highest semantic correlation degree as the representative image set through a dynamic threshold; among them, applying a weight coefficient of 1.3 to the images operated by the user in the last 5 minutes;

[0018] Among them, the lightweight attention network includes a multi-scale feature extraction layer, a semantic enhancement layer, a bilinear outer product network layer, a dynamic adaptation layer, and a feature storage optimization layer;

[0019] The multi-scale feature extraction layer adopts the EfficientNet-B4 architecture and is used for performing dynamic histogram equalization on the screened representative images and outputting grid local descriptors;

[0020] The semantic enhancement layer adopts a 4-head Transformer network layer to calculate the context dependence between grid local descriptors and outputs a context-aware semantic feature vector;

[0021] The bilinear outer product network layer is used for cross-modal fusion of grid local descriptors and context-aware semantic feature vectors and outputs a normalized image feature vector;

[0022] The dynamic adaptation layer is used to detect the editing operations of the user on the canvas page images and update the Triplet Loss function;

[0023] The feature storage optimization layer is used to establish a hierarchical inverted index according to the normalized image feature vectors.

[0024] In an alternative approach, the generation of the design context feature vector based on the context information in the current design canvas page further includes:

[0025] Convert the original RGB color space of all visible elements in the canvas page to the HSV space, and use the K-means clustering algorithm to identify 3 main color clusters and 2 secondary color clusters. Among them, the feature vector of each color cluster includes the hue mean and standard deviation, saturation weighted value, and spatial distribution entropy;

[0026] Perform text region detection and cropping on the bitmap data of all text elements in the canvas page through the EAST algorithm, calculate the font diversity index of the text region, and select the highest-frequency font vector;

[0027] Detect the element vectors in the layer structure data of the canvas page according to the YOLOv5s algorithm, where the elements include graphics, text, pictures, icons, buttons, backgrounds, and decorations; predict the style feature vector according to the StyleNet model for the element vectors;

[0028] Perform weighted fusion on the feature vectors of each color cluster, the highest-frequency font vector, the element vector, and the style feature vector through a gated attention mechanism to obtain the design context feature vector.

[0029] In an alternative approach, the retrieval of the multimodal feature vector in the image database through the graph neural network model to obtain potentially associated candidate images further includes:

[0030] Pre-construct the graph embedding space of the image database, where the nodes in the graph embedding space store the multimodal feature vectors corresponding to the material images; the edges in the graph embedding space are established according to the similarity between the image feature vectors;

[0031] Calculate the cosine similarity matrix between all nodes in the graph embedding space. For each node, establish bidirectional edges to connect its Top-k similar nodes, and add color similarity and style similarity to each edge;

[0032] Input the multimodal feature vector into the graph neural network model, and project it into the graph embedding space through the message passing layer and the fully connected layer; among them, after multiple rounds of message passing, each node obtains a new feature vector representation, and the feature vector representation fuses the node information and the information of the surrounding nodes;

[0033] Calculate the similarity between the updated feature vector of each node and the multimodal feature vector, and select the material images corresponding to several nodes with the highest similarity as potentially associated candidate images.

[0034] In an alternative manner, the graph neural network model includes an embedding layer, a message passing layer, a fully connected layer, a similarity calculation layer, and a selection layer;

[0035] Among them, the embedding layer is used to initialize the node feature vector representation according to the input multi-modal feature vector;

[0036] The message passing layer is used for message construction, message aggregation, and feature update;

[0037] The fully connected layer is used to map the node feature vector after message passing to the graph embedding space;

[0038] The similarity calculation layer is used to calculate the similarity between the updated feature vector of the node and the multi-modal feature vector to obtain the node in the graph embedding space that is most relevant to the user's needs;

[0039] The selection layer is used to select the top-k nodes with the highest similarity according to the similarity score.

[0040] In an alternative manner, the style transfer of the candidate image by the adversarial network model to make the color, texture, and composition of the candidate image consistent with the style of the user's current design canvas page further includes:

[0041] Construct a generative GAN adversarial network based on the Transformer architecture, where the GAN adversarial network includes multiple generators and multiple discriminators; the generator is used to transfer the style of the candidate image to a state consistent with the style of the canvas page, and the discriminator is used to distinguish the image generated by the generator from the real canvas page image;

[0042] Among them, the generator consists of multiple Transformer Encoder and Transformer Decoder modules. The Transformer Encoder module is used to extract the deep feature representation of the candidate image, and the Transformer Decoder module is used to fuse the extracted feature representation with the style information of the canvas page to generate the image after style transfer; the style information of the canvas page analyzes the color histogram, gray-level co-occurrence matrix, local binary pattern, edge direction histogram, and salient region of the canvas page through a style encoder to generate a style embedding vector.

[0043] In an alternative manner, the generation of a saliency map on the candidate image according to the saliency detection algorithm further includes:

[0044] Select the spectral residual method as the saliency detection algorithm, input the candidate image into the saliency detection algorithm, and calculate the saliency value of each pixel or region in the image;

[0045] Map the saliency value of each pixel or region to a grayscale value to generate a saliency map.

[0046] In an alternative approach, the formula for the Triplet Loss function is:

[0047]

[0048] where N is the number of triplets in the training set; ai is the feature vector of the anchor image; pi is the feature vector of the positive example image; ni is the feature vector of the negative example image; L(ai, pi, ni) is the distance metric between the anchor, positive example, and negative example, and L(ai, pi, ni) = d(ai, pi) - d(ai, ni), where d(ai, pi) and d(ai, ni) are the Euclidean distances between ai and pi, and ai and ni respectively; margin is the minimum margin by which the distance between the anchor and the positive example should be less than the distance between the anchor and the negative example.

[0049] According to another aspect of the present application, there is provided an image search and display device, including:

[0050] A feature vector generation module, configured to generate a text search feature vector according to a search keyword input by a user; analyze a representative image in the canvas page through an image encoder model and generate an image feature vector; generate a design context feature vector according to context information in the current design canvas page; where the context information includes the color distribution, font type, element type, and design style of the canvas page;

[0051] A multimodal fusion module, configured to perform weighted fusion on the text search feature vector, the image feature vector, and the design context feature vector to generate a multimodal feature vector;

[0052] An image retrieval module, configured to retrieve the multimodal feature vector in an image database through a graph neural network model to obtain potentially associated candidate images, where the image database contains material images and their corresponding multimodal feature vectors;

[0053] An image display module, configured to perform style transfer on the candidate images through an adversarial network model, so that the color, texture, and composition of the candidate images are consistent with the style of the user's current design canvas page; display the candidate images and decompose them into multiple editable objects, where the position, size, color, and transparency of each editable object can be adjusted individually.

[0054] The solution provided by the above embodiments of the present invention generates a text search feature vector according to the search keywords input by the user; analyzes the representative images in the canvas page through an image encoder model and generates an image feature vector; generates a design context feature vector according to the context information in the current design canvas page, where the context information includes the color distribution, font type, element type, and design style of the canvas page; performs weighted fusion on the text search feature vector, the image feature vector, and the design context feature vector to generate a multi-modal feature vector; retrieves the multi-modal feature vector in the image database through a graph neural network model to obtain candidate images with potential associations, where the image database contains material images and their corresponding multi-modal feature vectors; performs style transfer on the candidate images through an adversarial network model to make the colors, textures, and compositions of the candidate images consistent with the style of the user's current design canvas page; displays the candidate images and decomposes them into multiple editable objects, where the position, size, color, and transparency of each editable object can be adjusted individually. The present invention further improves the speed of image generation and display, improves the user experience in complex visual content search, and enables users to quickly obtain high-quality images.

[0055] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above description and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. Brief Description of the Drawings

[0056] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0057] Figure 1 Shows a schematic flow chart of the picture search and display method according to an embodiment of the present invention;

[0058] Figure 2 Shows a schematic diagram of the picture search and display according to an embodiment of the present invention;

[0059] Figure 3 Shows a schematic flow of the picture search and display method according to an embodiment of the present invention Figure 2 ;

[0060] Figure 4 Shows a schematic flow chart of the significance analysis according to an embodiment of the present invention;

[0061] Figure 5A schematic functional structure diagram of the picture search and display method and apparatus according to an embodiment of the present invention is shown. Detailed implementation manners

[0062] Hereinafter, exemplary embodiments of the present invention will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present invention can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.

[0063] The picture search and display method and apparatus proposed by the present invention will be described in detail below through specific embodiments.

[0064] Embodiment 1:

[0065] Figure 1 A schematic functional structure diagram of the picture search and display method according to an embodiment of the present invention is shown. Specifically, as Figure 1 shown, the following steps are included:

[0066] Step S101, generating a text search feature vector according to the search keyword input by the user; analyzing the representative image in the canvas page through an image encoder model and generating an image feature vector; generating a design context feature vector according to the context information in the current design canvas page; wherein, the context information includes the color distribution, font type, element type and design style of the canvas page.

[0067] In this embodiment, as Figure 2 shown, the text keyword input by the user, the representative image in the canvas page and the context design situation information (color, font, element, style) of the canvas page are fused, so that the search result is highly relevant to the actual needs and design scenarios of the user, reducing invalid or irrelevant search results. It overcomes the defect that traditional image search only depends on the content of the image itself or the text description input by the user and is difficult to capture the design intention. By the context information of the canvas page, materials that match the current design style and theme more precisely are selected, thereby improving the design efficiency. Searching based on the design context provides users with a richer selection of materials. Even if the user does not have a clear search target, materials can be searched according to the current design style, increasing the user's design inspiration. As Figure 2As shown, for the splash screen of a travel app, the designer hopes to find a "scenic place" as the background image. When the user enters the keyword: "scenic place", the keyword is converted into a 512-dimensional text feature vector. Analyze the content of the current app design draft (such as logo, search box, etc.) and extract a 256-dimensional image feature vector. For example, detect the main color of the app design draft (blue) and extract the color feature vector, detect that the layout is a centered layout and extract the layout feature vector, identify the shape and color of the logo and the position of the search box and extract the element feature vector, and the design style is flat and extract the corresponding style feature vector. As Figure 3 shown, the text feature vector, the image feature vector, and the design context feature vector are weighted and fused (text weight 0.4, image weight 0.3, context weight 0.3) to obtain a comprehensive multi-modal feature vector. Use the multi-modal feature vector to search for images in the image database that match the theme of "scenic place" and the style of the current app design draft. The search results obtained include scenic pictures that conform to the app style in terms of color, composition, etc.

[0068] In an optional manner, the analyzing the representative images in the canvas page by the image encoder model and generating the image feature vector further includes:

[0069] Calculate the semantic correlation degree between each image in the canvas page and the design theme through a lightweight attention network, and screen the images with the top 30% of the semantic correlation degree rankings as the representative image set through a dynamic threshold; among them, apply a weight coefficient of 1.3 to the images operated by the user within the last 5 minutes;

[0070] Among them, the lightweight attention network includes a multi-scale feature extraction layer, a semantic enhancement layer, a bilinear outer product network layer, a dynamic adaptation layer, and a feature storage optimization layer;

[0071] The multi-scale feature extraction layer adopts the EfficientNet-B4 architecture and is used for dynamic histogram equalization of the screened representative images, and outputs grid local descriptors;

[0072] The semantic enhancement layer adopts a 4-head Transformer network layer to calculate the context dependence between grid local descriptors and outputs a context-aware semantic feature vector;

[0073] The bilinear outer product network layer is used for cross-modal fusion of grid local descriptors and context-aware semantic feature vectors, and outputs a normalized image feature vector;

[0074] The dynamic adaptation layer is used to detect the user's editing operation on the canvas page image and update the Triplet Loss function;

[0075] The feature storage optimization layer is used to establish a hierarchical inverted index based on the standardized image feature vectors.

[0076] In this embodiment, by screening the dynamic threshold and calculating the semantic correlation degree, the selected representative images are highly relevant to the design theme, further improving the accuracy of the search results. For example, in the startup page of a travel App, it is desired to find a "scenic place" as the background image. The semantic correlation degree between each image on the canvas page and the theme of "scenic place" is calculated through a lightweight attention network, and a weight coefficient of 1.3 times is applied to the images that the user has operated on in the last 5 minutes. The images with the top 30% semantic correlation degree are selected as the representative image set through the dynamic threshold screening. The EfficientNet-B4 architecture is used to perform dynamic histogram equalization on the selected representative images, and grid local descriptors are extracted. A 4-head Transformer network layer is used to calculate the context dependence between the grid local descriptors, and a context-aware semantic feature vector is output (such as there are mountains and lakes in the image). Through the bilinear outer product network layer, cross-modal fusion of the grid local descriptors and the context-aware semantic feature vectors is performed, and a standardized image feature vector is output. The fused feature vector can contain both the visual features and semantic features of the image. Detect the user's editing operations on the canvas page images and update the TripletLoss loss function. For example, if the user makes adjustments to a certain image, the loss function is updated according to the adjustment results to ensure the continuous optimization of the model. A hierarchical inverted index is established based on the standardized image feature vectors to quickly retrieve images related to the theme of "scenic place" and improve the search efficiency. Therefore, users can find scenic materials that meet the design requirements more quickly and accurately, and can automatically adapt to the style of the App design draft, greatly improving the design efficiency and quality.

[0077] In an optional manner, the formula of the Triplet Loss loss function is:

[0078]

[0079] where N is the number of triplets in the training set; ai is the feature vector of the anchor image; pi is the feature vector of the positive example image; ni is the feature vector of the negative example image; L(ai, pi, ni) is the distance metric between the anchor, positive example, and negative example, and L(ai, pi, ni) = d(ai, pi) - d(ai, ni), where d(ai, pi) and d(ai, ni) are the Euclidean distances between ai and pi, ni respectively; margin is the minimum interval that the distance between the anchor and the positive example should be less than the distance between the anchor and the negative example.

[0080] In this embodiment, the Triplet Loss constructs triplets to force the distance between the anchor image and the positive example image to be less than the distance between the anchor image and the negative example image, enabling the learning of more discriminative feature vectors, thereby improving the accuracy of image search and matching. The margin ensures that the distance between the anchor and the positive example should be less than a preset threshold of the distance between the anchor and the negative example, preventing the model from stopping learning after only distinguishing between positive and negative examples. For example, for the splash screen of a travel app, it is desired to find a "scenic place" as the background image. The user has placed an image of a mountain on the canvas. The mountain image on the canvas is selected as the anchor through the anchor ai, and an image belonging to the same semantic category (mountain) or having a similar style to the anchor image is selected as the positive example through the positive example pi. If the user adjusts the mountain image (adjusts the brightness and contrast), the adjusted mountain image can be used as the positive example. An image belonging to a different semantic category (such as a beach) or having a different style from the anchor image is selected as the negative example ni. The multi-scale feature extraction layer, semantic enhancement layer, and bilinear outer product network layer are used to extract the feature vectors (ai, pi, ni) of the anchor, positive example, and negative example images. The Euclidean distance d(ai, pi) between the anchor and the positive example and the Euclidean distance d(ai, ni) between the anchor and the negative example are calculated. By continuously iterating and training, better feature vectors of mountains and beaches are learned, thereby improving the accuracy of image search and matching.

[0081] In an alternative manner, the generating of the design context feature vector according to the context information in the current design canvas page further includes:

[0082] Convert the original RGB color space of all visible elements in the canvas page to the HSV space, and use the K-means clustering algorithm to identify 3 main color clusters and 2 auxiliary color clusters. Among them, the feature vector of each color cluster includes the hue mean and standard deviation, saturation weighted value, and spatial distribution entropy;

[0083] Perform text region detection and cropping on the bitmap data of all text elements in the canvas page through the EAST algorithm, calculate the font diversity index of the text region, and select the highest-frequency font vector;

[0084] Detect the element vectors in the layer structure data of the canvas page according to the YOLOv5s algorithm, where the elements include graphics, text, pictures, icons, buttons, backgrounds, and decorations; predict the style feature vector according to the StyleNet model for the element vectors;

[0085] Perform weighted fusion on the feature vectors of each color cluster, the highest-frequency font vector, the element vector, and the style feature vector through the gated attention mechanism to obtain the design context feature vector.

[0086] In this embodiment, converting the RGB color space to the HSV space and using the K-means clustering algorithm to identify the main color cluster and the auxiliary color cluster can more accurately capture the color features of the canvas page. Using the EAST algorithm to detect, crop, and calculate the font diversity index of the text area can accurately identify and extract the features of text elements. Detecting graphics, text, pictures, icons, buttons, backgrounds, and decorations in the canvas page through the YOLOv5s algorithm can comprehensively capture the structural information of the canvas page. Using the StyleNet model to predict the style feature vector of elements can better understand and express the design style of the canvas page.

[0087] Step S102, perform weighted fusion on the text search feature vector, the image feature vector, and the design context feature vector to generate a multi-modal feature vector.

[0088] In this embodiment, if the information of a certain modality is missing (such as only having text description without images), the fused feature vector can still use the information of other modalities for inference, learn a more general representation, and has better generalization ability when facing new and unseen data.

[0089] Step S103, retrieve the multi-modal feature vector in the image database through a graph neural network model to obtain potentially associated candidate images, where the image database contains material images and their corresponding multi-modal feature vectors.

[0090] Traditional image retrieval methods usually only focus on the features of a single image and ignore the relationships between images. In this embodiment, the graph neural network model (GNN) can represent the image database as a graph, where nodes represent images and edges represent the relationships between images (such as similarity, co-occurrence relationship), and can capture the implicit associations between images, thereby improving the accuracy of retrieval. In addition, through techniques such as graph sampling and graph partitioning, GNN can process large-scale image databases (databases containing millions or even billions of images).

[0091] In an alternative manner, the retrieving the multi-modal feature vector in the image database through the graph neural network model to obtain potentially associated candidate images further includes:

[0092] Pre-construct the graph embedding space of the image database, where the nodes in the graph embedding space store the multi-modal feature vectors corresponding to the material images; the edges in the graph embedding space are established according to the similarity between the image feature vectors;

[0093] Calculate the cosine similarity matrix between all nodes in the graph embedding space, where a two-way edge is established for each node to connect its Top-k similar nodes, and color similarity and style similarity are added to each edge;

[0094] Input the multi-modal feature vector into the graph neural network model, and project it into the graph embedding space through the message passing layer and the fully connected layer; wherein, after multiple rounds of message passing, each node obtains a new feature vector representation, and the feature vector representation fuses the node information and the information of surrounding nodes;

[0095] Calculate the similarity between the updated feature vector of each node and the multi-modal feature vector, and select the material images corresponding to several nodes with the highest similarity as candidate images for potential association.

[0096] In this embodiment, in addition to the image feature vector similarity, color similarity and style similarity are also considered to make the retrieval result more in line with the user's visual preference. Through the message passing of the GNN, the features of the node itself and the surrounding nodes can be fused, so as to generate a more expressive node representation. In addition, the construction process of the graph embedding space and the similarity matrix is relatively transparent, so that it can explain how the model performs image retrieval. For example, store the multi-modal feature vectors of each landscape image in the nodes of the graph embedding space, calculate the cosine similarity matrix between the nodes, establish two-way edges of the top-k similar nodes and add color and style similarity. Input the target feature vector (such as the feature vector of "beautiful scenery place") into the graph neural network model, project the feature vector into the graph embedding space through the message passing layer and the fully connected layer for multiple rounds of message passing, and calculate the similarity between the updated feature vector of each node and the target feature vector. Select the landscape images corresponding to several nodes with the highest similarity as candidate background images, and the candidate images that most conform to the "beautiful scenery" feature can be efficiently found from a large number of landscape images for the startup page background of the travel App.

[0097] In an optional manner, the graph neural network model includes an embedding layer, a message passing layer, a fully connected layer, a similarity calculation layer and a selection layer;

[0098] Among them, the embedding layer is used to initialize the node feature vector representation according to the input multi-modal feature vector;

[0099] The message passing layer is used for message construction, message aggregation and feature update;

[0100] The fully connected layer is used to map the node feature vector after message passing to the graph embedding space;

[0101] The similarity calculation layer is used to calculate the similarity between the updated feature vector of the node and the multi-modal feature vector to obtain the nodes in the graph embedding space that are most relevant to the user's needs;

[0102] The selection layer is used to select the top-k nodes with the highest similarity according to the similarity score.

[0103] For example, the multi-modal feature vectors of a landscape image (color histogram (256 dimensions), SIFT features (128 dimensions), semantic label embedding (300 dimensions)) are input through an embedding layer, with a total of 684 dimensions. The 684-dimensional input is mapped to a 128-dimensional embedding space through a linear layer, and then input into a message passing layer (two layers). Each node i concatenates its own features with the features of adjacent node j, and then maps them to a new space through a linear layer. Average pooling is performed on the messages sent by all adjacent nodes, and a GRU cell is used to fuse the aggregated messages with the current node features. The fully connected layer contains two layers. The first layer maps 128 dimensions to 64 dimensions, and the second layer maps 64 dimensions to 32 dimensions. The 32-dimensional node feature vector in the graph embedding space is input into a similarity calculation layer to output the similarity score of each node. The similarity score of each node is input into a selection layer to output the indices of the top 5 nodes.

[0104] In step S104, style transfer is performed on the candidate image through an adversarial network model to make the color, texture, and composition of the candidate image consistent with the style of the user's current design canvas page; the candidate image is displayed and decomposed into multiple editable objects, where the position, size, color, and transparency of each editable object can be adjusted individually.

[0105] In this embodiment, style transfer is performed through an adversarial network model to make the color, texture, and composition of the candidate image consistent with the style of the user's current design canvas page. At the same time, the adversarial network model can perform efficient style transfer without paired data and has strong adaptability. The candidate image is decomposed into multiple editable objects, and the user can flexibly adjust the position, size, color, and transparency of each object. For example, the user hopes to add a "scenic place" as a background image to the design canvas page and hopes that the color, texture, and composition of the background image are consistent with the style of the current canvas page. The candidate image ("scenic place") is input into the trained CycleGAN model to generate an image with consistent style. The YOLO algorithm is used to perform object detection on the candidate image after style transfer to identify objects such as mountains, trees, and sky in the image. The Mask R-CNN algorithm is used to accurately segment the detected objects to obtain the mask of each object. Each object is extracted according to the mask and stored as an independent editable object. The candidate image after style transfer is displayed on the design canvas page and decomposed into multiple editable objects such as mountains, trees, and sky, and then the position, size, color, and transparency of each object can be edited and adjusted. For example, drag the position of the mountain, scale the size of the tree, use a color picker to adjust the color of the sky, and adjust the transparency of the object.

[0106] In an alternative approach, the style transfer of the candidate image by the adversarial network model to make the color, texture, and composition of the candidate image consistent with the style of the user's current design canvas page further includes:

[0107] Construct a generative GAN adversarial network based on the Transformer architecture, where the GAN adversarial network includes multiple generators and multiple discriminators; the generator is used to transfer the style of the candidate image to a state consistent with the style of the canvas page, and the discriminator is used to distinguish between the images generated by the generator and the real canvas page images;

[0108] Among them, the generator is composed of multiple Transformer Encoder and Transformer Decoder modules. The Transformer Encoder module is used to extract the deep feature representation of the candidate image, and the Transformer Decoder module is used to fuse the extracted feature representation with the style information of the canvas page to generate the style-transferred image; the style information of the canvas page analyzes the color histogram, gray-level co-occurrence matrix, local binary pattern, edge direction histogram, and salient region of the canvas page through a style encoder to generate a style embedding vector.

[0109] In this embodiment, using a style encoder to analyze the color histogram, gray-level co-occurrence matrix, local binary pattern, edge direction histogram, and salient region of the canvas page can generate an accurate style embedding vector to ensure the accuracy of style transfer.

[0110] In an alternative approach, the method further includes:

[0111] Generate a saliency map on the candidate image according to a saliency detection algorithm, where the saliency map indicates the region that is visually most attention-grabbing in the image;

[0112] Based on the saliency map, use an image segmentation algorithm to segment the candidate image into multiple objects, where the regions with high saliency are preferentially retained during the segmentation process;

[0113] Extract the feature vectors of each object and match them with the keywords input by the user;

[0114] Perform weighted fusion according to the matching results of each object with the keywords and the saliency scores of each object in the original image to generate a candidate image representation.

[0115] In this embodiment, as Figure 4As shown, through saliency detection and image segmentation, key regions and objects in an image can be identified more accurately. For example, for a travel App that selects a background and the candidate image is a landscape picture of the Alps, and the keyword entered by the user is "snow mountain". Through the saliency detection algorithm, the snow-capped peaks and blue sky of the Alps are identified as having high saliency. The image segmentation algorithm segments the image into objects such as snow-capped peaks, blue sky, vegetation, and rocks. During the segmentation process, the snow-capped peaks and blue sky are more precisely segmented due to their high saliency. Feature vectors are extracted for each object. For example, the feature vector of the snow-capped peak describes its white peaks and snow texture, and the feature vector of the blue sky describes its blue color and cloud shape. The feature vector of the snow-capped peak has the highest matching degree with the keyword "snow mountain", the matching degree of the blue sky is the second, and the matching degrees of vegetation and rocks are lower. Suppose the saliency score of the snow-capped peak is 0.9 and its matching degree with the keyword "snow mountain" is 0.8. The saliency score of the blue sky is 0.7 and its matching degree with the keyword "snow mountain" is 0.5. The saliency scores of vegetation and rocks are very low, and the matching degrees are also very low. The weighted weight of the snow-capped peak = 0.9 × 0.8 = 0.72, and the weighted weight of the blue sky = 0.7 × 0.5 = 0.35. Finally, the weighted and fused image representation highlights the features of the snow-capped peak more prominently, making it the main element of the App background. When the user opens the App, they can immediately see the magnificent snow scene of the Alps, thereby ensuring that the background of the App always displays the most eye-catching scenic elements and maintains a high correlation with the keywords entered by the user.

[0116] In an alternative approach, the generating a saliency map on the candidate image according to the saliency detection algorithm further includes:

[0117] Select the spectral residual method as the saliency detection algorithm, input the candidate image into the saliency detection algorithm, and calculate the saliency value of each pixel or region in the image;

[0118] Map the saliency value of each pixel or region to a gray value to generate a saliency map.

[0119] In this embodiment, the spectral residual method is an unsupervised algorithm that does not require a large amount of labeled data for training. The spectral residual method has good robustness to factors such as illumination changes and image noise, can stably extract significant regions in the image, and focuses on highlighting the global contrast in the image, and can effectively identify regions with large differences from the surrounding environment.

[0120] Specifically, perform a two-dimensional discrete Fourier transform on the grayscale image to obtain the spectrum of the image. Decompose the spectrum into an amplitude spectrum and a phase spectrum. Take the logarithm of the amplitude spectrum to obtain the logarithmic amplitude spectrum, and perform mean filtering on the logarithmic amplitude spectrum to obtain a smoothed logarithmic amplitude spectrum. Perform an inverse Fourier transform on the spectral residual to obtain a saliency map, and take the absolute value squared of the saliency map to obtain the final saliency map.

[0121] According to the solution provided in the above embodiments of the present invention, a text search feature vector is generated based on the search keywords input by the user; the representative images in the canvas page are analyzed by an image encoder model to generate image feature vectors; a design context feature vector is generated according to the context information in the current design canvas page; wherein, the context information includes the color distribution, font type, element type, and design style of the canvas page; the text search feature vector, the image feature vector, and the design context feature vector are weighted and fused to generate a multi-modal feature vector; the multi-modal feature vector is retrieved in the image database through a graph neural network model to obtain potentially associated candidate images, wherein the image database contains material images and their corresponding multi-modal feature vectors; the candidate images are subjected to style transfer through an adversarial network model to make the color, texture, and composition of the candidate images consistent with the style of the user's current design canvas page; the candidate images are displayed and decomposed into multiple editable objects, and the position, size, color, and transparency of each editable object can be adjusted individually. The present invention further improves the speed of image generation and display, improves the user experience in complex visual content search, and enables the user to quickly obtain high-quality images.

[0122] Embodiment 2:

[0123] Figure 5 The functional structure diagram of the image search and display method device according to the embodiments of the present invention is shown. As Figure 5 shown, the device includes:

[0124] A feature vector generation module 501, configured to generate a text search feature vector according to the search keywords input by the user; analyze the representative images in the canvas page through an image encoder model to generate image feature vectors; generate a design context feature vector according to the context information in the current design canvas page; wherein, the context information includes the color distribution, font type, element type, and design style of the canvas page;

[0125] A multi-modal fusion module 502, configured to perform weighted fusion on the text search feature vector, the image feature vector, and the design context feature vector to generate a multi-modal feature vector;

[0126] An image retrieval module 503, configured to retrieve the multi-modal feature vector in the image database through a graph neural network model to obtain potentially associated candidate images, wherein the image database contains material images and their corresponding multi-modal feature vectors;

[0127] An image display module 504, configured to perform style transfer on the candidate image through an adversarial network model, so that the color, texture, and composition of the candidate image are consistent with the style of the user's current design canvas page; display the candidate image and decompose it into multiple editable objects, where the position, size, color, and transparency of each editable object can be adjusted individually.

[0128] In an alternative approach, it further includes a saliency analysis module, which is further configured to:

[0129] Generate a saliency map on the candidate image according to a saliency detection algorithm, where the saliency map indicates the region that is visually most attention-grabbing in the image;

[0130] Based on the saliency map, use an image segmentation algorithm to segment the candidate image into multiple objects, where regions with high saliency are preferentially retained during the segmentation process;

[0131] Extract the feature vectors of each object and match them with the keywords input by the user;

[0132] Perform weighted fusion according to the matching results of each object with the keywords and the saliency scores of each object in the original image to generate a candidate image representation.

[0133] In an alternative approach, the feature vector generation module 501 is further configured to:

[0134] Calculate the semantic correlation degree between each image in the canvas page and the design theme through a lightweight attention network, and screen out the top 30% of the images with the highest semantic correlation degree as a representative image set through a dynamic threshold; where a weight coefficient of 1.3 is applied to the images operated by the user within the last 5 minutes;

[0135] Among them, the lightweight attention network includes a multi-scale feature extraction layer, a semantic enhancement layer, a bilinear outer product network layer, a dynamic adaptation layer, and a feature storage optimization layer;

[0136] The multi-scale feature extraction layer adopts the EfficientNet-B4 architecture and is used to perform dynamic histogram equalization on the screened representative images and output grid local descriptors;

[0137] The semantic enhancement layer adopts a 4-head Transformer network layer to calculate the context dependence between grid local descriptors and outputs context-aware semantic feature vectors;

[0138] The bilinear outer product network layer is used to perform cross-modal fusion on grid local descriptors and context-aware semantic feature vectors and outputs normalized image feature vectors;

[0139] The dynamic adaptation layer is used to detect the user's editing operations on the canvas page image and update the Triplet Loss function;

[0140] The feature storage optimization layer is used to establish a hierarchical inverted index based on the standardized image feature vectors.

[0141] In an alternative approach, the feature vector generation module 501 is further configured to:

[0142] Convert the original RGB color space of all visible elements in the canvas page to the HSV space, and use the K-means clustering algorithm to identify 3 main color clusters and 2 secondary color clusters. Among them, the feature vector of each color cluster includes the hue mean and standard deviation, saturation weighted value, and spatial distribution entropy;

[0143] Detect and crop the bitmap data of all text elements in the canvas page through the EAST algorithm, calculate the font diversity index of the text area, and select the highest-frequency font vector;

[0144] Detect the element vectors in the layer structure data of the canvas page according to the YOLOv5s algorithm, where the elements include graphics, text, pictures, icons, buttons, backgrounds, and decorations; predict the style feature vectors according to the StyleNet model for the element vectors;

[0145] Perform weighted fusion on the feature vectors of each color cluster, the highest-frequency font vector, the element vector, and the style feature vector through a gated attention mechanism to obtain a design context feature vector.

[0146] In an alternative approach, the image retrieval module 503 is further configured to:

[0147] Pre-construct the graph embedding space of the image database, where the nodes of the graph embedding space store the multi-modal feature vectors corresponding to the material images; the edges of the graph embedding space are established according to the similarity between the image feature vectors;

[0148] Calculate the cosine similarity matrix between all nodes in the graph embedding space, where a two-way edge is established for each node to connect its Top-k similar nodes, and color similarity and style similarity are added to each edge;

[0149] Input the multi-modal feature vectors into the graph neural network model, and project them into the graph embedding space through the message passing layer and the fully connected layer; among them, after multiple rounds of message passing, each node obtains a new feature vector representation, and the feature vector representation fuses the node information and the information of the surrounding nodes;

[0150] Calculate the similarity between the updated feature vector of each node and the multi-modal feature vector, and select the material images corresponding to several nodes with the highest similarity as candidate images for potential association.

[0151] In an alternative approach, the graph neural network model includes an embedding layer, a message passing layer, a fully connected layer, a similarity calculation layer, and a selection layer;

[0152] Among them, the embedding layer is used to initialize the node feature vector representation according to the input multi-modal feature vector;

[0153] The message passing layer is used for message construction, message aggregation, and feature update;

[0154] The fully connected layer is used to map the node feature vector after message passing to the graph embedding space;

[0155] The similarity calculation layer is used to calculate the similarity between the updated feature vector of the node and the multi-modal feature vector to obtain the nodes most relevant to the user's needs in the graph embedding space;

[0156] The selection layer is used to select the top-k nodes with the highest similarity according to the similarity scores.

[0157] In an alternative approach, the image display module 504 is further configured to:

[0158] Construct a generative GAN adversarial network based on the Transformer architecture, where the GAN adversarial network includes multiple generators and multiple discriminators; the generator is used to transfer the style of the candidate image to a state consistent with the style of the canvas page, and the discriminator is used to distinguish the image generated by the generator from the real canvas page image;

[0159] Among them, the generator consists of multiple Transformer Encoder and Transformer Decoder modules. The Transformer Encoder module is used to extract the deep feature representation of the candidate image, and the Transformer Decoder module is used to fuse the extracted feature representation with the style information of the canvas page to generate the image after style transfer; the style information of the canvas page analyzes the color histogram, gray-level co-occurrence matrix, local binary pattern, edge direction histogram, and salient region of the canvas page through a style encoder to generate a style embedding vector.

[0160] In an alternative approach, the saliency analysis module is further configured to:

[0161] Select the spectral residual method as the saliency detection algorithm, input the candidate image into the saliency detection algorithm, and calculate the saliency value of each pixel or region in the image;

[0162] Map the saliency value of each pixel or region to a grayscale value to generate a saliency map.

[0163] In an alternative approach, the formula for the Triplet Loss function is:

[0164]

[0165] where N is the number of triplets in the training set; ai is the feature vector of the anchor image; pi is the feature vector of the positive example image; ni is the feature vector of the negative example image; L(ai, pi, ni) is the distance metric between the anchor, positive example, and negative example, and L(ai, pi, ni) = d(ai, pi) - d(ai, ni), where d(ai, pi) and d(ai, ni) are the Euclidean distances between ai and pi, and ai and ni respectively; margin is the minimum margin by which the distance between the anchor and the positive example should be less than the distance between the anchor and the negative example.

[0166] The solution provided in the above embodiments of the present invention generates a text search feature vector according to the search keywords input by the user; analyzes the representative images in the canvas page through an image encoder model and generates image feature vectors; generates a design context feature vector according to the context information in the current design canvas page, where the context information includes the color distribution, font type, element type, and design style of the canvas page; performs weighted fusion on the text search feature vector, image feature vector, and design context feature vector to generate a multi-modal feature vector; retrieves the multi-modal feature vector in the image database through a graph neural network model to obtain potentially related candidate images, where the image database contains material images and their corresponding multi-modal feature vectors; performs style transfer on the candidate images through an adversarial network model to make the color, texture, and composition of the candidate images consistent with the style of the user's current design canvas page; displays the candidate images and decomposes them into multiple editable objects, where the position, size, color, and transparency of each editable object can be adjusted individually. The present invention further improves the speed of image generation and display, improves the user experience in complex visual content search, and enables the user to quickly obtain high-quality images.

[0167] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. A variety of general-purpose systems can also be used in conjunction with the teachings provided herein. The structure required to construct such systems will be apparent from the above description. Additionally, embodiments of the present invention are not directed to any particular programming language. It should be understood that the teachings of the present invention described herein can be implemented using a variety of programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present invention.

[0168] In the specification provided herein, numerous specific details are set forth. However, it can be understood that embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0169] Similarly, it should be understood that, in order to streamline the present invention and assist in understanding one or more of the various inventive aspects, in the above description of the exemplary embodiments of the present invention, the various features of the embodiments of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspects lie in less than all of the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate embodiment of the present invention.

[0170] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except for the fact that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature that provides the same, equivalent, or similar purpose.

[0171] In addition, those skilled in the art can understand that although some embodiments herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.

[0172] Each component embodiment of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or in any other form.

[0173] It should be noted that the above embodiments illustrate the present invention rather than limit the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. A method for image search and display, characterized in that, Including: Generating a text search feature vector according to the search keywords input by the user; Analyzing the representative images in the canvas page through an image encoder model and generating an image feature vector; generating a design context feature vector according to the context information in the current design canvas page; wherein, the context information includes the color distribution, font type, element type and design style of the canvas page; Performing weighted fusion on the text search feature vector, the image feature vector and the design context feature vector to generate a multi-modal feature vector; Retrieving the multi-modal feature vector in the image database through a graph neural network model to obtain candidate images with potential associations, wherein the image database contains material images and their corresponding multi-modal feature vectors; Performing style transfer on the candidate images through an adversarial network model to make the colors, textures and compositions of the candidate images consistent with the style of the user's current design canvas page; displaying the candidate images and decomposing them into multiple editable objects, wherein the position, size, color and transparency of each editable object can be adjusted separately.

2. The picture search and display method according to claim 1, wherein The method further includes: Generating a saliency map on the candidate images according to a saliency detection algorithm, wherein the saliency map indicates the region in the image that is visually most attention-grabbing; Based on the saliency map, using an image segmentation algorithm to segment the candidate images into multiple objects, wherein regions with high saliency are preferentially retained during the segmentation process; Extracting the feature vectors of each object and matching them with the keywords input by the user; Performing weighted fusion according to the matching results of each object and the keywords and the saliency scores of each object in the original image to generate a candidate image representation.

3. The method for picture search and display according to claim 1, characterized in that The further step of analyzing the representative images in the canvas page through an image encoder model and generating an image feature vector includes: Calculating the semantic association degree between each image in the canvas page and the design theme through a lightweight attention network, and screening the images with the top 30% of the semantic association degrees as the representative image set through a dynamic threshold; wherein, a weight coefficient of 1.3 is applied to the images operated by the user within the last 5 minutes; Wherein, the lightweight attention network includes a multi-scale feature extraction layer, a semantic enhancement layer, a bilinear outer product network layer, a dynamic adaptation layer and a feature storage optimization layer; The multi-scale feature extraction layer adopts the EfficientNet-B4 architecture to perform dynamic histogram equalization on the screened representative images and output grid local descriptors; The semantic enhancement layer adopts a 4-head Transformer network layer to calculate the context dependence between grid local descriptors and output a context-aware semantic feature vector; The bilinear outer product network layer is used for cross-modal fusion of grid local descriptors and context-aware semantic feature vectors and outputs a normalized image feature vector; The dynamic adaptation layer is used to detect the user's editing operations on the canvas page images and update the Triplet Loss function; The feature storage optimization layer is used to establish a hierarchical inverted index according to the normalized image feature vectors.

4. The method for picture search and display according to claim 1, wherein The further step of generating a design context feature vector according to the context information in the current design canvas page includes: Convert the original RGB color space of all visible elements in the canvas page to the HSV space, and use the K-means clustering algorithm to identify 3 main color clusters and 2 secondary color clusters. Among them, the feature vector of each color cluster includes the hue mean and standard deviation, saturation weighted value, and spatial distribution entropy; Detect and crop the bitmap data of all text elements in the canvas page through the EAST algorithm, calculate the font diversity index of the text area, and select the highest-frequency font vector; Detect the element vectors in the layer structure data of the canvas page according to the YOLOv5s algorithm, where the elements include graphics, text, pictures, icons, buttons, backgrounds, and decorations; predict the style feature vector according to the StyleNet model for the element vectors; Perform weighted fusion on the feature vector of each color cluster, the highest-frequency font vector, the element vector, and the style feature vector through a gated attention mechanism to obtain a design context feature vector.

5. The picture search and display method according to claim 1, characterized in that The step of retrieving the multimodal feature vector in the image database through the graph neural network model to obtain potentially associated candidate images further includes: Pre-construct the graph embedding space of the image database, where the nodes of the graph embedding space store the multimodal feature vectors corresponding to the material images; the edges of the graph embedding space are established according to the similarity between the image feature vectors; Calculate the cosine similarity matrix between all nodes in the graph embedding space, where a two-way edge is established for each node to connect its Top-k similar nodes, and color similarity and style similarity are added to each edge; Input the multimodal feature vector into the graph neural network model, and project it into the graph embedding space through the message passing layer and the fully connected layer; among them, after multiple rounds of message passing, each node obtains a new feature vector representation, and the feature vector representation fuses the node information and the information of surrounding nodes; Calculate the similarity between the updated feature vector of each node and the multimodal feature vector, and select the material images corresponding to several nodes with the highest similarity as potentially associated candidate images.

6. The method for picture search and display according to claim 1 or 5, characterized in that, The graph neural network model includes an embedding layer, a message passing layer, a fully connected layer, a similarity calculation layer, and a selection layer; Among them, the embedding layer is used to initialize the node feature vector representation according to the input multimodal feature vector; The message passing layer is used for message construction, message aggregation, and feature update; The fully connected layer is used to map the node feature vector after message passing to the graph embedding space; The similarity calculation layer is used to calculate the similarity between the updated feature vector of the node and the multimodal feature vector to obtain the node in the graph embedding space that is most relevant to the user's needs; The selection layer is used to select the Top-k nodes with the highest similarity according to the similarity score.

7. The method for picture search and display according to claim 1, characterized in that, The step of performing style transfer on the candidate images through the adversarial network model to make the color, texture, and composition of the candidate images consistent with the style of the user's current design canvas page further includes: Construct a generative GAN adversarial network based on the Transformer architecture, where the GAN adversarial network includes multiple generators and multiple discriminators; the generator is used to transfer the style of the candidate image to a state consistent with the style of the canvas page, and the discriminator is used to distinguish between the image generated by the generator and the real canvas page image; Among them, the generator consists of multiple Transformer Encoder and Transformer Decoder modules. The Transformer Encoder module is used to extract the deep feature representation of the candidate image, and the Transformer Decoder module is used to fuse the extracted feature representation with the style information of the canvas page to generate the image after style transfer; the style information of the canvas page analyzes the color histogram, gray-level co-occurrence matrix, local binary pattern, edge direction histogram, and salient region of the canvas page through the style encoder to generate a style embedding vector.

8. The method for picture search and display according to claim 2, characterized in that The generation of the saliency map on the candidate image according to the saliency detection algorithm further includes: Select the spectral residual method as the saliency detection algorithm, input the candidate image into the saliency detection algorithm, and calculate the saliency value of each pixel or region in the image; Map the saliency value of each pixel or region to a gray value to generate a saliency map.

9. The picture search and display method according to claim 3, wherein The formula of the Triplet Loss function is: Among them, N is the number of triplets in the training set; ai is the feature vector of the anchor image; pi is the feature vector of the positive example image; ni is the feature vector of the negative example image; L(ai, pi, ni) is the distance metric between the anchor, positive example, and negative example, and L(ai, pi, ni)=d(ai, pi)-d(ai, ni), where d(ai, pi) and d(ai, ni) are the Euclidean distances between ai and pi, ni respectively; margin is the minimum interval that the distance between the anchor and the positive example should be less than the distance between the anchor and the negative example.

10. An image search and display device, characterized in that, It includes: A feature vector generation module for generating a text search feature vector according to the search keywords input by the user; Analyze the representative images in the canvas page through the image encoder model and generate image feature vectors; generate design context feature vectors according to the context information in the current design canvas page; where the context information includes the color distribution, font type, element type, and design style of the canvas page.