Method, device and storage medium for searching art design images based on virtual reality

By constructing an image encoder to process distortion and parallax in VR image data, and extracting and fusing artistic design features, the problem of low accuracy in VR image data search is solved, resulting in high-quality search results and improved user experience.

CN120994865BActive Publication Date: 2026-02-10GUANGZHOU COLLEGE OF COMMERCE +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511517517.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-23
Publication Date
2026-02-10
Estimated Expiration
2045-10-23

AI Technical Summary

Technical Problem

In virtual reality scenarios, VR image data suffers from distortion and parallax, resulting in low accuracy of existing search methods. Directly using two-dimensional images for searching leads to biases.

Method used

An image encoder is constructed, including a first sub-encoder, a second sub-encoder, a third sub-encoder, and a fusion unit. Artistic design features are extracted through inverse distortion and parallax processing and then fused to generate high-quality fingerprint data for searching.

Benefits of technology

It improves the accuracy of VR image data search, integrates fragmented resources, expands the search space, and provides more search results in rare cases, thereby enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994865B_ABST
    Figure CN120994865B_ABST
Patent Text Reader

Abstract

The application provides a method and device for searching artistic design images based on virtual reality, and a storage medium. The method comprises the following steps: constructing first template image data and second template image data under virtual reality according to a search request; performing a difference operation on the first template image data and the second template image data to obtain third template image data; determining an image encoder; inputting the first template image data into a first sub-encoder to extract first artistic design features; inputting the second template image data into a second sub-encoder to extract second artistic design features; inputting the third template image data into a third sub-encoder to extract third artistic design features; inputting the first artistic design features, the second artistic design features and the third artistic design features into a fusion device to fuse the first artistic design features, the second artistic design features and the third artistic design features into fourth artistic design features; and searching target image data in artistic design according to the fourth artistic design features. The embodiment effectively improves the accuracy of VR search.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of computer vision, and in particular relates to a method, device and storage medium for searching art design images based on virtual reality. Background Technology

[0002] As people's living standards improve, their demands for entertainment experiences are increasing, and they crave more immersive and interactive entertainment. VR (Virtual Reality) technology provides an immersive experience that meets users' needs, allowing them to enter virtual game, movie, and other scenarios as if they were actually there.

[0003] In VR scenarios, some applications (such as search engines, LLM (Large Language Model), resource sharing websites, etc.) provide search functions, allowing users to search for VR image data.

[0004] Currently, searching for VR image data in VR scenes mainly involves using a deep learning-based encoder to encode the corresponding two-dimensional image data of the VR image data, obtaining fingerprint data of the two-dimensional image data, calculating the cosine similarity between the two-dimensional images corresponding to different VR image data based on the fingerprint data, and then using the cosine similarity to search for similar VR image data.

[0005] However, VR image data often contains distortion and parallax, presenting a three-dimensional effect to users. Directly using a two-dimensional perspective for searching will produce certain deviations, resulting in lower accuracy. Summary of the Invention

[0006] In view of this, embodiments of this application provide a method, device, and storage medium for searching art and design images based on virtual reality, in order to improve the accuracy of retrieving VR image data.

[0007] The first aspect of this application provides a method for searching art and design images based on virtual reality, including:

[0008] When a virtual reality-based image search request is received, first template image data and second template image data under virtual reality are constructed according to the search request.

[0009] Perform a difference operation on the first template image data and the second template image data to obtain the third template image data;

[0010] Determine the image encoder; the image encoder includes a first sub-encoder, a second sub-encoder, a third sub-encoder, and a fusion unit;

[0011] The first template image data is input into the first sub-encoder to extract the first artistic design feature from the first template image data for the purpose of inverse distortion;

[0012] The second template image data is input into the second sub-encoder to extract the second artistic design features from the second template image data for the purpose of inverse distortion;

[0013] The third template image data is input into the third sub-encoder to extract third art design features from the third template image data for classification purposes;

[0014] The first art design feature, the second art design feature, and the third art design feature are input into the fusion unit and fused into a fourth art design feature for the purpose of classification.

[0015] Based on the fourth art design feature, search for target image data that is similar to the first template image data and the second template image data in the art design.

[0016] A second aspect of this application provides an apparatus for searching art and design images based on virtual reality, comprising:

[0017] The template image data construction module is used to construct first template image data and second template image data under virtual reality based on the search request when a virtual reality-based image search request is received.

[0018] The template difference module is used to perform a difference operation on the first template image data and the second template image data to obtain the third template image data;

[0019] An image encoder determination module is used to determine an image encoder; the image encoder includes a first sub-encoder, a second sub-encoder, a third sub-encoder, and a fusion unit.

[0020] The first art design feature extraction module is used to input the first template image data into the first sub-encoder and extract the first art design features from the first template image data for the purpose of inverse distortion.

[0021] The second art design feature extraction module is used to input the second template image data into the second sub-encoder to extract the second art design features from the second template image data for the purpose of inverse distortion.

[0022] The third art design feature extraction module is used to input the third template image data into the third sub-encoder and extract the third art design features from the third template image data for classification purposes.

[0023] The fourth art design feature fusion module is used to input the first art design feature, the second art design feature and the third art design feature into the fusion unit and fuse them into a fourth art design feature for the purpose of classification.

[0024] The target image data search module is used to search for target image data that is similar to the first template image data and the second template image data in terms of art design based on the fourth art design feature.

[0025] A third aspect of this application provides a terminal device including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for searching art design images based on virtual reality as described in the first aspect above.

[0026] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method for searching art design images based on virtual reality as described in the first aspect above.

[0027] A fifth aspect of this application provides a computer program product that, when run on a computer, causes the computer to execute the method for searching art design images based on virtual reality as described in the first aspect.

[0028] In this embodiment, when a virtual reality-based image search request is received, first template image data and second template image data under virtual reality are constructed according to the search request; a difference operation is performed on the first template image data and the second template image data to obtain third template image data; an image encoder is determined; the image encoder includes a first sub-encoder, a second sub-encoder, a third sub-encoder, and a fusion unit; the first template image data is input into the first sub-encoder to extract a first artistic design feature from the first template image data for the purpose of inverse distortion; the second template image data is input into the second sub-encoder to extract a second artistic design feature from the second template image data for the purpose of inverse distortion; the third template image data is input into the third sub-encoder to extract a third artistic design feature from the third template image data for the purpose of classification; the first artistic design feature, the second artistic design feature, and the third artistic design feature are input into the fusion unit to fuse them into a fourth artistic design feature for the purpose of classification; and target image data similar to the first template image data and the second template image data are searched for in terms of artistic design based on the fourth artistic design feature. This embodiment reverses the distortion of VR image data and considers the parallax within the VR image data, thereby constructing high-quality fingerprint data for VR image data, enabling VR image data search, and ensuring that the representation of VR image data and fingerprint data are consistent, effectively improving the accuracy of VR search.

[0029] Furthermore, this embodiment provides a unified-scale encoding for VR image data of various formats during the search process, effectively integrating fragmented VR image data resources, expanding the search space, and further improving the accuracy of the search.

[0030] Furthermore, this embodiment provides a generalized search mode based on art design, offering users more search results and potential possibilities when VR image data is relatively scarce, thereby improving the user's VR search experience. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0032] Figure 1 This is a schematic diagram illustrating a method for searching art and design images based on virtual reality, as provided in an embodiment of this application.

[0033] Figure 2 This is a schematic diagram of an image encoder provided in an embodiment of this application;

[0034] Figure 3 This is a schematic diagram of a first generator-adversary and a second generator-adversary provided in an embodiment of this application;

[0035] Figure 4 This is a schematic diagram of an image classifier provided in an embodiment of this application;

[0036] Figure 5 This is a schematic diagram of a device for searching art and design images based on virtual reality, provided in an embodiment of this application;

[0037] Figure 6 This is a schematic diagram of a terminal device provided in an embodiment of this application. Detailed Implementation

[0038] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0039] The technical solution of this application will be described below through specific embodiments.

[0040] Reference Figure 1 The diagram illustrates a method for searching art and design images based on virtual reality, as provided in an embodiment of this application. Specifically, it may include the following steps:

[0041] Step 101: When a virtual reality-based image search request is received, construct the first template image data and the second template image data under virtual reality according to the search request.

[0042] Users wear VR devices (such as VR headsets, VR glasses, etc.), launch applications within the VR devices, and trigger image search requests using motion sensing or other methods, requesting to search for VR image data from an art and design perspective.

[0043] In practical applications, VR image data comes in many formats, such as polarized vertical arrangement, polarized horizontal arrangement, horizontal 3D arrangement, cube fitting method type A, cube fitting method type B, waist drum distortion, waist drum distortion polarized vertical arrangement, waist drum distortion polarized horizontal arrangement, concave and convex mirrors, concave and convex mirror horizontal 3D arrangement, etc., which makes VR image data resources fragmented.

[0044] In this embodiment, based on the instructions of the search request, the measurement scale of VR image data in various formats can be unified to construct the first template image data and the second template image data representing VR.

[0045] In one scenario, if the search request contains a frame of original image data from a virtual reality environment, the original image data is directly set as the first template image data and the second template image data, respectively.

[0046] In another scenario, if the search request contains two frames of original image data from virtual reality, then, based on the parallax, one frame of original image data is set as the first template image data, and the other frame of original image data is set as the second template image data.

[0047] Generally, image data located at the top, left, etc., can be set as the first template image data, and image data located at the bottom, right, etc., can be set as the second template image data.

[0048] Step 102: Perform a difference operation on the first template image data and the second template image data to obtain the third template image data.

[0049] In this embodiment, a difference operation (Diff) can be performed on the first template image data and the second template image data to obtain the third template image data.

[0050] In one scenario, if the search request contains a frame of original image data from a virtual reality environment, then the third template image data is empty.

[0051] In another scenario, if the search request contains two frames of original image data under virtual reality, and there is a parallax between the two frames of original image data, then the third template image data represents the parallax.

[0052] Step 103: Determine the image encoder.

[0053] In this embodiment, an image encoder can be built and trained based on deep learning. The image encoder is used to encode VR image data from a three-dimensional perspective (such as considering parallax, inverse distortion, etc.) to obtain fingerprint data of VR image data.

[0054] The image encoder includes a first sub-encoder Encoder_1, a second sub-encoder Encoder_2, a third sub-encoder Encoder_3, and a fusion unit Fusion.

[0055] The first sub-encoder, Encoder_1, is responsible for encoding one frame of VR image data.

[0056] The second sub-encoder, Encoder_2, is responsible for encoding another frame of VR image data.

[0057] The third sub-encoder, Encoder_3, is responsible for encoding the disparity map between two frames of VR image data.

[0058] Furthermore, the first sub-encoder Encoder_1, the second sub-encoder Encoder_2, and the third sub-encoder Encoder_3 can reuse some of the structures of third-party pre-trained image processing networks (such as VGG, ResNet, YOLO, etc.) to extract image features, thereby reducing development costs while maintaining feature quality.

[0059] The Fusion fusion unit includes structures such as convolutional layers, pooling layers, and attention mechanisms. It is responsible for fusing the characteristics of the outputs of the first sub-encoder (Encoder_1), the second sub-encoder (Encoder_2), and the third sub-encoder (Encoder_3) to obtain the fingerprint data of the VR image data.

[0060] In one embodiment of this application, step 103 may include the following steps:

[0061] Step 1031: Train the first sub-encoder and the second sub-encoder simultaneously in a generative adversarial mode so that the first sub-encoder and the second sub-encoder have the function of inverse distortion.

[0062] In this embodiment, the training of the image encoder is divided into two rounds: the first sub-encoder Encoder_1, the second sub-encoder Encoder_2, the third sub-encoder Encoder_3, and the fusion generator Fusion. The first round trains the first sub-encoder Encoder_1 and the second sub-encoder Encoder_2, and the second round trains the third sub-encoder Encoder_3 and the fusion generator Fusion.

[0063] In the first round of training, the first sub-encoder Encoder_1 and the second sub-encoder Encoder_2 are trained synchronously using a generative adversarial mode, so that the first sub-encoder Encoder_1 and the second sub-encoder Encoder_2 have the function of inverse distortion.

[0064] Generative adversarial mode refers to the game between generation (inversely distorted image data) and discrimination (whether the image data is distorted), so that the first sub-encoder Encoder_1 and the second sub-encoder Encoder_2 can learn the potential distribution of the samples, thereby converting the image data into image data without distortion while keeping the content of the image data unchanged.

[0065] In practical implementation, it is possible to collect sample image data with distortion and sample image data without distortion under virtual reality (VR).

[0066] Based on the first sub-encoder Encoder_1 and the second sub-encoder Encoder_2, the first generative adversarial network GAN_1 and the second generative adversarial network GAN_2 are determined.

[0067] The first generative adversarial network (GAN_1) includes a first generator (Generator_1) and a discriminator (Discriminator).

[0068] The second generative adversarial network, GAN_2, consists of a second generator (Generator_2) and a discriminator (Discriminator).

[0069] At this point, the first generative adversarial network GAN_1 and the second generative adversarial network GAN_2 share the same discriminator, which facilitates maintaining the balance between the first generator Generator_1 and the second generator Generator_2 during training.

[0070] The first generator, Generator_1, includes the first sub-encoder, Encoder_1, and the first decoder, Decoder_1.

[0071] The second generator, Generator_2, includes a second sub-encoder, Encoder_2, and a second decoder, Decoder_2.

[0072] Generally, the structure of the first decoder Decoder_1 is symmetrical to the structure of the first sub-encoder Encoder_1, and the structure of the second decoder Decoder_2 is symmetrical to the structure of the second sub-encoder Encoder_2.

[0073] The first sub-encoder, Encoder_1, is responsible for decoding the features output by the first sub-encoder, Encoder_1, to obtain new image data (especially image data without distortion). The second decoder, Decoder_2, is responsible for decoding the features output by the second sub-encoder, Encoder_2, to obtain new image data (especially image data without distortion).

[0074] The discriminator can reuse the discriminator in a third-party pre-trained Generative Adversarial Network (GAN), such as StackGAN, to reduce development costs while maintaining feature quality.

[0075] The generative adversarial mode is divided into two iterations:

[0076] In the first iteration, the discriminator is trained.

[0077] In the first iteration, a discriminator is trained based on the sample image data so that it can be used to determine whether there is distortion.

[0078] Furthermore, a batch of real image data x is sampled from undistorted sample image data, and the loss value of the discriminator for the real image data x is calculated. This is typically achieved using binary cross-entropy: D loss (x)=-log(D(x)).

[0079] Sample a batch of distorted image data z, generate image data G(z) using the first generator Generator_1 and / or the second generator Generator_2, and calculate the loss value of the discriminator on the generated image data G(z): D loss (G(z))=-log(1-D(G(z))).

[0080] The total loss value D of the discriminator total_loss =D loss (x)+D loss (G(z)), based on the total loss value D total_loss The parameters of the discriminator are updated using the backpropagation algorithm, enabling the discriminator to better distinguish between real image data and generated image data.

[0081] In the first iteration, the first generator Generator_1 and the second generator Generator_2 are trained.

[0082] On the one hand, with the assistance of the discriminator, the first generator Generator_1 is trained based on the sample image data so that the first generator Generator_1 can be used to recover the distortion of the sample image data.

[0083] On the other hand, with the assistance of the discriminator, a second generator, Generator_2, is trained based on the sample image data so that it can be used to recover the distortion of the sample image data.

[0084] Furthermore, a batch of distorted sample image data z is sampled, and image data G(z) is generated through the first generator Generator_1 and the second generator Generator_2. The loss values ​​G of the first generator Generator_1 and the second generator Generator_2 are then calculated. loss=-log(D(G(z))), the image data generated by the first generator Generator_1 and the second generator Generator_2 can deceive the discriminator, that is, make the discriminator think that the generated image data is real and has no distortion. Based on the loss value, the backpropagation algorithm is used to update the parameters of the first generator Generator_1 and the second generator Generator_2, so that the output of the image data generated by the first generator Generator_1 and the second generator Generator_2 on the discriminator is closer to 1.

[0085] The first and second iterations are repeated multiple times. During the training process, the discriminator competes with the first generator (Generator_1) and the second generator (Generator_2). The first generator (Generator_1) and the second generator (Generator_2) continuously learn to generate more realistic image data without distortion, while the discriminator continuously improves its ability to distinguish between real image data and generated image data.

[0086] Upon completion of training, the first sub-encoder Encoder_1 and the second sub-encoder Encoder_2 are retained, while the first decoder Decoder_1, the second decoder Decoder_2, and the discriminator are deleted.

[0087] Step 1032: If the training of the first sub-encoder and the second sub-encoder is completed, then while keeping the first sub-encoder and the second sub-encoder unchanged, train the third sub-encoder and the fusion machine in a classification mode so that the third sub-encoder and the fusion machine have the ability to classify.

[0088] If the first round of training is completed, that is, the training of the first sub-encoder Encoder_1 and the second sub-encoder Encoder_2 is completed, then the second round of training is started. While keeping the first sub-encoder Encoder_1 and the second sub-encoder Encoder_2 unchanged, the third sub-encoder Encoder_3 and the fusion generator are trained using a classification mode, so that the third sub-encoder Encoder_3 and the fusion generator have the ability to classify in art design.

[0089] In one embodiment of this application, step 1032 may further include the following steps:

[0090] Step 10321: Collect the first label image data and the second label image data under virtual reality.

[0091] In one scenario, if the VR image data contains a frame of virtual reality image data, then the image data is directly set as the first label image data and the second label image data respectively.

[0092] In another scenario, if the VR image data contains two frames of virtual reality image data, then based on the parallax, one frame of image data is set as the first label image data, and the other frame of image data is set as the second label image data.

[0093] For VR image data, multiple generalized categories can be divided in the dimension of art design, such as character design, scene design, storyboard design, etc. In this case, the category jointly labeled by the first label image data and the second label image data is used as a label, which is recorded as the real category. That is, the first label image data and the second label image data jointly label the real category representing the art design.

[0094] Step 10322: Perform a difference operation on the first label image data and the second label image data to obtain the third label image data.

[0095] In this embodiment, a difference operation (Diff) can be performed on the first label image data and the second label image data to obtain the third label image data.

[0096] In one scenario, if the VR image data contains a frame of virtual reality image data, then the third label image data is empty.

[0097] In another case, if there are two frames of virtual reality image data in the VR image data, and there is a parallax between the two frames, then the third labeled image data represents the parallax.

[0098] Step 10323: Determine the image classification network.

[0099] In this embodiment, a head structure can be cascaded on the image encoder to construct an image classification network. That is, the image classification network includes a first sub-encoder Encoder_1, a second sub-encoder Encoder_2, a third sub-encoder Encoder_3, a fusion unit Fusion, and a head structure.

[0100] The head structure includes fully connected layers (FC), activation functions (such as the sigmoid function), and other structures, and is used to perform multi-classification tasks in art and design.

[0101] Step 10324: Input the first label image data into the first sub-encoder to extract the first label image features from the first label image data for the purpose of inverse distortion.

[0102] The first label image data is input into the first sub-encoder Encoder_1 for encoding, and features are extracted from the first label image data for the purpose of inverse distortion, which are denoted as the first label image features.

[0103] Step 10325: Input the second label image data into the second sub-encoder to extract the second label image features representing the art design from the second label image data for the purpose of inverse distortion.

[0104] The second label image data is input into the second sub-encoder Encoder_2 for encoding, and features are extracted from the second label image data for the purpose of inverse distortion, which are denoted as the second label image features.

[0105] Step 10326: Input the third label image data into the third sub-encoder and extract the third label image features representing the art design from the third label image data.

[0106] The third-label image data is input into the third sub-encoder Encoder_3 for encoding. Features representing the art design are extracted from the third-label image data and denoted as third-label image features.

[0107] Step 10327: Input the first label image features, the second label image features and the third label image features into the fusion unit and fuse them into the fourth label image features.

[0108] The first label image features, the second label image features, and the third label image features are input into the Fusion fusion unit. The Fusion fusion unit integrates the first label image features, the second label image features, and the third label image features to enhance the expressive power of the features and fuse them into a fourth label image feature.

[0109] In one design, the Fusion fusion module includes a first depthwise separable convolutional layer DSC_1, a second depthwise separable convolutional layer DSC_2, a first convolutional module ConvBolck_1, a second convolutional module ConvBolck_2, a third convolutional module ConvBolck_3, a fourth convolutional module ConvBolck_4, and a fifth convolutional module ConvBolck_5.

[0110] Among them, the first depthwise separable convolutional layer DSC_1 and the second depthwise separable convolutional layer DSC_2 are both depthwise separable convolutional layers, which can perform a spatial convolution while keeping the channels independent and performing a depthwise convolution operation.

[0111] The first convolutional module ConvBolck_1, the second convolutional module ConvBolck_2, the third convolutional module ConvBolck_3, the fourth convolutional module ConvBolck_4, and the fifth convolutional module ConvBolck_5 each include a convolutional layer, a normalization layer (such as BN (Batch Normalization)), and an activation layer (ReLU (Rectified Linear Unit)). The convolutional layer provides convolution operations, the normalization layer provides normalization operations, and the activation layer provides activation operations.

[0112] In this design, the features of the third label image are input into the first depthwise separable convolutional layer DSC_1 to extract the features of the first label differential image.

[0113] The first label difference image features are input into the second depthwise separable convolutional layer DSC_2 to extract the second label difference image features.

[0114] On the one hand, functions such as Add are used to fuse the features of the first label image with the features of the first label difference image into the features of the first sample image.

[0115] The features of the first sample image are input into the first convolution module ConvBolck_1, and convolution, normalization and activation operations are performed sequentially to obtain the features of the second sample image.

[0116] The second sample image features and the second label difference image features are fused together using functions such as Add to form the third sample image features.

[0117] The features of the third sample image are input into the second convolution module ConvBolck_2, where convolution, normalization, and activation operations are performed sequentially to obtain the features of the fourth sample image.

[0118] On the other hand, functions such as Add are used to fuse the features of the second-labeled image with the features of the first-labeled differential image into the features of the fifth sample image.

[0119] The features of the fifth sample image are input into the third convolution module ConvBolck_3, where convolution, normalization, and activation operations are performed sequentially to obtain the features of the sixth sample image.

[0120] The sixth sample image features and the second label difference image features are fused together using functions such as Add to form the seventh sample image features.

[0121] The features of the seventh sample image are input into the fourth convolution module ConvBolck_4, where convolution, normalization, and activation operations are performed sequentially to obtain the features of the eighth sample image.

[0122] Combining the above two aspects, functions such as Concat are used to concatenate the features of the fourth sample image with the features of the eighth sample image to form the features of the ninth sample image.

[0123] The features of the ninth sample image are input into the fifth convolution module ConvBolck_5, where convolution, normalization, and activation operations are performed sequentially to obtain the features of the fourth label image.

[0124] Step 10328: Input the fourth label image features into the head structure to divide the predicted category in terms of art design.

[0125] The features of the fourth label image are input into the Head structure for processing, and the categories are divided in terms of art and design, which are denoted as the predicted categories.

[0126] Step 10329: During the training of the image classification network based on the loss value between the true category and the predicted category, keep the first sub-encoder and the second sub-encoder unchanged, and update the third sub-encoder, the fusion unit and the head structure.

[0127] In this embodiment, the true category and the predicted category are substituted into the loss function (such as multivariate cross-entropy) to calculate the loss value. Based on the loss value, the backpropagation algorithm is used to update the parameters of the image classification network. During this process, the parameters of the first sub-encoder Encoder_1 and the second sub-encoder Encoder_2 are kept unchanged, while the parameters of the third sub-encoder Encoder_3, the fusion unit, and the head structure are updated.

[0128] Upon completion of training, the first sub-encoder (Encoder_1), the second sub-encoder (Encoder_2), the third sub-encoder (Encoder_3), and the fusion generator (Fusion) are retained, while the head structure is deleted.

[0129] At this point, the image encoders (i.e., the first sub-encoder Encoder_1, the second sub-encoder Encoder_2, the third sub-encoder Encoder_3, and the fusion generator Fusion) can be deployed online.

[0130] Step 104: Input the first template image data into the first sub-encoder to extract the first artistic design feature from the first template image data for the purpose of inverse distortion.

[0131] The first template image data is input into the first sub-encoder Encoder_1 for encoding, and features are extracted from the first template image data for the purpose of inverse distortion, which are denoted as the first art design features.

[0132] Step 105: Input the second template image data into the second sub-encoder to extract the second artistic design features from the second template image data for the purpose of inverse distortion.

[0133] The second template image data is input into the second sub-encoder Encoder_2 for encoding, and features are extracted from the second template image data for the purpose of inverse distortion, which are denoted as the second art design features.

[0134] Step 106: Input the third template image data into the third sub-encoder to extract the third art design features from the third template image data for classification purposes.

[0135] The third template image data is input into the third sub-encoder Encoder_3 for encoding. Features representing the art design are extracted from the third template image data for classification purposes and are denoted as the third art design features.

[0136] Step 107: Input the first art design feature, the second art design feature and the third art design feature into the fusion device, and fuse them into a fourth art design feature for the purpose of classification.

[0137] The first, second, and third art design features are input into the Fusion fusion engine. The Fusion fusion engine integrates the first, second, and third art design features to enhance the expressive power of the features and merges them into a fourth art design feature for the purpose of classification.

[0138] In one design, the Fusion fusion module includes a first depthwise separable convolutional layer DSC_1, a second depthwise separable convolutional layer DSC_2, a first convolutional module ConvBolck_1, a second convolutional module ConvBolck_2, a third convolutional module ConvBolck_3, a fourth convolutional module ConvBolck_4, and a fifth convolutional module ConvBolck_5.

[0139] In this design, the third art design features are input into the first depthwise separable convolutional layer DSC_1 to extract the first template difference image features.

[0140] The first template difference image features are input into the second depth separable convolutional layer DSC_2 to extract the second template difference image features.

[0141] On the one hand, functions such as Add are used to fuse the first artistic design features with the first template difference image features into the first target image features.

[0142] The first target image features are input into the first convolution module ConvBolck_1, where convolution, normalization, and activation operations are performed sequentially to obtain the second target image features.

[0143] The second target image features and the second template difference image features are fused together using functions such as Add to form the third target image features.

[0144] The third target image features are input into the second convolution module ConvBolck_2, where convolution, normalization, and activation operations are performed sequentially to obtain the fourth target image features.

[0145] On the other hand, functions such as Add are used to fuse the second art design features with the first template difference image features into the fifth target image features.

[0146] The features of the fifth target image are input into the third convolution module ConvBolck_3, where convolution, normalization, and activation operations are performed sequentially to obtain the features of the sixth target image.

[0147] The sixth target image features and the second template difference image features are fused together using functions such as Add to form the seventh target image features.

[0148] The seventh target image features are input into the fourth convolution module ConvBolck_4, where convolution, normalization, and activation operations are performed sequentially to obtain the eighth target image features.

[0149] Combining the above two aspects, functions such as Concat are used to concatenate the features of the fourth target image with the features of the eighth target image to form the features of the ninth target image.

[0150] The ninth target image features are input into the fifth convolution module ConvBolck_5, where convolution, normalization, and activation operations are performed sequentially to obtain the fourth art design features.

[0151] Step 108: Based on the fourth art design feature, search for target image data that is similar to the first template image data and the second template image data in the art design.

[0152] In this embodiment, the fourth artistic design feature is used as the fingerprint data of the current VR image data. Other VR image data similar to the first template image data and the second template image data are searched in the database and used as target image data that is similar to the current VR image data in terms of artistic design. The target image data is then returned to the VR device and displayed to the user for browsing.

[0153] In the specific implementation, the query history uses reference image features encoded by an image encoder for the target image data (VR image data) under virtual reality.

[0154] The similarity between the fourth art design features and the reference image features is calculated using methods such as the cosine algorithm.

[0155] The target image data with the highest similarity is determined to be similar to the first template image data and the second template image data in terms of artistic design.

[0156] In this embodiment, when a virtual reality-based image search request is received, first template image data and second template image data under virtual reality are constructed according to the search request; a difference operation is performed on the first template image data and the second template image data to obtain third template image data; an image encoder is determined; the image encoder includes a first sub-encoder, a second sub-encoder, a third sub-encoder, and a fusion unit; the first template image data is input into the first sub-encoder to extract a first artistic design feature from the first template image data for the purpose of inverse distortion; the second template image data is input into the second sub-encoder to extract a second artistic design feature from the second template image data for the purpose of inverse distortion; the third template image data is input into the third sub-encoder to extract a third artistic design feature from the third template image data for the purpose of classification; the first artistic design feature, the second artistic design feature, and the third artistic design feature are input into the fusion unit to fuse them into a fourth artistic design feature for the purpose of classification; and target image data similar to the first template image data and the second template image data are searched for in terms of artistic design based on the fourth artistic design feature. This embodiment reverses the distortion of VR image data and considers the parallax within the VR image data, thereby constructing high-quality fingerprint data for VR image data, enabling VR image data search, and ensuring that the representation of VR image data and fingerprint data are consistent, effectively improving the accuracy of VR search.

[0157] Furthermore, this embodiment provides a unified-scale encoding for VR image data of various formats during the search process, effectively integrating fragmented VR image data resources, expanding the search space, and further improving the accuracy of the search.

[0158] Furthermore, this embodiment provides a generalized search mode based on art design, offering users more search results and potential possibilities when VR image data is relatively scarce, thereby improving the user's VR search experience.

[0159] It should be noted that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0160] Reference Figure 5 The diagram illustrates a device for searching art and design images based on virtual reality, as provided in an embodiment of this application. Specifically, it may include the following modules:

[0161] The template image data construction module 501 is used to construct first template image data and second template image data under virtual reality according to the search request when a virtual reality-based image search request is received.

[0162] Template difference module 502 is used to perform a difference operation on the first template image data and the second template image data to obtain the third template image data;

[0163] Image encoder determination module 503 is used to determine the image encoder; the image encoder includes a first sub-encoder, a second sub-encoder, a third sub-encoder and a fusion unit;

[0164] The first art design feature extraction module 504 is used to input the first template image data into the first sub-encoder to extract the first art design features from the first template image data for the purpose of inverse distortion.

[0165] The second art design feature extraction module 505 is used to input the second template image data into the second sub-encoder to extract the second art design features from the second template image data for the purpose of inverse distortion.

[0166] The third art design feature extraction module 506 is used to input the third template image data into the third sub-encoder and extract the third art design features from the third template image data for classification purposes.

[0167] The fourth art design feature fusion module 507 is used to input the first art design feature, the second art design feature and the third art design feature into the fusion device and fuse them into a fourth art design feature for the purpose of classification.

[0168] The target image data search module 508 is used to search for target image data that is similar to the first template image data and the second template image data in terms of art design based on the fourth art design feature.

[0169] In one embodiment of this application, the template image data construction module 501 is further configured to:

[0170] If the search request contains a frame of original image data under virtual reality, then the original image data is set as the first template image data and the second template image data respectively;

[0171] If the search request contains two frames of original image data under virtual reality, then one frame of the original image data is set as the first template image data, and the other frame of the original image data is set as the second template image data.

[0172] In one embodiment of this application, the image encoder determination module 503 includes:

[0173] The first-stage training module is used to synchronously train the first sub-encoder and the second sub-encoder in a generative adversarial mode so that the first sub-encoder and the second sub-encoder have the function of inverse distortion.

[0174] The second-stage training module is used to train the third sub-encoder and the fusion machine in a classification mode while keeping the first and second sub-encoders unchanged, once the training of the first and second sub-encoders is completed, so that the third sub-encoder and the fusion machine have the ability to classify in artistic design.

[0175] In one embodiment of this application, the first-stage training module is further configured to:

[0176] Collect sample image data with and without distortion under virtual reality;

[0177] A first generative adversarial network and a second generative adversarial network are determined; wherein, the first generative adversarial network includes a first generator and a discriminator, and the second generative adversarial network includes a second generator and a discriminator; the first generator includes a first sub-encoder and a first decoder, and the second generator includes a second sub-encoder and a second decoder;

[0178] The discriminator is trained based on the sample image data so that it can be used to determine whether distortion exists.

[0179] With the assistance of the discriminator, the first generator is trained based on the sample image data so that the first generator can be used to recover the distortion of the sample image data.

[0180] With the assistance of the discriminator, the second generator is trained based on the sample image data so that the second generator can be used to recover the distortion of the sample image data.

[0181] In one embodiment of this application, the second-stage training module includes:

[0182] The label image data acquisition module is used to acquire first label image data and second label image data under virtual reality; the first label image data and the second label image data jointly label and represent the real category of the art design.

[0183] The label difference module is used to perform a difference operation on the first label image data and the second label image data to obtain the third label image data;

[0184] An image classification network determination module is used to determine an image classification network; the image classification network includes a first sub-encoder, a second sub-encoder, a third sub-encoder, a fusion unit, and a head structure;

[0185] The first label image feature extraction module is used to input the first label image data into the first sub-encoder and extract the first label image features from the first label image data for the purpose of inverse distortion.

[0186] The second label image feature extraction module is used to input the second label image data into the second sub-encoder to extract the second label image features from the second label image data for the purpose of inverse distortion.

[0187] The third label image feature extraction module is used to input the third label image data into the third sub-encoder and extract the third label image features representing the art design from the third label image data;

[0188] The fourth label image feature fusion module is used to input the first label image features, the second label image features and the third label image features into the fusion unit and fuse them into a fourth label image feature;

[0189] The prediction category segmentation module is used to input the features of the fourth label image into the head structure to segment the predicted category in terms of artistic design.

[0190] A partial update module is used to maintain the first sub-encoder and the second sub-encoder unchanged, and update the third sub-encoder, the fusion unit and the head structure while training the image classification network based on the loss value between the true category and the predicted category.

[0191] In one embodiment of this application, the fusion unit includes a first depthwise separable convolutional layer, a second depthwise separable convolutional layer, a first convolutional module, a second convolutional module, a third convolutional module, a fourth convolutional module, and a fifth convolutional module;

[0192] The fourth label image feature fusion module is also used for:

[0193] The third label image features are input into the first depthwise separable convolutional layer to extract the first label difference image features;

[0194] The first label difference image features are input into the second depthwise separable convolutional layer to extract the second label difference image features;

[0195] The features of the first label image and the features of the first label difference image are fused together to form the features of the first sample image;

[0196] The first sample image features are input into the first convolution module, and convolution, normalization and activation operations are performed sequentially to obtain the second sample image features.

[0197] The features of the second sample image are fused with the features of the second label difference image to form the features of the third sample image;

[0198] The third sample image features are input into the second convolution module to perform convolution, normalization and activation operations in sequence to obtain the fourth sample image features.

[0199] The features of the second label image are fused with the features of the first label difference image to form the features of the fifth sample image;

[0200] The features of the fifth sample image are input into the third convolution module, and convolution, normalization and activation operations are performed sequentially to obtain the features of the sixth sample image.

[0201] The features of the sixth sample image are fused with the features of the second label difference image to form the features of the seventh sample image;

[0202] The features of the seventh sample image are input into the fourth convolution module, and convolution, normalization and activation operations are performed sequentially to obtain the features of the eighth sample image.

[0203] The features of the fourth sample image are concatenated with the features of the eighth sample image to form the features of the ninth sample image.

[0204] The features of the ninth sample image are input into the fifth convolution module, where convolution, normalization, and activation operations are performed sequentially to obtain the features of the fourth label image.

[0205] In one embodiment of this application, the fourth art design feature fusion module 507 is further configured to:

[0206] The third art design feature is input into the first depthwise separable convolutional layer to extract the first template difference image feature;

[0207] The first template difference image features are input into the second depthwise separable convolutional layer to extract the second template difference image features;

[0208] The first artistic design feature and the first template difference image feature are fused together to form the first target image feature;

[0209] The first target image features are input into the first convolution module, and convolution, normalization and activation operations are performed sequentially to obtain the second target image features.

[0210] The second target image features are fused with the second template difference image features to form a third target image feature;

[0211] The third target image features are input into the second convolution module, where convolution, normalization and activation operations are performed sequentially to obtain the fourth target image features.

[0212] The second artistic design feature is fused with the first template difference image feature to form the fifth target image feature;

[0213] The fifth target image features are input into the third convolution module, where convolution, normalization and activation operations are performed sequentially to obtain the sixth target image features.

[0214] The sixth target image features are fused with the second template difference image features to form the seventh target image features;

[0215] The seventh target image features are input into the fourth convolution module, where convolution, normalization and activation operations are performed sequentially to obtain the eighth target image features.

[0216] The fourth target image feature is concatenated with the eighth target image feature to form the ninth target image feature;

[0217] The ninth target image features are input into the fifth convolution module, where convolution, normalization, and activation operations are performed sequentially to obtain the fourth art design features.

[0218] In one embodiment of this application, the target image data search module 508 is further configured to:

[0219] The query history uses the reference image features encoded by the image encoder for target image data in virtual reality;

[0220] Calculate the similarity between the fourth artistic design feature and the reference image feature;

[0221] The target image data with the highest similarity is determined to be similar to the first template image data and the second template image data in terms of artistic design.

[0222] This application provides an apparatus for searching art and design images based on virtual reality. By using this apparatus, the steps in the aforementioned method embodiments can be implemented.

[0223] As the apparatus embodiments are basically similar to the method embodiments, they are described in a relatively simple manner. For relevant details, please refer to the description in the method embodiment section.

[0224] Reference Figure 6 The diagram illustrates a terminal device provided in an embodiment of this application. Figure 6 As shown, the terminal device 600 in this embodiment includes: a processor 610, a memory 620, and a computer program 621 stored in the memory 620 and executable on the processor 610. When the processor 610 executes the computer program 621, it implements the steps in the various embodiments of the method for searching art design images based on virtual reality. Alternatively, when the processor 610 executes the computer program 621, it implements the functions of each module / unit in the various device embodiments described above.

[0225] For example, the computer program 621 may be divided into one or more modules / units, which are stored in the memory 620 and executed by the processor 610 to complete this application. The one or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which may be used to describe the execution process of the computer program 621 in the terminal device 600.

[0226] The terminal device 600 may include, but is not limited to, a processor 610 and a memory 620. Those skilled in the art will understand that... Figure 6 This is merely one example of terminal device 600 and does not constitute a limitation on terminal device 600. It may include more or fewer components than shown, or combine certain components, or different components. For example, terminal device 600 may also include input / output devices, network access devices, buses, etc.

[0227] The processor 610 can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0228] The memory 620 can be an internal storage unit of the terminal device 600, such as a hard disk or memory of the terminal device 600. The memory 620 can also be an external storage device of the terminal device 600, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD) card, flash card, etc., equipped on the terminal device 600. Furthermore, the memory 620 can include both internal and external storage units of the terminal device 600. The memory 620 is used to store the computer program 621 and other programs and data required by the terminal device 600. The memory 620 can also be used to temporarily store data that has been output or will be output.

[0229] This application also discloses a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the method for searching art design images based on virtual reality as described in the foregoing embodiments.

[0230] This application also discloses a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method for searching art design images based on virtual reality as described in the foregoing embodiments.

[0231] This application also discloses a computer program product that, when run on a computer, causes the computer to execute the method for searching art design images based on virtual reality as described in the foregoing embodiments.

[0232] The embodiments described above are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for searching art and design images based on virtual reality, characterized in that, include: When a virtual reality-based image search request is received, first template image data and second template image data under virtual reality are constructed according to the search request. Perform a difference operation on the first template image data and the second template image data to obtain the third template image data; Determine the image encoder; the image encoder includes a first sub-encoder, a second sub-encoder, a third sub-encoder, and a fusion unit; The first template image data is input into the first sub-encoder to extract the first artistic design feature from the first template image data for the purpose of inverse distortion; The second template image data is input into the second sub-encoder to extract the second artistic design features from the second template image data for the purpose of inverse distortion; The third template image data is input into the third sub-encoder to extract third art design features from the third template image data for classification purposes; The first art design feature, the second art design feature, and the third art design feature are input into the fusion unit and fused into a fourth art design feature for the purpose of classification. Based on the fourth art design feature, search for target image data that is similar to the first template image data and the second template image data in the art design; The determination of the image encoder includes: The first sub-encoder and the second sub-encoder are trained synchronously in a generative adversarial mode so that the first sub-encoder and the second sub-encoder have the function of inverse distortion. If the training of the first sub-encoder and the second sub-encoder is completed, then while keeping the first sub-encoder and the second sub-encoder unchanged, the third sub-encoder and the fusion machine are trained in a classification mode so that the third sub-encoder and the fusion machine have the ability to classify in artistic design. The step of synchronously training the first sub-encoder and the second sub-encoder in a generative adversarial mode to enable the first sub-encoder and the second sub-encoder to have inverse distortion functionality includes: Collect sample image data with and without distortion under virtual reality; A first generative adversarial network and a second generative adversarial network are determined; wherein, the first generative adversarial network includes a first generator and a discriminator, and the second generative adversarial network includes a second generator and a discriminator; the first generator includes a first sub-encoder and a first decoder, and the second generator includes a second sub-encoder and a second decoder; The discriminator is trained based on the sample image data so that it can be used to determine whether distortion exists. With the assistance of the discriminator, the first generator is trained based on the sample image data so that the first generator can be used to recover the distortion of the sample image data. With the assistance of the discriminator, the second generator is trained based on the sample image data so that the second generator can be used to recover the distortion of the sample image data.

2. The method according to claim 1, characterized in that, The construction of the first template image data and the second template image data under virtual reality based on the search request includes: If the search request contains a frame of original image data under virtual reality, then the original image data is set as the first template image data and the second template image data respectively; If the search request contains two frames of original image data under virtual reality, then one frame of the original image data is set as the first template image data, and the other frame of the original image data is set as the second template image data.

3. The method according to claim 1, characterized in that, The step of training the third sub-encoder and the fusion generator in a classification mode while keeping the first and second sub-encoders unchanged, so that the third sub-encoder and the fusion generator have the ability to classify in artistic design, includes: Collect first-label image data and second-label image data under virtual reality; the first-label image data and the second-label image data are jointly labeled to represent the real category of the art design; Perform a difference operation on the first label image data and the second label image data to obtain the third label image data; An image classification network is defined; the image classification network includes a first sub-encoder, a second sub-encoder, a third sub-encoder, a fusion unit, and a head structure; The first label image data is input into the first sub-encoder to extract the first label image features from the first label image data for the purpose of inverse distortion; The second label image data is input into the second sub-encoder to extract the second label image features from the second label image data for the purpose of inverse distortion; The third label image data is input into the third sub-encoder, and the third label image features representing the art design are extracted from the third label image data; The first label image features, the second label image features, and the third label image features are input into the fusion unit and fused into a fourth label image feature; The fourth label image features are input into the head structure to classify and predict categories in terms of art and design. During the training of the image classification network based on the loss value between the true category and the predicted category, the first sub-encoder and the second sub-encoder are kept unchanged, while the third sub-encoder, the fusion unit, and the head structure are updated.

4. The method according to claim 3, characterized in that, The fusion unit includes a first depthwise separable convolutional layer, a second depthwise separable convolutional layer, a first convolutional module, a second convolutional module, a third convolutional module, a fourth convolutional module, and a fifth convolutional module; The step of inputting the first label image features, the second label image features, and the third label image features into the fusion unit to fuse them into a fourth label image feature includes: The third label image features are input into the first depthwise separable convolutional layer to extract the first label difference image features; The first label difference image features are input into the second depthwise separable convolutional layer to extract the second label difference image features; The features of the first label image and the features of the first label difference image are fused together to form the features of the first sample image; The first sample image features are input into the first convolution module, and convolution, normalization and activation operations are performed sequentially to obtain the second sample image features. The features of the second sample image are fused with the features of the second label difference image to form the features of the third sample image; The third sample image features are input into the second convolution module to perform convolution, normalization and activation operations in sequence to obtain the fourth sample image features. The features of the second label image are fused with the features of the first label difference image to form the features of the fifth sample image; The features of the fifth sample image are input into the third convolution module, and convolution, normalization and activation operations are performed sequentially to obtain the features of the sixth sample image. The features of the sixth sample image are fused with the features of the second label difference image to form the features of the seventh sample image; The features of the seventh sample image are input into the fourth convolution module, and convolution, normalization and activation operations are performed sequentially to obtain the features of the eighth sample image. The features of the fourth sample image are concatenated with the features of the eighth sample image to form the features of the ninth sample image. The features of the ninth sample image are input into the fifth convolution module, where convolution, normalization, and activation operations are performed sequentially to obtain the features of the fourth label image.

5. The method according to claim 4, characterized in that, The step of inputting the first art design feature, the second art design feature, and the third art design feature into the fusion unit and fusing them into a fourth art design feature for classification purposes includes: The third art design feature is input into the first depthwise separable convolutional layer to extract the first template difference image feature; The first template difference image features are input into the second depthwise separable convolutional layer to extract the second template difference image features; The first artistic design feature and the first template difference image feature are fused together to form the first target image feature; The first target image features are input into the first convolution module, and convolution, normalization and activation operations are performed sequentially to obtain the second target image features. The second target image features are fused with the second template difference image features to form a third target image feature; The third target image features are input into the second convolution module, where convolution, normalization and activation operations are performed sequentially to obtain the fourth target image features. The second artistic design feature is fused with the first template difference image feature to form the fifth target image feature; The fifth target image features are input into the third convolution module, where convolution, normalization and activation operations are performed sequentially to obtain the sixth target image features. The sixth target image features are fused with the second template difference image features to form the seventh target image features; The seventh target image features are input into the fourth convolution module, where convolution, normalization and activation operations are performed sequentially to obtain the eighth target image features. The fourth target image feature is concatenated with the eighth target image feature to form the ninth target image feature; The ninth target image features are input into the fifth convolution module, where convolution, normalization, and activation operations are performed sequentially to obtain the fourth art design features.

6. The method according to any one of claims 1-5, characterized in that, The step of searching for target image data similar to the first template image data and the second template image data in terms of art design based on the fourth art design feature includes: The query history uses the reference image features encoded by the image encoder for target image data in virtual reality; Calculate the similarity between the fourth artistic design feature and the reference image feature; The target image data with the highest similarity is determined to be similar to the first template image data and the second template image data in terms of artistic design.

7. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for searching art and design images based on virtual reality as described in any one of claims 1-6.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method for searching art and design images based on virtual reality as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Living body detection method, training method of living body detection model and corresponding device

    CN115482591A

  • Building facade image material matching method and device, equipment and storage medium

    CN116012626A