Device and method

The apparatus and method address the challenge of user preference-based style conversion by extracting trends from NFT images, ensuring style conversion aligns with user preferences and enhances image data value.

WO2025169471A1PCT designated stage Publication Date: 2025-08-14NTT DOCOMO INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/004589
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-09
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing image conversion devices fail to perform style conversion based on user preferences, making the process difficult and time-consuming.

Method used

An apparatus and method that extracts user image preference trends from owned NFT images using vector clustering and conversion models like img2img, generating output image data that aligns with the user's preferences.

Benefits of technology

Enables style conversion in the metaverse that accurately reflects user preferences, enhancing user satisfaction and asset value of image data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024004589_14082025_PF_FP_ABST
    Figure JP2024004589_14082025_PF_FP_ABST
Patent Text Reader

Abstract

The purpose of the present invention is to provide a device and a method which are capable of performing style conversion on the basis of the tendency of a user's preference. In a metaverse space provision device 100, a style extraction unit 102 extracts the tendency of a user's preference for image data on the basis of owned image data having an NFT corresponding to transaction history information (ownership information) of the user. Then, a style conversion unit 103 functions as a generation unit for generating, on the basis of the tendency of the user's preference, output image data from an input image to be converted, and generates, on the basis of a vector of the possessed image data, output image data in a style matched with the tendency of the user's preference.
Need to check novelty before this filing date? Find Prior Art

Description

Apparatus and method

[0001] The present invention relates to an apparatus and method for performing style transfer.

[0002] Patent document 1 describes a device that can increase a user's interest in image conversion by selecting any style from a number of styles in response to the user's intuitive actions and generating a painterly image by converting the original image into that style.

[0003] JP 2012-238086 A

[0004] However, the device described in Patent Document 1 cannot perform style conversion based on the user's preferences.

[0005] Therefore, an object of the present invention is to provide an apparatus and method that can perform style conversion based on the user's preferences.

[0006] The device of the present invention includes an extraction unit that extracts a user's image preference trends based on owned image data corresponding to the user's ownership information, and a generation unit that generates output image data from input image data based on the trends.

[0007] According to the present invention, it is possible to perform style conversion based on the user's preferences.

[0008] 1 is a diagram showing the system configuration of a metaverse space providing system that provides a metaverse space including the metaverse space providing device 100 of the present disclosure to a user terminal 100. FIG. 2 is a diagram schematically showing information stored in a wallet 501, a memory 401, and a blockchain 300. FIG. 3 is a diagram showing a transaction history of NFT images. FIG. 4 is a block diagram showing the functional configuration of the metaverse space providing device 200. FIG. 5 is a diagram showing an art style extraction process. FIG. 6 is an explanatory diagram illustrating the art style extraction process and the art style conversion process. FIG. 7 is a flowchart showing the operation of the metaverse space providing device 100 of the present disclosure. FIG. 8 is a flowchart showing the art style extraction and art style conversion processes of process S103 and process S104. FIG. 9 is a diagram schematically showing the art style suggestion process. FIG. 10 is a diagram showing an example of the hardware configuration of the metaverse space providing device 100 and a user terminal 500 according to an embodiment of the present disclosure.

[0009] The present disclosure will be described with reference to the accompanying drawings. Whenever possible, the same parts are designated by the same reference numerals and redundant description will be omitted.

[0010] 1 is a diagram showing the system configuration of a metaverse space providing system that provides a metaverse space including a metaverse space providing device 100 of the present disclosure to a user terminal 500. As shown in the figure, the metaverse space providing system includes the metaverse space providing device 100, an NFT issuing device 200, a blockchain 300, and a server 400.

[0011] The user terminal 500 is provided with the metaverse space and an avatar for acting therein by the metaverse space providing device 100 .

[0012] A user of the user terminal 500 can use an avatar as his or her own alter ego in the metaverse provided by the metaverse space providing device 100. In addition, famous characters (NPCs or avatars) may appear at events in this metaverse. A user may wish to take a memorable photo of their avatar together with a famous character, and in such a case, may wish to edit the photo to make their avatar look better (so-called enhancement). There may also be a desire to make the photo permanent (permanently preserved on a blockchain) as evidence that they were actually present at the location.

[0013] The metaverse space providing device 100 of the present disclosure can modify a user's avatar by converting the style of the avatar to meet the user's needs. Conventionally, adjusting the style conversion can be difficult and time-consuming. Therefore, the metaverse space providing device 100 of the present disclosure performs style conversion based on NFT images owned (or previously owned) by the user that are likely to be the user's preferences. The use of NFT images can more accurately emphasize the user's preferences. NFT images have asset value to the user and are considered to be more important to the user than general content. Naturally, the style conversion is not limited to NFT images, and general content (image data) stored in the user terminal 500 may also be used.

[0014] A user can operate the user terminal 500 to request the NFT issuing device 200 to issue an NFT image and obtain the NFT image. To this end, the user terminal 500 is equipped with a wallet 501. The wallet 501 stores a private key, which is a code required to extract the NFT image. The NFT image and metadata are stored in the memory 401 of the server 400.

[0015] The NFT issuing device 200 is a device that issues an NFT (Non-Fungible Token) image in response to a request from a user terminal 500. This NFT image is stored in memory 401 constructed in the server 400, and transaction history information (creator, owner, etc.) of the NFT image is stored in multiple computers (PCs) that make up the blockchain 300. As described above, the private key required to extract the NFT image is stored in the wallet 501 of the user terminal 500. In the present disclosure, the NFT image is assumed to be image data purchased by the user, but may be any other NFT image.

[0016] Here, we will explain how to extract an NFT image. FIG. 2 is a diagram schematically showing the information stored in the wallet 501, memory 401, and blockchain 300. FIG. 2(a) is a diagram showing the wallet 501. As shown in the diagram, the wallet 501 stores address A and a private key. Address A is an address for referencing transaction history information in the blockchain 300. The private key is used to prove ownership of a specific blockchain address and the assets (including NFT images) associated with that address. A user who possesses the private key is considered to be the legitimate owner of that address.

[0017] The NFT acquisition unit 101 can determine what is held or has been held by referring to the transaction history from the wallet 501 address.

[0018] For example, if the address of wallet 501 is address A, referring to the transaction history reveals that it holds an NFT image identified by address B and token ID (1).

[0019] Then, the NFT acquisition unit 101 queries the off-chain (server 400) side for the metadata specified by address B and token ID (1), and acquires the NFT image from the URI described in the metadata.

[0020] Furthermore, as described above, the transaction history of NFT images can be obtained from the blockchain 300. FIG. 3 is a diagram showing the transaction history of NFT images. This transaction is information obtained from the transaction history information stored in the blockchain. As shown in this diagram, the owner transition, contract address, token ID, and transition time (timestamp) are obtained from the transaction history information. In this diagram, "method" indicates the method of issuing or transferring an NFT image, and "from" and "to" indicate the transfer of ownership. While the transaction history is omitted in FIG. 2, the information shown in FIG. 3 is managed by the blockchain 300 so that it can be obtained. The accuracy of user preferences can be improved by performing filtering such as: (1) calculating the user's preference trends (features) using only NFT images that the user still owns; or (2) calculating the user's preference trends (features) using only NFT images that the user has held for a certain period of time.

[0021] 4 is a block diagram showing the functional configuration of the metaverse space providing device 100. As shown in the figure, the metaverse space providing device 100 includes an NFT acquisition unit 101, an art style extraction unit 102, an art style conversion unit 103, a metaverse space processing unit 104, a learning unit 105, and a metaverse space DB 106.

[0022] The NFT acquisition unit 101 is a part that acquires one or more NFT images from the memory 401 of the server 400. When the NFT acquisition unit 101 receives a request to edit an avatar from the user terminal 500, it acquires an NFT image as information for the editing. For example, when a user takes a photo of their own avatar, the instruction to take a photo serves as a trigger, and the NFT acquisition unit 101 acquires the NFT image.

[0023] The style extraction unit 102 extracts an art style from the acquired NFT image. In the present disclosure, an NFT image is image data. The style extraction unit 102 converts the image data into a real vector and performs clustering on the result of the conversion into a real vector. The style extraction unit 102 then extracts preferred image data from the obtained clustering results according to a certain rule (e.g., in descending order of the number of images belonging to the cluster).

[0024] The style extraction process will now be explained using diagrams. Figure 5(a) is a schematic diagram showing vector conversion in the style extraction unit 102. As shown in the figure, img2vec is a module that vectorizes image data, and the style extraction unit 102 is equipped with this module. In the figure, img2vec is equipped with VGG-16 or ResNet. VGG-16 is a convolutional neural network with a depth of 16 layers that extracts features of image data and outputs them as vectors. ResNet (Residual Neural Networks) is a model that trains a highly accurate deep CNN by serially connecting multiple residual blocks that utilize residual connections to model the residual sequence. This also inputs image data and outputs vectors representing its features.

[0025] FIG. 5(b) is a diagram illustrating the clustering of vectorized NFT images. In the diagram, the vertical and horizontal axes represent each vector, and the circles represent the vectors of each image data. Note that the vector axes are omitted in the diagram for ease of explanation, but naturally, there are more than two vector axes. By clustering these vectorized NFT images, they can be divided into several groups (clusters). It can be determined that the group containing the most NFT images contains the user's preferred image data. Note that the criteria for this determination are not limited to the number of image data items, but can also be based, for example, on the time of purchase as an NFT image, the purchase price, etc. NFT images belonging to a cluster containing NFT images purchased at a high price, or NFT images belonging to a cluster containing NFT images purchased recently, can be determined to be the user's preferred image data. In this case, the blockchain 300 may include price information.

[0026] The style extraction unit 102 averages the vectors of several acquired NFT images to obtain a single NFT image vector. This single NFT image vector can be considered a vector of image data preferred by the user, and can be considered a vector representing the user's preferred style. Note that this is not limited to averaging, and a single vector can be calculated or obtained from the vectors of multiple NFT images using a median or other statistical method.

[0027] The style extraction unit 102 then performs style extraction processing based on the NFT images belonging to the cluster determined to be preferred image data. Based on the NFT image group determined to be a personal favorite, the style extraction unit 102 compares the vector of the obtained preferred NFT image with the vector of a predetermined teacher, and obtains text (e.g., watercolor style) corresponding to a similar teacher drawing. This text corresponds to the style.

[0028] The style conversion unit 103 converts input image data into output image data that meets the user's preferences. This style conversion unit 103 is configured with image conversion models such as img2img. Details will be described later.

[0029] The metaverse space processing unit 104 provides the user terminal 500 with a metaverse space and performs image processing to reflect operations by the user terminal 500 in the metaverse space. This metaverse space includes the user's avatar and other characters (NPCs, avatars of other users). Here, for example, in response to the user's photographing operation, the style conversion unit 103 converts the style of the user's avatar in the metaverse space, and the metaverse space processing unit 104 provides the style-converted avatar in the metaverse. That is, the style conversion unit 103 extracts character data constituting the user's avatar that has been registered in advance, inputs it into img2img, and executes style conversion processing. The metaverse space processing unit 104 then photographs the converted avatar and acquires image data of the converted avatar.

[0030] Furthermore, the metaverse space processing unit 104 requests the NFT issuing device 200 to issue an NFT image in response to a user operation from the user terminal 500. For example, when a user's avatar performs a photographing process in the metaverse space through a user operation, the NFT issuing device 200 issues the captured image data as an NFT image. To this end, the metaverse space processing unit 104 transmits the image data captured by the user's avatar to the NFT issuing device 200. The NFT issuing device 200 stores the image data in the wallet 501 and stores the private key in the wallet 501 and transaction information such as the creator and owner of the image data in the blockchain 300. This ensures that the NFT is unique data and has been acquired (including the image) in the metaverse.

[0031] Although the style of the avatar itself is changed in the above, it is also possible to change the style of only the avatar portion in the image data obtained by photographing.

[0032] In the present disclosure, an NFT image of a user's preference can be generated by taking a picture of an avatar with a converted style. Furthermore, if the user can take a picture with a famous character, the asset value of the image data can be increased.

[0033] The learning unit 105 is a part that trains the style conversion unit 103 (img2img) based on acquired NFT images (proprietary image data). Learning is not limited to NFT images, and general image data may also be used. In this disclosure, an NFT image is defined as proprietary image data that is an image indicated by non-fungible information written on a blockchain.

[0034] The metaverse space DB 106 is a part that stores character data such as avatars and NPCs provided in the metaverse space, as well as other image data that constitutes the metaverse.

[0035] Next, detailed processing of the art style extraction process and art style conversion process in the metaverse space providing device 100 of the present disclosure will be described. Figure 6(a) is a diagram showing the correspondence between the vector of the teacher's drawing and the image data of the style. As shown in the figure, teacher drawing vector 1 is a vector of image data showing a watercolor style, and teacher drawing vector 2 is a vector of image data showing an anime style or a brighter style. Each image data vector is associated with text such as watercolor style or anime style, and the text is stored as a conversion table in the art style extraction unit 102.

[0036] 6B is a block diagram showing the detailed configuration of the style extraction unit 102, the style conversion unit 103, and the learning unit 105. The style extraction unit 102 includes a conversion table, and the style conversion unit 103 includes img2img. The learning unit 105 includes a fine-tuning unit and LoRA.

[0037] The conversion table compares the vector of the obtained preferred NFT image (for example, the average value of the vectors of multiple NFT images) with the vector of a predetermined teacher's painting to obtain text (e.g., watercolor style) corresponding to the teacher's painting. This text is a so-called prompt, and is information that determines the conversion content (painting style) of the input image to be converted.

[0038] In this figure, processes such as vector averaging of the NFT image input to the conversion table are omitted.

[0039] img2img is an image conversion model that, by inputting input image data and text, converts the style of the input image data according to the style indicated in the text derived from a conversion table. For example, Stable Diffusion is known as such an image conversion model. This allows the style conversion unit 103 (img2img) to convert the style of the input image data.

[0040] The learning unit 105 is a part that fine-tunes img2img using the finetuning unit based on the obtained preferred owned image data. The learning unit 105 may also perform LoRA on the obtained preferred owned image data. LoRA (Low-Rank Adaptation) is a technique used when fine-tuning to customize output to suit a specific field. Specifically, while keeping the parameters of img2img fixed, additional parameters are added and only those parts are fine-tuned.

[0041] The learning unit 105 is not necessarily a required component. The conversion table is also not necessarily required; an img2img table adjusted through learning may be used. In this case, a predetermined character string set in advance is used as the prompt.

[0042] Next, a description will be given of the operation of the metaverse space providing device 100 configured as described above. Fig. 7 is a flowchart showing the operation of the metaverse space providing device 100 of the present disclosure.

[0043] When the metaverse spatial processing unit 104 receives an instruction to photograph an avatar in the metaverse (S101), the NFT acquisition unit 101 acquires an NFT image (owned image data) that is owned (or was owned) by the user operating the avatar, using the private key of the wallet 501 and the blockchain 300 (S102). The metaverse spatial processing unit 104 can confirm that the NFT image belongs to the user based on the owner information in the blockchain 300.

[0044] When the NFT acquisition unit 101 acquires an NFT image, the style extraction unit 102 acquires vectors of the acquired multiple NFT images (owned image data) and performs clustering processing.The style extraction unit 102 then acquires an NFT image (owned image data) that belongs to one group from several groups of NFT images (owned image data) obtained by clustering, and obtains the vector of that NFT image (owned image data).

[0045] The style conversion unit 103 performs style conversion processing based on the obtained vectors (S104). That is, the style conversion unit 103 compares the obtained vectors with the vectors of the teacher image to obtain text information. The text information and the input image (here, the user's avatar) are input to img2img to obtain output image data with the style converted. This output image data is image data with the style converted.

[0046] The metaverse spatial processing unit 104 then performs a photographing process of the style-converted avatar in accordance with the user's operation, and obtains output image data (S105). The metaverse spatial processing unit 104 then sends the output image data and the user's identification information to the NFT issuing device 200, requesting the issuance of an NFT image. The NFT issuing device 200 then converts the output image data into an NFT image (S106). That is, the NFT issuing device 200 stores the output image data in the server 400 (memory 401), stores the encryption key in the wallet 501, and stores the user's identification information as the owner in the blockchain 300, thereby converting the output image data into an NFT image. Note that in the above process, the style determination is triggered by photography, but this is not limited to this. To determine the style in advance, an NFT image (owned image data) based on ownership information may be acquired, a style based on the acquired image data may be stored, and a conversion process based on the style may be performed when the image is captured.

[0047] Next, the style extraction and style conversion processes of steps S103 and S104 will be described. FIG. 8 is a flowchart showing these processes. The style extraction unit 102 vectorizes the NFT images (S103a) and performs clustering based on the vectors (S103b). The style extraction unit 102 then selects the most common style (group of vectors) of the NFT images (S103c). The style conversion unit 103 then performs style conversion processing using the vectors (S104a).

[0048] That is, as described above, the style extraction unit 102 extracts a style (vector) based on the acquired vector and the vector of the teacher's drawing, and the style conversion unit 103 performs style conversion processing based on the extracted style (vector). The style conversion processing uses img2img, but other conversion models may also be used.

[0049] In step S103c, the style extraction unit 102 selects the group with the largest number of vectors from the multiple NFT image vectors obtained by clustering, but this is not limited to this. For example, the style extraction unit 102 may select groups in descending order of the number of vectors included in other groups, and a suggestion unit (not shown) may suggest the selected group to the user, allowing the user to select one. In this case, it is preferable to use a conversion table to present the user with information indicating the style of the vectors of the image data.

[0050] Fig. 9 is a diagram showing the proposed process. As shown in the figure, there are three NFT images (vectors) in cluster 1. There are two image data (vectors) in cluster 2, and one image data (vector) in cluster 3.

[0051] The image feature represents a vector obtained by averaging the vectors of the images contained in each cluster. In Figure 9, this image feature 1 is designated as user suggestion 1, but since the image feature represents a vector of image data, it is incomprehensible to the user. Therefore, the style extraction unit 102 should use a conversion table in the style conversion unit 103 to compare the input image with each of multiple teacher images and present the user with the style indicated by the closest teacher image. Alternatively, style conversion may actually be applied to the input image and presented to the user.

[0052] Next, the effects of the metaverse space providing device 100 of the present disclosure will be described. In the metaverse space providing device 100, the style extraction unit 102 extracts a user's preference for image data based on NFT images (owned image data) corresponding to the user's transaction history information (ownership information). The style conversion unit 103 functions as a generator that generates output image data from the input image to be converted based on the user's preference, and generates output image data in a style that matches the user's preference based on vectors of NFT images (owned image data) that the user owns or has owned.

[0053] This allows the system to determine a user's preferences based on the NFT images the user owns (or has owned), and convert image data into a style that matches those preferences. This allows the user to obtain image data that suits their preferences. For example, by converting their own avatar in the metaverse into a style that matches their preferences and photographing it, they can obtain image data that gives them satisfaction.

[0054] In the present disclosure, the metaverse space providing device 100 includes an NFT acquisition unit 101 that functions as a data acquisition unit that acquires owned image data, which is non-fungible image data (NFT images). The art style extraction unit 102 extracts a user's preference tendency based on the owned image data corresponding to the ownership information.

[0055] By using NFT images, it is possible to extract trends in user preferences.

[0056] The NFT acquisition unit 101 acquires an NFT image that is currently owned or has been owned in the past as owned image data. As described above, the ownership history of the image data can be determined based on the transaction history information stored in the blockchain 300, and the NFT acquisition unit 101 can acquire the specified NFT image. This NFT image (owned image data) may be image data based on the period of ownership by the user. The period of ownership can be determined from the transaction history information.

[0057] In the present disclosure, the art style extraction unit 102 acquires vectors from NFT images (proprietary image data) and extracts user preference trends based on the vectors and reference data. Here, the reference data is vectors of pre-prepared teacher images. This allows the user's preference trends to be grasped.

[0058] The style extraction unit 102 clusters the vectors of multiple NFT images (proprietary image data) to obtain multiple vectors of multiple NFT images (proprietary image data) that belong to a group that satisfies a condition. Here, the condition is, for example, the group that contains the largest number of vectors. The style extraction unit 102 then extracts the user's preference tendency based on the multiple vectors (for example, their average values) and reference data.

[0059] The art style extraction unit 102 has a conversion table that associates reference data with text information. The art style extraction unit 102 then refers to the conversion table and extracts text information from a vector obtained by aggregating (e.g., averaging) multiple vectors, as a trend in the user's preferences. This text information is information indicating an art style, such as watercolor or anime style.

[0060] The style conversion unit 103 generates output image data from input image data using an image conversion model such as img2img. The parameters of this image conversion model are learned by a learning unit 105 based on the user's preferences.

[0061] This allows the style conversion unit 103 to convert the image into a style that matches the preferences of the user.

[0062] The learning unit 105 may fine-tune the img2img (image transformation model) based on the user's preferences or may learn it using LoRA.

[0063] The apparatus and method of the present invention have the following configuration.

[0064] [1] A device comprising: an extraction unit that extracts a user's image preference tendency based on owned image data corresponding to the user's ownership information; and a generation unit that generates output image data from input image data based on the tendency.

[0065] [2] The device described in [1], further comprising a data acquisition unit that acquires owned image data, which is an image based on information written on a non-fungible blockchain, and the extraction unit extracts the trend based on the owned image data corresponding to the owned information.

[0066] [3] The device according to [2], wherein the data acquisition unit acquires image data that is currently owned or image data that has been owned in the past as the owned image data.

[0067] [4] The device according to [3], wherein the owned image data is image data based on a period of time that the image data has been owned by the user.

[0068] [5] The device according to [4], wherein the extraction unit acquires vector data from the owned image data and extracts user preference trends based on the vector data and reference data.

[0069] [6] The device described in [5], wherein the extraction unit clusters the vector data of the plurality of owned image data to obtain a plurality of vector data of the plurality of owned image data belonging to a group that satisfies a condition, and extracts the user's preference tendency based on the plurality of vector data and the reference data.

[0070] [7] The device described in [6], wherein the extraction unit has a conversion table that associates reference data with text information, and refers to the conversion table to extract text information from a single vector data set that aggregates the multiple vector data sets as a preference trend of the user.

[0071] [8] The device described in any one of [1] to [7], wherein the generation unit generates converted image data from source image data using an image transformation model, and further includes a learning unit that learns parameters of the image transformation model based on the user's preference trends. [9] The device described in [8], wherein the learning unit fine-tunes the image transformation model based on the user's preference trends.

[0072]

[10] A method comprising: an extraction step of extracting a user's image preference tendency based on owned image data corresponding to the user's ownership information; and a generation step of generating converted image data from source image data based on the tendency.

[0073] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (e.g., via wire, wirelessly, etc.) and these multiple devices. The functional block may also be realized by combining the single device or multiple devices with software.

[0074] Functions include, but are not limited to, judgment, determination, assessment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.

[0075] For example, the metaverse space providing device 100 and the user terminal 500 according to an embodiment of the present disclosure may function as a computer that performs processing of the image processing method of the present disclosure. Fig. 10 is a diagram illustrating an example of the hardware configuration of the metaverse space providing device 100 and the user terminal 500 according to an embodiment of the present disclosure. The metaverse space providing device 100 and the user terminal 500 described above may be physically configured as a computer device including a processor 1001, a memory 1002, a storage 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like.

[0076] In the following description, the term "device" can be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the metaverse space providing device 100 and the user terminal 500 may be configured to include one or more of the devices shown in the figure, or may be configured to exclude some of the devices.

[0077] Each function of the metaverse space providing device 100 and the user terminal 500 is realized by loading specified software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via the communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.

[0078] The processor 1001 controls the entire computer by running, for example, an operating system. The processor 1001 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control unit, an arithmetic unit, a register, etc. For example, the above-mentioned style extraction unit 102, style conversion unit 103, etc. may be realized by the processor 1001.

[0079] The processor 1001 also loads programs (program codes), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with these programs. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the style extraction unit 102 and the style conversion unit 103 may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may be used for other functional blocks. While the above-described various processes have been described as being executed by a single processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may also be transmitted from a network via a telecommunications line.

[0080] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be called a register, a cache, a main memory (primary storage device), etc. The memory 1002 can store executable programs (program codes), software modules, etc. for implementing an image processing method according to an embodiment of the present disclosure.

[0081] Storage 1003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003.

[0082] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, or a communication module. The communication device 1004 may be configured to include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. to realize at least one of frequency division duplex (FDD) and time division duplex (TDD). For example, the above-mentioned NFT acquisition unit 101 may be realized by the communication device 1004. The communication device 1004 may be implemented with a transmitter and a receiver that are physically or logically separated.

[0083] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. Note that the input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).

[0084] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.

[0085] Furthermore, the metaverse space providing device 100 and the user terminal 500 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a field programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware.

[0086] The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed by physical layer signaling (e.g., Downlink Control Information (DCI) and Uplink Control Information (UCI)), higher layer signaling (e.g., Radio Resource Control (RRC) signaling, Medium Access Control (MAC) signaling, broadcast information (Master Information Block (MIB) and System Information Block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, and may be, for example, an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.

[0087] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.

[0088] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.

[0089] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).

[0090] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).

[0091] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.

[0092] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0093] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.

[0094] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0095] Note that terms described in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Furthermore, a signal may be a message. Furthermore, a component carrier (CC) may be called a carrier frequency, a cell, a frequency carrier, etc.

[0096] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed using absolute values, relative values ​​from a predetermined value, or other corresponding information. For example, a radio resource may be indicated by an index.

[0097] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.

[0098] In this disclosure, the terms "Mobile Station (MS)," "user terminal," "User Equipment (UE)," "terminal," and the like may be used interchangeably.

[0099] A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other suitable terminology.

[0100] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.

[0101] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.

[0102] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."

[0103] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.

[0104] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.

[0105] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.

[0106] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."

[0107] 100...Metaverse space providing device, 200...NFT issuing device, 300...Blockchain, 400...Server, 401...Memory, 500...User terminal, 501...Wallet, 101...NFT acquisition unit, 102...Art style extraction unit, 103...Art style conversion unit, 104...Metaverse space processing unit, 105...Learning unit, 106...Metaverse space DB.

Claims

1. An apparatus comprising: an extraction unit that extracts a user's image preference trends based on owned image data corresponding to the user's ownership information; and a generation unit that generates output image data from input image data based on the trends.

2. The device according to claim 1, further comprising a data acquisition unit that acquires proprietary image data, which is an image based on information recorded on a non-fungible blockchain, and the extraction unit extracts the tendency based on the proprietary image data corresponding to the proprietary information.

3. The device according to claim 2, wherein the data acquisition unit acquires image data that is currently owned or that has been owned in the past as the owned image data.

4. The device according to claim 3, wherein the owned image data is image data based on a period of time that the image data has been owned by the user.

5. The device according to claim 4, wherein the extraction unit acquires vector data from the owned image data and extracts user preference trends based on the vector data and reference data.

6. The device according to claim 5, wherein the extraction unit clusters the vector data of the plurality of owned image data to obtain a plurality of vector data of the plurality of owned image data that belong to a group that satisfies a condition, and extracts the user's preference trends based on the plurality of vector data and the reference data.

7. The device according to claim 6, wherein the extraction unit has a conversion table that associates reference data with text information, and refers to the conversion table to extract text information from a single vector data set that aggregates the plurality of vector data sets as a preference trend of the user.

8. The device according to claim 1, wherein the generation unit generates converted image data from source image data using an image transformation model, and further comprises a learning unit that learns parameters of the image transformation model based on the user's preferences.

9. The device according to claim 8, wherein the learning unit fine-tunes the image transformation model based on the user's preference trends.

10. A method comprising: an extraction step of extracting a user's image preference tendency based on owned image data corresponding to the user's ownership information; and a generation step of generating converted image data from source image data based on the tendency.

Citation Information

Patent Citations

  • Picture style conversion method, device and equipment and computer readable storage medium

    CN110599393A

  • User preference prediction method and device, equipment and storage medium

    CN111784069A

  • Method, apparatus and computer program for providing modularized artificial intelligence model platform service

    KR1020220000491A