System

A system allows users to input preferences in natural language, using AI to generate and apply custom smartphone designs, addressing the limitations of conventional customization methods by simplifying the process and enhancing user experience.

JP2026024001APending Publication Date: 2026-02-13SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024126322
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-01
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Conventional methods for designing and customizing mobile phones require users to have a high level of design sense and skill, and existing design templates and customization options are limited, making it difficult to fully address individual user needs.

Method used

A system that allows users to input preferences and images in natural language, uses keyword extraction to understand user intentions, formats the data for a generative AI model, generates original design data, and applies it to the user's terminal without special operations.

Benefits of technology

Enables users to easily create original smartphone designs by intuitively inputting preferences, with AI generating professional designs that are instantly applied, enhancing user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026024001000001_ABST
    Figure 2026024001000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for acquiring a preference or an image input by a user in a natural language; means for analyzing the acquired preference or image and extracting a keyword; means for formatting input data for a generative artificial intelligence model using the extracted keyword; means for generating original design data based on the formatted input data by the generative artificial intelligence model; means for transmitting the generated design data to a terminal of the user; and means for applying the received design data to the terminal.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional methods for designing and customizing mobile phones require users to have a high level of design sense and skill, making it difficult to easily create a design that reflects their individual preferences. Furthermore, existing design templates and customization options are limited, making it difficult to fully address individual user needs. The present invention aims to solve these problems by providing a system that allows users to easily create original designs and apply them to their mobile phones. [Means for solving the problem]

[0005] The present invention solves the above problems by the following means.

[0006] The system provides a means for acquiring preferences and images entered by a user in natural language, allowing the user to intuitively enter the design they desire. It also provides a means for analyzing the acquired preferences and images and extracting keywords, accurately grasping the user's intentions through this keyword extraction. It also provides a means for formatting input data to a generative artificial intelligence model using the extracted keywords, allowing the generative artificial intelligence model to appropriately generate design data. It also provides a means for the generative artificial intelligence model to generate original design data based on the formatted input data, automatically creating a design exclusive to the user. It also provides a means for transmitting the generated design data to the user's terminal, allowing the generated design to be quickly provided to the user. Finally, it provides a means for applying the received design data to the terminal, allowing the user to apply the design without performing any special operations.

[0007] A "user" is a person who uses the system to customize the design of a mobile phone based on their own preferences and image.

[0008] A "natural language" is a language that humans use on a daily basis, and is not a standardized command or programming language, but rather generates sentences from a combination of letters and words.

[0009] "Means for acquiring" refers to hardware and software for capturing natural language data entered by a user and incorporating it into the system.

[0010] The "analyzing means" is a part of the system that includes natural language processing technology for analyzing acquired natural language data and understanding the user's intent.

[0011] "Keywords" are important words and phrases extracted from user input, and serve as the basis for the generative AI model to generate designs.

[0012] A "generative artificial intelligence model" is a technology that includes machine learning algorithms and neural networks to generate new designs and content based on input data.

[0013] "Formatting means" refers to the process or algorithm used to convert the analyzed keywords into a format that can be processed by the generative artificial intelligence model.

[0014] "Design data" refers to digital content such as wallpapers, ringtones, and icons generated by generative artificial intelligence models.

[0015] "Transmitting means" refers to the Internet communication protocol and related hardware and software for transmitting the generated design data to the user's terminal.

[0016] The "means for applying" is a part of the system for reflecting the received design data on the user's terminal, and includes a function for setting a new design on the user's mobile phone. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] This invention is a system for customizing the design of a mobile phone based on preferences and images input by the user in natural language. This system is carried out by the cooperation of the user, the terminal, and the server.

[0039] Program processing explanation

[0040] Getting user-entered data

[0041] Using a dedicated app, the user inputs content in natural language, such as "I want a wallpaper with a spring cherry blossom theme." When the user presses the "Send" button, the device receives the input data and proceeds to the next step.

[0042] Sending data

[0043] The terminal converts the natural language data entered by the user into JSON format, and then transmits the converted JSON data to the server using the HTTPS protocol, ensuring accurate transmission of the data.

[0044] Data reception and analysis

[0045] The server receives the HTTPS request, parses the JSON data sent, and retrieves the user's input. The server uses a natural language processing (NLP) engine to extract meaning from the input text and identify keywords such as "spring," "cherry blossoms," and "wallpaper."

[0046] Data Formatting

[0047] The server then uses the extracted keywords to format the data required by the generative AI model. Specifically, it prepares keywords such as "spring," "cherry blossoms," and "wallpaper" to be input into the model in array format.

[0048] Data generation

[0049] The server inputs the formatted data into a generative AI model, which then generates original design data, such as a wallpaper with a "spring cherry blossom" theme.

[0050] Sending generated data

[0051] The server converts the generated wallpaper image data back into JSON format and prepares to send it to the device. The wallpaper image data is sent to the user's device using the HTTPS protocol.

[0052] Receiving and applying data

[0053] The device receives the wallpaper image data and parses the JSON data to obtain the image data. The device then executes a system call to immediately apply the image data as the smartphone's wallpaper. As a result, the user's smartphone wallpaper is set to the new original "Spring Cherry Blossom" theme.

[0054] Specific examples

[0055] For example, consider a case where a user inputs, "I want a summer beach-themed icon set." When the user's input is transmitted and analyzed, the keywords "summer," "beach," and "icon" are extracted. The server uses a generative AI model to generate a new icon set based on this input and transmits the generated icon set to the device. The device then applies the received icon set to the user's smartphone.

[0056] This invention allows users to easily enjoy creating original smartphone designs simply by intuitively inputting their preferences and ideas using natural language. AI automatically generates professional designs, independent of the user's taste or skills, and instantly applies them, making everyday life more appealing.

[0057] The processing flow will be explained below.

[0058] Step 1:

[0059] User:

[0060] Users start up a dedicated smartphone app and enter their preferences and images in natural language into the input field that appears, such as "I want a wallpaper with an autumn leaf theme." Once they've finished entering the information, they press the "Send" button.

[0061] Step 2:

[0062] Device:

[0063] The terminal receives input data from the user, converts it into JSON format, and then sends the converted JSON data to the server using the HTTPS protocol.

[0064] Step 3:

[0065] server:

[0066] The server receives the HTTPS request, parses the JSON data sent, and extracts the natural language input from the parsed data.

[0067] Step 4:

[0068] server:

[0069] The server uses a natural language processing (NLP) engine to analyze the meaning of the natural language data it receives. Specifically, it extracts keywords such as "autumn," "autumn leaves," and "wallpaper" from the input "I want a wallpaper with an autumn foliage theme."

[0070] Step 5:

[0071] server:

[0072] The server then uses the extracted keywords to format the data appropriately for use as input for a generative AI model. This formatting process results in a dataset containing keywords such as "autumn," "autumn leaves," and "wallpaper."

[0073] Step 6:

[0074] server:

[0075] The formatted data is input into a generative AI model, which then generates original wallpaper data based on this data.

[0076] Step 7:

[0077] server:

[0078] The generated wallpaper data is received, converted to JSON format, and prepared for transmission to the user's device. This wallpaper data is sent to the device using the HTTPS protocol.

[0079] Step 8:

[0080] Device:

[0081] The device receives the wallpaper data sent from the server, parses the JSON data, and obtains the original wallpaper image data.

[0082] Step 9:

[0083] Device:

[0084] The acquired original wallpaper image data is immediately set as the smartphone wallpaper. This setting changes the user's smartphone wallpaper to an original one with an "Autumn Leaves" theme.

[0085] Example 1

[0086] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0087] Today's smartphone users want to be able to easily customize their devices to suit their tastes and preferences. However, existing services often make customization difficult, relying on the user's taste and skills. Furthermore, the interface for users to specify designs is often not intuitive, making it difficult to achieve high-quality customization easily. Furthermore, for users who are not technically savvy, the complex and time-consuming operation poses a problem.

[0088] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0089] In this invention, the server includes means for acquiring preferences and images input by a user in natural language, means for converting the acquired preferences and images into JSON format, means for transmitting the converted JSON format data using the HTTPS protocol, means for receiving and analyzing the transmitted JSON format data, means for extracting keywords from the analyzed data, means for formatting input data for a generative artificial intelligence model using the extracted keywords, means for generating original design data based on the formatted input data using the generative artificial intelligence model, means for converting the generated design data back into JSON format and transmitting it using the HTTPS protocol, and means for analyzing the received design data and applying it to the terminal. This allows a user to intuitively and quickly customize the design of their smartphone simply by inputting their preferences in natural language.

[0090] "Means for acquiring preferences and images input by the user in natural language" refers to a function that allows the user to input their wishes and images in natural language and receives the input content from the system.

[0091] "Means for converting to JSON format" is a function that converts natural language data entered by the user into a standardized data format called JavaScript Object Notation (JSON).

[0092] "Means for sending using the HTTPS protocol" refers to a function that uses a protocol called HTTP Secure (HTTPS) to send data to a server in a secure and encrypted manner.

[0093] "Means for receiving and analyzing transmitted JSON format data" refers to a function in which a server receives JSON data transmitted via the HTTPS protocol and analyzes the contents of that data.

[0094] "Means for extracting keywords from analyzed data" refers to a function that analyzes received data using natural language processing and identifies and extracts important keywords.

[0095] The "means for formatting input data for a generative artificial intelligence model" is a function that converts extracted keywords into data in a format that is easy for the generative artificial intelligence model to understand.

[0096] "Means for generating original design data based on input data generated by a generative artificial intelligence model" refers to a function in which a generative artificial intelligence model automatically creates design data that meets the user's wishes based on formatted input data.

[0097] "Means for converting the generated design data back into JSON format and transmitting it using the HTTPS protocol" refers to a function for converting the generated design data back into JSON format and transmitting it to the terminal using the HTTPS protocol.

[0098] The "means for analyzing received design data and applying it to the terminal" is a function that analyzes the JSON format design data received by the terminal and applies the specified design to the terminal.

[0099] This invention is a system that customizes a mobile information terminal based on preferences and images input by the user in natural language. This system performs processing through collaboration between the user, the terminal, and the server. The user uses a dedicated app to input specific wishes and images in natural language. For example, the user might input, "I want a wallpaper with a spring cherry blossom theme."

[0100] The device receives natural language data entered by the user and converts it into JSON format. The converted data is then sent to the server using the HTTPS protocol. The server analyzes the received JSON data and extracts keywords using a natural language processing engine (libraries such as TensorFlow and PyTorch). The extracted keywords are then used to format the input data for a generative AI model, which then generates original design data.

[0101] The generated design data is converted back to JSON format and sent to the device using the HTTPS protocol. The device then analyzes the received data and extracts the image data. Finally, the device applies the extracted image data as wallpaper or icons on the smartphone.

[0102] For example, if a user types "I want a summer beach-themed icon set," the system will extract the keywords "summer," "beach," and "icon," and the AI ​​model will generate a new icon set, resulting in a summer beach-themed icon set that will be applied to the user's smartphone.

[0103] Examples of specific prompts include:

[0104] "I want a wallpaper with a spring cherry blossom theme."

[0105] "I want a summer beach themed icon set."

[0106] "Create an autumn leaf-themed lock screen"

[0107] Based on these prompts, the generative AI model creates a design that meets the user's wishes. Users can easily enjoy creating their own original smartphone design by simply inputting their wishes in natural language.

[0108] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0109] Step 1:

[0110] The user opens the app and enters a prompt in natural language, such as "I want a wallpaper with a spring cherry blossom theme," into the input field. When the user presses the "Send" button, the input data is sent to the device.

[0111] Input: A natural language prompt entered by the user.

[0112] Output: The prompt is sent to the terminal.

[0113] Step 2:

[0114] The device receives the natural language data entered by the user and converts it into JavaScript Object Notation (JSON) format, such as "message: 'I want a wallpaper with a spring cherry blossom theme'".

[0115] Input: A prompt sentence typed in natural language.

[0116] Output: Data converted to JSON format.

[0117] Step 3:

[0118] The device sends the converted JSON data to the server using the HTTPS protocol, which may also include data integrity checks and retransmission mechanisms.

[0119] Input: Data converted to JSON format.

[0120] Output: Data sent to the server using the HTTPS protocol.

[0121] Step 4:

[0122] The server receives the HTTPS request and analyzes the JSON data sent. Specifically, it parses the received data and obtains the contents of the "message" field.

[0123] Input: JSON data sent over HTTPS.

[0124] Output: Parsed content (prompt sentence).

[0125] Step 5:

[0126] The server uses a natural language processing engine (such as TensorFlow or PyTorch) to extract keywords from the prompt. In this example, "spring," "cherry blossoms," and "wallpaper" are extracted.

[0127] Input: The parsed prompt sentence.

[0128] Output: Extracted keywords.

[0129] Step 6:

[0130] The server formats the data based on the extracted keywords into a format that the generative AI model can understand, such as a JSON object with the format "keywords: ['spring', 'cherry blossoms', 'wallpaper']".

[0131] Input: Extracted keywords.

[0132] Output: The formatted data.

[0133] Step 7:

[0134] The server inputs the formatted data into a generative AI model, which then generates original design data (e.g., wallpaper images) based on the data.

[0135] Input: Formatted data.

[0136] Output: Generated design data (wallpaper image).

[0137] Step 8:

[0138] The server converts the generated design data into JSON format and sends it to the terminal using the HTTPS protocol.

[0139] Input: Generated design data.

[0140] Output: Design data converted to JSON format.

[0141] Step 9:

[0142] The device receives the sent JSON data, parses it, and obtains the image data. Specifically, it extracts the "image" field from the JSON data and decodes the Base64-encoded image.

[0143] Input: Design data submitted in JSON format.

[0144] Output: Decoded image data.

[0145] Step 10:

[0146] The device executes a system call to apply the acquired image data as the smartphone's wallpaper. Ultimately, the user's smartphone wallpaper is changed to the new original "Spring Cherry Blossom" theme.

[0147] Input: Decoded image data.

[0148] Output: The new wallpaper applied.

[0149] (Application example 1)

[0150] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0151] In virtual stores, there are inconvenient ways to quickly and intuitively reflect the preferences and images that users input in natural language and customize the store design, so there is a need for technology that can generate attractive designs in real time based on user input and apply them immediately.Furthermore, there is also a need to easily save and share the generated designs.

[0152] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0153] In this invention, the server includes means for acquiring preferences and images input by a user in natural language, means for analyzing the acquired preferences and images and extracting keywords, means for shaping input data to a generative artificial intelligence model using the extracted keywords, means for generating original design data based on the shaped input data using the generative artificial intelligence model, means for transmitting the generated design data to the user's terminal, means for applying the received design data to the terminal, means for applying the generated design in real time in a virtual space, and means for saving and sharing the generated design. This allows a user to intuitively customize the design of a virtual store in real time based on input in natural language, and to save and share the generated design.

[0154] "Natural language" refers to the language that a user normally uses in conversation or writing, without being converted into a form that is easy for the system to parse.

[0155] "Preferences and images" refers to the user's personal hobbies and visual concepts / themes.

[0156] "Keyword extraction" refers to the process of extracting meaningful words and phrases from a user's natural language input.

[0157] A "generative artificial intelligence model" refers to an artificial intelligence that has the ability to automatically generate new designs and content based on input data and conditions.

[0158] "Formatting input data" refers to the process of converting data into a format that is easy for a generative artificial intelligence model to understand.

[0159] "Original design data" refers to unique design information that is newly generated by a generative artificial intelligence model and does not exist anywhere else.

[0160] "Sending to the user's device" refers to sending the generated design data to the device used by the user using a communication means such as the Internet.

[0161] "Applying design data" refers to reflecting the received design data on the user's device and making it actually displayable and usable.

[0162] "Virtual space" refers to a 3D space or virtual store generated by a computer.

[0163] "Real-time application" refers to the process of updating and displaying the design so that it reflects user input almost immediately.

[0164] "Saving a design" refers to saving the generated design data to a storage medium for later use.

[0165] "Means of sharing" refers to the methods by which the generated design can be published and transferred to other users or platforms.

[0166] This invention is a system that customizes the design of a virtual store based on preferences and images input by the user in natural language. This system executes processing in cooperation with the user, the terminal, and the server.

[0167] Using a smartphone, smart glasses, or head-mounted display, users can request a virtual store design in natural language, for example, by typing, "I want a store with an autumn leaf theme."

[0168] The device converts the user's input into JSON format and sends it to the server using the HTTPS protocol. The server analyzes the received JSON data and uses a natural language processing engine to extract keywords such as "autumn," "autumn leaves," and "store." It then formats these keywords into the format required by the generative artificial intelligence model and inputs them into the model.

[0169] The generative AI model generates original design data based on the given keywords. This design data is then converted back to JSON format and sent to the user's device via HTTPS. The device parses and retrieves the received design data and immediately applies it to the virtual store.

[0170] Furthermore, the generated designs can be saved on the device and shared with other users and platforms, allowing users to intuitively make design requests in natural language, customize the design of their virtual store in real time, and even save and share the designs.

[0171] The hardware requires a smartphone, smart glasses, or a head-mounted display to receive user input, and the software uses Python 3.x, the Requests library, and specific API endpoints to apply the virtual store design.

[0172] For example, if a user inputs "I want a Halloween-themed store," the server extracts the keywords "Halloween," "theme," and "store" and inputs them into a generative AI model. Based on this prompt, the AI ​​model generates a virtual store design, which is then instantly applied.

[0173] Example prompt sentence:

[0174] Theme: Halloween

[0175] Design: Pumpkins, bats, dark colors

[0176] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0177] Step 1:

[0178] The user inputs a design request for the virtual store in natural language. For example, they might input, "I want a store with an autumn foliage theme." The input is sent to the terminal.

[0179] Step 2:

[0180] The terminal receives the user's input and converts it into JSON format. Specifically, it converts it into a JSON object with the natural language input as a key. The converted JSON data is then sent to the server.

[0181] Step 3:

[0182] The server receives the JSON data sent from the device and analyzes it using a natural language processing engine. Keywords such as "autumn," "autumn leaves," and "store" are extracted from the input data. The results of this analysis become the input for the next processing step.

[0183] Step 4:

[0184] Based on the extracted keywords, the server formats the data to be input into the generative AI model. Specifically, it converts the keywords into a list format and generates a prompt sentence. The keywords "autumn," "autumn leaves," and "store" become the elements of the prompt sentence.

[0185] Step 5:

[0186] The server inputs the formatted data into a generative AI model, which generates original design data based on the prompt. This generated design data becomes the input for the next processing step.

[0187] Step 6:

[0188] The server converts the generated design data back into JSON format and sends it to the terminal using the HTTPS protocol. Specifically, it encodes the design data into a JSON object and sends a POST request to the endpoint.

[0189] Step 7:

[0190] The device receives the JSON data sent from the server and parses it to obtain the design data. The parsing process extracts the JSON format data as specific design data.

[0191] Step 8:

[0192] The terminal immediately applies the acquired design data to the virtual store by sending a request to the API endpoint of the virtual space to reflect the design data.

[0193] Step 9:

[0194] The device provides functions for saving the generated design and sharing it with other users or platforms. Specifically, it saves the design data in local storage and generates a link or QR code for sharing it.

[0195] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0196] This invention relates to a system that realizes advanced design customization by combining preferences and images input by the user in natural language with an emotion engine that recognizes the user's emotions. This system performs processing in cooperation with the user, the terminal, and the server.

[0197] Program processing explanation

[0198] Getting user-entered data

[0199] Using a dedicated app, users input content in natural language, such as "I want a wallpaper with a spring cherry blossom theme." When the user presses the "send" button, the device receives the input data and proceeds to the next step. The device is also equipped with a camera and microphone, and the emotion engine analyzes the user's facial expressions and tone of voice, capturing their emotions as data.

[0200] Sending data

[0201] The device converts the input data from the user into JSON format. At the same time, the emotion data recognized by the emotion engine is also added to the JSON format. The converted JSON data is then sent to the server using the HTTPS protocol.

[0202] Data reception and analysis

[0203] The server receives the HTTPS request and parses the JSON data sent. From the parsed data, it extracts the natural language content and emotion data of the user's input. Using a natural language processing (NLP) engine, it analyzes the semantics of the natural language data and identifies keywords such as "spring," "cherry blossoms," and "wallpaper." The emotion data is also analyzed to extract the user's emotional state (e.g., joy, sadness, surprise, etc.) at the time of input.

[0204] Data Formatting

[0205] The server then uses the extracted keywords and recognized emotions to format the data appropriately for use as input for a generative AI model. For example, it generates a dataset that combines keywords such as "spring," "cherry blossoms," and "wallpaper" with the user's emotion of "happy."

[0206] Data generation

[0207] The server inputs the formatted data into a generative AI model, which then generates original design data, such as wallpaper data with the theme of "joyful spring cherry blossoms."

[0208] Sending generated data

[0209] The server receives the generated wallpaper data, converts it back to JSON format, and prepares it for transmission to the user's device. This wallpaper data is sent to the device using the HTTPS protocol.

[0210] Receiving and applying data

[0211] The device receives the wallpaper data sent from the server, parses the JSON data, and obtains the original wallpaper image data. The device immediately sets the image data as the smartphone's wallpaper. This setting changes the user's smartphone wallpaper to the original wallpaper with the theme "Joyful Spring Cherry Blossoms."

[0212] Specific examples

[0213] For example, consider the case where a user inputs "I want an icon set with a summer beach theme," and at the same time, the emotion engine recognizes the user's "excitement." The acquired data includes the keywords "summer," "beach," and "icon set," as well as the emotion "excitement." Based on this data, the server uses a generative AI model to generate a new icon set with the theme of "exciting summer beach" and sends it to the device. The device immediately applies the received icon set, and the user's smartphone is updated with the new icon set.

[0214] This invention allows users to intuitively input their preferences and ideas using natural language, and easily experience original smartphone designs that reflect their emotions. AI and emotion recognition technology work together to generate professional designs that can be instantly applied, without relying on the user's taste or skills, making everyday life more fulfilling.

[0215] The processing flow will be explained below.

[0216] Program processing explanation

[0217] Step 1:

[0218] User:

[0219] The user launches a dedicated smartphone app and enters their preferences and image in natural language, such as "I want a wallpaper with an autumn leaf theme," into the input field that appears. When the user presses the "Send" button, the emotion engine begins analyzing the user's emotions through their voice and facial expressions. The emotion engine generates emotion data such as "happy" or "excited" from the user's facial expressions and tone of voice.

[0220] Step 2:

[0221] Device:

[0222] The device receives input data from the user and converts it into JSON format. At the same time, the emotion data output by the emotion engine is also added to the JSON format. The converted JSON data is structured to include the user's input and emotion data. It is then sent to the server using the HTTPS protocol.

[0223] Step 3:

[0224] server:

[0225] The server receives the HTTPS request and parses the JSON data. From the parsed data, it extracts the natural language content and emotion data entered by the user. Specifically, it extracts the text data "I want a wallpaper with an autumn leaf theme" and the emotion data "I'm excited."

[0226] Step 4:

[0227] server:

[0228] The server uses a natural language processing (NLP) engine to analyze the meaning of the acquired natural language data. Keywords such as "autumn," "autumn leaves," and "wallpaper" are extracted from the text data. Emotion data is analyzed in the same way, and the keyword "excitement" is extracted.

[0229] Step 5:

[0230] server:

[0231] Based on the extracted keywords and recognized emotions, the server formats the data into an appropriate format for use as input for the generative AI model. Specifically, it generates a dataset consisting of the keywords "autumn," "autumn leaves," and "wallpaper" and the emotion data "excitement." This dataset is used as input for the generative AI model.

[0232] Step 6:

[0233] server:

[0234] The formatted data set is then input into a generative AI model, which then generates original design data. For example, it might generate an exciting wallpaper image featuring vivid autumn leaves.

[0235] Step 7:

[0236] server:

[0237] The generated wallpaper data is received, converted back to JSON format, and prepared for transmission to the user's device. This wallpaper data is sent to the device using the HTTPS protocol.

[0238] Step 8:

[0239] Device:

[0240] The device receives the wallpaper data sent from the server, parses the JSON data, and obtains the original wallpaper image data.

[0241] Step 9:

[0242] Device:

[0243] The acquired original wallpaper image data is instantly set as the smartphone wallpaper, so that the user's smartphone wallpaper is set to an original design with an "Autumn Leaves" theme that reflects the user's emotion of "Excitement."

[0244] Specific examples

[0245] For example, if a user inputs "I want a summer beach-themed icon set," and the emotion engine recognizes the user's "relaxation," the retrieved data includes the keywords "summer," "beach," and "icon set" along with the emotion "relaxation." This data is sent to the server, which then inputs the formatted data into a generative AI model to generate a new icon set with the theme of "relaxing summer beach." The generated icon set is then sent to the device and instantly applied to the user's smartphone.

[0246] The present invention allows users to easily enjoy original designs that reflect their own tastes and feelings, making their daily lives more enriching.

[0247] Example 2

[0248] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0249] In the field of modern digital design, there is a growing need for users to easily input their preferences and preferences in natural language and generate customized, original designs based on those inputs. However, current systems have difficulty appropriately reflecting users' emotions, extracting keywords from the input natural language, and generating designs that take emotions into account. Furthermore, the process of immediately applying the generated designs to the user's device is not efficient enough. This can sometimes result in a poor user experience.

[0250] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0251] In this invention, the server includes means for acquiring preferences and images input by a user in natural language, means for analyzing the acquired preferences and images to extract keywords, means for integrating the extracted keywords and user emotion data to format input data for a generative AI model, means for analyzing the user's facial expressions and voice using a camera and microphone to acquire emotion data, means for generating original design data based on the formatted input data using the generative AI model, means for transmitting the generated design data to the user's terminal, and means for applying the received design data to the terminal. This enables the generation of original designs that reflect the user's preferences and emotions and their immediate application.

[0252] "User" refers to an individual who uses the system to input preferences and ideas in natural language.

[0253] A "dedicated app" refers to a software application that allows users to input their preferences and images and communicate with a server.

[0254] "Terminal" refers to a device used by a user, such as a smartphone or tablet.

[0255] An "emotion engine" refers to a software module that analyzes a user's facial expressions and tone of voice to recognize their emotional state.

[0256] "JSON" stands for JavaScript Object Notation and refers to a lightweight data interchange format for structuring data and communicating it over the Internet.

[0257] "HTTPS protocol" stands for Hypertext Transfer Protocol Secure and refers to a communication protocol for securely transmitting data over the Internet.

[0258] "Server" refers to a computer system that receives, analyzes, and processes data sent by a user and sends the generated data to the user's terminal.

[0259] A "natural language processing (NLP) engine" refers to a software module that analyzes natural language input by a user and extracts meaning and keywords.

[0260] A "generative artificial intelligence model" refers to a machine learning algorithm for generating original design data based on input data.

[0261] "Design Data" refers to original visual or functional designs generated based on user input and emotional data.

[0262] This invention is a system that realizes more advanced design customization by combining preferences and images input by the user in natural language with an emotion engine that recognizes the user's emotions. This system is carried out by the cooperation of the user, the terminal, and the server.

[0263] First, a user launches a dedicated app on a device such as a smartphone or tablet and inputs the desired design in natural language. For example, they might input something like, "I want a wallpaper with a spring cherry blossom theme." When the user presses the "send" button, the device captures this text data. At the same time, an emotion engine is activated on the device, which has a built-in camera and microphone, and analyzes the user's facial expressions and tone of voice to capture emotional data in real time. This emotional data includes emotional states such as "joy" and "excitement."

[0264] The device converts the acquired data (natural language data and emotion data) into JSON format, and then transmits this JSON-formatted dataset to the server using the HTTPS protocol, which ensures data security and confidentiality.

[0265] The server receives the HTTPS request and parses the JSON data sent. The analysis module extracts the natural language and sentiment data entered by the user from the parsed data. Using a natural language processing (NLP) engine, it identifies keywords such as "spring," "cherry blossoms," and "wallpaper," and then uses a sentiment analysis algorithm to extract the emotional state (e.g., "joy").

[0266] The server then combines the extracted keywords and emotion data and formats them into a format appropriate for the generative AI model. For example, it combines the keywords "spring," "cherry blossoms," and "wallpaper" with "joy" to create a prompt sentence to be input into the generative AI model.

[0267] Based on the input prompts, the generative AI model generates original design data that reflects the user's preferences and emotions. Specifically, wallpaper data with the theme of "joyful spring cherry blossoms" is generated.

[0268] The generated design data is converted back to JSON format by the server and sent to the user's device using the HTTPS protocol. The device receives this data, parses the JSON data, and obtains the original wallpaper image data. The device then immediately sets the obtained image data as the smartphone's wallpaper. As a result, the user's smartphone wallpaper is changed to the original wallpaper with the theme "Joyful Spring Cherry Blossoms."

[0269] As a concrete example, consider the case where a user inputs "I want a summer beach-themed icon set," and the emotion engine recognizes the user's "excitement." The acquired dataset contains the keywords "summer," "beach," and "icon set," as well as the emotion "excitement." The server inputs this data into the generative AI model and generates a new icon set with the theme of "exciting summer beach." This icon set is sent to the device, which immediately applies the received icon set, and the user's smartphone is updated with the new icon set.

[0270] In this way, users can intuitively input their preferences and images using natural language, and easily experience original smartphone designs that reflect their emotions.This system makes it possible to generate professional designs and apply them immediately, without relying on the user's taste or skills.

[0271] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0272] Step 1: Getting user-entered data

[0273] Specific explanation

[0274] Using a dedicated app, users input their preferences and images in natural language, such as "I want a wallpaper with a spring cherry blossom theme." When the user presses the "Send" button, the device acquires this text data.

[0275] input

[0276] User natural language input (e.g., "I want a wallpaper with a spring cherry blossom theme")

[0277] output

[0278] The text data is saved on the device.

[0279] concrete action

[0280] The terminal captures the user input in text format and stores it in a variable.

[0281] Step 2: Acquire emotion data using a camera or microphone

[0282] Specific explanation

[0283] While the user is typing, the device's camera and microphone are activated to analyze the user's facial expressions and tone of voice in real time. An emotion engine processes this data to determine the user's emotional state.

[0284] input

[0285] User video and audio data

[0286] output

[0287] Analyzed emotion data (e.g., "joy," "excitement")

[0288] concrete action

[0289] The camera and microphone are activated to capture video and audio in real time, and the emotion engine analyzes the data and outputs a specific emotional state.

[0290] Step 3: Convert the data to JSON format

[0291] Specific explanation

[0292] The natural language data and emotion data acquired by the device is converted into JSON format.

[0293] input

[0294] Natural language input data, emotion data

[0295] output

[0296] JSON formatted dataset

[0297] concrete action

[0298] Natural language data and sentiment data are integrated and converted into JSON format programmatically.

[0299] Step 4: Sending data to the server

[0300] Specific explanation

[0301] The terminal sends the converted data set to the server using the HTTPS protocol.

[0302] input

[0303] JSON formatted dataset

[0304] output

[0305] The server receives the JSON-formatted dataset.

[0306] concrete action

[0307] The device generates an HTTPS request and sends data to the server's URL.

[0308] Step 5: Receiving and analyzing data on the server

[0309] Specific explanation

[0310] The server parses the received JSON data and extracts natural language and sentiment data. The NLP engine identifies keywords from the natural language data, and the sentiment analysis algorithm extracts the emotional state.

[0311] input

[0312] JSON formatted dataset

[0313] output

[0314] Identified keywords (e.g., "spring," "cherry blossoms," "wallpaper"), emotional states (e.g., "joy")

[0315] concrete action

[0316] The server parses the JSON data and analyzes it using an NLP engine and sentiment analysis algorithms.

[0317] Step 6: Format the data

[0318] Specific explanation

[0319] The server integrates the extracted keywords and sentiment data to generate appropriate prompts for the generative AI model.

[0320] input

[0321] Identified keywords, emotional states

[0322] output

[0323] Prompt statement (e.g., "Spring cherry blossom wallpaper that inspires joy")

[0324] concrete action

[0325] The server combines the keywords and emotion data and formats them into a prompt sentence.

[0326] Step 7: Generate original design data

[0327] Specific explanation

[0328] The server inputs the prompt sentence into the generative AI model and generates design data.

[0329] input

[0330] Prompt statement

[0331] output

[0332] Generated original design data (e.g., "Joyful Spring Cherry Blossom Wallpaper")

[0333] concrete action

[0334] The generative AI model generates design data that fits the theme based on the prompt text.

[0335] Step 8: Convert your design data to JSON format and prepare it for submission

[0336] Specific explanation

[0337] The generated design data is converted back into JSON format and prepared for transmission to the user's device using the HTTPS protocol.

[0338] input

[0339] Generated design data

[0340] output

[0341] Design data in JSON format

[0342] concrete action

[0343] The server converts the generated design data into JSON format and prepares it for transmission via HTTPS protocol.

[0344] Step 9: Sending and receiving design data to and from your device

[0345] Specific explanation

[0346] The server sends JSON format design data to the device, which then receives it.

[0347] input

[0348] Design data in JSON format

[0349] output

[0350] The device receives the design data

[0351] concrete action

[0352] The server sends the design data to the terminal via an HTTPS request, and the terminal receives it.

[0353] Step 10: Applying design data

[0354] Specific explanation

[0355] The device parses the received design data and obtains the original wallpaper image data, which the device then instantly sets as the smartphone's wallpaper.

[0356] input

[0357] Design data in JSON format

[0358] output

[0359] The device wallpaper has been changed

[0360] concrete action

[0361] The device parses the received JSON data, obtains the wallpaper image data, and immediately sets it as the wallpaper.

[0362] (Application example 2)

[0363] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0364] When workers use industrial equipment, they need specific work procedures and guidance. However, there are no systems that can flexibly respond to the worker's emotions and situation. In particular, emotions such as tension and anxiety can affect work efficiency, so appropriate support is needed.

[0365] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing preferences, images, and emotions input by the user in natural language and extracting keywords, means for shaping input data to a generative AI model using the extracted keywords and emotions, and means for generating original design data based on the shaped input data by the generative AI model. This makes it possible to generate and display customized work guides in real time according to the instructions and emotions input by the worker in natural language.

[0366] A "natural language" is a human language that is used by users in normal conversation and writing, and that can be converted into a form that a computer can parse.

[0367] "Emotion" refers to a psychological state that can be recognized from a user's facial expression, tone of voice, etc., and includes, for example, states such as joy, sadness, surprise, and tension.

[0368] "Keywords" refer to important information or characteristic terms extracted from a user's natural language input and sentiment data.

[0369] A "generative artificial intelligence model" refers to an algorithm or system that generates new content or designs based on input data.

[0370] "Original design data" refers to new and unique design data generated by a generative artificial intelligence model based on a user's input and emotional state.

[0371] "Terminal" refers to electronic devices used by users, such as smartphones, tablets, and personal computers.

[0372] "Industrial equipment" refers to machinery and equipment used in manufacturing and production processes, including equipment operated by workers.

[0373] "Work guide" means visual or textual instructions that show a worker how to perform a particular work task.

[0374] MODE FOR CARRYING OUT THE INVENTION

[0375] The system embodying this invention is configured to provide customized operation guides based on the user's instructions and emotions. It has the function of analyzing preferences, images, and emotion data input by the user in natural language, and generating and displaying appropriate operation guides in real time.

[0376] Hardware and Software Configuration

[0377] The system mainly consists of the following hardware and software:

[0378] Hardware

[0379] Smart glasses or head-mounted display (HMD): A device that allows users to provide natural language instructions and emotional input.

[0380] Camera and microphone: Devices for analyzing the user's facial expressions and tone of voice to obtain emotional data.

[0381] software

[0382] Emotion recognition engine (e.g., Microsoft Azure Cognitive Services Emotion API): Analyzes emotions from the user's facial expressions and tone of voice and obtains data.

[0383] Natural language processing (NLP) engine (e.g., Google Cloud Natural Language API): Analyzes the user's natural language input and extracts keywords.

[0384] Generative AI models (e.g., OpenAI GPT-4): Generate customized work guides and design data based on user keywords and emotional data.

[0385] System Operation

[0386] The system operates as follows.

[0387] User interface: The user wears smart glasses and inputs specific instructions for the task in natural language. For example, they might give instructions such as, "Tell me how to tighten a bolt." At this time, the emotion recognition engine analyzes the user's facial expressions and tone of voice via a camera and microphone, and obtains emotional data such as tension or anxiety.

[0388] Data analysis: The natural language data and emotion data entered by the user are acquired and converted into JSON format. This allows the data sent to the server to be processed in a unified format. The server receives the JSON data, analyzes it using a natural language processing engine, and extracts keywords and emotion data.

[0389] Design data generation: The server uses a generative AI model to generate original work guides and design data based on the extracted keywords and emotion data. This generated data includes appropriate procedures and guides that take emotion into consideration.

[0390] Data application: The generated work guide data is converted back to JSON format and sent to the user's device. The smart glasses parse the received design data and apply it to the user's vision. For example, the bolt tightening procedure is displayed as an illustrated guide in the user's field of vision.

[0391] Specific examples

[0392] For example, if a worker says in natural language, "Tell me how to tighten a bolt," and the emotion recognition engine detects that the user is nervous, the server analyzes this data and uses a generative AI model to generate "bolt tightening procedures to relieve tension." The worker's smart glasses display easy-to-understand procedures and illustrated guides that instill a sense of security.

[0393] Prompt Sentence Examples

[0394] User Instructions: What is the procedure for tightening the bolts?

[0395] User Emotion: Tension

[0396] This prompt enables the generative AI model to provide optimal task guidance.

[0397] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0398] Step 1:

[0399] The user puts on the smart glasses and inputs specific work instructions in natural language. For example, they might say, "Tell me how to tighten a bolt." At this time, the user's facial expressions and tone of voice are recorded via a camera and microphone. The input is captured as natural language data and sent to an emotion recognition engine. The acquired natural language data, facial expression data, and voice data are then passed to the respective processing modules, where analysis begins.

[0400] Step 2:

[0401] The device's emotion recognition engine analyzes the acquired facial and voice data to identify the user's emotions. For example, if the user is nervous, the device identifies that state of tension and records it as emotion data. The input facial and voice data is converted into emotion data. The output emotion data is "tension."

[0402] Step 3:

[0403] The device converts the user's natural language input data and the acquired emotion data into JSON format. A JSON object is generated with the natural language data and emotion data stored in their respective fields. The generated JSON data is ready to be sent to the server in the next step.

[0404] Step 4:

[0405] The device sends the generated JSON data to the server using the HTTPS protocol. The JSON data, including the user's input amount and emotion data, is securely sent to the server. The server receives the HTTPS request and proceeds to the next step.

[0406] Step 5:

[0407] The server parses the received JSON data and extracts natural language data and emotion data separately. Keywords such as "bolt," "tightening," and "procedure" and the emotion "tension" are obtained from the parsed data. The NLP engine on the server analyzes the meaning and organizes the keywords.

[0408] Step 6:

[0409] The server creates a prompt for the generative AI model based on the extracted keywords and emotion data. The prompt is formatted, for example, as "Tell me how to tighten a bolt," including the word "tension." The prompt is sent to the generative AI model, which then begins generating design data.

[0410] Step 7:

[0411] Based on the prompt received, the generative AI model generates original work guides and design data that take emotions into account. For example, it generates "bolt tightening procedures to relieve tension." The generated design data is then returned to the server.

[0412] Step 8:

[0413] The server converts the generated design data back into JSON format and prepares it for transmission to the device. The data organized into JSON objects is sent to the device via HTTPS protocol.

[0414] Step 9:

[0415] The device parses the JSON data received from the server and extracts the generated work guide and design data, which the device receives, parses, and converts into visual information such as an illustration guide.

[0416] Step 10:

[0417] The device then displays the received design data on the smart glasses display, and the user is shown a "bolt tightening procedure to ease tension," allowing them to proceed with the work with peace of mind.

[0418] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0419] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0420] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0421] [Second embodiment]

[0422] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0423] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0424] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0425] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0426] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0427] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0428] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0429] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0430] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0431] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0432] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0433] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0434] This invention is a system for customizing the design of a mobile phone based on preferences and images input by the user in natural language. This system is carried out by the cooperation of the user, the terminal, and the server.

[0435] Program processing explanation

[0436] Getting user-entered data

[0437] Using a dedicated app, the user inputs content in natural language, such as "I want a wallpaper with a spring cherry blossom theme." When the user presses the "Send" button, the device receives the input data and proceeds to the next step.

[0438] Sending data

[0439] The terminal converts the natural language data entered by the user into JSON format, and then transmits the converted JSON data to the server using the HTTPS protocol, ensuring accurate transmission of the data.

[0440] Data reception and analysis

[0441] The server receives the HTTPS request, parses the JSON data sent, and retrieves the user's input. The server uses a natural language processing (NLP) engine to extract meaning from the input text and identify keywords such as "spring," "cherry blossoms," and "wallpaper."

[0442] Data Formatting

[0443] The server then uses the extracted keywords to format the data required by the generative AI model. Specifically, it prepares keywords such as "spring," "cherry blossoms," and "wallpaper" to be input into the model in array format.

[0444] Data generation

[0445] The server inputs the formatted data into a generative AI model, which then generates original design data, such as a wallpaper with a "spring cherry blossom" theme.

[0446] Sending generated data

[0447] The server converts the generated wallpaper image data back into JSON format and prepares to send it to the device. The wallpaper image data is sent to the user's device using the HTTPS protocol.

[0448] Receiving and applying data

[0449] The device receives the wallpaper image data and parses the JSON data to obtain the image data. The device then executes a system call to immediately apply the image data as the smartphone's wallpaper. As a result, the user's smartphone wallpaper is set to the new original "Spring Cherry Blossom" theme.

[0450] Specific examples

[0451] For example, consider a case where a user inputs, "I want a summer beach-themed icon set." When the user's input is transmitted and analyzed, the keywords "summer," "beach," and "icon" are extracted. The server uses a generative AI model to generate a new icon set based on this input and transmits the generated icon set to the device. The device then applies the received icon set to the user's smartphone.

[0452] This invention allows users to easily enjoy creating original smartphone designs simply by intuitively inputting their preferences and ideas using natural language. AI automatically generates professional designs, independent of the user's taste or skills, and instantly applies them, making everyday life more appealing.

[0453] The processing flow will be explained below.

[0454] Step 1:

[0455] User:

[0456] Users start up a dedicated smartphone app and enter their preferences and images in natural language into the input field that appears, such as "I want a wallpaper with an autumn leaf theme." Once they've finished entering the information, they press the "Send" button.

[0457] Step 2:

[0458] Device:

[0459] The terminal receives input data from the user, converts it into JSON format, and then sends the converted JSON data to the server using the HTTPS protocol.

[0460] Step 3:

[0461] server:

[0462] The server receives the HTTPS request, parses the JSON data sent, and extracts the natural language input from the parsed data.

[0463] Step 4:

[0464] server:

[0465] The server uses a natural language processing (NLP) engine to analyze the meaning of the natural language data it receives. Specifically, it extracts keywords such as "autumn," "autumn leaves," and "wallpaper" from the input "I want a wallpaper with an autumn foliage theme."

[0466] Step 5:

[0467] server:

[0468] The server then uses the extracted keywords to format the data appropriately for use as input for a generative AI model. This formatting process results in a dataset containing keywords such as "autumn," "autumn leaves," and "wallpaper."

[0469] Step 6:

[0470] server:

[0471] The formatted data is input into a generative AI model, which then generates original wallpaper data based on this data.

[0472] Step 7:

[0473] server:

[0474] The generated wallpaper data is received, converted to JSON format, and prepared for transmission to the user's device. This wallpaper data is sent to the device using the HTTPS protocol.

[0475] Step 8:

[0476] Device:

[0477] The device receives the wallpaper data sent from the server, parses the JSON data, and obtains the original wallpaper image data.

[0478] Step 9:

[0479] Device:

[0480] The acquired original wallpaper image data is immediately set as the smartphone wallpaper. This setting changes the user's smartphone wallpaper to an original one with an "Autumn Leaves" theme.

[0481] Example 1

[0482] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0483] Today's smartphone users want to be able to easily customize their devices to suit their tastes and preferences. However, existing services often make customization difficult, relying on the user's taste and skills. Furthermore, the interface for users to specify designs is often not intuitive, making it difficult to achieve high-quality customization easily. Furthermore, for users who are not technically savvy, the complex and time-consuming operation poses a problem.

[0484] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0485] In this invention, the server includes means for acquiring preferences and images input by a user in natural language, means for converting the acquired preferences and images into JSON format, means for transmitting the converted JSON format data using the HTTPS protocol, means for receiving and analyzing the transmitted JSON format data, means for extracting keywords from the analyzed data, means for formatting input data for a generative artificial intelligence model using the extracted keywords, means for generating original design data based on the formatted input data using the generative artificial intelligence model, means for converting the generated design data back into JSON format and transmitting it using the HTTPS protocol, and means for analyzing the received design data and applying it to the terminal. This allows a user to intuitively and quickly customize the design of their smartphone simply by inputting their preferences in natural language.

[0486] "Means for acquiring preferences and images input by the user in natural language" refers to a function that allows the user to input their wishes and images in natural language and receives the input content from the system.

[0487] "Means for converting to JSON format" is a function that converts natural language data entered by the user into a standardized data format called JavaScript Object Notation (JSON).

[0488] "Means for sending using the HTTPS protocol" refers to a function that uses a protocol called HTTP Secure (HTTPS) to send data to a server in a secure and encrypted manner.

[0489] "Means for receiving and analyzing transmitted JSON format data" refers to a function in which a server receives JSON data transmitted via the HTTPS protocol and analyzes the contents of that data.

[0490] "Means for extracting keywords from analyzed data" refers to a function that analyzes received data using natural language processing and identifies and extracts important keywords.

[0491] "Means for formatting input data for a generative artificial intelligence model" is a function that converts extracted keywords into data in a format that is easy for the generative artificial intelligence model to understand.

[0492] "Means for generating original design data based on input data generated by a generative artificial intelligence model" refers to a function in which a generative artificial intelligence automatically creates design data that meets the user's wishes based on formatted input data.

[0493] "Means for converting the generated design data back into JSON format and transmitting it using the HTTPS protocol" refers to a function for converting the generated design data back into JSON format and transmitting it to the terminal using the HTTPS protocol.

[0494] The "means for analyzing received design data and applying it to the terminal" is a function that analyzes the JSON format design data received by the terminal and applies the specified design to the terminal.

[0495] This invention is a system that customizes a mobile information terminal based on preferences and images input by the user in natural language. This system performs processing through collaboration between the user, the terminal, and the server. The user uses a dedicated app to input specific wishes and images in natural language. For example, the user might input, "I want a wallpaper with a spring cherry blossom theme."

[0496] The device receives natural language data entered by the user and converts it into JSON format. The converted data is then sent to the server using the HTTPS protocol. The server analyzes the received JSON data and extracts keywords using a natural language processing engine (libraries such as TensorFlow and PyTorch). The extracted keywords are then used to format the input data for a generative AI model, which then generates original design data.

[0497] The generated design data is converted back to JSON format and sent to the device using the HTTPS protocol. The device then analyzes the received data and extracts the image data. Finally, the device applies the extracted image data as the smartphone's wallpaper or icon.

[0498] For example, if a user types "I want a summer beach-themed icon set," the system will extract the keywords "summer," "beach," and "icon," and the AI ​​model will generate a new icon set, resulting in a summer beach-themed icon set that will be applied to the user's smartphone.

[0499] Examples of specific prompts include:

[0500] "I want a wallpaper with a spring cherry blossom theme."

[0501] "I want a summer beach themed icon set."

[0502] "Create an autumn leaf-themed lock screen"

[0503] Based on these prompts, the generative AI model creates a design that meets the user's wishes. Users can easily enjoy creating their own original smartphone design by simply inputting their wishes in natural language.

[0504] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0505] Step 1:

[0506] The user opens the app and enters a prompt in natural language, such as "I want a spring cherry blossom themed wallpaper," into the input field. When the user presses the "Send" button, the input data is sent to the device.

[0507] Input: A natural language prompt entered by the user.

[0508] Output: The prompt is sent to the terminal.

[0509] Step 2:

[0510] The device receives the natural language data entered by the user and converts it into JavaScript Object Notation (JSON) format, such as "message: 'I want a wallpaper with a spring cherry blossom theme'".

[0511] Input: A prompt sentence typed in natural language.

[0512] Output: Data converted to JSON format.

[0513] Step 3:

[0514] The device sends the converted JSON data to the server using the HTTPS protocol, which may also include data integrity checks and retransmission mechanisms.

[0515] Input: Data converted to JSON format.

[0516] Output: Data sent to the server using the HTTPS protocol.

[0517] Step 4:

[0518] The server receives the HTTPS request and analyzes the JSON data sent. Specifically, it parses the received data and obtains the contents of the "message" field.

[0519] Input: JSON data sent over HTTPS.

[0520] Output: Parsed content (prompt sentence).

[0521] Step 5:

[0522] The server uses a natural language processing engine (such as TensorFlow or PyTorch) to extract keywords from the prompt. In this example, "spring," "cherry blossoms," and "wallpaper" are extracted.

[0523] Input: The parsed prompt sentence.

[0524] Output: Extracted keywords.

[0525] Step 6:

[0526] The server formats the data based on the extracted keywords into a format that the generative AI model can understand, such as a JSON object with the format "keywords: ['spring', 'cherry blossoms', 'wallpaper']".

[0527] Input: Extracted keywords.

[0528] Output: The formatted data.

[0529] Step 7:

[0530] The server inputs the formatted data into a generative AI model, which then generates original design data (e.g., wallpaper images) based on the data.

[0531] Input: Formatted data.

[0532] Output: Generated design data (wallpaper image).

[0533] Step 8:

[0534] The server converts the generated design data into JSON format and sends it to the terminal using the HTTPS protocol.

[0535] Input: Generated design data.

[0536] Output: Design data converted to JSON format.

[0537] Step 9:

[0538] The device receives the sent JSON data, parses it, and obtains the image data. Specifically, it extracts the "image" field from the JSON data and decodes the Base64-encoded image.

[0539] Input: Design data submitted in JSON format.

[0540] Output: Decoded image data.

[0541] Step 10:

[0542] The device executes a system call to apply the acquired image data as the smartphone's wallpaper. Ultimately, the user's smartphone wallpaper is changed to the new original "Spring Cherry Blossom" theme.

[0543] Input: Decoded image data.

[0544] Output: The new wallpaper applied.

[0545] (Application example 1)

[0546] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0547] In virtual stores, there are inconvenient ways to quickly and intuitively reflect the preferences and images that users input in natural language and customize the store design, so there is a need for technology that can generate attractive designs in real time based on user input and apply them immediately.Furthermore, there is also a need to easily save and share the generated designs.

[0548] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0549] In this invention, the server includes means for acquiring preferences and images input by a user in natural language, means for analyzing the acquired preferences and images and extracting keywords, means for shaping input data to a generative artificial intelligence model using the extracted keywords, means for generating original design data based on the shaped input data using the generative artificial intelligence model, means for transmitting the generated design data to the user's terminal, means for applying the received design data to the terminal, means for applying the generated design in real time in a virtual space, and means for saving and sharing the generated design. This allows a user to intuitively customize the design of a virtual store in real time based on input in natural language, and to save and share the generated design.

[0550] "Natural language" refers to the language that a user normally uses in conversation or writing, without being converted into a form that is easy for the system to parse.

[0551] "Preferences and images" refers to the user's personal hobbies and visual concepts / themes.

[0552] "Keyword extraction" refers to the process of extracting meaningful words and phrases from a user's natural language input.

[0553] A "generative artificial intelligence model" refers to an artificial intelligence that has the ability to automatically generate new designs and content based on input data and conditions.

[0554] "Formatting input data" refers to the process of converting data into a format that is easy for a generative artificial intelligence model to understand.

[0555] "Original design data" refers to unique design information that is newly generated by a generative artificial intelligence model and does not exist anywhere else.

[0556] "Sending to the user's device" refers to sending the generated design data to the device used by the user using a communication means such as the Internet.

[0557] "Applying design data" refers to reflecting the received design data on the user's device and making it actually displayable and usable.

[0558] "Virtual space" refers to a 3D space or virtual store generated by a computer.

[0559] "Real-time application" refers to the process of updating and displaying the design so that it reflects user input almost immediately.

[0560] "Saving a design" refers to saving the generated design data to a storage medium for later use.

[0561] "Means of sharing" refers to the methods by which the generated design can be published and transferred to other users or platforms.

[0562] This invention is a system that customizes the design of a virtual store based on preferences and images input by the user in natural language. This system executes processing in cooperation with the user, the terminal, and the server.

[0563] Using a smartphone, smart glasses, or head-mounted display, users can request a virtual store design in natural language, for example, by typing, "I want a store with an autumn leaf theme."

[0564] The device converts the user's input into JSON format and sends it to the server using the HTTPS protocol. The server analyzes the received JSON data and uses a natural language processing engine to extract keywords such as "autumn," "autumn leaves," and "store." It then formats these keywords into the format required by the generative artificial intelligence model and inputs them into the model.

[0565] The generative AI model generates original design data based on the given keywords. This design data is then converted back to JSON format and sent to the user's device via HTTPS. The device parses and retrieves the received design data and immediately applies it to the virtual store.

[0566] Furthermore, the generated designs can be saved on the device and shared with other users and platforms, allowing users to intuitively make design requests in natural language, customize the design of their virtual store in real time, and even save and share the designs.

[0567] The hardware requires a smartphone, smart glasses, or a head-mounted display to receive user input, and the software uses Python 3.x, the Requests library, and specific API endpoints to apply the virtual store design.

[0568] For example, if a user inputs "I want a Halloween-themed store," the server extracts the keywords "Halloween," "theme," and "store" and inputs them into a generative AI model. Based on this prompt, the AI ​​model generates a virtual store design, which is then instantly applied.

[0569] Example prompt sentence:

[0570] Theme: Halloween

[0571] Design: Pumpkins, bats, dark colors

[0572] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0573] Step 1:

[0574] The user inputs a design request for the virtual store in natural language. For example, they might input, "I want a store with an autumn foliage theme." The input is sent to the terminal.

[0575] Step 2:

[0576] The terminal receives the user's input and converts it into JSON format. Specifically, it converts it into a JSON object with the natural language input as a key. The converted JSON data is then sent to the server.

[0577] Step 3:

[0578] The server receives the JSON data sent from the device and analyzes it using a natural language processing engine. Keywords such as "autumn," "autumn leaves," and "store" are extracted from the input data. The results of this analysis become the input for the next processing step.

[0579] Step 4:

[0580] Based on the extracted keywords, the server formats the data to be input into the generative AI model. Specifically, it converts the keywords into a list format and generates a prompt sentence. The keywords "autumn," "autumn leaves," and "store" become the elements of the prompt sentence.

[0581] Step 5:

[0582] The server inputs the formatted data into a generative AI model, which generates original design data based on the prompt. This generated design data becomes the input for the next processing step.

[0583] Step 6:

[0584] The server converts the generated design data back into JSON format and sends it to the terminal using the HTTPS protocol. Specifically, it encodes the design data into a JSON object and sends a POST request to the endpoint.

[0585] Step 7:

[0586] The device receives the JSON data sent from the server and parses it to obtain the design data. The parsing process extracts the JSON format data as specific design data.

[0587] Step 8:

[0588] The terminal immediately applies the acquired design data to the virtual store by sending a request to the API endpoint of the virtual space to reflect the design data.

[0589] Step 9:

[0590] The device provides functions for saving the generated design and sharing it with other users or platforms. Specifically, it saves the design data in local storage and generates a link or QR code for sharing it.

[0591] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0592] This invention relates to a system that realizes advanced design customization by combining preferences and images input by the user in natural language with an emotion engine that recognizes the user's emotions. This system performs processing in cooperation with the user, the terminal, and the server.

[0593] Program processing explanation

[0594] Getting user-entered data

[0595] Using a dedicated app, users input content in natural language, such as "I want a wallpaper with a spring cherry blossom theme." When the user presses the "send" button, the device receives the input data and proceeds to the next step. The device is also equipped with a camera and microphone, and the emotion engine analyzes the user's facial expressions and tone of voice, capturing their emotions as data.

[0596] Sending data

[0597] The device converts the input data from the user into JSON format. At the same time, the emotion data recognized by the emotion engine is also added to the JSON format. The converted JSON data is then sent to the server using the HTTPS protocol.

[0598] Data reception and analysis

[0599] The server receives the HTTPS request and parses the JSON data sent. From the parsed data, it extracts the natural language content and emotion data of the user's input. Using a natural language processing (NLP) engine, it analyzes the semantics of the natural language data and identifies keywords such as "spring," "cherry blossoms," and "wallpaper." The emotion data is also analyzed to extract the user's emotional state (e.g., joy, sadness, surprise, etc.) at the time of input.

[0600] Data Formatting

[0601] The server then uses the extracted keywords and recognized emotions to format the data appropriately for use as input for a generative AI model. For example, it generates a dataset that combines keywords such as "spring," "cherry blossoms," and "wallpaper" with the user's emotion of "happy."

[0602] Data generation

[0603] The server inputs the formatted data into a generative AI model, which then generates original design data, such as wallpaper data with the theme of "joyful spring cherry blossoms."

[0604] Sending generated data

[0605] The server receives the generated wallpaper data, converts it back to JSON format, and prepares it for transmission to the user's device. This wallpaper data is sent to the device using the HTTPS protocol.

[0606] Receiving and applying data

[0607] The device receives the wallpaper data sent from the server, parses the JSON data, and obtains the original wallpaper image data. The device immediately sets the image data as the smartphone's wallpaper. This setting changes the user's smartphone wallpaper to the original wallpaper with the theme "Joyful Spring Cherry Blossoms."

[0608] Specific examples

[0609] For example, consider the case where a user inputs "I want an icon set with a summer beach theme," and at the same time, the emotion engine recognizes the user's "excitement." The acquired data includes the keywords "summer," "beach," and "icon set," as well as the emotion "excitement." Based on this data, the server uses a generative AI model to generate a new icon set with the theme of "exciting summer beach" and sends it to the device. The device immediately applies the received icon set, and the user's smartphone is updated with the new icon set.

[0610] This invention allows users to intuitively input their preferences and ideas using natural language, and easily experience original smartphone designs that reflect their emotions. AI and emotion recognition technology work together to generate professional designs that can be instantly applied, without relying on the user's taste or skills, making everyday life more fulfilling.

[0611] The processing flow will be explained below.

[0612] Program processing explanation

[0613] Step 1:

[0614] User:

[0615] The user launches a dedicated smartphone app and enters their preferences and image in natural language, such as "I want a wallpaper with an autumn leaf theme," into the input field that appears. When the user presses the "Send" button, the emotion engine begins analyzing the user's emotions through their voice and facial expressions. The emotion engine generates emotion data such as "happy" or "excited" from the user's facial expressions and tone of voice.

[0616] Step 2:

[0617] Device:

[0618] The device receives input data from the user and converts it into JSON format. At the same time, the emotion data output by the emotion engine is also added to the JSON format. The converted JSON data is structured to include the user's input and emotion data. It is then sent to the server using the HTTPS protocol.

[0619] Step 3:

[0620] server:

[0621] The server receives the HTTPS request and parses the JSON data. From the parsed data, it extracts the natural language content and emotion data entered by the user. Specifically, it extracts the text data "I want a wallpaper with an autumn leaf theme" and the emotion data "I'm excited."

[0622] Step 4:

[0623] server:

[0624] The server uses a natural language processing (NLP) engine to analyze the meaning of the acquired natural language data. Keywords such as "autumn," "autumn leaves," and "wallpaper" are extracted from the text data. Emotion data is analyzed in the same way, and the keyword "excitement" is extracted.

[0625] Step 5:

[0626] server:

[0627] Based on the extracted keywords and recognized emotions, the server formats the data into an appropriate format for use as input for the generative AI model. Specifically, it generates a dataset consisting of the keywords "autumn," "autumn leaves," and "wallpaper" and the emotion data "excitement." This dataset is used as input for the generative AI model.

[0628] Step 6:

[0629] server:

[0630] The formatted data set is then input into a generative AI model, which then generates original design data. For example, it generates an exciting wallpaper image featuring vivid autumn leaves.

[0631] Step 7:

[0632] server:

[0633] The generated wallpaper data is received, converted back to JSON format, and prepared for transmission to the user's device. This wallpaper data is sent to the device using the HTTPS protocol.

[0634] Step 8:

[0635] Device:

[0636] The device receives the wallpaper data sent from the server, parses the JSON data, and obtains the original wallpaper image data.

[0637] Step 9:

[0638] Device:

[0639] The acquired original wallpaper image data is immediately set as the smartphone wallpaper, so that the user's smartphone wallpaper is set to an original design with an "Autumn Leaves" theme that reflects the user's emotion of "Excitement."

[0640] Specific examples

[0641] For example, if a user inputs "I want a summer beach-themed icon set," and the emotion engine recognizes the user's "relaxation," the retrieved data includes the keywords "summer," "beach," and "icon set" along with the emotion "relaxation." This data is sent to the server, which then inputs the formatted data into a generative AI model to generate a new icon set with the theme of "relaxing summer beach." The generated icon set is then sent to the device and instantly applied to the user's smartphone.

[0642] The present invention allows users to easily enjoy original designs that reflect their own tastes and feelings, making their daily lives more enriching.

[0643] Example 2

[0644] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0645] In the field of modern digital design, there is a growing need for users to easily input their preferences and preferences in natural language and generate customized, original designs based on those inputs. However, current systems have difficulty appropriately reflecting users' emotions, extracting keywords from the input natural language, and generating designs that take emotions into account. Furthermore, the process of immediately applying the generated designs to the user's device is not efficient enough. This can sometimes result in a poor user experience.

[0646] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0647] In this invention, the server includes means for acquiring preferences and images input by a user in natural language, means for analyzing the acquired preferences and images to extract keywords, means for integrating the extracted keywords and user emotion data to format input data for a generative AI model, means for analyzing the user's facial expressions and voice using a camera and microphone to acquire emotion data, means for generating original design data based on the formatted input data using the generative AI model, means for transmitting the generated design data to the user's terminal, and means for applying the received design data to the terminal. This enables the generation of original designs that reflect the user's preferences and emotions and their immediate application.

[0648] "User" refers to an individual who uses the system to input preferences and ideas in natural language.

[0649] A "dedicated app" refers to a software application that allows users to input their preferences and images and communicate with a server.

[0650] "Terminal" refers to a device used by a user, such as a smartphone or tablet.

[0651] An "emotion engine" refers to a software module that analyzes a user's facial expressions and tone of voice to recognize their emotional state.

[0652] "JSON" stands for JavaScript Object Notation and refers to a lightweight data interchange format for structuring data and communicating it over the Internet.

[0653] "HTTPS protocol" stands for Hypertext Transfer Protocol Secure and refers to a communication protocol for securely transmitting data over the Internet.

[0654] "Server" refers to a computer system that receives, analyzes, and processes data sent by a user and sends the generated data to the user's terminal.

[0655] A "natural language processing (NLP) engine" refers to a software module that analyzes natural language input by a user and extracts meaning and keywords.

[0656] A "generative artificial intelligence model" refers to a machine learning algorithm for generating original design data based on input data.

[0657] "Design Data" refers to original visual or functional designs generated based on user input and emotional data.

[0658] This invention is a system that realizes more advanced design customization by combining preferences and images input by the user in natural language with an emotion engine that recognizes the user's emotions. This system is carried out by the cooperation of the user, the terminal, and the server.

[0659] First, a user launches a dedicated app on a device such as a smartphone or tablet and inputs the desired design in natural language. For example, they might input something like, "I want a wallpaper with a spring cherry blossom theme." When the user presses the "send" button, the device captures this text data. At the same time, an emotion engine is activated on the device, which has a built-in camera and microphone, and analyzes the user's facial expressions and tone of voice to capture emotional data in real time. This emotional data includes emotional states such as "joy" and "excitement."

[0660] The device converts the acquired data (natural language data and emotion data) into JSON format, and then transmits this JSON-formatted dataset to the server using the HTTPS protocol, which ensures data security and confidentiality.

[0661] The server receives the HTTPS request and parses the JSON data sent. The analysis module extracts the natural language and sentiment data entered by the user from the parsed data. Using a natural language processing (NLP) engine, it identifies keywords such as "spring," "cherry blossoms," and "wallpaper," and then uses a sentiment analysis algorithm to extract the emotional state (e.g., "joy").

[0662] The server then combines the extracted keywords and emotion data and formats them into a format appropriate for the generative AI model. For example, it combines the keywords "spring," "cherry blossoms," and "wallpaper" with "joy" to create a prompt sentence to be input into the generative AI model.

[0663] Based on the input prompts, the generative AI model generates original design data that reflects the user's preferences and emotions. Specifically, wallpaper data with the theme of "joyful spring cherry blossoms" is generated.

[0664] The generated design data is converted back to JSON format by the server and sent to the user's device using the HTTPS protocol. The device receives this data, parses the JSON data, and obtains the original wallpaper image data. The device then immediately sets the obtained image data as the smartphone's wallpaper. As a result, the user's smartphone wallpaper is changed to the original wallpaper with the theme "Joyful Spring Cherry Blossoms."

[0665] As a concrete example, consider the case where a user inputs "I want a summer beach-themed icon set," and the emotion engine recognizes the user's "excitement." The acquired dataset contains the keywords "summer," "beach," and "icon set," as well as the emotion "excitement." The server inputs this data into the generative AI model and generates a new icon set with the theme of "exciting summer beach." This icon set is sent to the device, which immediately applies the received icon set, and the user's smartphone is updated with the new icon set.

[0666] In this way, users can intuitively input their preferences and images using natural language, and easily experience original smartphone designs that reflect their emotions.This system makes it possible to generate professional designs and apply them immediately, without relying on the user's taste or skills.

[0667] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0668] Step 1: Getting user-entered data

[0669] Specific explanation

[0670] Using a dedicated app, users input their preferences and images in natural language, such as "I want a wallpaper with a spring cherry blossom theme." When the user presses the "Send" button, the device acquires this text data.

[0671] input

[0672] User natural language input (e.g., "I want a wallpaper with a spring cherry blossom theme")

[0673] output

[0674] The text data is saved on the device.

[0675] concrete action

[0676] The terminal captures the user input in text format and stores it in a variable.

[0677] Step 2: Acquire emotion data using a camera or microphone

[0678] Specific explanation

[0679] While the user is typing, the device's camera and microphone are activated to analyze the user's facial expressions and tone of voice in real time. An emotion engine processes this data to determine the user's emotional state.

[0680] input

[0681] User video and audio data

[0682] output

[0683] Analyzed emotion data (e.g., "joy," "excitement")

[0684] concrete action

[0685] The camera and microphone are activated to capture video and audio in real time, and the emotion engine analyzes the data and outputs a specific emotional state.

[0686] Step 3: Convert the data to JSON format

[0687] Specific explanation

[0688] The natural language data and emotion data acquired by the device is converted into JSON format.

[0689] input

[0690] Natural language input data, emotion data

[0691] output

[0692] JSON formatted dataset

[0693] concrete action

[0694] Natural language data and sentiment data are integrated and converted into JSON format programmatically.

[0695] Step 4: Sending data to the server

[0696] Specific explanation

[0697] The terminal sends the converted data set to the server using the HTTPS protocol.

[0698] input

[0699] JSON formatted dataset

[0700] output

[0701] The server receives the JSON-formatted dataset.

[0702] concrete action

[0703] The device generates an HTTPS request and sends data to the server's URL.

[0704] Step 5: Receiving and analyzing data on the server

[0705] Specific explanation

[0706] The server parses the received JSON data and extracts natural language and sentiment data. The NLP engine identifies keywords from the natural language data, and the sentiment analysis algorithm extracts the emotional state.

[0707] input

[0708] JSON formatted dataset

[0709] output

[0710] Identified keywords (e.g., "spring," "cherry blossoms," "wallpaper"), emotional states (e.g., "joy")

[0711] concrete action

[0712] The server parses the JSON data and analyzes it using an NLP engine and sentiment analysis algorithms.

[0713] Step 6: Format the data

[0714] Specific explanation

[0715] The server integrates the extracted keywords and sentiment data to generate appropriate prompts for the generative AI model.

[0716] input

[0717] Identified keywords, emotional states

[0718] output

[0719] Prompt statement (e.g., "Spring cherry blossom wallpaper that inspires joy")

[0720] concrete action

[0721] The server combines the keywords and emotion data and formats them into a prompt sentence.

[0722] Step 7: Generate original design data

[0723] Specific explanation

[0724] The server inputs the prompt sentence into the generative AI model and generates design data.

[0725] input

[0726] Prompt statement

[0727] output

[0728] Generated original design data (e.g., "Joyful Spring Cherry Blossom Wallpaper")

[0729] concrete action

[0730] The generative AI model generates design data that fits the theme based on the prompt text.

[0731] Step 8: Convert your design data to JSON format and prepare it for submission

[0732] Specific explanation

[0733] The generated design data is converted back into JSON format and prepared for transmission to the user's device using the HTTPS protocol.

[0734] input

[0735] Generated design data

[0736] output

[0737] Design data in JSON format

[0738] concrete action

[0739] The server converts the generated design data into JSON format and prepares it for transmission via HTTPS protocol.

[0740] Step 9: Sending and receiving design data to and from your device

[0741] Specific explanation

[0742] The server sends JSON format design data to the device, which then receives it.

[0743] input

[0744] Design data in JSON format

[0745] output

[0746] The device receives the design data

[0747] concrete action

[0748] The server sends the design data to the terminal via an HTTPS request, and the terminal receives it.

[0749] Step 10: Applying design data

[0750] Specific explanation

[0751] The device parses the received design data and obtains the original wallpaper image data, which the device then instantly sets as the smartphone's wallpaper.

[0752] input

[0753] Design data in JSON format

[0754] output

[0755] The device wallpaper has been changed

[0756] concrete action

[0757] The device parses the received JSON data, obtains the wallpaper image data, and immediately sets it as the wallpaper.

[0758] (Application example 2)

[0759] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0760] When workers use industrial equipment, they need specific work procedures and guidance. However, there are no systems that can flexibly respond to the worker's emotions and situation. In particular, emotions such as tension and anxiety can affect work efficiency, so appropriate support is needed.

[0761] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing preferences, images, and emotions input by the user in natural language and extracting keywords, means for shaping input data to a generative AI model using the extracted keywords and emotions, and means for generating original design data based on the shaped input data by the generative AI model. This makes it possible to generate and display customized work guides in real time according to the instructions and emotions input by the worker in natural language.

[0762] A "natural language" is a human language that is used by users in normal conversation and writing, and that can be converted into a form that a computer can parse.

[0763] "Emotion" refers to a psychological state that can be recognized from a user's facial expression, tone of voice, etc., and includes, for example, states such as joy, sadness, surprise, and tension.

[0764] "Keywords" refer to important information or characteristic terms extracted from a user's natural language input and sentiment data.

[0765] A "generative artificial intelligence model" refers to an algorithm or system that generates new content or designs based on input data.

[0766] "Original design data" refers to new and unique design data generated by a generative artificial intelligence model based on a user's input and emotional state.

[0767] "Terminal" refers to electronic devices used by users, such as smartphones, tablets, and personal computers.

[0768] "Industrial equipment" refers to machinery and equipment used in manufacturing and production processes, including equipment operated by workers.

[0769] "Work guide" means visual or textual instructions that show a worker how to perform a particular work task.

[0770] MODE FOR CARRYING OUT THE INVENTION

[0771] The system embodying this invention is configured to provide customized operation guides based on the user's instructions and emotions. It has the function of analyzing preferences, images, and emotion data input by the user in natural language, and generating and displaying appropriate operation guides in real time.

[0772] Hardware and Software Configuration

[0773] The system mainly consists of the following hardware and software:

[0774] Hardware

[0775] Smart glasses or head-mounted display (HMD): A device that allows users to provide natural language instructions and emotional input.

[0776] Camera and microphone: Devices for analyzing the user's facial expressions and tone of voice to obtain emotional data.

[0777] software

[0778] Emotion recognition engine (e.g., Microsoft Azure Cognitive Services Emotion API): Analyzes emotions from the user's facial expressions and tone of voice and obtains data.

[0779] Natural language processing (NLP) engine (e.g., Google Cloud Natural Language API): Analyzes the user's natural language input and extracts keywords.

[0780] Generative AI models (e.g., OpenAI GPT-4): Generate customized work guides and design data based on user keywords and emotional data.

[0781] System Operation

[0782] The system operates as follows.

[0783] User interface: The user wears smart glasses and inputs specific instructions for the task in natural language. For example, they might give instructions such as, "Tell me how to tighten a bolt." At this time, the emotion recognition engine analyzes the user's facial expressions and tone of voice via a camera and microphone, and obtains emotional data such as tension or anxiety.

[0784] Data analysis: The natural language data and emotion data entered by the user are acquired and converted into JSON format. This allows the data sent to the server to be processed in a unified format. The server receives the JSON data, analyzes it using a natural language processing engine, and extracts keywords and emotion data.

[0785] Design data generation: The server uses a generative AI model to generate original work guides and design data based on the extracted keywords and emotion data. This generated data includes appropriate procedures and guides that take emotion into consideration.

[0786] Data application: The generated work guide data is converted back to JSON format and sent to the user's device. The smart glasses parse the received design data and apply it to the user's vision. For example, the bolt tightening procedure is displayed as an illustrated guide in the user's field of vision.

[0787] Specific examples

[0788] For example, if a worker says in natural language, "Tell me how to tighten a bolt," and the emotion recognition engine detects that the user is nervous, the server analyzes this data and uses a generative AI model to generate "bolt tightening procedures to relieve tension." The worker's smart glasses display easy-to-understand procedures and illustrated guides that instill a sense of security.

[0789] Prompt Sentence Examples

[0790] User Instructions: What is the procedure for tightening the bolts?

[0791] User Emotion: Tension

[0792] This prompt enables the generative AI model to provide optimal task guidance.

[0793] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0794] Step 1:

[0795] The user puts on the smart glasses and inputs specific work instructions in natural language. For example, they might say, "Tell me how to tighten a bolt." At this time, the user's facial expressions and tone of voice are recorded via a camera and microphone. The input is captured as natural language data and sent to an emotion recognition engine. The acquired natural language data, facial expression data, and voice data are then passed to the respective processing modules, where analysis begins.

[0796] Step 2:

[0797] The device's emotion recognition engine analyzes the acquired facial and voice data to identify the user's emotions. For example, if the user is nervous, the device identifies that state of tension and records it as emotion data. The input facial and voice data is converted into emotion data. The output emotion data is "tension."

[0798] Step 3:

[0799] The device converts the user's natural language input data and the acquired emotion data into JSON format. A JSON object is generated with the natural language data and emotion data stored in their respective fields. The generated JSON data is ready to be sent to the server in the next step.

[0800] Step 4:

[0801] The device sends the generated JSON data to the server using the HTTPS protocol. The JSON data, including the user's input amount and emotion data, is securely sent to the server. The server receives the HTTPS request and proceeds to the next step.

[0802] Step 5:

[0803] The server parses the received JSON data and extracts natural language data and emotion data separately. Keywords such as "bolt," "tightening," and "procedure" and the emotion "tension" are obtained from the parsed data. The NLP engine on the server analyzes the meaning and organizes the keywords.

[0804] Step 6:

[0805] The server creates a prompt for the generative AI model based on the extracted keywords and emotion data. The prompt is formatted, for example, as "Tell me how to tighten a bolt," including the word "tension." The prompt is sent to the generative AI model, which then begins generating design data.

[0806] Step 7:

[0807] Based on the prompt received, the generative AI model generates original work guides and design data that take emotions into account. For example, it generates "bolt tightening procedures to relieve tension." The generated design data is then returned to the server.

[0808] Step 8:

[0809] The server converts the generated design data back into JSON format and prepares it for transmission to the device. The data organized into JSON objects is sent to the device via HTTPS.

[0810] Step 9:

[0811] The device parses the JSON data received from the server and extracts the generated work guide and design data, which the device receives, parses, and converts into visual information such as an illustration guide.

[0812] Step 10:

[0813] The device then displays the received design data on the smart glasses display, and the user is shown a "bolt tightening procedure to ease tension," allowing them to proceed with the work with peace of mind.

[0814] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0815] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0816] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0817] [Third embodiment]

[0818] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0819] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0820] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0821] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0822] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0823] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0824] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0825] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0826] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0827] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0828] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0829] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0830] This invention is a system for customizing the design of a mobile phone based on preferences and images input by the user in natural language. This system is carried out by the cooperation of the user, the terminal, and the server.

[0831] Program processing explanation

[0832] Getting user-entered data

[0833] Using a dedicated app, the user inputs content in natural language, such as "I want a wallpaper with a spring cherry blossom theme." When the user presses the "Send" button, the device receives the input data and proceeds to the next step.

[0834] Sending data

[0835] The terminal converts the natural language data entered by the user into JSON format, and then transmits the converted JSON data to the server using the HTTPS protocol, ensuring accurate transmission of the data.

[0836] Data reception and analysis

[0837] The server receives the HTTPS request, parses the JSON data sent, and retrieves the user's input. The server uses a natural language processing (NLP) engine to extract meaning from the input text and identify keywords such as "spring," "cherry blossoms," and "wallpaper."

[0838] Data Formatting

[0839] The server then uses the extracted keywords to format the data required by the generative AI model. Specifically, it prepares keywords such as "spring," "cherry blossoms," and "wallpaper" to be input into the model in array format.

[0840] Data generation

[0841] The server inputs the formatted data into a generative AI model, which then generates original design data, such as a wallpaper with a "spring cherry blossom" theme.

[0842] Sending generated data

[0843] The server converts the generated wallpaper image data back into JSON format and prepares to send it to the device. The wallpaper image data is sent to the user's device using the HTTPS protocol.

[0844] Receiving and applying data

[0845] The device receives the wallpaper image data and parses the JSON data to obtain the image data. The device then executes a system call to immediately apply the image data as the smartphone's wallpaper. As a result, the user's smartphone wallpaper is set to the new original "Spring Cherry Blossom" theme.

[0846] Specific examples

[0847] For example, consider a case where a user inputs, "I want a summer beach-themed icon set." When the user's input is transmitted and analyzed, the keywords "summer," "beach," and "icon" are extracted. The server uses a generative AI model to generate a new icon set based on this input and transmits the generated icon set to the device. The device then applies the received icon set to the user's smartphone.

[0848] This invention allows users to easily enjoy creating original smartphone designs simply by intuitively inputting their preferences and ideas using natural language. AI automatically generates professional designs, independent of the user's taste or skills, and instantly applies them, making everyday life more appealing.

[0849] The processing flow will be explained below.

[0850] Step 1:

[0851] User:

[0852] Users start up a dedicated smartphone app and enter their preferences and images in natural language into the input field that appears, such as "I want a wallpaper with an autumn leaf theme." Once they've finished entering the information, they press the "Send" button.

[0853] Step 2:

[0854] Device:

[0855] The terminal receives input data from the user, converts it into JSON format, and then sends the converted JSON data to the server using the HTTPS protocol.

[0856] Step 3:

[0857] server:

[0858] The server receives the HTTPS request, parses the JSON data sent, and extracts the natural language input from the parsed data.

[0859] Step 4:

[0860] server:

[0861] The server uses a natural language processing (NLP) engine to analyze the meaning of the natural language data it receives. Specifically, it extracts keywords such as "autumn," "autumn leaves," and "wallpaper" from the input "I want a wallpaper with an autumn foliage theme."

[0862] Step 5:

[0863] server:

[0864] The server then uses the extracted keywords to format the data appropriately for use as input for a generative AI model. This formatting process results in a dataset containing keywords such as "autumn," "autumn leaves," and "wallpaper."

[0865] Step 6:

[0866] server:

[0867] The formatted data is input into a generative AI model, which then generates original wallpaper data based on this data.

[0868] Step 7:

[0869] server:

[0870] The generated wallpaper data is received, converted to JSON format, and prepared for transmission to the user's device. This wallpaper data is sent to the device using the HTTPS protocol.

[0871] Step 8:

[0872] Device:

[0873] The device receives the wallpaper data sent from the server, parses the JSON data, and obtains the original wallpaper image data.

[0874] Step 9:

[0875] Device:

[0876] The acquired original wallpaper image data is immediately set as the smartphone wallpaper. This setting changes the user's smartphone wallpaper to an original one with an "Autumn Leaves" theme.

[0877] Example 1

[0878] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0879] Today's smartphone users want to be able to easily customize their devices to suit their tastes and preferences. However, existing services often make customization difficult, relying on the user's taste and skills. Furthermore, the interface for users to specify designs is often not intuitive, making it difficult to achieve high-quality customization easily. Furthermore, for users who are not technically savvy, the complex and time-consuming operation poses a problem.

[0880] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0881] In this invention, the server includes means for acquiring preferences and images input by a user in natural language, means for converting the acquired preferences and images into JSON format, means for transmitting the converted JSON format data using the HTTPS protocol, means for receiving and analyzing the transmitted JSON format data, means for extracting keywords from the analyzed data, means for formatting input data for a generative artificial intelligence model using the extracted keywords, means for generating original design data based on the formatted input data using the generative artificial intelligence model, means for converting the generated design data back into JSON format and transmitting it using the HTTPS protocol, and means for analyzing the received design data and applying it to the terminal. This allows a user to intuitively and quickly customize the design of their smartphone simply by inputting their preferences in natural language.

[0882] "Means for acquiring preferences and images input by the user in natural language" refers to a function that allows the user to input their wishes and images in natural language and receives the input content from the system.

[0883] "Means for converting to JSON format" is a function that converts natural language data entered by the user into a standardized data format called JavaScript Object Notation (JSON).

[0884] "Means for sending using the HTTPS protocol" refers to a function that uses a protocol called HTTP Secure (HTTPS) to send data to a server in a secure and encrypted manner.

[0885] "Means for receiving and analyzing transmitted JSON format data" refers to a function in which a server receives JSON data transmitted via the HTTPS protocol and analyzes the contents of that data.

[0886] "Means for extracting keywords from analyzed data" refers to a function that analyzes received data using natural language processing and identifies and extracts important keywords.

[0887] "Means for formatting input data for a generative artificial intelligence model" is a function that converts extracted keywords into data in a format that is easy for the generative artificial intelligence model to understand.

[0888] "Means for generating original design data based on input data generated by a generative artificial intelligence model" refers to a function in which a generative artificial intelligence automatically creates design data that meets the user's wishes based on formatted input data.

[0889] "Means for converting the generated design data back into JSON format and transmitting it using the HTTPS protocol" refers to a function for converting the generated design data back into JSON format and transmitting it to the terminal using the HTTPS protocol.

[0890] The "means for analyzing received design data and applying it to the terminal" is a function that analyzes the JSON format design data received by the terminal and applies the specified design to the terminal.

[0891] This invention is a system that customizes a mobile information terminal based on preferences and images input by the user in natural language. This system performs processing through collaboration between the user, the terminal, and the server. The user uses a dedicated app to input specific wishes and images in natural language. For example, the user might input, "I want a wallpaper with a spring cherry blossom theme."

[0892] The device receives natural language data entered by the user and converts it into JSON format. The converted data is then sent to the server using the HTTPS protocol. The server analyzes the received JSON data and extracts keywords using a natural language processing engine (libraries such as TensorFlow and PyTorch). The extracted keywords are then used to format the input data for a generative AI model, which then generates original design data.

[0893] The generated design data is converted back to JSON format and sent to the device using the HTTPS protocol. The device then analyzes the received data and extracts the image data. Finally, the device applies the extracted image data as the smartphone's wallpaper or icon.

[0894] For example, if a user types "I want a summer beach-themed icon set," the system will extract the keywords "summer," "beach," and "icon," and the AI ​​model will generate a new icon set, resulting in a summer beach-themed icon set that will be applied to the user's smartphone.

[0895] Examples of specific prompts include:

[0896] "I want a wallpaper with a spring cherry blossom theme."

[0897] "I want a summer beach themed icon set."

[0898] "Create an autumn leaf-themed lock screen"

[0899] Based on these prompts, the generative AI model creates a design that meets the user's wishes. Users can easily enjoy creating their own original smartphone design by simply inputting their wishes in natural language.

[0900] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0901] Step 1:

[0902] The user opens the app and enters a prompt in natural language, such as "I want a spring cherry blossom themed wallpaper," into the input field. When the user presses the "Send" button, the input data is sent to the device.

[0903] Input: A natural language prompt entered by the user.

[0904] Output: The prompt is sent to the terminal.

[0905] Step 2:

[0906] The device receives the natural language data entered by the user and converts it into JavaScript Object Notation (JSON) format, such as "message: 'I want a wallpaper with a spring cherry blossom theme'".

[0907] Input: A prompt sentence typed in natural language.

[0908] Output: Data converted to JSON format.

[0909] Step 3:

[0910] The device sends the converted JSON data to the server using the HTTPS protocol, which may also include data integrity checks and retransmission mechanisms.

[0911] Input: Data converted to JSON format.

[0912] Output: Data sent to the server using the HTTPS protocol.

[0913] Step 4:

[0914] The server receives the HTTPS request and analyzes the JSON data sent. Specifically, it parses the received data and obtains the contents of the "message" field.

[0915] Input: JSON data sent over HTTPS.

[0916] Output: Parsed content (prompt sentence).

[0917] Step 5:

[0918] The server uses a natural language processing engine (such as TensorFlow or PyTorch) to extract keywords from the prompt. In this example, "spring," "cherry blossoms," and "wallpaper" are extracted.

[0919] Input: The parsed prompt sentence.

[0920] Output: Extracted keywords.

[0921] Step 6:

[0922] The server formats the data based on the extracted keywords into a format that the generative AI model can understand, such as a JSON object with the format "keywords: ['spring', 'cherry blossoms', 'wallpaper']".

[0923] Input: Extracted keywords.

[0924] Output: The formatted data.

[0925] Step 7:

[0926] The server inputs the formatted data into a generative AI model, which then generates original design data (e.g., wallpaper images) based on the data.

[0927] Input: Formatted data.

[0928] Output: Generated design data (wallpaper image).

[0929] Step 8:

[0930] The server converts the generated design data into JSON format and sends it to the terminal using the HTTPS protocol.

[0931] Input: Generated design data.

[0932] Output: Design data converted to JSON format.

[0933] Step 9:

[0934] The device receives the sent JSON data, parses it, and obtains the image data. Specifically, it extracts the "image" field from the JSON data and decodes the Base64-encoded image.

[0935] Input: Design data submitted in JSON format.

[0936] Output: Decoded image data.

[0937] Step 10:

[0938] The device executes a system call to apply the acquired image data as the smartphone's wallpaper. Ultimately, the user's smartphone wallpaper is changed to the new original "Spring Cherry Blossom" theme.

[0939] Input: Decoded image data.

[0940] Output: The new wallpaper applied.

[0941] (Application example 1)

[0942] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0943] In virtual stores, there are inconvenient ways to quickly and intuitively reflect the preferences and images that users input in natural language and customize the store design, so there is a need for technology that can generate attractive designs in real time based on user input and apply them immediately.Furthermore, there is also a need to easily save and share the generated designs.

[0944] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0945] In this invention, the server includes means for acquiring preferences and images input by a user in natural language, means for analyzing the acquired preferences and images and extracting keywords, means for shaping input data to a generative artificial intelligence model using the extracted keywords, means for generating original design data based on the shaped input data using the generative artificial intelligence model, means for transmitting the generated design data to the user's terminal, means for applying the received design data to the terminal, means for applying the generated design in real time in a virtual space, and means for saving and sharing the generated design. This allows a user to intuitively customize the design of a virtual store in real time based on input in natural language, and to save and share the generated design.

[0946] "Natural language" refers to the language that a user normally uses in conversation or writing, without being converted into a form that is easy for the system to parse.

[0947] "Preferences and images" refers to the user's personal hobbies and visual concepts / themes.

[0948] "Keyword extraction" refers to the process of extracting meaningful words and phrases from a user's natural language input.

[0949] A "generative artificial intelligence model" refers to an artificial intelligence that has the ability to automatically generate new designs and content based on input data and conditions.

[0950] "Formatting input data" refers to the process of converting data into a format that is easy for a generative artificial intelligence model to understand.

[0951] "Original design data" refers to unique design information that is newly generated by a generative artificial intelligence model and does not exist anywhere else.

[0952] "Sending to the user's device" refers to sending the generated design data to the device used by the user using a communication means such as the Internet.

[0953] "Applying design data" refers to reflecting the received design data on the user's device and making it actually displayable and usable.

[0954] "Virtual space" refers to a 3D space or virtual store generated by a computer.

[0955] "Real-time application" refers to the process of updating and displaying the design so that it reflects user input almost immediately.

[0956] "Saving a design" refers to saving the generated design data to a storage medium for later use.

[0957] "Means of sharing" refers to the methods by which the generated design can be published and transferred to other users or platforms.

[0958] This invention is a system that customizes the design of a virtual store based on preferences and images input by the user in natural language. This system executes processing in cooperation with the user, the terminal, and the server.

[0959] Using a smartphone, smart glasses, or head-mounted display, users can request a virtual store design in natural language, for example, by typing, "I want a store with an autumn leaf theme."

[0960] The device converts the user's input into JSON format and sends it to the server using the HTTPS protocol. The server analyzes the received JSON data and uses a natural language processing engine to extract keywords such as "autumn," "autumn leaves," and "store." It then formats these keywords into the format required by the generative artificial intelligence model and inputs them into the model.

[0961] The generative AI model generates original design data based on the given keywords. This design data is then converted back to JSON format and sent to the user's device via HTTPS. The device parses and retrieves the received design data and immediately applies it to the virtual store.

[0962] Furthermore, the generated designs can be saved on the device and shared with other users and platforms, allowing users to intuitively make design requests in natural language, customize the design of their virtual store in real time, and even save and share the designs.

[0963] The hardware requires a smartphone, smart glasses, or a head-mounted display to receive user input, and the software uses Python 3.x, the Requests library, and specific API endpoints to apply the virtual store design.

[0964] For example, if a user inputs "I want a Halloween-themed store," the server extracts the keywords "Halloween," "theme," and "store" and inputs them into a generative AI model. Based on this prompt, the AI ​​model generates a virtual store design, which is then instantly applied.

[0965] Example prompt sentence:

[0966] Theme: Halloween

[0967] Design: Pumpkins, bats, dark colors

[0968] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0969] Step 1:

[0970] The user inputs a design request for the virtual store in natural language. For example, they might input, "I want a store with an autumn foliage theme." The input is sent to the terminal.

[0971] Step 2:

[0972] The terminal receives the user's input and converts it into JSON format. Specifically, it converts it into a JSON object with the natural language input as a key. The converted JSON data is then sent to the server.

[0973] Step 3:

[0974] The server receives the JSON data sent from the device and analyzes it using a natural language processing engine. Keywords such as "autumn," "autumn leaves," and "store" are extracted from the input data. The results of this analysis become the input for the next processing step.

[0975] Step 4:

[0976] Based on the extracted keywords, the server formats the data to be input into the generative AI model. Specifically, it converts the keywords into a list format and generates a prompt sentence. The keywords "autumn," "autumn leaves," and "store" become the elements of the prompt sentence.

[0977] Step 5:

[0978] The server inputs the formatted data into a generative AI model, which generates original design data based on the prompt. This generated design data becomes the input for the next processing step.

[0979] Step 6:

[0980] The server converts the generated design data back into JSON format and sends it to the terminal using the HTTPS protocol. Specifically, it encodes the design data into a JSON object and sends a POST request to the endpoint.

[0981] Step 7:

[0982] The device receives the JSON data sent from the server and parses it to obtain the design data. The parsing process extracts the JSON format data as specific design data.

[0983] Step 8:

[0984] The terminal immediately applies the acquired design data to the virtual store by sending a request to the API endpoint of the virtual space to reflect the design data.

[0985] Step 9:

[0986] The device provides functions for saving the generated design and sharing it with other users or platforms. Specifically, it saves the design data in local storage and generates a link or QR code for sharing it.

[0987] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0988] This invention relates to a system that realizes advanced design customization by combining preferences and images input by the user in natural language with an emotion engine that recognizes the user's emotions. This system performs processing in cooperation with the user, the terminal, and the server.

[0989] Program processing explanation

[0990] Getting user-entered data

[0991] Using a dedicated app, users input content in natural language, such as "I want a wallpaper with a spring cherry blossom theme." When the user presses the "send" button, the device receives the input data and proceeds to the next step. The device is also equipped with a camera and microphone, and the emotion engine analyzes the user's facial expressions and tone of voice, capturing their emotions as data.

[0992] Sending data

[0993] The device converts the input data from the user into JSON format. At the same time, the emotion data recognized by the emotion engine is also added to the JSON format. The converted JSON data is then sent to the server using the HTTPS protocol.

[0994] Data reception and analysis

[0995] The server receives the HTTPS request and parses the JSON data sent. From the parsed data, it extracts the natural language content and emotion data of the user's input. Using a natural language processing (NLP) engine, it analyzes the semantics of the natural language data and identifies keywords such as "spring," "cherry blossoms," and "wallpaper." The emotion data is also analyzed to extract the user's emotional state (e.g., joy, sadness, surprise, etc.) at the time of input.

[0996] Data Formatting

[0997] The server then uses the extracted keywords and recognized emotions to format the data appropriately for use as input for a generative AI model. For example, it generates a dataset that combines keywords such as "spring," "cherry blossoms," and "wallpaper" with the user's emotion of "happy."

[0998] Data generation

[0999] The server inputs the formatted data into a generative AI model, which then generates original design data, such as wallpaper data with the theme of "joyful spring cherry blossoms."

[1000] Sending generated data

[1001] The server receives the generated wallpaper data, converts it back to JSON format, and prepares it for transmission to the user's device. This wallpaper data is sent to the device using the HTTPS protocol.

[1002] Receiving and applying data

[1003] The device receives the wallpaper data sent from the server, parses the JSON data, and obtains the original wallpaper image data. The device immediately sets the image data as the smartphone's wallpaper. This setting changes the user's smartphone wallpaper to the original wallpaper with the theme "Joyful Spring Cherry Blossoms."

[1004] Specific examples

[1005] For example, consider the case where a user inputs "I want an icon set with a summer beach theme," and at the same time, the emotion engine recognizes the user's "excitement." The acquired data includes the keywords "summer," "beach," and "icon set," as well as the emotion "excitement." Based on this data, the server uses a generative AI model to generate a new icon set with the theme of "exciting summer beach" and sends it to the device. The device immediately applies the received icon set, and the user's smartphone is updated with the new icon set.

[1006] This invention allows users to intuitively input their preferences and ideas using natural language, and easily experience original smartphone designs that reflect their emotions. AI and emotion recognition technology work together to generate professional designs that can be instantly applied, without relying on the user's taste or skills, making everyday life more fulfilling.

[1007] The processing flow will be explained below.

[1008] Program processing explanation

[1009] Step 1:

[1010] User:

[1011] The user launches a dedicated smartphone app and enters their preferences and image in natural language, such as "I want a wallpaper with an autumn leaf theme," into the input field that appears. When the user presses the "Send" button, the emotion engine begins analyzing the user's emotions through their voice and facial expressions. The emotion engine generates emotion data such as "happy" or "excited" from the user's facial expressions and tone of voice.

[1012] Step 2:

[1013] Device:

[1014] The device receives input data from the user and converts it into JSON format. At the same time, the emotion data output by the emotion engine is also added to the JSON format. The converted JSON data is structured to include the user's input and emotion data. It is then sent to the server using the HTTPS protocol.

[1015] Step 3:

[1016] server:

[1017] The server receives the HTTPS request and parses the JSON data. From the parsed data, it extracts the natural language content and emotion data entered by the user. Specifically, it extracts the text data "I want a wallpaper with an autumn leaf theme" and the emotion data "I'm excited."

[1018] Step 4:

[1019] server:

[1020] The server uses a natural language processing (NLP) engine to analyze the meaning of the acquired natural language data. Keywords such as "autumn," "autumn leaves," and "wallpaper" are extracted from the text data. Emotion data is analyzed in the same way, and the keyword "excitement" is extracted.

[1021] Step 5:

[1022] server:

[1023] Based on the extracted keywords and recognized emotions, the server formats the data into an appropriate format for use as input for the generative AI model. Specifically, it generates a dataset consisting of the keywords "autumn," "autumn leaves," and "wallpaper" and the emotion data "excitement." This dataset is used as input for the generative AI model.

[1024] Step 6:

[1025] server:

[1026] The formatted data set is then input into a generative AI model, which then generates original design data. For example, it generates an exciting wallpaper image featuring vivid autumn leaves.

[1027] Step 7:

[1028] server:

[1029] The generated wallpaper data is received, converted back to JSON format, and prepared for transmission to the user's device. This wallpaper data is sent to the device using the HTTPS protocol.

[1030] Step 8:

[1031] Device:

[1032] The device receives the wallpaper data sent from the server, parses the JSON data, and obtains the original wallpaper image data.

[1033] Step 9:

[1034] Device:

[1035] The acquired original wallpaper image data is immediately set as the smartphone wallpaper, so that the user's smartphone wallpaper is set to an original design with an "Autumn Leaves" theme that reflects the user's emotion of "Excitement."

[1036] Specific examples

[1037] For example, if a user inputs "I want a summer beach-themed icon set," and the emotion engine recognizes the user's "relaxation," the retrieved data includes the keywords "summer," "beach," and "icon set" along with the emotion "relaxation." This data is sent to the server, which then inputs the formatted data into a generative AI model to generate a new icon set with the theme of "relaxing summer beach." The generated icon set is then sent to the device and instantly applied to the user's smartphone.

[1038] The present invention allows users to easily enjoy original designs that reflect their own tastes and feelings, making their daily lives more enriching.

[1039] Example 2

[1040] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1041] In the field of modern digital design, there is a growing need for users to easily input their preferences and preferences in natural language and generate customized, original designs based on those inputs. However, current systems have difficulty appropriately reflecting users' emotions, extracting keywords from the input natural language, and generating designs that take emotions into account. Furthermore, the process of immediately applying the generated designs to the user's device is not efficient enough. This can sometimes result in a poor user experience.

[1042] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1043] In this invention, the server includes means for acquiring preferences and images input by a user in natural language, means for analyzing the acquired preferences and images to extract keywords, means for integrating the extracted keywords and user emotion data to format input data for a generative AI model, means for analyzing the user's facial expressions and voice using a camera and microphone to acquire emotion data, means for generating original design data based on the formatted input data using the generative AI model, means for transmitting the generated design data to the user's terminal, and means for applying the received design data to the terminal. This enables the generation of original designs that reflect the user's preferences and emotions and their immediate application.

[1044] "User" refers to an individual who uses the system to input preferences and ideas in natural language.

[1045] A "dedicated app" refers to a software application that allows users to input their preferences and images and communicate with a server.

[1046] "Terminal" refers to a device used by a user, such as a smartphone or tablet.

[1047] An "emotion engine" refers to a software module that analyzes a user's facial expressions and tone of voice to recognize their emotional state.

[1048] "JSON" stands for JavaScript Object Notation and refers to a lightweight data interchange format for structuring data and communicating it over the Internet.

[1049] "HTTPS protocol" stands for Hypertext Transfer Protocol Secure and refers to a communication protocol for securely transmitting data over the Internet.

[1050] "Server" refers to a computer system that receives, analyzes, and processes data sent by a user and sends the generated data to the user's terminal.

[1051] A "natural language processing (NLP) engine" refers to a software module that analyzes natural language input by a user and extracts meaning and keywords.

[1052] A "generative artificial intelligence model" refers to a machine learning algorithm for generating original design data based on input data.

[1053] "Design Data" refers to original visual or functional designs generated based on user input and emotional data.

[1054] This invention is a system that realizes more advanced design customization by combining preferences and images input by the user in natural language with an emotion engine that recognizes the user's emotions. This system is carried out by the cooperation of the user, the terminal, and the server.

[1055] First, a user launches a dedicated app on a device such as a smartphone or tablet and inputs the desired design in natural language. For example, they might input something like, "I want a wallpaper with a spring cherry blossom theme." When the user presses the "send" button, the device captures this text data. At the same time, an emotion engine is activated on the device, which has a built-in camera and microphone, and analyzes the user's facial expressions and tone of voice to capture emotional data in real time. This emotional data includes emotional states such as "joy" and "excitement."

[1056] The device converts the acquired data (natural language data and emotion data) into JSON format, and then transmits this JSON-formatted dataset to the server using the HTTPS protocol, which ensures data security and confidentiality.

[1057] The server receives the HTTPS request and parses the JSON data sent. The analysis module extracts the natural language and sentiment data entered by the user from the parsed data. Using a natural language processing (NLP) engine, it identifies keywords such as "spring," "cherry blossoms," and "wallpaper," and then uses a sentiment analysis algorithm to extract the emotional state (e.g., "joy").

[1058] The server then combines the extracted keywords and emotion data and formats them into a format appropriate for the generative AI model. For example, it combines the keywords "spring," "cherry blossoms," and "wallpaper" with "joy" to create a prompt sentence to be input into the generative AI model.

[1059] Based on the input prompts, the generative AI model generates original design data that reflects the user's preferences and emotions. Specifically, wallpaper data with the theme of "joyful spring cherry blossoms" is generated.

[1060] The generated design data is converted back to JSON format by the server and sent to the user's device using the HTTPS protocol. The device receives this data, parses the JSON data, and obtains the original wallpaper image data. The device then immediately sets the obtained image data as the smartphone's wallpaper. As a result, the user's smartphone wallpaper is changed to the original wallpaper with the theme "Joyful Spring Cherry Blossoms."

[1061] As a concrete example, consider the case where a user inputs "I want a summer beach-themed icon set," and the emotion engine recognizes the user's "excitement." The acquired dataset contains the keywords "summer," "beach," and "icon set," as well as the emotion "excitement." The server inputs this data into the generative AI model and generates a new icon set with the theme of "exciting summer beach." This icon set is sent to the device, which immediately applies the received icon set, and the user's smartphone is updated with the new icon set.

[1062] In this way, users can intuitively input their preferences and images using natural language, and easily experience original smartphone designs that reflect their emotions.This system makes it possible to generate professional designs and apply them immediately, without relying on the user's taste or skills.

[1063] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1064] Step 1: Getting user-entered data

[1065] Specific explanation

[1066] Using a dedicated app, users input their preferences and images in natural language, such as "I want a wallpaper with a spring cherry blossom theme." When the user presses the "Send" button, the device acquires this text data.

[1067] input

[1068] User natural language input (e.g., "I want a wallpaper with a spring cherry blossom theme")

[1069] output

[1070] The text data is saved on the device.

[1071] concrete action

[1072] The terminal captures the user input in text format and stores it in a variable.

[1073] Step 2: Acquire emotion data using a camera or microphone

[1074] Specific explanation

[1075] While the user is typing, the device's camera and microphone are activated to analyze the user's facial expressions and tone of voice in real time. An emotion engine processes this data to determine the user's emotional state.

[1076] input

[1077] User video and audio data

[1078] output

[1079] Analyzed emotion data (e.g., "joy," "excitement")

[1080] concrete action

[1081] The camera and microphone are activated to capture video and audio in real time, and the emotion engine analyzes the data and outputs a specific emotional state.

[1082] Step 3: Convert the data to JSON format

[1083] Specific explanation

[1084] The natural language data and emotion data acquired by the device is converted into JSON format.

[1085] input

[1086] Natural language input data, emotion data

[1087] output

[1088] JSON formatted dataset

[1089] concrete action

[1090] Natural language data and sentiment data are integrated and converted into JSON format programmatically.

[1091] Step 4: Sending data to the server

[1092] Specific explanation

[1093] The terminal sends the converted data set to the server using the HTTPS protocol.

[1094] input

[1095] JSON formatted dataset

[1096] output

[1097] The server receives the JSON-formatted dataset.

[1098] concrete action

[1099] The device generates an HTTPS request and sends data to the server's URL.

[1100] Step 5: Receiving and analyzing data on the server

[1101] Specific explanation

[1102] The server parses the received JSON data and extracts natural language and sentiment data. The NLP engine identifies keywords from the natural language data, and the sentiment analysis algorithm extracts the emotional state.

[1103] input

[1104] JSON formatted dataset

[1105] output

[1106] Identified keywords (e.g., "spring," "cherry blossoms," "wallpaper"), emotional states (e.g., "joy")

[1107] concrete action

[1108] The server parses the JSON data and analyzes it using an NLP engine and sentiment analysis algorithms.

[1109] Step 6: Format the data

[1110] Specific explanation

[1111] The server integrates the extracted keywords and sentiment data to generate appropriate prompts for the generative AI model.

[1112] input

[1113] Identified keywords, emotional states

[1114] output

[1115] Prompt statement (e.g., "Spring cherry blossom wallpaper that inspires joy")

[1116] concrete action

[1117] The server combines the keywords and emotion data and formats them into a prompt sentence.

[1118] Step 7: Generate original design data

[1119] Specific explanation

[1120] The server inputs the prompt sentence into the generative AI model and generates design data.

[1121] input

[1122] Prompt statement

[1123] output

[1124] Generated original design data (e.g., "Joyful Spring Cherry Blossom Wallpaper")

[1125] concrete action

[1126] The generative AI model generates design data that fits the theme based on the prompt text.

[1127] Step 8: Convert your design data to JSON format and prepare it for submission

[1128] Specific explanation

[1129] The generated design data is converted back into JSON format and prepared for transmission to the user's device using the HTTPS protocol.

[1130] input

[1131] Generated design data

[1132] output

[1133] Design data in JSON format

[1134] concrete action

[1135] The server converts the generated design data into JSON format and prepares it for transmission via HTTPS protocol.

[1136] Step 9: Sending and receiving design data to and from your device

[1137] Specific explanation

[1138] The server sends JSON format design data to the device, which then receives it.

[1139] input

[1140] Design data in JSON format

[1141] output

[1142] The device receives the design data

[1143] concrete action

[1144] The server sends the design data to the terminal via an HTTPS request, and the terminal receives it.

[1145] Step 10: Applying design data

[1146] Specific explanation

[1147] The device parses the received design data and obtains the original wallpaper image data, which the device then instantly sets as the smartphone's wallpaper.

[1148] input

[1149] Design data in JSON format

[1150] output

[1151] The device wallpaper has been changed

[1152] concrete action

[1153] The device parses the received JSON data, obtains the wallpaper image data, and immediately sets it as the wallpaper.

[1154] (Application example 2)

[1155] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1156] When workers use industrial equipment, they need specific work procedures and guidance. However, there are no systems that can flexibly respond to the worker's emotions and situation. In particular, emotions such as tension and anxiety can affect work efficiency, so appropriate support is needed.

[1157] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing preferences, images, and emotions input by the user in natural language and extracting keywords, means for shaping input data to a generative AI model using the extracted keywords and emotions, and means for generating original design data based on the shaped input data by the generative AI model. This makes it possible to generate and display customized work guides in real time according to the instructions and emotions input by the worker in natural language.

[1158] A "natural language" is a human language that is used by users in normal conversation and writing, and that can be converted into a form that a computer can parse.

[1159] "Emotion" refers to a psychological state that can be recognized from a user's facial expression, tone of voice, etc., and includes, for example, states such as joy, sadness, surprise, and tension.

[1160] "Keywords" refer to important information or characteristic terms extracted from a user's natural language input and sentiment data.

[1161] A "generative artificial intelligence model" refers to an algorithm or system that generates new content or designs based on input data.

[1162] "Original design data" refers to new and unique design data generated by a generative artificial intelligence model based on a user's input and emotional state.

[1163] "Terminal" refers to electronic devices used by users, such as smartphones, tablets, and personal computers.

[1164] "Industrial equipment" refers to machinery and equipment used in manufacturing and production processes, including equipment operated by workers.

[1165] "Work guide" means visual or textual instructions that show a worker how to perform a particular work task.

[1166] MODE FOR CARRYING OUT THE INVENTION

[1167] The system embodying this invention is configured to provide customized operation guides based on the user's instructions and emotions. It has the function of analyzing preferences, images, and emotion data input by the user in natural language, and generating and displaying appropriate operation guides in real time.

[1168] Hardware and Software Configuration

[1169] The system mainly consists of the following hardware and software:

[1170] Hardware

[1171] Smart glasses or head-mounted display (HMD): A device that allows users to provide natural language instructions and emotional input.

[1172] Camera and microphone: Devices for analyzing the user's facial expressions and tone of voice to obtain emotional data.

[1173] software

[1174] Emotion recognition engine (e.g., Microsoft Azure Cognitive Services Emotion API): Analyzes emotions from the user's facial expressions and tone of voice and obtains data.

[1175] Natural language processing (NLP) engine (e.g., Google Cloud Natural Language API): Analyzes the user's natural language input and extracts keywords.

[1176] Generative AI models (e.g., OpenAI GPT-4): Generate customized work guides and design data based on user keywords and emotional data.

[1177] System Operation

[1178] The system operates as follows.

[1179] User interface: The user wears smart glasses and inputs specific instructions for the task in natural language. For example, they might give instructions such as, "Tell me how to tighten a bolt." At this time, the emotion recognition engine analyzes the user's facial expressions and tone of voice via a camera and microphone, and obtains emotional data such as tension or anxiety.

[1180] Data analysis: The natural language data and emotion data entered by the user are acquired and converted into JSON format. This allows the data sent to the server to be processed in a unified format. The server receives the JSON data, analyzes it using a natural language processing engine, and extracts keywords and emotion data.

[1181] Design data generation: The server uses a generative AI model to generate original work guides and design data based on the extracted keywords and emotion data. This generated data includes appropriate procedures and guides that take emotion into consideration.

[1182] Data application: The generated work guide data is converted back to JSON format and sent to the user's device. The smart glasses parse the received design data and apply it to the user's vision. For example, the bolt tightening procedure is displayed as an illustrated guide in the user's field of vision.

[1183] Specific examples

[1184] For example, if a worker says in natural language, "Tell me how to tighten a bolt," and the emotion recognition engine detects that the user is nervous, the server analyzes this data and uses a generative AI model to generate "bolt tightening procedures to relieve tension." The worker's smart glasses display easy-to-understand procedures and illustrated guides that instill a sense of security.

[1185] Prompt Sentence Examples

[1186] User Instructions: What is the procedure for tightening the bolts?

[1187] User Emotion: Tension

[1188] This prompt enables the generative AI model to provide optimal task guidance.

[1189] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1190] Step 1:

[1191] The user puts on the smart glasses and inputs specific work instructions in natural language. For example, they might say, "Tell me how to tighten a bolt." At this time, the user's facial expressions and tone of voice are recorded via a camera and microphone. The input is captured as natural language data and sent to an emotion recognition engine. The acquired natural language data, facial expression data, and voice data are then passed to the respective processing modules, where analysis begins.

[1192] Step 2:

[1193] The device's emotion recognition engine analyzes the acquired facial and voice data to identify the user's emotions. For example, if the user is nervous, the device identifies that state of tension and records it as emotion data. The input facial and voice data is converted into emotion data. The output emotion data is "tension."

[1194] Step 3:

[1195] The device converts the user's natural language input data and the acquired emotion data into JSON format. A JSON object is generated with the natural language data and emotion data stored in their respective fields. The generated JSON data is ready to be sent to the server in the next step.

[1196] Step 4:

[1197] The device sends the generated JSON data to the server using the HTTPS protocol. The JSON data, including the user's input amount and emotion data, is securely sent to the server. The server receives the HTTPS request and proceeds to the next step.

[1198] Step 5:

[1199] The server parses the received JSON data and extracts natural language data and emotion data separately. Keywords such as "bolt," "tightening," and "procedure" and the emotion "tension" are obtained from the parsed data. The NLP engine on the server analyzes the meaning and organizes the keywords.

[1200] Step 6:

[1201] The server creates a prompt for the generative AI model based on the extracted keywords and emotion data. The prompt is formatted, for example, as "Tell me how to tighten a bolt," including the word "tension." The prompt is sent to the generative AI model, which then begins generating design data.

[1202] Step 7:

[1203] Based on the prompt received, the generative AI model generates original work guides and design data that take emotions into account. For example, it generates "bolt tightening procedures to relieve tension." The generated design data is then returned to the server.

[1204] Step 8:

[1205] The server converts the generated design data back into JSON format and prepares it for transmission to the device. The data organized into JSON objects is sent to the device via HTTPS.

[1206] Step 9:

[1207] The device parses the JSON data received from the server and extracts the generated work guide and design data, which the device receives, parses, and converts into visual information such as an illustration guide.

[1208] Step 10:

[1209] The device then displays the received design data on the smart glasses display, and the user is shown a "bolt tightening procedure to ease tension," allowing them to proceed with the work with peace of mind.

[1210] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1211] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1212] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1213] [Fourth embodiment]

[1214] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1215] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1216] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1217] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1218] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1219] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1220] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1221] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1222] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1223] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1224] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1225] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1226] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1227] This invention is a system for customizing the design of a mobile phone based on preferences and images input by the user in natural language. This system is carried out by the cooperation of the user, the terminal, and the server.

[1228] Program processing explanation

[1229] Getting user-entered data

[1230] Using a dedicated app, the user inputs content in natural language, such as "I want a wallpaper with a spring cherry blossom theme." When the user presses the "Send" button, the device receives the input data and proceeds to the next step.

[1231] Sending data

[1232] The terminal converts the natural language data entered by the user into JSON format, and then transmits the converted JSON data to the server using the HTTPS protocol, ensuring accurate transmission of the data.

[1233] Data reception and analysis

[1234] The server receives the HTTPS request, parses the JSON data sent, and retrieves the user's input. The server uses a natural language processing (NLP) engine to extract meaning from the input text and identify keywords such as "spring," "cherry blossoms," and "wallpaper."

[1235] Data Formatting

[1236] The server then uses the extracted keywords to format the data required by the generative AI model. Specifically, it prepares keywords such as "spring," "cherry blossoms," and "wallpaper" to be input into the model in array format.

[1237] Data generation

[1238] The server inputs the formatted data into a generative AI model, which then generates original design data, such as a wallpaper with a "spring cherry blossom" theme.

[1239] Sending generated data

[1240] The server converts the generated wallpaper image data back into JSON format and prepares to send it to the device. The wallpaper image data is sent to the user's device using the HTTPS protocol.

[1241] Receiving and applying data

[1242] The device receives the wallpaper image data and parses the JSON data to obtain the image data. The device then executes a system call to immediately apply the image data as the smartphone's wallpaper. As a result, the user's smartphone wallpaper is set to the new original "Spring Cherry Blossom" theme.

[1243] Specific examples

[1244] For example, consider a case where a user inputs, "I want a summer beach-themed icon set." When the user's input is transmitted and analyzed, the keywords "summer," "beach," and "icon" are extracted. The server uses a generative AI model to generate a new icon set based on this input and transmits the generated icon set to the device. The device then applies the received icon set to the user's smartphone.

[1245] This invention allows users to easily enjoy creating original smartphone designs simply by intuitively inputting their preferences and ideas using natural language. AI automatically generates professional designs, independent of the user's taste or skills, and instantly applies them, making everyday life more appealing.

[1246] The processing flow will be explained below.

[1247] Step 1:

[1248] User:

[1249] Users start up a dedicated smartphone app and enter their preferences and images in natural language into the input field that appears, such as "I want a wallpaper with an autumn leaf theme." Once they've finished entering the information, they press the "Send" button.

[1250] Step 2:

[1251] Device:

[1252] The terminal receives input data from the user, converts it into JSON format, and then sends the converted JSON data to the server using the HTTPS protocol.

[1253] Step 3:

[1254] server:

[1255] The server receives the HTTPS request, parses the JSON data sent, and extracts the natural language input from the parsed data.

[1256] Step 4:

[1257] server:

[1258] The server uses a natural language processing (NLP) engine to analyze the meaning of the natural language data it receives. Specifically, it extracts keywords such as "autumn," "autumn leaves," and "wallpaper" from the input "I want a wallpaper with an autumn foliage theme."

[1259] Step 5:

[1260] server:

[1261] The server then uses the extracted keywords to format the data appropriately for use as input for a generative AI model. This formatting process results in a dataset containing keywords such as "autumn," "autumn leaves," and "wallpaper."

[1262] Step 6:

[1263] server:

[1264] The formatted data is input into a generative AI model, which then generates original wallpaper data based on this data.

[1265] Step 7:

[1266] server:

[1267] The generated wallpaper data is received, converted to JSON format, and prepared for transmission to the user's device. This wallpaper data is sent to the device using the HTTPS protocol.

[1268] Step 8:

[1269] Device:

[1270] The device receives the wallpaper data sent from the server, parses the JSON data, and obtains the original wallpaper image data.

[1271] Step 9:

[1272] Device:

[1273] The acquired original wallpaper image data is immediately set as the smartphone wallpaper. This setting changes the user's smartphone wallpaper to an original one with an "Autumn Leaves" theme.

[1274] Example 1

[1275] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1276] Today's smartphone users want to be able to easily customize their devices to suit their tastes and preferences. However, existing services often make customization difficult, relying on the user's taste and skills. Furthermore, the interface for users to specify designs is often not intuitive, making it difficult to achieve high-quality customization easily. Furthermore, for users who are not technically savvy, the complex and time-consuming operation poses a problem.

[1277] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1278] In this invention, the server includes means for acquiring preferences and images input by a user in natural language, means for converting the acquired preferences and images into JSON format, means for transmitting the converted JSON format data using the HTTPS protocol, means for receiving and analyzing the transmitted JSON format data, means for extracting keywords from the analyzed data, means for formatting input data for a generative artificial intelligence model using the extracted keywords, means for generating original design data based on the formatted input data using the generative artificial intelligence model, means for converting the generated design data back into JSON format and transmitting it using the HTTPS protocol, and means for analyzing the received design data and applying it to the terminal. This allows a user to intuitively and quickly customize the design of their smartphone simply by inputting their preferences in natural language.

[1279] "Means for acquiring preferences and images input by the user in natural language" refers to a function that allows the user to input their wishes and images in natural language and receives the input content from the system.

[1280] "Means for converting to JSON format" is a function that converts natural language data entered by the user into a standardized data format called JavaScript Object Notation (JSON).

[1281] "Means for sending using the HTTPS protocol" refers to a function that uses a protocol called HTTP Secure (HTTPS) to send data to a server in a secure and encrypted manner.

[1282] "Means for receiving and analyzing transmitted JSON format data" refers to a function in which a server receives JSON data transmitted via the HTTPS protocol and analyzes the contents of that data.

[1283] "Means for extracting keywords from analyzed data" refers to a function that analyzes received data using natural language processing and identifies and extracts important keywords.

[1284] "Means for formatting input data for a generative artificial intelligence model" is a function that converts extracted keywords into data in a format that is easy for the generative artificial intelligence model to understand.

[1285] "Means for generating original design data based on input data generated by a generative artificial intelligence model" refers to a function in which a generative artificial intelligence automatically creates design data that meets the user's wishes based on formatted input data.

[1286] "Means for converting the generated design data back into JSON format and transmitting it using the HTTPS protocol" refers to a function for converting the generated design data back into JSON format and transmitting it to the terminal using the HTTPS protocol.

[1287] The "means for analyzing received design data and applying it to the terminal" is a function that analyzes the JSON format design data received by the terminal and applies the specified design to the terminal.

[1288] This invention is a system that customizes a mobile information terminal based on preferences and images input by the user in natural language. This system performs processing through collaboration between the user, the terminal, and the server. The user uses a dedicated app to input specific wishes and images in natural language. For example, the user might input, "I want a wallpaper with a spring cherry blossom theme."

[1289] The device receives natural language data entered by the user and converts it into JSON format. The converted data is then sent to the server using the HTTPS protocol. The server analyzes the received JSON data and extracts keywords using a natural language processing engine (libraries such as TensorFlow and PyTorch). The extracted keywords are then used to format the input data for a generative AI model, which then generates original design data.

[1290] The generated design data is converted back to JSON format and sent to the device using the HTTPS protocol. The device then analyzes the received data and extracts the image data. Finally, the device applies the extracted image data as the smartphone's wallpaper or icon.

[1291] For example, if a user types "I want a summer beach-themed icon set," the system will extract the keywords "summer," "beach," and "icon," and the AI ​​model will generate a new icon set, resulting in a summer beach-themed icon set that will be applied to the user's smartphone.

[1292] Examples of specific prompts include:

[1293] "I want a wallpaper with a spring cherry blossom theme."

[1294] "I want a summer beach themed icon set."

[1295] "Create an autumn leaf-themed lock screen"

[1296] Based on these prompts, the generative AI model creates a design that meets the user's wishes. Users can easily enjoy creating their own original smartphone design by simply inputting their wishes in natural language.

[1297] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1298] Step 1:

[1299] The user opens the app and enters a prompt in natural language, such as "I want a spring cherry blossom themed wallpaper," into the input field. When the user presses the "Send" button, the input data is sent to the device.

[1300] Input: A natural language prompt entered by the user.

[1301] Output: The prompt is sent to the terminal.

[1302] Step 2:

[1303] The device receives the natural language data entered by the user and converts it into JavaScript Object Notation (JSON) format, such as "message: 'I want a wallpaper with a spring cherry blossom theme'".

[1304] Input: A prompt sentence typed in natural language.

[1305] Output: Data converted to JSON format.

[1306] Step 3:

[1307] The device sends the converted JSON data to the server using the HTTPS protocol, which may also include data integrity checks and retransmission mechanisms.

[1308] Input: Data converted to JSON format.

[1309] Output: Data sent to the server using the HTTPS protocol.

[1310] Step 4:

[1311] The server receives the HTTPS request and analyzes the JSON data sent. Specifically, it parses the received data and obtains the contents of the "message" field.

[1312] Input: JSON data sent over HTTPS.

[1313] Output: Parsed content (prompt sentence).

[1314] Step 5:

[1315] The server uses a natural language processing engine (such as TensorFlow or PyTorch) to extract keywords from the prompt. In this example, "spring," "cherry blossoms," and "wallpaper" are extracted.

[1316] Input: The parsed prompt sentence.

[1317] Output: Extracted keywords.

[1318] Step 6:

[1319] The server formats the data based on the extracted keywords into a format that the generative AI model can understand, such as a JSON object with the format "keywords: ['spring', 'cherry blossoms', 'wallpaper']".

[1320] Input: Extracted keywords.

[1321] Output: The formatted data.

[1322] Step 7:

[1323] The server inputs the formatted data into a generative AI model, which then generates original design data (e.g., wallpaper images) based on the data.

[1324] Input: Formatted data.

[1325] Output: Generated design data (wallpaper image).

[1326] Step 8:

[1327] The server converts the generated design data into JSON format and sends it to the terminal using the HTTPS protocol.

[1328] Input: Generated design data.

[1329] Output: Design data converted to JSON format.

[1330] Step 9:

[1331] The device receives the sent JSON data, parses it, and obtains the image data. Specifically, it extracts the "image" field from the JSON data and decodes the Base64-encoded image.

[1332] Input: Design data submitted in JSON format.

[1333] Output: Decoded image data.

[1334] Step 10:

[1335] The device executes a system call to apply the acquired image data as the smartphone's wallpaper. Ultimately, the user's smartphone wallpaper is changed to the new original "Spring Cherry Blossom" theme.

[1336] Input: Decoded image data.

[1337] Output: The new wallpaper applied.

[1338] (Application example 1)

[1339] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1340] In virtual stores, there are inconvenient ways to quickly and intuitively reflect the preferences and images that users input in natural language and customize the store design, so there is a need for technology that can generate attractive designs in real time based on user input and apply them immediately.Furthermore, there is also a need to easily save and share the generated designs.

[1341] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1342] In this invention, the server includes means for acquiring preferences and images input by a user in natural language, means for analyzing the acquired preferences and images and extracting keywords, means for shaping input data to a generative artificial intelligence model using the extracted keywords, means for generating original design data based on the shaped input data using the generative artificial intelligence model, means for transmitting the generated design data to the user's terminal, means for applying the received design data to the terminal, means for applying the generated design in real time in a virtual space, and means for saving and sharing the generated design. This allows a user to intuitively customize the design of a virtual store in real time based on input in natural language, and to save and share the generated design.

[1343] "Natural language" refers to the language that a user normally uses in conversation or writing, without being converted into a form that is easy for the system to parse.

[1344] "Preferences and images" refers to the user's personal hobbies and visual concepts / themes.

[1345] "Keyword extraction" refers to the process of extracting meaningful words and phrases from a user's natural language input.

[1346] A "generative artificial intelligence model" refers to an artificial intelligence that has the ability to automatically generate new designs and content based on input data and conditions.

[1347] "Formatting input data" refers to the process of converting data into a format that is easy for a generative artificial intelligence model to understand.

[1348] "Original design data" refers to unique design information that is newly generated by a generative artificial intelligence model and does not exist anywhere else.

[1349] "Sending to the user's device" refers to sending the generated design data to the device used by the user using a communication means such as the Internet.

[1350] "Applying design data" refers to reflecting the received design data on the user's device and making it actually displayable and usable.

[1351] "Virtual space" refers to a 3D space or virtual store generated by a computer.

[1352] "Real-time application" refers to the process of updating and displaying the design so that it reflects user input almost immediately.

[1353] "Saving a design" refers to saving the generated design data to a storage medium for later use.

[1354] "Means of sharing" refers to the methods by which the generated design can be published and transferred to other users or platforms.

[1355] This invention is a system that customizes the design of a virtual store based on preferences and images input by the user in natural language. This system executes processing in cooperation with the user, the terminal, and the server.

[1356] Using a smartphone, smart glasses, or head-mounted display, users can request a virtual store design in natural language, for example, by typing, "I want a store with an autumn leaf theme."

[1357] The device converts the user's input into JSON format and sends it to the server using the HTTPS protocol. The server analyzes the received JSON data and uses a natural language processing engine to extract keywords such as "autumn," "autumn leaves," and "store." It then formats these keywords into the format required by the generative artificial intelligence model and inputs them into the model.

[1358] The generative AI model generates original design data based on the given keywords. This design data is then converted back to JSON format and sent to the user's device via HTTPS. The device parses and retrieves the received design data and immediately applies it to the virtual store.

[1359] Furthermore, the generated designs can be saved on the device and shared with other users and platforms, allowing users to intuitively make design requests in natural language, customize the design of their virtual store in real time, and even save and share the designs.

[1360] The hardware requires a smartphone, smart glasses, or a head-mounted display to receive user input, and the software uses Python 3.x, the Requests library, and specific API endpoints to apply the virtual store design.

[1361] For example, if a user inputs "I want a Halloween-themed store," the server extracts the keywords "Halloween," "theme," and "store" and inputs them into a generative AI model. Based on this prompt, the AI ​​model generates a virtual store design, which is then instantly applied.

[1362] Example prompt sentence:

[1363] Theme: Halloween

[1364] Design: Pumpkins, bats, dark colors

[1365] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1366] Step 1:

[1367] The user inputs a design request for the virtual store in natural language. For example, they might input, "I want a store with an autumn foliage theme." The input is sent to the terminal.

[1368] Step 2:

[1369] The terminal receives the user's input and converts it into JSON format. Specifically, it converts it into a JSON object with the natural language input as a key. The converted JSON data is then sent to the server.

[1370] Step 3:

[1371] The server receives the JSON data sent from the device and analyzes it using a natural language processing engine. Keywords such as "autumn," "autumn leaves," and "store" are extracted from the input data. The results of this analysis become the input for the next processing step.

[1372] Step 4:

[1373] Based on the extracted keywords, the server formats the data to be input into the generative AI model. Specifically, it converts the keywords into a list format and generates a prompt sentence. The keywords "autumn," "autumn leaves," and "store" become the elements of the prompt sentence.

[1374] Step 5:

[1375] The server inputs the formatted data into a generative AI model, which generates original design data based on the prompt. This generated design data becomes the input for the next processing step.

[1376] Step 6:

[1377] The server converts the generated design data back into JSON format and sends it to the terminal using the HTTPS protocol. Specifically, it encodes the design data into a JSON object and sends a POST request to the endpoint.

[1378] Step 7:

[1379] The device receives the JSON data sent from the server and parses it to obtain the design data. The parsing process extracts the JSON format data as specific design data.

[1380] Step 8:

[1381] The terminal immediately applies the acquired design data to the virtual store by sending a request to the API endpoint of the virtual space to reflect the design data.

[1382] Step 9:

[1383] The device provides functions for saving the generated design and sharing it with other users or platforms. Specifically, it saves the design data in local storage and generates a link or QR code for sharing it.

[1384] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1385] This invention relates to a system that realizes advanced design customization by combining preferences and images input by the user in natural language with an emotion engine that recognizes the user's emotions. This system performs processing in cooperation with the user, the terminal, and the server.

[1386] Program processing explanation

[1387] Getting user-entered data

[1388] Using a dedicated app, users input content in natural language, such as "I want a wallpaper with a spring cherry blossom theme." When the user presses the "send" button, the device receives the input data and proceeds to the next step. The device is also equipped with a camera and microphone, and the emotion engine analyzes the user's facial expressions and tone of voice, capturing their emotions as data.

[1389] Sending data

[1390] The device converts the input data from the user into JSON format. At the same time, the emotion data recognized by the emotion engine is also added to the JSON format. The converted JSON data is then sent to the server using the HTTPS protocol.

[1391] Data reception and analysis

[1392] The server receives the HTTPS request and parses the JSON data sent. From the parsed data, it extracts the natural language content and emotion data of the user's input. Using a natural language processing (NLP) engine, it analyzes the semantics of the natural language data and identifies keywords such as "spring," "cherry blossoms," and "wallpaper." The emotion data is also analyzed to extract the user's emotional state (e.g., joy, sadness, surprise, etc.) at the time of input.

[1393] Data Formatting

[1394] The server then uses the extracted keywords and recognized emotions to format the data appropriately for use as input for a generative AI model. For example, it generates a dataset that combines keywords such as "spring," "cherry blossoms," and "wallpaper" with the user's emotion of "happy."

[1395] Data generation

[1396] The server inputs the formatted data into a generative AI model, which then generates original design data, such as wallpaper data with the theme of "joyful spring cherry blossoms."

[1397] Sending generated data

[1398] The server receives the generated wallpaper data, converts it back to JSON format, and prepares it for transmission to the user's device. This wallpaper data is sent to the device using the HTTPS protocol.

[1399] Receiving and applying data

[1400] The device receives the wallpaper data sent from the server, parses the JSON data, and obtains the original wallpaper image data. The device immediately sets the image data as the smartphone's wallpaper. This setting changes the user's smartphone wallpaper to the original wallpaper with the theme "Joyful Spring Cherry Blossoms."

[1401] Specific examples

[1402] For example, consider the case where a user inputs "I want an icon set with a summer beach theme," and at the same time, the emotion engine recognizes the user's "excitement." The acquired data includes the keywords "summer," "beach," and "icon set," as well as the emotion "excitement." Based on this data, the server uses a generative AI model to generate a new icon set with the theme of "exciting summer beach" and sends it to the device. The device immediately applies the received icon set, and the user's smartphone is updated with the new icon set.

[1403] This invention allows users to intuitively input their preferences and ideas using natural language, and easily experience original smartphone designs that reflect their emotions. AI and emotion recognition technology work together to generate professional designs that can be instantly applied, without relying on the user's taste or skills, making everyday life more fulfilling.

[1404] The processing flow will be explained below.

[1405] Program processing explanation

[1406] Step 1:

[1407] User:

[1408] The user launches a dedicated smartphone app and enters their preferences and image in natural language, such as "I want a wallpaper with an autumn leaf theme," into the input field that appears. When the user presses the "Send" button, the emotion engine begins analyzing the user's emotions through their voice and facial expressions. The emotion engine generates emotion data such as "happy" or "excited" from the user's facial expressions and tone of voice.

[1409] Step 2:

[1410] Device:

[1411] The device receives input data from the user and converts it into JSON format. At the same time, the emotion data output by the emotion engine is also added to the JSON format. The converted JSON data is structured to include the user's input and emotion data. It is then sent to the server using the HTTPS protocol.

[1412] Step 3:

[1413] server:

[1414] The server receives the HTTPS request and parses the JSON data. From the parsed data, it extracts the natural language content and emotion data entered by the user. Specifically, it extracts the text data "I want a wallpaper with an autumn leaf theme" and the emotion data "I'm excited."

[1415] Step 4:

[1416] server:

[1417] The server uses a natural language processing (NLP) engine to analyze the meaning of the acquired natural language data. Keywords such as "autumn," "autumn leaves," and "wallpaper" are extracted from the text data. Emotion data is analyzed in the same way, and the keyword "excitement" is extracted.

[1418] Step 5:

[1419] server:

[1420] Based on the extracted keywords and recognized emotions, the server formats the data into an appropriate format for use as input for the generative AI model. Specifically, it generates a dataset consisting of the keywords "autumn," "autumn leaves," and "wallpaper" and the emotion data "excitement." This dataset is used as input for the generative AI model.

[1421] Step 6:

[1422] server:

[1423] The formatted data set is then input into a generative AI model, which then generates original design data. For example, it generates an exciting wallpaper image featuring vivid autumn leaves.

[1424] Step 7:

[1425] server:

[1426] The generated wallpaper data is received, converted back to JSON format, and prepared for transmission to the user's device. This wallpaper data is sent to the device using the HTTPS protocol.

[1427] Step 8:

[1428] Device:

[1429] The device receives the wallpaper data sent from the server, parses the JSON data, and obtains the original wallpaper image data.

[1430] Step 9:

[1431] Device:

[1432] The acquired original wallpaper image data is immediately set as the smartphone wallpaper, so that the user's smartphone wallpaper is set to an original design with an "Autumn Leaves" theme that reflects the user's emotion of "Excitement."

[1433] Specific examples

[1434] For example, if a user inputs "I want a summer beach-themed icon set," and the emotion engine recognizes the user's "relaxation," the retrieved data includes the keywords "summer," "beach," and "icon set" along with the emotion "relaxation." This data is sent to the server, which then inputs the formatted data into a generative AI model to generate a new icon set with the theme of "relaxing summer beach." The generated icon set is then sent to the device and instantly applied to the user's smartphone.

[1435] The present invention allows users to easily enjoy original designs that reflect their own tastes and feelings, making their daily lives more enriching.

[1436] Example 2

[1437] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1438] In the field of modern digital design, there is a growing need for users to easily input their preferences and preferences in natural language and generate customized, original designs based on those inputs. However, current systems have difficulty appropriately reflecting users' emotions, extracting keywords from the input natural language, and generating designs that take emotions into account. Furthermore, the process of immediately applying the generated designs to the user's device is not efficient enough. This can sometimes result in a poor user experience.

[1439] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1440] In this invention, the server includes means for acquiring preferences and images input by a user in natural language, means for analyzing the acquired preferences and images to extract keywords, means for integrating the extracted keywords and user emotion data to format input data for a generative AI model, means for analyzing the user's facial expressions and voice using a camera and microphone to acquire emotion data, means for generating original design data based on the formatted input data using the generative AI model, means for transmitting the generated design data to the user's terminal, and means for applying the received design data to the terminal. This enables the generation of original designs that reflect the user's preferences and emotions and their immediate application.

[1441] "User" refers to an individual who uses the system to input preferences and ideas in natural language.

[1442] A "dedicated app" refers to a software application that allows users to input their preferences and images and communicate with a server.

[1443] "Terminal" refers to a device used by a user, such as a smartphone or tablet.

[1444] An "emotion engine" refers to a software module that analyzes a user's facial expressions and tone of voice to recognize their emotional state.

[1445] "JSON" stands for JavaScript Object Notation and refers to a lightweight data interchange format for structuring data and communicating it over the Internet.

[1446] "HTTPS protocol" stands for Hypertext Transfer Protocol Secure and refers to a communication protocol for securely transmitting data over the Internet.

[1447] "Server" refers to a computer system that receives, analyzes, and processes data sent by a user and sends the generated data to the user's terminal.

[1448] A "natural language processing (NLP) engine" refers to a software module that analyzes natural language input by a user and extracts meaning and keywords.

[1449] A "generative artificial intelligence model" refers to a machine learning algorithm for generating original design data based on input data.

[1450] "Design Data" refers to original visual or functional designs generated based on user input and emotional data.

[1451] This invention is a system that realizes more advanced design customization by combining preferences and images input by the user in natural language with an emotion engine that recognizes the user's emotions. This system is carried out by the cooperation of the user, the terminal, and the server.

[1452] First, a user launches a dedicated app on a device such as a smartphone or tablet and inputs the desired design in natural language. For example, they might input something like, "I want a wallpaper with a spring cherry blossom theme." When the user presses the "send" button, the device captures this text data. At the same time, an emotion engine is activated on the device, which has a built-in camera and microphone, and analyzes the user's facial expressions and tone of voice to capture emotional data in real time. This emotional data includes emotional states such as "joy" and "excitement."

[1453] The device converts the acquired data (natural language data and emotion data) into JSON format, and then transmits this JSON-formatted dataset to the server using the HTTPS protocol, which ensures data security and confidentiality.

[1454] The server receives the HTTPS request and parses the JSON data sent. The analysis module extracts the natural language and sentiment data entered by the user from the parsed data. Using a natural language processing (NLP) engine, it identifies keywords such as "spring," "cherry blossoms," and "wallpaper," and then uses a sentiment analysis algorithm to extract the emotional state (e.g., "joy").

[1455] The server then combines the extracted keywords and emotion data and formats them into a format appropriate for the generative AI model. For example, it combines the keywords "spring," "cherry blossoms," and "wallpaper" with "joy" to create a prompt sentence to be input into the generative AI model.

[1456] Based on the input prompts, the generative AI model generates original design data that reflects the user's preferences and emotions. Specifically, wallpaper data with the theme of "joyful spring cherry blossoms" is generated.

[1457] The generated design data is converted back to JSON format by the server and sent to the user's device using the HTTPS protocol. The device receives this data, parses the JSON data, and obtains the original wallpaper image data. The device then immediately sets the obtained image data as the smartphone's wallpaper. As a result, the user's smartphone wallpaper is changed to the original wallpaper with the theme "Joyful Spring Cherry Blossoms."

[1458] As a concrete example, consider the case where a user inputs "I want a summer beach-themed icon set," and the emotion engine recognizes the user's "excitement." The acquired dataset contains the keywords "summer," "beach," and "icon set," as well as the emotion "excitement." The server inputs this data into the generative AI model and generates a new icon set with the theme of "exciting summer beach." This icon set is sent to the device, which immediately applies the received icon set, and the user's smartphone is updated with the new icon set.

[1459] In this way, users can intuitively input their preferences and images using natural language, and easily experience original smartphone designs that reflect their emotions.This system makes it possible to generate professional designs and apply them immediately, without relying on the user's taste or skills.

[1460] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1461] Step 1: Getting user-entered data

[1462] Specific explanation

[1463] Using a dedicated app, users input their preferences and images in natural language, such as "I want a wallpaper with a spring cherry blossom theme." When the user presses the "Send" button, the device acquires this text data.

[1464] input

[1465] User natural language input (e.g., "I want a wallpaper with a spring cherry blossom theme")

[1466] output

[1467] The text data is saved on the device.

[1468] concrete action

[1469] The terminal captures the user input in text format and stores it in a variable.

[1470] Step 2: Acquire emotion data using a camera or microphone

[1471] Specific explanation

[1472] While the user is typing, the device's camera and microphone are activated to analyze the user's facial expressions and tone of voice in real time. An emotion engine processes this data to determine the user's emotional state.

[1473] input

[1474] User video and audio data

[1475] output

[1476] Analyzed emotion data (e.g., "joy," "excitement")

[1477] concrete action

[1478] The camera and microphone are activated to capture video and audio in real time, and the emotion engine analyzes the data and outputs a specific emotional state.

[1479] Step 3: Convert the data to JSON format

[1480] Specific explanation

[1481] The natural language data and emotion data acquired by the device is converted into JSON format.

[1482] input

[1483] Natural language input data, emotion data

[1484] output

[1485] JSON formatted dataset

[1486] concrete action

[1487] Natural language data and sentiment data are integrated and converted into JSON format programmatically.

[1488] Step 4: Sending data to the server

[1489] Specific explanation

[1490] The terminal sends the converted data set to the server using the HTTPS protocol.

[1491] input

[1492] JSON formatted dataset

[1493] output

[1494] The server receives the JSON-formatted dataset.

[1495] concrete action

[1496] The device generates an HTTPS request and sends data to the server's URL.

[1497] Step 5: Receiving and analyzing data on the server

[1498] Specific explanation

[1499] The server parses the received JSON data and extracts natural language and sentiment data. The NLP engine identifies keywords from the natural language data, and the sentiment analysis algorithm extracts the emotional state.

[1500] input

[1501] JSON formatted dataset

[1502] output

[1503] Identified keywords (e.g., "spring," "cherry blossoms," "wallpaper"), emotional states (e.g., "joy")

[1504] concrete action

[1505] The server parses the JSON data and analyzes it using an NLP engine and sentiment analysis algorithms.

[1506] Step 6: Format the data

[1507] Specific explanation

[1508] The server integrates the extracted keywords and sentiment data to generate appropriate prompts for the generative AI model.

[1509] input

[1510] Identified keywords, emotional states

[1511] output

[1512] Prompt statement (e.g., "Spring cherry blossom wallpaper that inspires joy")

[1513] concrete action

[1514] The server combines the keywords and emotion data and formats them into a prompt sentence.

[1515] Step 7: Generate original design data

[1516] Specific explanation

[1517] The server inputs the prompt sentence into the generative AI model and generates design data.

[1518] input

[1519] Prompt statement

[1520] output

[1521] Generated original design data (e.g., "Joyful Spring Cherry Blossom Wallpaper")

[1522] concrete action

[1523] The generative AI model generates design data that fits the theme based on the prompt text.

[1524] Step 8: Convert your design data to JSON format and prepare it for submission

[1525] Specific explanation

[1526] The generated design data is converted back into JSON format and prepared for transmission to the user's device using the HTTPS protocol.

[1527] input

[1528] Generated design data

[1529] output

[1530] Design data in JSON format

[1531] concrete action

[1532] The server converts the generated design data into JSON format and prepares it for transmission via HTTPS protocol.

[1533] Step 9: Sending and receiving design data to and from your device

[1534] Specific explanation

[1535] The server sends JSON format design data to the device, which then receives it.

[1536] input

[1537] Design data in JSON format

[1538] output

[1539] The device receives the design data

[1540] concrete action

[1541] The server sends the design data to the terminal via an HTTPS request, and the terminal receives it.

[1542] Step 10: Applying design data

[1543] Specific explanation

[1544] The device parses the received design data and obtains the original wallpaper image data, which the device then instantly sets as the smartphone's wallpaper.

[1545] input

[1546] Design data in JSON format

[1547] output

[1548] The device wallpaper has been changed

[1549] concrete action

[1550] The device parses the received JSON data, obtains the wallpaper image data, and immediately sets it as the wallpaper.

[1551] (Application example 2)

[1552] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1553] When workers use industrial equipment, they need specific work procedures and guidance. However, there are no systems that can flexibly respond to the worker's emotions and situation. In particular, emotions such as tension and anxiety can affect work efficiency, so appropriate support is needed.

[1554] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for analyzing preferences, images, and emotions input by the user in natural language and extracting keywords, means for shaping input data to a generative AI model using the extracted keywords and emotions, and means for generating original design data based on the shaped input data by the generative AI model. This makes it possible to generate and display customized work guides in real time according to the instructions and emotions input by the worker in natural language.

[1555] A "natural language" is a human language that is used by users in normal conversation and writing, and that can be converted into a form that a computer can parse.

[1556] "Emotion" refers to a psychological state that can be recognized from a user's facial expression, tone of voice, etc., and includes, for example, states such as joy, sadness, surprise, and tension.

[1557] "Keywords" refer to important information or characteristic terms extracted from a user's natural language input and sentiment data.

[1558] A "generative artificial intelligence model" refers to an algorithm or system that generates new content or designs based on input data.

[1559] "Original design data" refers to new and unique design data generated by a generative artificial intelligence model based on a user's input and emotional state.

[1560] "Terminal" refers to electronic devices used by users, such as smartphones, tablets, and personal computers.

[1561] "Industrial equipment" refers to machinery and equipment used in manufacturing and production processes, including equipment operated by workers.

[1562] "Work guide" means visual or textual instructions that show a worker how to perform a particular work task.

[1563] MODE FOR CARRYING OUT THE INVENTION

[1564] The system embodying this invention is configured to provide customized operation guides based on the user's instructions and emotions. It has the function of analyzing preferences, images, and emotion data input by the user in natural language, and generating and displaying appropriate operation guides in real time.

[1565] Hardware and Software Configuration

[1566] The system mainly consists of the following hardware and software:

[1567] Hardware

[1568] Smart glasses or head-mounted display (HMD): A device that allows users to provide natural language instructions and emotional input.

[1569] Camera and microphone: Devices for analyzing the user's facial expressions and tone of voice to obtain emotional data.

[1570] software

[1571] Emotion recognition engine (e.g., Microsoft Azure Cognitive Services Emotion API): Analyzes emotions from the user's facial expressions and tone of voice and obtains data.

[1572] Natural language processing (NLP) engine (e.g., Google Cloud Natural Language API): Analyzes the user's natural language input and extracts keywords.

[1573] Generative AI models (e.g., OpenAI GPT-4): Generate customized work guides and design data based on user keywords and emotional data.

[1574] System Operation

[1575] The system operates as follows.

[1576] User interface: The user wears smart glasses and inputs specific instructions for the task in natural language. For example, they might give instructions such as, "Tell me how to tighten a bolt." At this time, the emotion recognition engine analyzes the user's facial expressions and tone of voice via a camera and microphone, and obtains emotional data such as tension or anxiety.

[1577] Data analysis: The natural language data and emotion data entered by the user are acquired and converted into JSON format. This allows the data sent to the server to be processed in a unified format. The server receives the JSON data, analyzes it using a natural language processing engine, and extracts keywords and emotion data.

[1578] Design data generation: The server uses a generative AI model to generate original work guides and design data based on the extracted keywords and emotion data. This generated data includes appropriate procedures and guides that take emotion into consideration.

[1579] Data application: The generated work guide data is converted back to JSON format and sent to the user's device. The smart glasses parse the received design data and apply it to the user's vision. For example, the bolt tightening procedure is displayed as an illustrated guide in the user's field of vision.

[1580] Specific examples

[1581] For example, if a worker says in natural language, "Tell me how to tighten a bolt," and the emotion recognition engine detects that the user is nervous, the server analyzes this data and uses a generative AI model to generate "bolt tightening procedures to relieve tension." The worker's smart glasses display easy-to-understand procedures and illustrated guides that instill a sense of security.

[1582] Prompt Sentence Examples

[1583] User Instructions: What is the procedure for tightening the bolts?

[1584] User Emotion: Tension

[1585] This prompt enables the generative AI model to provide optimal task guidance.

[1586] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1587] Step 1:

[1588] The user puts on the smart glasses and inputs specific work instructions in natural language. For example, they might say, "Tell me how to tighten a bolt." At this time, the user's facial expressions and tone of voice are recorded via a camera and microphone. The input is captured as natural language data and sent to an emotion recognition engine. The acquired natural language data, facial expression data, and voice data are then passed to the respective processing modules, where analysis begins.

[1589] Step 2:

[1590] The device's emotion recognition engine analyzes the acquired facial and voice data to identify the user's emotions. For example, if the user is nervous, the device identifies that state of tension and records it as emotion data. The input facial and voice data is converted into emotion data. The output emotion data is "tension."

[1591] Step 3:

[1592] The device converts the user's natural language input data and the acquired emotion data into JSON format. A JSON object is generated with the natural language data and emotion data stored in their respective fields. The generated JSON data is ready to be sent to the server in the next step.

[1593] Step 4:

[1594] The device sends the generated JSON data to the server using the HTTPS protocol. The JSON data, including the user's input amount and emotion data, is securely sent to the server. The server receives the HTTPS request and proceeds to the next step.

[1595] Step 5:

[1596] The server parses the received JSON data and extracts natural language data and emotion data separately. Keywords such as "bolt," "tightening," and "procedure" and the emotion "tension" are obtained from the parsed data. The NLP engine on the server analyzes the meaning and organizes the keywords.

[1597] Step 6:

[1598] The server creates a prompt for the generative AI model based on the extracted keywords and emotion data. The prompt is formatted, for example, as "Tell me how to tighten a bolt," including the word "tension." The prompt is sent to the generative AI model, which then begins generating design data.

[1599] Step 7:

[1600] Based on the prompt received, the generative AI model generates original work guides and design data that take emotions into account. For example, it generates "bolt tightening procedures to relieve tension." The generated design data is then returned to the server.

[1601] Step 8:

[1602] The server converts the generated design data back into JSON format and prepares it for transmission to the device. The data organized into JSON objects is sent to the device via HTTPS.

[1603] Step 9:

[1604] The device parses the JSON data received from the server and extracts the generated work guide and design data, which the device receives, parses, and converts into visual information such as an illustration guide.

[1605] Step 10:

[1606] The device then displays the received design data on the smart glasses display, and the user is shown a "bolt tightening procedure to ease tension," allowing them to proceed with the work with peace of mind.

[1607] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1608] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1609] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1610] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1611] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1612] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1613] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1614] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1615] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1616] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1617] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1618] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1619] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1620] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1621] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1622] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1623] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1624] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1625] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1626] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1627] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1628] The following is further disclosed regarding the above embodiment.

[1629] (Claim 1)

[1630] A means for acquiring preferences and images input by a user in natural language;

[1631] A means of analyzing acquired preferences and images to extract keywords,

[1632] A means for formatting input data into a generative artificial intelligence model using the extracted keywords;

[1633] A means for generating original design data based on the formatted input data using a generative artificial intelligence model;

[1634] means for transmitting the generated design data to a user terminal;

[1635] means for applying the received design data to the terminal;

[1636] A system including:

[1637] (Claim 2)

[1638] 2. The system according to claim 1, wherein natural language data input by a user is converted into JSON format and transmitted.

[1639] (Claim 3)

[1640] 10. The system of claim 1, wherein the system parses and analyzes data received in JSON format.

[1641] "Example 1"

[1642] (Claim 1)

[1643] A means for acquiring preferences and images input by a user in natural language;

[1644] A way to convert the acquired preferences and images into JSON format,

[1645] A means for transmitting the converted JSON format data using the HTTPS protocol;

[1646] A means for receiving and parsing the transmitted JSON format data;

[1647] A means for extracting keywords from the analyzed data;

[1648] A means for formatting input data into a generative artificial intelligence model using the extracted keywords;

[1649] A means for generating original design data based on the formatted input data using a generative artificial intelligence model;

[1650] A means to convert the generated design data back into JSON format and transmit it using the HTTPS protocol;

[1651] means for analyzing and applying the received design data to the terminal;

[1652] A system including:

[1653] (Claim 2)

[1654] 2. The system according to claim 1, wherein the generated design data is applied as wallpaper on the user's terminal.

[1655] (Claim 3)

[1656] 10. The system of claim 1, including a natural language processing engine for analyzing natural language data.

[1657] "Application Example 1"

[1658] (Claim 1)

[1659] A means for acquiring preferences and images input by a user in natural language;

[1660] A means of analyzing acquired preferences and images to extract keywords,

[1661] A means for formatting input data into a generative artificial intelligence model using the extracted keywords;

[1662] A means for generating original design data based on the formatted input data using a generative artificial intelligence model;

[1663] means for transmitting the generated design data to a user terminal;

[1664] means for applying the received design data to the terminal;

[1665] A means for applying the generated design in real time in the virtual space;

[1666] A way to save and share the generated designs;

[1667] A system including:

[1668] (Claim 2)

[1669] 2. The system according to claim 1, wherein natural language data input by a user is converted into JSON format and transmitted.

[1670] (Claim 3)

[1671] 10. The system of claim 1, wherein the system parses and analyzes data received in JSON format.

[1672] "Example 2: Combining Emotion Engines"

[1673] (Claim 1)

[1674] A means for acquiring preferences and images input by a user in natural language;

[1675] A means of analyzing acquired preferences and images to extract keywords,

[1676] A means for integrating the extracted keywords and user emotion data to form input data into a generative artificial intelligence model;

[1677] A means for analyzing the user's facial expressions and voice using a camera and microphone to obtain emotional data;

[1678] A means for generating original design data based on the formatted input data using a generative artificial intelligence model;

[1679] means for transmitting the generated design data to a user terminal;

[1680] means for applying the received design data to the terminal;

[1681] A system including:

[1682] (Claim 2)

[1683] 2. The system according to claim 1, wherein natural language data input by a user is converted into JSON format and transmitted.

[1684] (Claim 3)

[1685] 10. The system of claim 1, wherein the system parses and analyzes data received in JSON format.

[1686] "Application example 2 when combining emotion engines"

[1687] (Claim 1)

[1688] A means for acquiring preferences and images input by a user in natural language;

[1689] A means for extracting keywords by analyzing the acquired preferences, images, and user emotions;

[1690] A means for shaping input data into a generative artificial intelligence model using the extracted keywords and emotions;

[1691] A means for generating original design data based on the formatted input data using a generative artificial intelligence model;

[1692] means for transmitting the generated design data to a user terminal;

[1693] means for applying the received design data to the terminal;

[1694] means for generating and displaying a work guide based on work instructions and emotion data for an industrial device;

[1695] A system including:

[1696] (Claim 2)

[1697] 2. The system according to claim 1, wherein the natural language data and emotion data input by the user are converted into a JSON format and transmitted.

[1698] (Claim 3)

[1699] 10. The system of claim 1, wherein the system parses and analyzes data received in JSON format. [Explanation of symbols]

[1700] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for acquiring preferences and images input by a user in natural language; A means of analyzing acquired preferences and images to extract keywords, A means for formatting input data into a generative artificial intelligence model using the extracted keywords; A means for generating original design data based on the formatted input data using a generative artificial intelligence model; means for transmitting the generated design data to a user terminal; means for applying the received design data to the terminal; A system including:

2. The system according to claim 1, wherein natural language data input by a user is converted into JSON format and transmitted.

3. The system of claim 1 , wherein the system parses and analyzes data received in JSON format.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A